Tells implemented apart from verified: a requirement contract from your own words, a Mirage Detector for fake completeness, evidence from real test/build/lint…

<img src="docs/media/banner.svg" alt="Claude Integrity: a verified requirement, an unverified one, a fake save flagged by the Mirage Detector, and a contradicted 'All tests pass' claim" width="720">
<h1 align="center">Claude Integrity</h1>
<strong>"All tests pass." Did they run? Claude Integrity checks what your coding agent claims against what it actually did.</strong>
<sub>A Claude Code mod that keeps a requirement contract from your own words, flags code that only looks finished, counts a requirement as verified only after a real passing check, and audits completion claims. Everything stays on your machine.</sub>
<a href="https://github.com/NMenzel/claude-integrity-mod/stargazers"><img src="https://img.shields.io/github/stars/NMenzel/claude-integrity-mod?style=flat-square&color=yellow&label=stars" alt="GitHub stars"></a> <a href="https://github.com/NMenzel/claude-integrity-mod/releases/latest"><img src="https://img.shields.io/github/v/release/NMenzel/claude-integrity-mod?style=flat-square&label=version&color=blue" alt="Latest release"></a> <a href="LICENSE"><img src="https://img.shields.io/badge/license-MIT-green?style=flat-square" alt="License: MIT"></a> <a href="https://claude.com/blog/claude-code-mods"><img src="https://img.shields.io/badge/Claude%20Code-2.1.294%2B%20mod-d97757?style=flat-square" alt="Claude Code 2.1.294+ mod"></a>
<a href="#install">Install</a> · <a href="#use-it">Use it</a> · <a href="#the-dashboard">Dashboard</a> · <a href="#what-it-checks">What it checks</a> · <a href="#commands">Commands</a> · <a href="#options">Options</a> · <a href="docs/SECURITY.md">Security</a>
Claude Code ends a task with a confident summary. "Implemented the save flow. All tests pass." Often that is true. Sometimes:
| tail.{ ok: true } and writes nothing. A button's onClick is () => {}. The "live" list still reads mockEvents. A service throws new Error("Not implemented").app/discover-v2/page.tsx. You said to leave app/legacy/** alone, and it was edited.Checking all of that by hand, after every turn, is the work you wanted to hand off.
Claude Integrity compares four things while Claude works, built on Claude Code's native Mods API:
| In Claude Code today | With Claude Integrity |
|---|---|
| "All tests pass." in the final answer | A note beneath the answer: ? "All tests pass.": no test run was observed for this task |
| A test run that scrolled past three edits ago | Evidence tied to the source state. Edits after the run mark it ⧗ stale |
return Response.json({ ok: true }) and nothing saved | MIR-B1: a POST handler returns success with no write, query, fetch or service call |
| "Do not create another Discover page" in your prompt | DRIFT-N1 when a new file is named after a second Discover page |
| Re-reading the diff to see what is really done | A status line, Integrity ✓ 2/5 verified · ? 3 unverified · ! 1 finding · ⧗ 1 stale, and a dashboard per requirement |
| Hoping review catches it | Strict mode refuses edits to excluded paths, and claude -p "/integrity-check" exits 1 on violations |
Nothing is verified because a model said so. A requirement is verified only by a real, passing, current check of a kind that can prove it, or by your own confirmation.
Native mod. No wrappers, no API keys, no network calls. No model calls unless you turn on /integrity-suggest. Two commands to install.
It sees what the Mods API exposes: your prompts, Claude's tool calls and their results, and the files Claude edits. It does not see tests you run in your own terminal or in CI (confirm those with /integrity-verify). Its rules read source text, not a syntax tree, so every finding that is not literal is marked potential. It never approves a tool call and never rewrites Claude's answer.
[!TIP] If Claude Integrity catches a "done" that wasn't, a ⭐ on the repo helps other developers find it.
Claude Integrity is a community project. It is not made or endorsed by Anthropic.
claude -p the hooks still run, and every /integrity-* command answers as text.Inside Claude Code:
/plugin marketplace add NMenzel/claude-integrity-mod
/plugin install integrity@claude-integrity-mod
/reload-plugins
/integrity-help
Or from a terminal, in one command (it adds the marketplace too):
claude plugin install integrity --marketplace NMenzel/claude-integrity-mod
Then run /reload-plugins in an open session, or start a new one.
The installer may say config options aren't set. The defaults are fine, and /config changes them. Mods are an early-access Claude Code feature, and their API can change between releases.
Check that it loaded: run /plugin and look for integrity in the mods active line. You can also run /integrity-help.
To update: claude plugin update integrity@claude-integrity-mod. To remove: claude plugin uninstall integrity@claude-integrity-mod.
Paste this into Claude Code, or any coding agent that can run shell commands:
Install the Claude Integrity mod for Claude Code from
https://github.com/NMenzel/claude-integrity-mod. Follow the
"Instructions for agents" in its README, then tell me what to run next.
If you are an AI agent installing Claude Integrity for a user, follow these steps.
claude --version. Claude Code 2.1.287 or later supports mods, and Claude Integrity is tested on 2.1.294. If the version is older, stop and tell the user to update Claude Code (claude update).sh claude plugin install integrity --marketplace NMenzel/claude-integrity-mod --json ` This adds the marketplace to the user's settings and installs at user scope. Exit code 0 means it is installed. To share it with everyone working in the current repository, add --scope project (it is then recorded in .claude/settings.json`). Ask the user before using a scope other than the default.--config <option>=<value> on the install command. Examples: --config notify=all, --config autoCapture=false. All options are listed under Options.claude plugin list --json includes integrity@claude-integrity-mod./reload-plugins, then /integrity-help. A newly started session loads the mod by itself.Notes for agents:
sudo. Do not clone the repository to install it.claude -p "/integrity-help" prints the help without a model call. In Git Bash on Windows, prefix MSYS_NO_PATHCONV=1 so the shell does not turn /integrity-help into a path./integrity-start <task>./integrity-contract (short: /igc). Edit it in place: ``text /integrity-contract add Events are fetched from /api/events /integrity-contract edit R2 Keep the existing /discover route /integrity-contract crit R2 should # must | should | optional /integrity-contract check R1 npm test # this command proves R1 /integrity-contract exclude app/legacy/** # never touch these paths /integrity-contract intent prototype # fixtures are expected, not findings ``/integrity-approve (/iga). That freezes the baseline. Later changes become versioned revisions.Integrity ✓ 2/5 verified · ? 3 unverified · ! 1 finding · ⧗ 1 stale./integrity, or /ig. Keys 1-7 switch between Overview, Requirements, Evidence, Findings, Claims, Activity and Report. Press a row for details. k acknowledges a finding, w waives it (a reason is required), c re-checks, e/j export./integrity-report export (/igr export). It writes .claude/integrity/<task>.md and .json./ig opens it. It is drawn in the style of Claude Flightdeck, and docked beside a fullscreen transcript it asks for the same 66 columns:
INTEGRITY · STRICT MODE · APPROVED V2, over a color legend.▰▰▰▱▱ 3/5 verified · must 2/3 gauge, and the open problems.i inspects it), a approve, c re-check, and the activity log.From 110 columns the cards sit side by side in two columns. Inline above the prompt it is a three-line summary. With the color option off, the glyphs (✓ ◐ ? ⧗ ✗) carry every state on their own.
Each requirement is your own sentence, with a criticality (must, should, optional), a category (functional, constraint, non-functional, test, preservation) and how it can be proved (static, test, runtime, manual). Exclusions (exclude <glob>), scope (scope <glob>) and intent (production, prototype) shape how strict the checks are.
The contract lives in the mod's own store, keyed per repository, outside the conversation. Compaction cannot drop it, and a new session picks it up.
Twelve rules for JavaScript, TypeScript, React and Next.js, each with its confidence:
throw new Error("Not implemented") (confirmed), TODOs that defer the real work, example.com endpoints and YOUR_API_KEY.setTimeout delay followed by a success state.onClick={() => {}}, log-only handlers, href="#".mockEvents in a production view (expected in a prototype)./api/x that no observed route serves, a save that only sets state, a fetch with no error handling, role-gated UI over unguarded handlers.And seven intent drift rules: edits to excluded paths, edits outside the scope, a new file named after something you said not to recreate, a new route while routes must be preserved, deleted code files, near-total rewrites, and new dependencies the contract never mentions.
Every finding is confirmed, potential (it says what to check) or expected (allowed by the contract, shown but never counted). Acknowledged and waived findings stay quiet across later scans. The full list, with each rule's false-positive guard, is in docs/DETECTION-RULES.md.
Test, build, lint, typecheck, migration and browser-test runs that Claude makes become evidence: the command, the outcome, the counts the runner printed (jest, vitest, mocha, pytest, cargo, go, node:test), and the git HEAD they ran against.
|| true is never a pass./integrity-verify R3 submitted the form and reloaded: saved.When Claude finishes, Integrity reads the final answer for completion claims ("all tests pass", "it's implemented", "uses live data", "production-ready") and checks each against the evidence: supported, partial, unsupported, contradicted or not assessable. When one is unsupported or contradicted, a short note appears beneath the answer:
Integrity · 2 completion claims checked: 1 partial · 1 unsupported
? "All tests pass.": no test run was observed for this task
details: /integrity claims
The answer itself is never changed. Hedged ("should pass") and negated sentences are skipped on purpose: a missed claim is better than a false alarm.
/integrity-check exit 1 on policy violations. It cannot stop a turn from completing, because the Mods API has no such hook. It says so instead of pretending.Set it per project with /integrity-mode observe|review|strict.
To use the check as a gate in a script or a git hook on the machine where the contract lives:
claude -p "/integrity-check" # exit 1 under strict mode when a policy is violated
| Command | Does | | - | - | | /integrity [view] | Open the pane (in a headless session: print the status). /integrity <subcommand> runs /integrity-<subcommand> | | /integrity-start [text] | Start a contract from the text, or from your last prompt | | /integrity-contract [action] | View or edit the contract (add, edit, remove, crit, category, verify-by, check, exclude, scope, assume, intent, summary) | | /integrity-approve | Approve the baseline or a revision | | /integrity-check | Re-read tracked files, re-scan and evaluate policies. Exits 1 under strict mode on violations | | /integrity-findings [all\|ack <id>\|reopen <id>] | List or triage findings | | /integrity-evidence [link E# R#] | List evidence, or link a check to a requirement | | /integrity-verify R# <what you checked> | Record your own confirmation | | /integrity-report [md\|json] [export] | Print or export the delivery report | | /integrity-mode observe\|review\|strict | Set the mode for this project | | /integrity-waive <id\|R#> <reason> | Waive a finding or requirement, with a reason | | /integrity-reset [all] [confirm] | Delete the active task (or all of this project's data) after confirmation | | /integrity-ui [setting value ...] | Status hint, notifications, color, findings per page, open or close the pane | | /integrity-layout auto\|mini\|compact\|wide | Pane layout | | /integrity-suggest | Opt-in AI requirement suggestions: one model call, results marked proposed | | /integrity-help | The workflow in short | | /ig [view\|subcommand] | Short for /integrity: /ig check, /ig start <task> and every other subcommand work too | | /igc · /iga · /igv · /igf · /igr | Short for /integrity-contract, /integrity-approve, /integrity-verify, /integrity-findings and /integrity-report |
Set them in /config, or at install time with claude plugin install ... --config <option>=<value>:
| Option | Default | Meaning | | - | - | - | | notify | important | Hints and notes beneath answers: off, important (only unsupported or contradicted claims) or all | | statusHint | true | The one-line status under the prompt | | paneAutoOpen | false | Open the pane when a session starts (only where it docks as a sidebar) | | layout | auto | Pane layout: auto, mini, compact or wide | | maxFindings | 8 | Findings per pane page, 1 to 50 | | color | true | Color the status glyphs (the glyphs always carry the meaning) | | autoCapture | true | Draft a contract from an implementation prompt when no task is active. Never approves it | | retentionDays | 30 | Idle tasks older than this are deleted, 1 to 3650 | | gitProvenance | true | Record the git HEAD with each piece of evidence (git rev-parse, read-only) | | aiAssist | false | Allow /integrity-suggest to make one model call per request |
/integrity-ui and /integrity-layout change the display options without leaving the session.
Local only: no telemetry and no network calls. The only processes it starts are two read-only git commands for provenance (off with gitProvenance). It stores a hash of your prompt and a short redacted excerpt, never the full prompt and never raw tool output. Credentials, tokens and home-directory paths are redacted. Details: docs/SECURITY.md.
Covered by 66 automated tests (claude plugin test) and a real headless load (claude -p --plugin-dir):
Not covered by tests: how the terminal and Desktop paint the pane (the test kit checks the drawn trees, not the pixels). See docs/LIMITATIONS.md.
npm install # local TypeScript only; nothing global
claude plugin validate . # manifest, hooks, calls, state contract
claude plugin test . # 66 tests: pure engine + real hooks and pane through claude-code/testing
npx tsc -p . # type-check (after one load has laid .claude-plugin/types)
The engine writes .claude-plugin/types/ the first time it loads the folder. To lay it without starting an interactive session, run once: claude -p --plugin-dir . "/integrity-help".
See docs/ARCHITECTURE.md, docs/DETECTION-RULES.md, docs/API-COMPATIBILITY.md, docs/SECURITY.md and docs/LIMITATIONS.md. CONTRIBUTING.md has the ground rules; changes are listed in CHANGELOG.md.
hooks/register.tsx 780 lines1// Claude Integrity: the Claude Code Mod adapter. It observes prompts, tool
2// calls and turn completions, keeps the engine's state in `$.store`, and draws
3// the status line and the /integrity pane. All judgement lives in the pure
4// engine (../src/engine); this file does IO and nothing else.
5
6import { atom, read, update } from 'claude-code'
7import type { EngineInterface, PluginOptions, Register } from 'claude-code'
8
9import type { IntegrityUi } from '../types'
10import * as E from '../src/engine/index'
11import { renderPane, type PaneActions } from './pane'
12
13const PANE = 'integrity'
14const rev = atom({ plugin: 'integrity', key: 'rev' } as const, 0)
15const uiState = atom({ plugin: 'integrity', key: 'ui' } as const, { tab: 'overview', page: 0 } as IntegrityUi)
16
17export type Prefs = {
18 notify: E.NotifyLevel
19 statusHint: boolean
20 paneAutoOpen: boolean
21 layout: E.Layout
22 maxFindings: number
23 color: boolean
24 autoCapture: boolean
25 retentionDays: number
26 gitProvenance: boolean
27 aiAssist: boolean
28}
29
30function prefsFrom(options: PluginOptions, stored: Partial<Prefs> | undefined): Prefs {
31 const pick = <T,>(key: keyof Prefs, fallback: T, ok: (v: unknown) => boolean): T => {
32 const s = stored?.[key]
33 if (s !== undefined && ok(s)) return s as T
34 const o = options[key]
35 return o !== undefined && ok(o) ? (o as T) : fallback
36 }
37 const isBool = (v: unknown) => typeof v === 'boolean'
38 const isNum = (v: unknown) => typeof v === 'number' && Number.isFinite(v) && v > 0
39 return {
40 notify: pick('notify', 'important', v => v === 'off' || v === 'important' || v === 'all'),
41 statusHint: pick('statusHint', true, isBool),
42 paneAutoOpen: pick('paneAutoOpen', false, isBool),
43 layout: pick('layout', 'auto', v => v === 'auto' || v === 'mini' || v === 'compact' || v === 'wide'),
44 maxFindings: Math.min(50, Math.round(pick('maxFindings', 8, isNum))),
45 color: pick('color', true, isBool),
46 autoCapture: pick('autoCapture', true, isBool),
47 retentionDays: pick('retentionDays', 30, isNum),
48 gitProvenance: pick('gitProvenance', true, isBool),
49 aiAssist: pick('aiAssist', false, isBool),
50 }
51}
52
53const COMMANDS: readonly { name: string; description: string; argumentHint?: string }[] = [
54 { name: 'integrity', description: 'Claude Integrity: open the dashboard (or /integrity <subcommand>)', argumentHint: '[help|status|...]' },
55 { name: 'integrity-start', description: 'Integrity: start a task contract from text or your last prompt', argumentHint: '[task description]' },
56 { name: 'integrity-contract', description: 'Integrity: view or edit the contract (add, edit, remove, crit, check, exclude, scope, intent, ...)', argumentHint: '[add|edit|remove|crit|check|exclude|scope|assume|intent|implemented] ...' },
57 { name: 'integrity-approve', description: 'Integrity: approve the contract as the baseline' },
58 { name: 'integrity-check', description: 'Integrity: re-check tracked files and evaluate policies (exit 1 in strict mode on violations)' },
59 { name: 'integrity-findings', description: 'Integrity: list findings, or ack/reopen one', argumentHint: '[all | ack <id> | reopen <id>]' },
60 { name: 'integrity-evidence', description: 'Integrity: list evidence, or link one to a requirement', argumentHint: '[link <E#> <R#>]' },
61 { name: 'integrity-verify', description: 'Integrity: record your own manual confirmation of a requirement', argumentHint: '<R#> <what you checked>' },
62 { name: 'integrity-report', description: 'Integrity: print the delivery report, or export it to files', argumentHint: '[md|json] [export]' },
63 { name: 'integrity-mode', description: 'Integrity: set the mode (observe, review, strict)', argumentHint: '[observe|review|strict]' },
64 { name: 'integrity-waive', description: 'Integrity: waive a finding or requirement, with a reason', argumentHint: '<finding-id|R#> <reason>' },
65 { name: 'integrity-reset', description: 'Integrity: delete the active task (or all stored data) after confirmation', argumentHint: '[all] [confirm]' },
66 { name: 'integrity-ui', description: 'Integrity: configure the status hint, pane and notifications', argumentHint: '[status on|off] [pane open|close] [autoopen on|off] [notify off|important|all] [color on|off] [max N]' },
67 { name: 'integrity-layout', description: 'Integrity: set the pane layout', argumentHint: '[auto|mini|compact|wide]' },
68 { name: 'integrity-suggest', description: 'Integrity: suggest requirements with one opt-in model call (proposed, needs approval)' },
69 { name: 'integrity-help', description: 'Integrity: how it works and every command' },
70]
71
72// Short aliases, like Claude DevTools' /bp family: /ig is /integrity (so /ig check, /ig approve ... work too).
73const SHORTCUTS: readonly { name: string; to: string; description: string; argumentHint?: string }[] = [
74 { name: 'ig', to: 'integrity', description: 'Integrity: open the dashboard (= /integrity); /ig check, /ig start ... run a subcommand', argumentHint: '[view|subcommand] ...' },
75 { name: 'igc', to: 'integrity-contract', description: 'Integrity: view or edit the contract (= /integrity-contract)', argumentHint: '[add|edit|remove|crit|check|exclude|scope|intent] ...' },
76 { name: 'iga', to: 'integrity-approve', description: 'Integrity: approve the contract (= /integrity-approve)' },
77 { name: 'igv', to: 'integrity-verify', description: 'Integrity: record your own confirmation (= /integrity-verify)', argumentHint: '<R#> <what you checked>' },
78 { name: 'igf', to: 'integrity-findings', description: 'Integrity: list or triage findings (= /integrity-findings)', argumentHint: '[all | ack <id> | reopen <id>]' },
79 { name: 'igr', to: 'integrity-report', description: 'Integrity: print or export the report (= /integrity-report)', argumentHint: '[md|json] [export]' },
80]
81const ALIASES: Readonly<Record<string, string>> = Object.fromEntries(SHORTCUTS.map(s => [s.name, s.to]))
82
83const HELP = `**Claude Integrity**: never confuse AI-generated implementation with verified completion.
84
85It keeps a *contract* (what you asked for), watches *changes* and *checks* Claude runs, flags *mirages* (code that looks finished but isn't) and audits *completion claims* against evidence. Nothing is "verified" without a real, passing, current check, or your own confirmation.
86
871. Describe a task, or \`/integrity-start <task>\`: a draft contract is captured from your words (no model call).
882. \`/integrity-contract\` to review; \`add\`, \`edit R2 ...\`, \`crit R2 should\`, \`check R1 npm test\`, \`exclude app/legacy/**\`, \`intent prototype\`.
893. \`/integrity-approve\` freezes the baseline; later edits become versioned revisions.
904. Work normally. Test/build/lint/typecheck runs become evidence; later edits make it stale.
915. \`/integrity\` opens the dashboard (keys 1-7 switch views). \`/integrity-report export\` writes Markdown + JSON.
92
93Other: \`/integrity-verify R3 <what you checked>\` · \`/integrity-waive <id> <reason>\` · \`/integrity-findings ack <id>\` · \`/integrity-evidence link E4 R2\` · \`/integrity-mode observe|review|strict\` · \`/integrity-check\` · \`/integrity-ui\` · \`/integrity-layout\` · \`/integrity-suggest\` (opt-in AI) · \`/integrity-reset\`.
94
95Short aliases: \`/ig\` (= \`/integrity\`; \`/ig check\`, \`/ig start ...\`) · \`/igc\` contract · \`/iga\` approve · \`/igv\` verify · \`/igf\` findings · \`/igr\` report.
96
97Modes: **observe** (default) records and notes; **review** adds a summary of open must-requirements and high findings after each answer; **strict** also refuses edits to paths the approved contract excludes and makes \`/integrity-check\` exit 1 on policy violations. Nothing blocks a turn from completing; the Mods API cannot, so Integrity says so instead of pretending.`
98
99type Runtime = {
100 options: PluginOptions
101 project?: E.ProjectRecord
102 loading?: Promise<E.ProjectRecord>
103 key: string
104 root: string
105 session: string
106 prefs: Prefs
107 storedPrefs?: Partial<Prefs>
108 lastStatus?: string
109 dirty: boolean
110 writing?: Promise<void>
111 deleted: Set<string>
112 lastPrompt?: string
113 suggested: Map<string, string[]>
114}
115
116
117function freshRuntime(options: PluginOptions): Runtime {
118 return { options, key: '', root: '', session: 'session', prefs: prefsFrom(options, undefined), dirty: false, deleted: new Set(), suggested: new Map() }
119}
120
121/** Module state, renewed each time the engine loads the module (a hot reload is a fresh load). */
122let rt: Runtime = freshRuntime({})
123
124function now($: EngineInterface): Promise<number> {
125 return $.clock.now()
126}
127
128function rel(path: string): string {
129 return E.normalizePath(path, rt.root)
130}
131
132function ensure($: EngineInterface): Promise<E.ProjectRecord> {
133 if (rt.project !== undefined) return Promise.resolve(rt.project)
134 rt.loading ??= (async () => {
135 const repo = await $.session.repo().catch(() => null)
136 const root = repo?.root ?? (await $.session.root().catch(() => $.session.cwd()))
137 rt.root = root
138 rt.key = E.projectKey(root, repo?.remote)
139 rt.session = await $.session.id().catch(() => 'session')
140 rt.storedPrefs = ((await $.store.get('prefs')) ?? undefined) as Partial<Prefs> | undefined
141 rt.prefs = prefsFrom(rt.options, rt.storedPrefs)
142 const stored = (await $.store.get(rt.key)) as E.ProjectRecord | undefined
143 const label = repo?.remote ?? root.replace(/\\/g, '/').split('/').slice(-2).join('/')
144 rt.project = stored?.schema === 1 ? stored : E.newProject(rt.key, label, rt.session, await now($))
145 return rt.project
146 })()
147 return rt.loading
148}
149
150/** Re-reads, merges and writes; writes coalesce, so a burst of changes is one or two writes. */
151async function writeOnce($: EngineInterface): Promise<void> {
152 const mine = rt.project
153 if (mine === undefined) return
154 const t = await now($)
155 const stored = (await $.store.get(rt.key)) as E.ProjectRecord | undefined
156 const merged = E.mergeProjects(mine, stored, t)
157 for (const id of rt.deleted) delete merged.tasks[id]
158 if (merged.activeTaskId !== undefined && merged.tasks[merged.activeTaskId] === undefined) delete merged.activeTaskId
159 merged.writer = rt.session
160 E.trimProject(merged, t, rt.prefs.retentionDays, 600_000)
161 await $.store.set(rt.key, merged)
162 rt.project = merged
163}
164
165function save($: EngineInterface): Promise<void> {
166 rt.dirty = true
167 rt.writing ??= (async () => {
168 try {
169 while (rt.dirty) {
170 rt.dirty = false
171 await writeOnce($)
172 }
173 } catch (err) {
174 $.ui.log(`integrity: could not save state (${String(err)})`, { to: 'debug' })
175 } finally {
176 rt.writing = undefined
177 }
178 })()
179 return rt.writing
180}
181
182async function refreshStatus($: EngineInterface): Promise<void> {
183 const p = await ensure($)
184 const line = rt.prefs.statusHint ? E.statusLine(E.summarize(E.activeTask(p), p.mode)) : undefined
185 if (line === rt.lastStatus) return
186 rt.lastStatus = line
187 $.ui.status(line)
188}
189
190function showHints($: EngineInterface, hints: readonly E.Hint[]): void {
191 for (const h of hints) {
192 if (rt.prefs.notify === 'off' || (rt.prefs.notify === 'important' && h.level !== 'important')) continue
193 $.ui.toast(h.text, { timeoutMs: 6000 })
194 }
195}
196
197/** After any state change: redraw subscribers, refresh the status line, show hints, persist. */
198async function changed($: EngineInterface, hints: readonly E.Hint[] = []): Promise<void> {
199 await update($, rev, n => n + 1)
200 await refreshStatus($)
201 showHints($, hints)
202 await save($)
203}
204
205/** Docked it asks Flightdeck's width beside the transcript; inline above the prompt, a short block. */
206function openPane($: EngineInterface) {
207 return $.ui.open({ id: PANE, title: 'Integrity', columns: 66, rows: 10 })
208}
209
210async function gitHead($: EngineInterface): Promise<string | undefined> {
211 if (!rt.prefs.gitProvenance) return undefined
212 const run = await $.process.run(['git', 'rev-parse', '--short', 'HEAD'], { timeoutMs: 3000 }).catch(() => undefined)
213 if (run === undefined || run.exitCode !== 0) return undefined
214 const dirty = await $.process.run(['git', 'status', '--porcelain', '--untracked-files=no'], { timeoutMs: 3000 }).catch(() => undefined)
215 return `${run.stdout.trim()}${dirty !== undefined && dirty.stdout.trim() !== '' ? `+dirty:${E.hash(dirty.stdout).slice(0, 6)}` : ''}`
216}
217
218async function readText($: EngineInterface, path: string): Promise<string | undefined> {
219 const text = await $.fs.read(path).catch(() => undefined)
220 return typeof text === 'string' ? text : undefined
221}
222
223function filePath(e: { tool: string } & Record<string, unknown>): string | undefined {
224 const path = e['file_path'] ?? e['notebook_path']
225 return (e.tool === 'Edit' || e.tool === 'Write' || e.tool === 'NotebookEdit' || e.tool === 'MultiEdit') && typeof path === 'string' ? path : undefined
226}
227
228// ---- Commands -------------------------------------------------------------------
229
230function words(args: string): string[] {
231 return args.trim().split(/\s+/).filter(Boolean)
232}
233
234async function requireTask($: EngineInterface): Promise<E.Task | string> {
235 const task = E.activeTask(await ensure($))
236 return task ?? 'No active task. Describe what you want built, or run `/integrity-start <task>`.'
237}
238
239function contractText(task: E.Task, mode: E.Mode): string {
240 const c = task.contract
241 const s = E.summarize(task, mode)
242 const lines = [
243 `**Contract** ${c.summary} · ${s.approval} v${c.version} · intent ${c.intent} · task \`${task.id}\``,
244 `Prompt: "${c.prompt.excerpt}" (hash ${c.prompt.hash.slice(0, 10)})`,
245 '',
246 ...(c.requirements.length === 0 ? ['_No requirements yet: `/integrity-contract add <requirement>`_'] : c.requirements.map(r => {
247 const st = E.requirementState(r, task)
248 return `- **${r.id}** [${r.criticality}/${r.category}/${r.verification}${r.origin === 'proposed' ? '/proposed' : ''}] ${r.description} \n → ${st.state}: ${st.reason}${r.checks.length ? ` · checks: ${r.checks.join(', ')}` : ''}`
249 })),
250 ]
251 if (c.exclusions.length) lines.push('', `Exclusions: ${c.exclusions.map(x => `\`${x}\``).join(', ')}`)
252 if (c.scope.length) lines.push(`Scope: ${c.scope.map(x => `\`${x}\``).join(', ')}`)
253 if (c.assumptions.length) lines.push(`Assumptions: ${c.assumptions.join('; ')}`)
254 if (c.revisions.length) lines.push(`Revisions: ${c.revisions.slice(-5).map(r => `v${r.version} ${r.change}`).join(' · ')}`)
255 lines.push('', s.approval === 'approved' ? 'Approved baseline kept; edits create revisions.' : 'Not approved yet: `/integrity-approve` when it reads right.')
256 return lines.join('\n')
257}
258
259async function editContract($: EngineInterface, args: string): Promise<string> {
260 const task = await requireTask($)
261 if (typeof task === 'string') return task
262 const p = rt.project as E.ProjectRecord
263 const [verb = '', ...rest] = words(args)
264 const t = await now($)
265 const c = task.contract
266 if (verb === '') return contractText(task, p.mode)
267 const id = rest[0] ?? ''
268 const text = rest.slice(1).join(' ')
269 const req = E.findRequirement(c, id)
270 const needReq = (): string | undefined => (req === undefined ? `No requirement "${id}". Requirements: ${c.requirements.map(r => r.id).join(', ') || 'none'}.` : undefined)
271 let done: string
272 switch (verb) {
273 case 'add': {
274 const description = rest.join(' ')
275 if (description === '') return 'Usage: `/integrity-contract add <requirement>`'
276 let added: E.Requirement | undefined
277 E.amend(c, t, `added requirement: ${description}`, x => { added = E.makeRequirement(x.requirements, description); x.requirements.push(added) })
278 done = `Added ${added?.id}.`
279 break
280 }
281 case 'edit':
282 if (needReq() || text === '') return needReq() ?? 'Usage: `/integrity-contract edit R2 <new text>`'
283 E.amend(c, t, `${req!.id} reworded`, () => { req!.description = E.redact(text).slice(0, 300); if (req!.origin === 'proposed') req!.origin = 'manual' })
284 done = `Updated ${req!.id}.`
285 break
286 case 'remove':
287 if (needReq()) return needReq() as string
288 E.amend(c, t, `${req!.id} removed: ${req!.description}`, x => { x.requirements = x.requirements.filter(r => r !== req) })
289 done = `Removed ${req!.id}.`
290 break
291 case 'crit':
292 case 'criticality':
293 if (needReq() || !['must', 'should', 'optional'].includes(text)) return needReq() ?? 'Usage: `/integrity-contract crit R2 must|should|optional`'
294 E.amend(c, t, `${req!.id} criticality ${req!.criticality} → ${text}`, () => { req!.criticality = text as E.Criticality })
295 done = `${req!.id} is now ${text}.`
296 break
297 case 'category':
298 if (needReq() || !['functional', 'constraint', 'non-functional', 'test', 'preservation'].includes(text)) return needReq() ?? 'Usage: `/integrity-contract category R2 functional|constraint|non-functional|test|preservation`'
299 E.amend(c, t, `${req!.id} category → ${text}`, () => { req!.category = text as E.RequirementCategory })
300 done = `${req!.id} category is now ${text}.`
301 break
302 case 'verify-by':
303 if (needReq() || !['static', 'test', 'runtime', 'manual', 'unknown'].includes(text)) return needReq() ?? 'Usage: `/integrity-contract verify-by R2 static|test|runtime|manual`'
304 E.amend(c, t, `${req!.id} verification → ${text}`, () => { req!.verification = text as E.VerificationKind })
305 done = `${req!.id} is verified by ${text} evidence.`
306 break
307 case 'check':
308 if (needReq() || text === '') return needReq() ?? 'Usage: `/integrity-contract check R2 <command substring, e.g. npm test -- events>`'
309 E.amend(c, t, `${req!.id} check: ${text}`, () => { req!.checks = [...new Set([...req!.checks, text])] })
310 done = `${req!.id} is verified by a passing run of a command containing "${text}".`
311 break
312 case 'implemented':
313 if (needReq()) return needReq() as string
314 E.amend(c, t, `${req!.id} marked implemented by the developer`, () => { req!.implemented = true }, false)
315 done = `${req!.id} marked implemented (not verified: that still needs evidence).`
316 break
317 case 'exclude':
318 case 'scope':
319 case 'assume': {
320 const value = rest.join(' ')
321 if (value === '') return `Usage: \`/integrity-contract ${verb} <${verb === 'assume' ? 'text' : 'glob, e.g. app/legacy/**'}>\``
322 const list = verb === 'exclude' ? 'exclusions' : verb === 'scope' ? 'scope' : 'assumptions'
323 E.amend(c, t, `${verb}: ${value}`, x => { x[list] = [...new Set([...x[list], E.redact(value)])] }, verb !== 'assume')
324 done = `Added to ${list}: ${value}`
325 break
326 }
327 case 'intent':
328 if (!['production', 'prototype', 'unknown'].includes(id)) return 'Usage: `/integrity-contract intent production|prototype|unknown`'
329 E.amend(c, t, `intent ${c.intent} → ${id}`, x => { x.intent = id as E.ContractIntent })
330 done = `Intent is now ${id}. Mirage rules read mock data as ${id === 'prototype' ? 'expected' : 'suspicious'}.`
331 break
332 case 'summary':
333 if (rest.length === 0) return 'Usage: `/integrity-contract summary <one line>`'
334 E.amend(c, t, 'summary changed', x => { x.summary = E.redact(rest.join(' ')).slice(0, 140) }, false)
335 done = 'Summary updated.'
336 break
337 default:
338 return `Unknown contract action "${verb}". Try: add, edit, remove, crit, category, verify-by, check, implemented, exclude, scope, assume, intent, summary.`
339 }
340 E.log(task, 'contract', done, t)
341 await changed($)
342 const note = c.baseline !== undefined && c.approval.status !== 'approved' ? ' Contract amended to v' + c.version + ': re-approve with `/integrity-approve`.' : ''
343 return done + note
344}
345
346async function findingsText($: EngineInterface, args: string): Promise<string> {
347 const task = await requireTask($)
348 if (typeof task === 'string') return task
349 const [verb = '', id = ''] = words(args)
350 if (verb === 'ack' || verb === 'reopen') {
351 const f = E.setFindingStatus(task, id, verb === 'ack' ? 'acknowledged' : 'open', await now($))
352 if (f === undefined) return `No finding "${id}".`
353 await changed($)
354 return `${f.ruleId} ${f.file}:${f.line} ${verb === 'ack' ? 'acknowledged' : 'reopened'}.`
355 }
356 const order: Record<E.Severity, number> = { high: 0, medium: 1, low: 2, info: 3 }
357 const list = task.findings
358 .filter(f => verb === 'all' || f.status === 'open')
359 .sort((a, b) => order[a.severity] - order[b.severity])
360 if (list.length === 0) return verb === 'all' ? 'No findings recorded.' : 'No open findings. `/integrity-findings all` lists resolved, waived and acknowledged ones.'
361 return list.slice(0, 40).map(f =>
362 `- \`${f.id.slice(0, 8)}\` **${f.severity}/${f.confidence}** ${f.ruleId} \`${f.file}:${f.line}\` (${f.status}): ${f.explanation}\n Verify: ${f.verify}`,
363 ).join('\n') + (list.length > 40 ? `\n…and ${list.length - 40} more` : '')
364}
365
366async function evidenceText($: EngineInterface, args: string): Promise<string> {
367 const task = await requireTask($)
368 if (typeof task === 'string') return task
369 const [verb = '', eid = '', rid = ''] = words(args)
370 if (verb === 'link') {
371 const ev = task.evidence.find(x => x.id.toLowerCase() === eid.toLowerCase())
372 const req = E.findRequirement(task.contract, rid)
373 if (ev === undefined || req === undefined) return 'Usage: `/integrity-evidence link E4 R2` (both must exist).'
374 req.evidenceIds = [...new Set([...req.evidenceIds, ev.id])]
375 E.log(task, 'requirement', `${ev.id} linked to ${req.id} by the developer`, await now($))
376 await changed($)
377 return `${ev.id} linked to ${req.id}: now ${E.requirementState(req, task).state}.`
378 }
379 if (task.evidence.length === 0) return 'No checks observed yet. Test, build, lint and typecheck runs Claude makes are recorded automatically.'
380 return task.evidence.slice(-30).map(ev => {
381 const fresh = E.freshness(ev, task)
382 const why = fresh === 'stale' ? ` (stale: ${E.staleBecause(ev, task).slice(0, 3).join(', ')} changed)` : ''
383 return `- **${ev.id}** ${ev.check} ${ev.outcome} · ${fresh}${why} · ${ev.command ? `\`${ev.command}\`` : ev.source} · ${ev.basis}${ev.revision ? ` @ ${ev.revision}` : ''}`
384 }).join('\n')
385}
386
387async function check($: EngineInterface): Promise<{ text: string; exitCode?: number }> {
388 const task = await requireTask($)
389 if (typeof task === 'string') return { text: task }
390 const p = rt.project as E.ProjectRecord
391 const current: Record<string, string | null> = {}
392 for (const path of Object.keys(task.files).slice(0, 200)) {
393 const abs = /^([A-Za-z]:)?[\\/]/.test(path) ? path : `${rt.root.replace(/[\\/]$/, '')}/${path}`
394 current[path] = (await $.fs.exists(abs).catch(() => false)) ? (await readText($, abs)) ?? null : null
395 }
396 const t = await now($)
397 const hints = E.reconcileFiles(task, current, t)
398 // Re-scan with the current contract (intent or wording may have changed since).
399 for (const [path, text] of Object.entries(current)) {
400 if (text !== null && /\.(c|m)?(t|j)sx?$/i.test(path)) {
401 const asked = [task.contract.summary, task.contract.prompt.excerpt, ...task.contract.requirements.map(r => r.description)].join(' ')
402 hints.push(...E.applyFindings(task, path, 'mirage', E.scanSource(path, text, task.contract.intent, asked), t))
403 }
404 }
405 const violations = E.policyViolations(task)
406 E.log(task, 'evidence', `re-check: ${Object.keys(current).length} tracked file(s), ${violations.length} policy violation(s)`, t)
407 await changed($, hints)
408 const s = E.summarize(task, p.mode)
409 const lines = [
410 E.statusLine(s, 200) ?? 'Integrity: no status',
411 `Re-read ${Object.keys(current).length} tracked file(s); ${hints.length} new hint(s).`,
412 ...(violations.length ? ['', `**Policy (${p.mode}):**`, ...violations.map(v => `- ✗ ${v.text}`)] : ['Policies: none violated.']),
413 ]
414 return { text: lines.join('\n'), ...(p.mode === 'strict' && violations.length > 0 ? { exitCode: 1 } : {}) }
415}
416
417async function exportReport($: EngineInterface, format: 'md' | 'json' | 'both'): Promise<string> {
418 const task = await requireTask($)
419 if (typeof task === 'string') return task
420 const p = rt.project as E.ProjectRecord
421 const t = await now($)
422 const dir = `${rt.root.replace(/[\\/]$/, '')}/.claude/integrity`
423 const written: string[] = []
424 if (format !== 'json') {
425 await $.fs.write(`${dir}/${task.id}.md`, E.toMarkdown(task, p.mode, p.label, t))
426 written.push(`.claude/integrity/${task.id}.md`)
427 }
428 if (format !== 'md') {
429 await $.fs.write(`${dir}/${task.id}.json`, JSON.stringify(E.toJson(task, p.mode, p.label, t), null, 2))
430 written.push(`.claude/integrity/${task.id}.json`)
431 }
432 E.log(task, 'task', `report exported: ${written.join(', ')}`, t)
433 await changed($)
434 return `Exported ${written.join(' and ')}.`
435}
436
437async function suggest($: EngineInterface): Promise<string> {
438 const task = await requireTask($)
439 if (typeof task === 'string') return task
440 if (!rt.prefs.aiAssist) {
441 return 'AI suggestions are off. Turn on "AI requirement suggestions" for integrity in /config, then run `/integrity-suggest` again. Each run makes one model call (about 1-2k tokens) on this session\'s own model provider; nothing goes to any other service.'
442 }
443 const c = task.contract
444 const source = rt.lastPrompt !== undefined && E.hash(rt.lastPrompt) === c.prompt.hash ? rt.lastPrompt : c.prompt.excerpt
445 const cacheKey = E.hash(source)
446 let lines = rt.suggested.get(cacheKey)
447 let cost = 'cached, no model call'
448 if (lines === undefined) {
449 const r = await $.model.complete({
450 model: 'haiku',
451 system: 'You list testable software requirements. Output 3 to 7 lines, each starting with "- ". No commentary. Do not invent features the request does not imply.',
452 prompt: `Developer request (secrets redacted):\n${E.redact(source).slice(0, 4000)}\n\nAlready recorded:\n${c.requirements.map(r => `- ${r.description}`).join('\n') || '(none)'}\n\nList missing acceptance criteria.`,
453 maxTokens: 500,
454 effort: 'low',
455 timeoutMs: 30_000,
456 })
457 if (!r.isAnswered) return `No suggestions: the model call ended with ${r.reason}. Nothing was changed.`
458 lines = r.text.split('\n').map(l => l.replace(/^\s*[-*•]\s*/, '').trim()).filter(l => l.length > 8).slice(0, 7)
459 rt.suggested.set(cacheKey, lines)
460 cost = `${r.usage.input_tokens} input + ${r.usage.output_tokens} output tokens`
461 }
462 if (lines.length === 0) return 'The model proposed nothing new.'
463 const t = await now($)
464 E.amend(c, t, `${lines.length} proposed requirement(s) added (model suggestion)`, x => {
465 for (const l of lines!) x.requirements.push(E.makeRequirement(x.requirements, l, { origin: 'proposed', criticality: 'should' }))
466 })
467 E.log(task, 'contract', `${lines.length} model-proposed requirement(s) added; awaiting approval`, t)
468 await changed($)
469 return `Added ${lines.length} **proposed** requirement(s) (${cost}). They count only once you approve the contract; remove any with \`/integrity-contract remove R#\`.\n\n${lines.map(l => `- ${l}`).join('\n')}`
470}
471
472async function savePrefs($: EngineInterface, change: Partial<Prefs>): Promise<void> {
473 rt.storedPrefs = { ...rt.storedPrefs, ...change }
474 await $.store.set('prefs', rt.storedPrefs)
475 rt.prefs = prefsFrom(rt.options, rt.storedPrefs)
476 rt.lastStatus = '\u0000'
477 await changed($)
478}
479
480async function run($: EngineInterface, command: string, args: string, presentation: { isFullscreen: boolean }): Promise<{ text: string; exitCode?: number }> {
481 const p = await ensure($)
482 const t = await now($)
483 switch (command) {
484 case 'integrity': {
485 const [sub = '', ...rest] = words(args)
486 if (sub !== '' && COMMANDS.some(c => c.name === `integrity-${sub}`)) return run($, `integrity-${sub}`, rest.join(' '), presentation)
487 if (sub === 'status' || sub === 'claims' || sub === '' || ['overview', 'requirements', 'evidence', 'findings', 'activity', 'report'].includes(sub)) {
488 const tab = (['overview', 'requirements', 'evidence', 'findings', 'claims', 'activity', 'report'].includes(sub) ? sub : 'overview') as E.Tab
489 if (sub !== 'status') {
490 await update($, uiState, u => ({ ...u, tab, page: 0 }))
491 void openPane($)
492 }
493 const s = E.summarize(E.activeTask(p), p.mode)
494 if (!s.hasTask) return { text: 'Claude Integrity: no task contract yet. Describe what you want built, or `/integrity-start <task>`. `/integrity-help` explains.' }
495 const claims = s.claims.latest.map(c => `- ${E.GLYPH[c.verdict]} ${c.verdict}: "${c.text}" (${c.reason})`)
496 return { text: [E.statusLine(s, 200), s.concern ? `Top concern: ${s.concern.text}` : 'No outstanding concern.', ...(sub === 'claims' ? ['', ...claims] : [])].join('\n') }
497 }
498 return { text: `Unknown subcommand "${sub}". \`/integrity-help\` lists them.` }
499 }
500 case 'integrity-start': {
501 const text = args.trim() !== '' ? args : rt.lastPrompt ?? ''
502 if (text === '') return { text: 'Usage: `/integrity-start <what you want built>` (or run it right after describing the task).' }
503 const old = E.activeTask(p)
504 const task = E.startTask(p, text, args.trim() !== '' ? 'manual' : 'prompt', t)
505 if (old) E.log(task, 'task', `previous task ${old.id} kept in history`, t)
506 await changed($)
507 return { text: `Started task \`${task.id}\`.\n\n${contractText(task, p.mode)}` }
508 }
509 case 'integrity-contract':
510 return { text: await editContract($, args) }
511 case 'integrity-approve': {
512 const task = await requireTask($)
513 if (typeof task === 'string') return { text: task }
514 if (task.contract.approval.status === 'approved') return { text: `Already approved (v${task.contract.version}).` }
515 if (task.contract.requirements.length === 0) return { text: 'Nothing to approve: add at least one requirement (`/integrity-contract add ...`).' }
516 const first = task.contract.baseline === undefined
517 E.approve(task.contract, t)
518 E.log(task, 'approval', `contract ${first ? 'approved as baseline' : 're-approved'} at v${task.contract.version}`, t)
519 await changed($)
520 const proposed = task.contract.requirements.filter(r => r.origin === 'proposed').length
521 return { text: `Contract ${first ? 'approved: baseline frozen' : 're-approved'} at v${task.contract.version} (${task.contract.requirements.length} requirements${proposed ? `, ${proposed} model-proposed` : ''}).` }
522 }
523 case 'integrity-check':
524 return check($)
525 case 'integrity-findings':
526 return { text: await findingsText($, args) }
527 case 'integrity-evidence':
528 return { text: await evidenceText($, args) }
529 case 'integrity-verify': {
530 const task = await requireTask($)
531 if (typeof task === 'string') return { text: task }
532 const [id = '', ...note] = words(args)
533 const req = E.findRequirement(task.contract, id)
534 if (req === undefined || note.length === 0) return { text: 'Usage: `/integrity-verify R3 <what you checked, e.g. "submitted the form and reloaded: saved">`' }
535 const ev = E.recordManual(task, req.id, note.join(' '), t)
536 await changed($)
537 return { text: `${ev.id}: your confirmation of ${req.id} recorded (${E.requirementState(req, task).state}). It goes stale if related code changes later.` }
538 }
539 case 'integrity-report': {
540 const task = await requireTask($)
541 if (typeof task === 'string') return { text: task }
542 const w = words(args)
543 const format = w.includes('json') ? 'json' : w.includes('md') ? 'md' : 'both'
544 if (w.includes('export')) return { text: await exportReport($, format) }
545 return { text: format === 'json' ? '```json\n' + JSON.stringify(E.toJson(task, p.mode, p.label, t), null, 2).slice(0, 60_000) + '\n```' : E.toMarkdown(task, p.mode, p.label, t) }
546 }
547 case 'integrity-mode': {
548 const mode = words(args)[0]
549 if (mode === undefined) return { text: `Mode: **${p.mode}**. Options: observe (default, never blocks), review (summary after answers), strict (refuses edits to excluded paths; /integrity-check exits 1 on violations).` }
550 if (mode !== 'observe' && mode !== 'review' && mode !== 'strict') return { text: 'Usage: `/integrity-mode observe|review|strict`' }
551 p.mode = mode
552 const task = E.activeTask(p)
553 if (task) E.log(task, 'mode', `mode → ${mode}`, t)
554 await changed($)
555 const strictNote = mode === 'strict' && (task?.contract.baseline === undefined) ? ' Strict policies apply once a contract is approved.' : ''
556 return { text: `Mode set to **${mode}** for this project.${strictNote}` }
557 }
558 case 'integrity-waive': {
559 const task = await requireTask($)
560 if (typeof task === 'string') return { text: task }
561 const [id = '', ...reason] = words(args)
562 if (id === '' || reason.length === 0) return { text: 'Usage: `/integrity-waive <finding-id|R#> <reason>`. A reason is required.' }
563 const req = E.findRequirement(task.contract, id)
564 if (req !== undefined) {
565 E.amend(task.contract, t, `${req.id} waived: ${reason.join(' ')}`, () => { req.waiver = { reason: E.redact(reason.join(' ')), at: t } }, false)
566 E.log(task, 'waiver', `${req.id} waived: ${reason.join(' ')}`, t)
567 await changed($)
568 return { text: `${req.id} waived: ${reason.join(' ')}` }
569 }
570 const f = E.setFindingStatus(task, id, 'waived', t, reason.join(' '))
571 if (f === undefined) return { text: `No finding or requirement "${id}". \`/integrity-findings\` lists ids.` }
572 await changed($)
573 return { text: `Waived ${f.ruleId} at ${f.file}:${f.line}: ${reason.join(' ')}. The same finding stays waived on later scans.` }
574 }
575 case 'integrity-reset': {
576 const w = words(args)
577 const all = w.includes('all')
578 const task = E.activeTask(p)
579 if (!all && task === undefined) return { text: 'No active task to reset.' }
580 const what = all ? `all Integrity data for this project (${Object.keys(p.tasks).length} task(s))` : `task \`${task!.id}\` (${task!.contract.summary})`
581 let confirmed = w.includes('confirm')
582 if (!confirmed) {
583 const answer = await $.ui.ask(`Delete ${what}? This cannot be undone.`, ['Delete', 'Keep']).catch(() => undefined)
584 if (answer === undefined) return { text: `This deletes ${what}. Run \`/integrity-reset ${all ? 'all ' : ''}confirm\` to proceed.` }
585 confirmed = answer === 'Delete'
586 }
587 if (!confirmed) return { text: 'Kept. Nothing was deleted.' }
588 const ids = all ? Object.keys(p.tasks) : [task!.id]
589 for (const id of ids) { rt.deleted.add(id); delete p.tasks[id] }
590 delete p.activeTaskId
591 await changed($)
592 if (all) { await $.store.delete(rt.key); rt.project = E.newProject(rt.key, p.label, rt.session, t); rt.deleted.clear() }
593 return { text: `Deleted ${what}.` }
594 }
595 case 'integrity-ui': {
596 const w = words(args)
597 if (w.length === 0) return { text: `Status hint ${rt.prefs.statusHint ? 'on' : 'off'} · pane auto-open ${rt.prefs.paneAutoOpen ? 'on' : 'off'} · notify ${rt.prefs.notify} · color ${rt.prefs.color ? 'on' : 'off'} · max findings ${rt.prefs.maxFindings} · layout ${rt.prefs.layout}` }
598 const change: Partial<Prefs> = {}
599 for (let i = 0; i < w.length; i += 2) {
600 const [k, v] = [w[i], w[i + 1]]
601 if (k === 'status') change.statusHint = v === 'on'
602 else if (k === 'autoopen') change.paneAutoOpen = v === 'on'
603 else if (k === 'color') change.color = v === 'on'
604 else if (k === 'notify' && (v === 'off' || v === 'important' || v === 'all')) change.notify = v
605 else if (k === 'max' && Number(v) > 0) change.maxFindings = Number(v)
606 else if (k === 'pane' && v === 'open') void openPane($)
607 else if (k === 'pane' && v === 'close') await $.ui.close({ id: PANE }).catch(() => undefined)
608 else return { text: `Unknown setting "${k} ${v ?? ''}". \`/integrity-ui\` with no arguments shows the current ones.` }
609 }
610 if (Object.keys(change).length > 0) await savePrefs($, change)
611 return { text: `Saved: ${Object.entries(change).map(([k, v]) => `${k}=${v}`).join(', ') || 'pane toggled'}.` }
612 }
613 case 'integrity-layout': {
614 const l = words(args)[0]
615 if (l !== 'auto' && l !== 'mini' && l !== 'compact' && l !== 'wide') return { text: `Layout: ${rt.prefs.layout}. Usage: \`/integrity-layout auto|mini|compact|wide\`` }
616 await savePrefs($, { layout: l })
617 return { text: `Layout set to ${l}.` }
618 }
619 case 'integrity-suggest':
620 return { text: await suggest($) }
621 case 'integrity-help':
622 return { text: HELP }
623 }
624 return { text: `Unknown command ${command}` }
625}
626
627// ---- The pane -----------------------------------------------------------------
628
629function actions($: EngineInterface): PaneActions {
630 return {
631 setUi: fn => update($, uiState, fn),
632 approve: async () => (await run($, 'integrity-approve', '', { isFullscreen: false })).text,
633 recheck: async () => (await check($)).text.split('\n')[1] ?? 'Re-checked.',
634 exportReport: format => exportReport($, format),
635 setFinding: async (id, status, reason) => {
636 const task = E.activeTask(await ensure($))
637 if (task === undefined) return 'No active task.'
638 const f = E.setFindingStatus(task, id, status, await now($), reason)
639 await changed($)
640 return f ? `${f.ruleId} ${status}.` : 'Finding not found.'
641 },
642 waiveRequirement: async (id, reason) => (await run($, 'integrity-waive', `${id} ${reason}`, { isFullscreen: false })).text,
643 close: () => $.ui.close({ id: PANE }),
644 }
645}
646
647export const register: Register = (on, options) => {
648 rt = freshRuntime(options)
649
650 // ---- lifecycle -------------------------------------------------------------
651
652 on('session.start', async ($, e, next) => {
653 await Promise.all(COMMANDS.map(c => $.command.register(c)))
654 await Promise.all(SHORTCUTS.map(({ to: _, ...c }) => $.command.register(c)))
655 await ensure($)
656 await refreshStatus($)
657 if (rt.prefs.paneAutoOpen && e.isInteractive && E.activeTask(rt.project as E.ProjectRecord) !== undefined) {
658 void openPane($)
659 }
660 return next(e)
661 })
662
663 on('session.end', async ($, e, next) => {
664 if (rt.dirty || rt.writing !== undefined) await save($)
665 return next(e)
666 })
667
668 // ---- Intent Fingerprint: capture --------------------------------------------
669
670 on('prompt.submit', async ($, e, next) => {
671 const human = e.origin.kind === 'composer' || e.origin.kind === 'bridge' || e.origin.kind === 'sdk'
672 if (human && !e.text.trim().startsWith('/')) {
673 rt.lastPrompt = e.text
674 const p = await ensure($)
675 if (rt.prefs.autoCapture && E.activeTask(p) === undefined && E.isMeaningfulTask(e.text)) {
676 const task = E.startTask(p, e.text, 'prompt', await now($))
677 await changed($, rt.prefs.notify === 'all' ? [{ key: `draft:${task.id}`, level: 'info', text: `Integrity: drafted a contract (${task.contract.requirements.length} explicit requirement(s)) · /integrity-contract` }] : [])
678 }
679 }
680 return next(e)
681 })
682
683 // ---- The Witness: observe tool calls ----------------------------------------
684
685 on('tool.call', async ($, e, next) => {
686 const args = e as unknown as { tool: string } & Record<string, unknown>
687 const p = await ensure($)
688 const task = E.activeTask(p)
689 const target = filePath(args)
690 if (task !== undefined && target !== undefined) {
691 const glob = E.deniedPath(task, p.mode, rel(target))
692 if (glob !== undefined) {
693 E.log(task, 'finding', `strict: refused ${args.tool} of ${rel(target)} (excluded by "${glob}")`, await now($))
694 await changed($)
695 return { deny: `Claude Integrity (strict mode): ${rel(target)} is excluded by the approved contract ("${glob}"). Amend the contract (/integrity-contract) or change mode (/integrity-mode) to edit it.` }
696 }
697 }
698
699 const res = await next(e)
700 if (task === undefined) return res
701 try {
702 const t = await now($)
703 const ranInCore = res.ref !== undefined
704 const result = res.result as Record<string, unknown> | undefined
705 const hints: E.Hint[] = []
706
707 if (target !== undefined) {
708 // A refused, failed, staged or plugin-answered write changed nothing core can vouch for.
709 if (res.deny === undefined && res.isError !== true && ranInCore && result?.['staged'] !== true) {
710 const path = rel(target)
711 const after = args.tool === 'Write' && typeof args['content'] === 'string' ? args['content'] : await readText($, target)
712 const before = typeof result?.['originalFile'] === 'string' ? (result['originalFile'] as string) : result?.['originalFile'] === null ? null : undefined
713 const kind = args.tool === 'Write' && result?.['type'] === 'create' ? 'create' : 'update'
714 hints.push(...E.recordChange(task, { path, kind, ...(after === undefined ? {} : { after }), ...(before === undefined ? {} : { before }) }, t))
715 }
716 } else if (args.tool === 'Bash' || /^mcp__.*(playwright|chrome|browser|puppeteer)/i.test(args.tool)) {
717 // File changes a shell command made come first: a check in the same line ran after them.
718 const diff = result?.['bashEditDiff'] as { files?: { filePath: string; created?: true; deleted?: true }[] } | undefined
719 for (const f of (ranInCore ? diff?.files ?? [] : []).slice(0, 30)) {
720 const after = f.deleted ? undefined : await readText($, f.filePath)
721 hints.push(...E.recordChange(task, { path: rel(f.filePath), kind: f.deleted ? 'delete' : f.created ? 'create' : 'update', ...(after === undefined ? {} : { after }) }, t))
722 }
723 const obs: E.ToolObservation = {
724 tool: args.tool, input: args, ranInCore, isError: res.isError === true, text: res.text ?? '', result: res.result,
725 ...(res.deny === undefined ? {} : { denied: res.deny }),
726 ...(e.agentId === undefined ? {} : { agentId: e.agentId }),
727 cwd: await $.session.cwd().catch(() => rt.root),
728 }
729 const recorded = E.recordTool(task, obs, t)
730 if (recorded.evidence.length > 0) {
731 const head = ranInCore ? await gitHead($) : undefined
732 if (head !== undefined) for (const ev of recorded.evidence) ev.revision = head
733 }
734 hints.push(...recorded.hints)
735 if (recorded.evidence.length === 0 && hints.length === 0 && (diff?.files ?? []).length === 0) return res
736 } else return res
737 await changed($, hints)
738 } catch (err) {
739 $.ui.log(`integrity: could not record ${args.tool} (${String(err)})`, { to: 'debug' })
740 }
741 return res
742 })
743
744 // ---- Completion Claim Auditor -------------------------------------------------
745
746 on('turn.complete', async ($, e, next) => {
747 const res = await next(e)
748 if (e.agentId !== undefined) return res
749 const p = await ensure($)
750 const task = E.activeTask(p)
751 if (task === undefined) return res
752 const t = await now($)
753 if (e.reason === 'aborted' || e.isAborted) {
754 E.log(task, 'turn', 'turn interrupted: its claims were not audited', t)
755 await changed($)
756 return res
757 }
758 if (e.reason !== 'answer') return res
759 const { notice } = E.recordAudit(task, e.answer, e.turnId, t, p.mode, rt.prefs.notify)
760 await changed($)
761 if (p.mode === 'review' && notice !== undefined) void openPane($)
762 if (notice === undefined) return res
763 // Another plugin's line beneath the answer stays; ours follows it. The answer itself is never rewritten.
764 const theirs = res.text !== e.answer ? `${res.text}\n` : ''
765 return { ...res, text: `${theirs}${notice}` }
766 })
767
768 // ---- Commands and pane ---------------------------------------------------------
769
770 on('command.run', { command: [...COMMANDS, ...SHORTCUTS].map(c => c.name) }, async ($, e) => run($, ALIASES[e.command] ?? e.command, e.args, e.presentation))
771
772 on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
773 await read($, rev)
774 const ui = await read($, uiState)
775 const p = await ensure($)
776 const view = { surface: e.surface, bodyColumns: e.props.bodyColumns, columns: e.viewport?.columns, placement: e.props.placement }
777 return renderPane($.ui.resolve(e), view, { project: p, task: E.activeTask(p), prefs: rt.prefs, ui, actions: actions($) })
778 })
779}
780src/engine/index.ts 283 lines1// The engine's state transitions. Pure: every function takes the state and the
2// time, mutates the task in place, and returns the hints to show. The Mod
3// adapter does all IO; a CLI or CI job can drive the same functions.
4
5import type {
6 Activity, ActivityKind, Evidence, FileChange, FileState, Finding, Hint, Mode, ProjectRecord, Task, ToolObservation,
7} from './types'
8import { createContract } from './contract'
9import { evidenceFromBash, evidenceFromBrowserTool, freshness, isProof, staleBecause } from './evidence'
10import { apiCalls, routeMatches, routeOf, scanSource, type RawFinding } from './mirage'
11import { driftFindings } from './drift'
12import { auditAnswer } from './claims'
13import { requirementState } from './status'
14import { clip, hash, isScannable, newId, redact } from './util'
15
16export * from './types'
17export { summarize, statusLine, type Summary, type Tab } from './summary'
18export { requirementState, policyViolations, deniedPath, linkedEvidence } from './status'
19export { freshness, staleBecause, isProof, classifyCommand, parseCounts } from './evidence'
20export { createContract, approve, amend, makeRequirement, findRequirement, isMeaningfulTask, detectIntent } from './contract'
21export { scanSource, RULES, apiCalls, routeOf, routeMatches } from './mirage'
22export { driftFindings } from './drift'
23export { auditAnswer, extractClaims } from './claims'
24export { toMarkdown, toJson } from './report'
25export { mergeProjects, trimProject } from './persist'
26export { redact, normalizePath, hash, projectKey } from './util'
27
28const LIMITS = { activity: 200, notified: 300 }
29
30export function newProject(key: string, label: string, writer: string, now: number): ProjectRecord {
31 return { schema: 1, key, label: redact(label), rev: 0, writer, updatedAt: now, mode: 'observe', tasks: {} }
32}
33
34export const activeTask = (p: ProjectRecord): Task | undefined =>
35 p.activeTaskId === undefined ? undefined : p.tasks[p.activeTaskId]
36
37export function startTask(p: ProjectRecord, prompt: string, source: 'prompt' | 'manual', now: number): Task {
38 const id = newId('t', now)
39 const task: Task = {
40 id, createdAt: now, updatedAt: now, contract: createContract(id, prompt, now, source),
41 evidence: [], files: {}, findings: [], claims: [], activity: [], seq: 0, notified: [], nextEvidence: 1,
42 }
43 p.tasks[id] = task
44 p.activeTaskId = id
45 log(task, 'task', `task started (${source === 'prompt' ? 'captured from prompt' : 'manual'}): ${task.contract.summary}`, now)
46 return task
47}
48
49export function log(task: Task, kind: ActivityKind, text: string, now: number): void {
50 const entry: Activity = { at: now, kind, text: clip(redact(text), 200) }
51 task.activity.push(entry)
52 if (task.activity.length > LIMITS.activity) task.activity.splice(0, task.activity.length - LIMITS.activity)
53 task.updatedAt = now
54}
55
56/** A hint at most once per key, ever, for this task. */
57function hint(task: Task, key: string, level: Hint['level'], text: string): Hint[] {
58 if (task.notified.includes(key)) return []
59 task.notified.push(key)
60 if (task.notified.length > LIMITS.notified) task.notified.splice(0, task.notified.length - LIMITS.notified)
61 return [{ key, level, text }]
62}
63
64const nextEvidenceId = (task: Task) => (): string => `E${task.nextEvidence++}`
65
66function contractText(task: Task): string {
67 const c = task.contract
68 return [c.summary, c.prompt.excerpt, ...c.requirements.map(r => r.description), ...c.assumptions].join(' ')
69}
70
71/**
72 * Merges a fresh scan of one file into the findings: an id seen before keeps
73 * its waiver or acknowledgement; one no longer found is resolved.
74 */
75export function applyFindings(
76 task: Task, file: string, system: Finding['system'], fresh: (RawFinding & { id: string; requirementIds?: string[] })[], now: number,
77 replaceRules?: (ruleId: string) => boolean,
78): Hint[] {
79 const hints: Hint[] = []
80 const ids = new Set(fresh.map(f => f.id))
81 for (const old of task.findings) {
82 if (old.file !== file || old.system !== system || ids.has(old.id)) continue
83 if (replaceRules !== undefined && !replaceRules(old.ruleId)) continue
84 if (old.status === 'open' || old.status === 'acknowledged') {
85 old.status = 'resolved'
86 log(task, 'resolved', `${old.ruleId} no longer found in ${file}`, now)
87 }
88 }
89 for (const f of fresh) {
90 const existing = task.findings.find(x => x.id === f.id)
91 if (existing) {
92 Object.assign(existing, { line: f.line, snippet: redact(f.snippet), lastSeq: task.seq, severity: f.severity, confidence: f.confidence })
93 if (existing.status === 'resolved') existing.status = 'open'
94 continue
95 }
96 task.findings.push({
97 id: f.id, ruleId: f.ruleId, system, category: f.category, file, line: f.line, snippet: redact(f.snippet),
98 explanation: f.explanation, evidence: f.evidence, severity: f.severity, confidence: f.confidence,
99 requirementIds: f.requirementIds ?? [], verify: f.verify, status: 'open', firstSeq: task.seq, lastSeq: task.seq,
100 })
101 log(task, 'finding', `${f.ruleId} ${f.confidence} in ${file}:${f.line}`, now)
102 if (f.confidence === 'expected' || f.severity === 'low' || f.severity === 'info') continue
103 const text = system === 'drift'
104 ? `Integrity: possible scope drift in ${file} (${f.category.toLowerCase()}) · /integrity findings`
105 : `Integrity: ${f.confidence === 'confirmed' ? '' : 'possible '}${f.category.replace(/^\w\.\s*/, '').toLowerCase()} in ${file}:${f.line} · /integrity findings`
106 hints.push(...hint(task, `finding:${f.id}`, f.severity === 'high' ? 'important' : 'info', text))
107 }
108 if (task.findings.length > 300) {
109 const drop = task.findings.filter(f => f.status === 'resolved').slice(0, task.findings.length - 300)
110 task.findings = task.findings.filter(f => !drop.includes(f))
111 }
112 return hints
113}
114
115/** One observed source change: bumps the change sequence, then drift and mirage checks, then staleness. */
116export function recordChange(task: Task, change: FileChange, now: number): Hint[] {
117 const proofBefore = task.evidence.filter(e => isProof(e, task))
118 task.seq += 1
119 const prev = task.files[change.path]
120 task.files[change.path] = {
121 path: change.path,
122 created: prev?.created ?? change.kind === 'create',
123 deleted: change.kind === 'delete',
124 firstSeq: prev?.firstSeq ?? task.seq,
125 lastSeq: task.seq,
126 changes: (prev?.changes ?? 0) + 1,
127 ...(change.after === undefined ? {} : { hash: hash(change.after) }),
128 }
129 if (Object.keys(task.files).length > 500) {
130 const oldest = Object.values(task.files).sort((a, b) => a.lastSeq - b.lastSeq)[0]
131 if (oldest) delete task.files[oldest.path]
132 }
133 log(task, 'evidence', `${change.kind} ${change.path}`, now)
134
135 const hints: Hint[] = []
136 // Drift is history: a later edit does not undo a creation or deletion, so nothing auto-resolves.
137 hints.push(...applyFindings(task, change.path, 'drift', driftFindings(change, task.contract), now, () => false))
138 if (change.kind === 'delete') hints.push(...applyFindings(task, change.path, 'mirage', [], now))
139 else if (change.after !== undefined && isScannable(change.path)) {
140 const file = task.files[change.path] as FileState
141 file.calls = apiCalls(change.after)
142 const route = routeOf(change.path)
143 if (route !== undefined) file.route = route
144 hints.push(...applyFindings(task, change.path, 'mirage', scanSource(change.path, change.after, task.contract.intent, contractText(task)), now, id => id !== 'MIR-D2'))
145 }
146 hints.push(...endpointCheck(task, now))
147
148 const invalidated = proofBefore.filter(e => freshness(e, task) === 'stale')
149 for (const e of invalidated) log(task, 'invalidated', `${e.id} ${e.check} evidence stale after ${change.path} changed`, now)
150 if (invalidated.length > 0) {
151 const checks = [...new Set(invalidated.map(e => e.check))].join('/')
152 hints.push(...hint(task, `stale:${invalidated.map(e => e.id).join(',')}`, 'important', `Integrity: ${checks} evidence is stale after ${change.path} changed · /integrity evidence`))
153 }
154 task.updatedAt = now
155 return hints
156}
157
158/**
159 * MIR-D2, across files: a client calls an `/api/...` endpoint that no observed
160 * route file serves. Runs only once the task has seen at least one route file.
161 */
162function endpointCheck(task: Task, now: number): Hint[] {
163 const files = Object.values(task.files).filter(f => !f.deleted)
164 const routes = files.map(f => f.route).filter((r): r is string => r !== undefined)
165 if (routes.length === 0) return []
166 const hints: Hint[] = []
167 for (const f of files) {
168 if (f.calls === undefined) continue
169 const missing = f.calls.filter(c => !routes.some(r => routeMatches(r, c.path)))
170 hints.push(...applyFindings(task, f.path, 'mirage', missing.map(c => ({
171 id: hash(`MIR-D2|${f.path}|${c.path}`), ruleId: 'MIR-D2', category: 'D. Incomplete data integration', line: c.line,
172 snippet: `request to ${c.path}`, severity: 'high' as const, confidence: 'potential' as const,
173 explanation: `${f.path} calls ${c.path}, but no route file seen in this task serves it (seen: ${routes.slice(0, 4).join(', ')}).`,
174 evidence: 'Literal endpoint compared with the route files observed in this task; routes outside the task are not seen.',
175 verify: `Confirm ${c.path} exists, or point the client at the route that does.`,
176 })), now, id => id === 'MIR-D2'))
177 }
178 return hints
179}
180
181/** One observed tool call that may be evidence: a check run, a browser session. */
182export function recordTool(task: Task, obs: ToolObservation, now: number): { evidence: Evidence[]; hints: Hint[] } {
183 const verifiedBefore = new Set(task.contract.requirements.filter(r => requirementState(r, task).state === 'verified').map(r => r.id))
184 let evidence: Evidence[] = []
185 if (obs.tool === 'Bash') evidence = evidenceFromBash(obs, task.seq, now, nextEvidenceId(task))
186 else {
187 const browser = evidenceFromBrowserTool(obs, task.seq, now, `E${task.nextEvidence}`)
188 if (browser) {
189 task.nextEvidence++
190 evidence = [browser]
191 }
192 }
193 if (evidence.length === 0) return { evidence, hints: [] }
194
195 for (const e of evidence) {
196 e.requirementIds = task.contract.requirements
197 .filter(r => r.checks.length > 0 && e.command !== undefined && r.checks.some(c => e.command!.toLowerCase().includes(c.toLowerCase())))
198 .map(r => r.id)
199 task.evidence.push(e)
200 log(task, 'evidence', `${e.id} ${e.summary} [${e.basis}]`, now)
201 }
202 if (task.evidence.length > 200) task.evidence.splice(0, task.evidence.length - 200)
203
204 const hints: Hint[] = []
205 const failed = evidence.find(e => e.outcome === 'fail')
206 if (failed) hints.push(...hint(task, `fail:${failed.id}`, 'important', `Integrity: ${failed.check} failed (${failed.basis}) · /integrity evidence`))
207 const newlyVerified = task.contract.requirements.filter(r => !verifiedBefore.has(r.id) && requirementState(r, task).state === 'verified')
208 if (newlyVerified.length > 0) {
209 hints.push(...hint(task, `verified:${newlyVerified.map(r => r.id).join(',')}:${task.seq}`, 'info', `Integrity: ${newlyVerified.map(r => r.id).join(', ')} verified by fresh evidence`))
210 }
211 task.updatedAt = now
212 return { evidence, hints }
213}
214
215/** A person's own confirmation: the only way a manual requirement is verified. */
216export function recordManual(task: Task, requirementId: string, note: string, now: number): Evidence {
217 const e: Evidence = {
218 id: `E${task.nextEvidence++}`, requirementIds: [requirementId], kind: 'manual-confirmation', check: 'manual', outcome: 'pass',
219 timestamp: now, source: 'person', summary: `confirmed by the developer: ${clip(redact(note), 120)}`, basis: 'explicit user verification',
220 scope: [], seq: task.seq,
221 }
222 task.evidence.push(e)
223 log(task, 'evidence', `${e.id} manual confirmation of ${requirementId}`, now)
224 return e
225}
226
227export type AuditResult = { notice?: string; hints: Hint[] }
228
229/** Audits a final answer and composes the notice shown beneath it, or none per the notification level. */
230export function recordAudit(task: Task, answer: string, turnId: string, now: number, mode: Mode, level: 'off' | 'important' | 'all'): AuditResult {
231 const claims = auditAnswer(answer, task, turnId, now)
232 const known = new Set(task.claims.map(c => c.id))
233 task.claims.push(...claims.filter(c => !known.has(c.id)))
234 if (task.claims.length > 100) task.claims.splice(0, task.claims.length - 100)
235 if (claims.length > 0) log(task, 'claims', `${claims.length} claim(s) audited: ${claims.map(c => c.verdict).join(', ')}`, now)
236
237 const bad = claims.filter(c => c.verdict === 'contradicted' || c.verdict === 'unsupported')
238 const lines: string[] = []
239 if (level !== 'off' && (bad.length > 0 || (level === 'all' && claims.length > 0))) {
240 const counts = (['supported', 'partial', 'unsupported', 'contradicted', 'not-assessable'] as const)
241 .map(v => [v, claims.filter(c => c.verdict === v).length] as const).filter(([, n]) => n > 0).map(([v, n]) => `${n} ${v}`)
242 lines.push(`Integrity · ${claims.length} completion claim${claims.length === 1 ? '' : 's'} checked: ${counts.join(' · ')}`)
243 for (const c of (level === 'all' ? claims : bad).slice(0, 4)) lines.push(` ${GLYPH[c.verdict]} "${clip(c.text, 70)}": ${c.reason}`)
244 }
245 if (mode !== 'observe' && level !== 'off') {
246 const high = task.findings.filter(f => f.status === 'open' && f.severity === 'high' && f.confidence !== 'expected')
247 const mustOpen = task.contract.requirements.filter(r => r.criticality === 'must' && r.waiver === undefined && requirementState(r, task).state !== 'verified')
248 if (high.length > 0 || mustOpen.length > 0) {
249 lines.push(`Integrity ${mode}: ${mustOpen.length} must requirement(s) not verified · ${high.length} high finding(s) open`)
250 }
251 }
252 if (lines.length > 0) lines.push(' details: /integrity claims')
253 return { ...(lines.length > 0 ? { notice: lines.join('\n') } : {}), hints: [] }
254}
255
256export const GLYPH = { supported: '✓', partial: '◐', unsupported: '?', contradicted: '✗', 'not-assessable': '·' } as const
257
258export function setFindingStatus(task: Task, id: string, status: 'acknowledged' | 'waived' | 'open', now: number, reason?: string): Finding | undefined {
259 const f = task.findings.find(x => x.id === id || x.id.startsWith(id))
260 if (f === undefined) return undefined
261 f.status = status
262 if (status === 'waived') f.waiver = { reason: clip(redact(reason ?? ''), 200), at: now }
263 else delete f.waiver
264 log(task, status === 'waived' ? 'waiver' : 'finding', `${f.ruleId} ${f.file}:${f.line} ${status}${reason ? `: ${reason}` : ''}`, now)
265 return f
266}
267
268/** Re-reads tracked files (passed in by the adapter) and records out-of-band edits as changes. */
269export function reconcileFiles(task: Task, current: Record<string, string | null>, now: number): Hint[] {
270 const hints: Hint[] = []
271 for (const [path, text] of Object.entries(current)) {
272 const known = task.files[path]
273 if (known === undefined) continue
274 if (text === null && !known.deleted) hints.push(...recordChange(task, { path, kind: 'delete' }, now))
275 else if (text !== null && known.hash !== undefined && hash(text) !== known.hash) {
276 hints.push(...recordChange(task, { path, kind: 'update', after: text }, now))
277 log(task, 'invalidated', `${path} changed outside observed tools`, now)
278 }
279 }
280 return hints
281}
282
283hooks/pane.tsx 565 lines1// The /integrity pane, drawn the way Claude Flightdeck draws its own: a centered
2// title over a color legend, rounded cards per subsystem (title left, count
3// right), a strip of requirement states, and an activity log. Everything comes
4// from the engine's canonical summary, so its counts always match
5// /integrity-report. Keys 1-7 switch views while the pane holds the keyboard;
6// every action is also a command for headless use.
7
8import type { ElementTable, RenderElement, ThemeKey } from 'claude-code'
9
10import type { IntegrityTab, IntegrityUi } from '../types'
11import * as E from '../src/engine/index'
12
13export type PaneActions = {
14 setUi: (fn: (u: IntegrityUi) => IntegrityUi) => Promise<unknown>
15 approve: () => Promise<string>
16 recheck: () => Promise<string>
17 exportReport: (format: 'md' | 'json' | 'both') => Promise<string>
18 setFinding: (id: string, status: 'acknowledged' | 'waived' | 'open', reason?: string) => Promise<string>
19 waiveRequirement: (id: string, reason: string) => Promise<string>
20 close: () => Promise<void>
21}
22
23type Prefs = { layout: E.Layout; maxFindings: number; color: boolean }
24
25export type PaneContext = {
26 project: E.ProjectRecord
27 task: E.Task | undefined
28 prefs: Prefs
29 ui: IntegrityUi
30 actions: PaneActions
31}
32
33const TABS: readonly [IntegrityTab, string, string][] = [
34 ['overview', 'Overview', 'Ovw'], ['requirements', 'Requirements', 'Reqs'], ['evidence', 'Evidence', 'Evid'],
35 ['findings', 'Findings', 'Find'], ['claims', 'Claims', 'Clms'], ['activity', 'Activity', 'Log'], ['report', 'Report', 'Rpt'],
36]
37
38// Theme keys, like Flightdeck's default palette: they follow light, dark and color-blind themes.
39const C = {
40 main: 'claude', ok: 'success', evidence: 'suggestion', warn: 'warning', bad: 'error', claims: 'merged', dim: 'inactive', faint: 'subtle',
41} as const satisfies Record<string, ThemeKey>
42
43type ReqState = Exclude<keyof E.Summary['requirements'], 'total'>
44const REQ_GLYPH: Record<ReqState, string> = { verified: '✓', stale: '⧗', failed: '✗', waived: '–', implemented: '◐', unverified: '?' }
45const REQ_COLOR: Record<ReqState, ThemeKey> = { verified: C.ok, implemented: C.evidence, unverified: C.dim, stale: C.warn, failed: C.bad, waived: C.faint }
46const REQ_LABEL: Record<ReqState, string> = {
47 verified: 'Verified', stale: 'Stale', failed: 'Failed', waived: 'Waived', implemented: 'Implemented, not verified', unverified: 'Not verified',
48}
49const LEGEND: readonly ReqState[] = ['verified', 'implemented', 'unverified', 'stale', 'failed']
50const SEV_GLYPH: Record<E.Severity, string> = { high: '‼', medium: '!', low: '·', info: '○' }
51const KIND_COLOR: Record<E.Activity['kind'], ThemeKey> = {
52 task: C.main, contract: C.main, approval: C.main, requirement: C.main, mode: C.main, evidence: C.evidence,
53 invalidated: C.warn, finding: C.warn, resolved: C.ok, waiver: C.dim, claims: C.claims, turn: 'text',
54}
55
56const clip = (text: string, max: number): string => (text.length <= max ? text : `${text.slice(0, Math.max(1, max - 1))}…`)
57const time = (at: number): string => new Date(at).toTimeString().slice(0, 5)
58const plural = (n: number, word: string): string => `${n} ${word}${n === 1 ? '' : 's'}`
59/** Flightdeck's gauge: ▰ filled, ▱ empty. */
60const gauge = (ratio: number, cells: number): { on: string; off: string } => {
61 const n = Math.max(0, Math.min(cells, Math.round(ratio * cells)))
62 return { on: '▰'.repeat(n), off: '▱'.repeat(cells - n) }
63}
64
65/** Inline above the prompt the pane is a short summary, as Flightdeck's mini; docked it follows the width. */
66export function layoutFor(columns: number, pref: E.Layout, placement: 'dock' | 'inline' = 'dock'): Exclude<E.Layout, 'auto'> {
67 if (pref !== 'auto') return pref
68 if (placement === 'inline') return 'mini'
69 return columns < 56 ? 'mini' : columns < 110 ? 'compact' : 'wide'
70}
71
72/** What the pane is drawn into: the Pane site's body width and placement, and the viewport's width when measured. */
73export type PaneView = {
74 surface: 'terminal' | 'desktop' | 'mobile' | 'vscode'
75 bodyColumns: number
76 columns?: number | undefined
77 placement?: 'dock' | 'inline' | undefined
78}
79
80export function renderPane(els: ElementTable, view: PaneView, ctx: PaneContext): RenderElement {
81 const { Box, Text, Button } = els
82 // The mobile app draws no field yet: waivers there go through the command.
83 const Input = view.surface !== 'mobile' && 'Input' in els ? els.Input : undefined
84 const { task, prefs, ui, actions, project } = ctx
85 const W = Math.max(24, view.bodyColumns || view.columns || 80)
86 const layout = layoutFor(W, prefs.layout, view.placement)
87 const isMini = layout === 'mini'
88 // Color supplements the glyphs and words, never replaces them, so it can be switched off.
89 const fg = (c: ThemeKey): { color?: ThemeKey } => (prefs.color ? { color: c } : {})
90 const faint = prefs.color ? { color: C.faint } : { dimColor: true }
91 const s = E.summarize(task, project.mode)
92
93 const go = (tab: IntegrityTab, selected?: string) => (): void => {
94 void actions.setUi(u => ({ tab, page: 0, ...(selected === undefined ? {} : { selected: u.selected === selected && u.tab === tab ? undefined : selected }) }))
95 }
96 const flash = (p: Promise<string>): void => {
97 void p.then(msg => actions.setUi(u => ({ ...u, flash: msg, waiving: undefined })))
98 }
99
100 /** A card: a rounded border in the section's color, its title left and a count right. Mini drops the border. */
101 const card = (key: string, color: ThemeKey, title: string, right: string, w: number, body: (RenderElement | null | false)[], double = false): RenderElement =>
102 isMini ? (
103 <Box key={key} flexDirection="column" width={w}>
104 <Text key={`${key}:title`} wrap="truncate-end">
105 <Text bold {...fg(color)}>{title}</Text>
106 {right !== '' ? <Text dimColor>{` ${right}`}</Text> : null}
107 </Text>
108 {body}
109 </Box>
110 ) : (
111 <Box key={key} flexDirection="column" borderStyle={double ? 'double' : 'round'} {...(prefs.color ? { borderColor: color } : {})} paddingX={1} width={w}>
112 <Box key={`${key}:title`} justifyContent="space-between">
113 <Text bold wrap="truncate-end" {...fg(color)}>{title}</Text>
114 {right !== '' ? <Text dimColor>{right}</Text> : null}
115 </Box>
116 {body}
117 </Box>
118 )
119 /** Inside a card's border and padding. */
120 const inner = (w: number): number => (isMini ? w : w - 4)
121 const rail = (key: string, w: number): RenderElement => <Text key={key} {...faint}>{'─'.repeat(Math.max(1, w))}</Text>
122
123 const tabBar = (
124 <Box key="tabs" flexDirection="row" flexWrap="wrap" columnGap={1} {...(isMini ? {} : { justifyContent: 'center' as const })}>
125 {TABS.map(([id, long, short], i) => (
126 <Button key={`tab:${id}`} plain hotkey={String(i + 1)} label={layout === 'wide' ? long : short}
127 variant={ui.tab === id ? 'primary' : 'secondary'} dimColor={ui.tab !== id} onPress={go(id)} />
128 ))}
129 </Box>
130 )
131
132 if (!s.hasTask || task === undefined) {
133 return (
134 <Box flexDirection="column" width={W}>
135 {!isMini && (
136 <Box key="head" justifyContent="center">
137 <Text bold wrap="truncate-end">INTEGRITY<Text {...faint}> · </Text><Text dimColor>NO CONTRACT</Text></Text>
138 </Box>
139 )}
140 {card('c-empty', C.faint, 'CONTRACT · none yet', '', W, [
141 <Text key="none" dimColor wrap="wrap">◇ No task contract. Nothing is tracked as verified yet.</Text>,
142 <Text key="how" wrap="wrap">Describe what you want built (an implementation prompt drafts a contract, no model call), or run /integrity-start <task>.</Text>,
143 <Text key="help" dimColor wrap="wrap">/integrity-help explains the workflow · /ig is short for /integrity</Text>,
144 ])}
145 </Box>
146 )
147 }
148
149 const inForce = s.requirements.total - s.requirements.waived
150 const approval = s.approval === 'approved' ? `approved v${s.version}` : s.approval === 'amended' ? 'amended · re-approve' : `draft v${s.version}`
151
152 // ---- header: Flightdeck's centered title and legend ------------------------------------
153
154 const legend = (() => {
155 const out: ReqState[] = []
156 let used = 0
157 for (const st of LEGEND) {
158 const cell = st.length + 4
159 if (used + cell > W) break
160 out.push(st)
161 used += cell
162 }
163 return out
164 })()
165 const header = isMini ? null : (
166 <Box key="header" flexDirection="column">
167 <Box key="head" justifyContent="center">
168 <Text bold wrap="truncate-end">
169 <Text>INTEGRITY</Text>
170 <Text {...faint}> · </Text>
171 <Text {...fg(C.main)}>{s.mode.toUpperCase()}</Text>
172 <Text> MODE</Text>
173 <Text {...faint}> · </Text>
174 <Text {...fg(s.approval === 'approved' ? C.ok : C.warn)}>{approval.toUpperCase()}</Text>
175 </Text>
176 </Box>
177 <Box key="legend" justifyContent="center" columnGap={2}>
178 {legend.map(st => (
179 <Text key={`lg:${st}`}>
180 <Text {...fg(REQ_COLOR[st])}>{prefs.color ? '■' : REQ_GLYPH[st]}</Text>
181 <Text dimColor>{` ${st}`}</Text>
182 </Text>
183 ))}
184 </Box>
185 </Box>
186 )
187 const footer = ui.flash ? <Text key="flash" dimColor wrap="truncate-end">{ui.flash}</Text> : null
188
189 // ---- overview pieces --------------------------------------------------------------------
190
191 /** One cell per requirement, colored by state (its glyph when color is off), like Flightdeck's gate strip. */
192 const strip = (w: number): RenderElement => {
193 const reqs = task.contract.requirements
194 const shown = reqs.slice(0, Math.max(4, w - 10))
195 return (
196 <Text key="strip" wrap="truncate-end">
197 <Text dimColor>reqs </Text>
198 {shown.map(r => {
199 const st = E.requirementState(r, task).state
200 return <Text key={`st:${r.id}`} {...fg(REQ_COLOR[st])}>{prefs.color ? '■' : REQ_GLYPH[st]}</Text>
201 })}
202 {reqs.length > shown.length ? <Text dimColor>{` +${reqs.length - shown.length}`}</Text> : null}
203 {reqs.length === 0 ? <Text dimColor>none yet</Text> : null}
204 </Text>
205 )
206 }
207 const progress = (w: number): RenderElement => {
208 const g = gauge(s.progress, Math.max(5, Math.min(14, w - 34)))
209 return (
210 <Text key="progress" wrap="truncate-end">
211 <Text {...fg(C.ok)}>{g.on}</Text>
212 <Text {...faint}>{g.off}</Text>
213 <Text bold>{` ${s.requirements.verified}/${inForce} verified`}</Text>
214 <Text dimColor>{` · must ${s.must.verified}/${s.must.total}`}</Text>
215 </Text>
216 )
217 }
218 const notVerified = s.requirements.unverified + s.requirements.implemented
219 const badClaims = s.claims.contradicted + s.claims.unsupported
220 const parts: [string, ThemeKey][] = [
221 ...(notVerified > 0 ? [[`? ${notVerified} not verified`, C.dim] as [string, ThemeKey]] : []),
222 ...(s.requirements.stale > 0 ? [[`⧗ ${s.requirements.stale} stale`, C.warn] as [string, ThemeKey]] : []),
223 ...(s.requirements.failed > 0 ? [[`✗ ${s.requirements.failed} failed`, C.bad] as [string, ThemeKey]] : []),
224 ...(s.findings.open > 0 ? [[`! ${plural(s.findings.open, 'finding')}${s.findings.high ? ` (${s.findings.high} high)` : ''}`, C.warn] as [string, ThemeKey]] : []),
225 ...(badClaims > 0 ? [[`✗ ${badClaims} claim(s)`, C.bad] as [string, ThemeKey]] : []),
226 ]
227 const counts = (
228 <Text key="counts" wrap="truncate-end">
229 {parts.length === 0 ? <Text {...fg(C.ok)}>no open problems recorded</Text> : parts.map(([text, color], i) => (
230 <Text key={`n${i}`}>
231 {i > 0 ? <Text {...faint}> · </Text> : null}
232 <Text {...fg(color)}>{text}</Text>
233 </Text>
234 ))}
235 </Text>
236 )
237
238 const contractCard = (w: number): RenderElement =>
239 card('c-contract', C.main, `CONTRACT · ${approval}`, `intent ${s.intent}`, w, [
240 <Text key="title" wrap="truncate-end">◆ {s.title}</Text>,
241 strip(inner(w)),
242 progress(inner(w)),
243 counts,
244 ])
245
246 const evidenceCard = (w: number): RenderElement =>
247 card('c-evidence', C.evidence, 'EVIDENCE · witness', plural(s.evidence.total, 'check'), w,
248 s.evidence.latest.length === 0
249 ? [<Text key="none" dimColor wrap="wrap">Not observed: no test, build, lint or typecheck run yet.</Text>]
250 : s.evidence.latest.map(ev => {
251 const stale = ev.freshness === 'stale'
252 const color = ev.outcome === 'fail' ? C.bad : ev.outcome !== 'pass' ? C.dim : stale ? C.warn : C.ok
253 return (
254 <Text key={`latest:${ev.check}`} wrap="truncate-end">
255 <Text {...fg(color)}>{ev.outcome === 'pass' ? '✓' : ev.outcome === 'fail' ? '✗' : '·'}</Text>
256 {` ${ev.check} ${ev.outcome} · `}
257 <Text {...(stale ? fg(C.warn) : {})}>{stale ? '⧗ stale' : ev.freshness}</Text>
258 <Text dimColor>{` · ${time(ev.at)} ${ev.id}`}</Text>
259 </Text>
260 )
261 }))
262
263 const order: Record<E.Severity, number> = { high: 0, medium: 1, low: 2, info: 3 }
264 const sorted = task.findings.filter(f => f.status !== 'resolved')
265 .sort((a, b) => (a.status === 'open' ? 0 : 1) - (b.status === 'open' ? 0 : 1) || order[a.severity] - order[b.severity])
266 const findingLabel = (f: E.Finding): string =>
267 `${f.confidence === 'expected' ? '○' : SEV_GLYPH[f.severity]} ${f.ruleId} ${f.file}:${f.line} ${f.category.replace(/^\w\.\s*/, '')}${f.status !== 'open' ? ` (${f.status})` : f.confidence === 'confirmed' ? ' [confirmed]' : ''}`
268
269 const findingsCard = (w: number): RenderElement => {
270 const top = sorted.filter(f => f.status === 'open' && f.confidence !== 'expected').slice(0, 3)
271 return card('c-findings', C.warn, 'FINDINGS · mirage + drift', `${s.findings.open} open`, w, [
272 <Text key="sev" wrap="truncate-end">
273 <Text {...fg(C.bad)}>‼</Text>{` ${s.findings.high} high · `}
274 <Text {...fg(C.warn)}>!</Text>{` ${s.findings.medium} medium · · ${s.findings.low + s.findings.info} low · ○ ${s.findings.expected} expected · ${s.findings.waived} waived`}
275 </Text>,
276 ...(top.length === 0
277 ? [<Text key="none" dimColor wrap="truncate-end">none open · checked on every observed edit</Text>]
278 : top.map(f => <Button key={`ovf:${f.id}`} plain label={clip(findingLabel(f), inner(w))} onPress={go('findings', f.id)} />)),
279 ])
280 }
281
282 const claimsCard = (w: number): RenderElement => {
283 const worst = s.claims.latest.filter(c => c.verdict === 'contradicted' || c.verdict === 'unsupported').slice(0, 2)
284 return card('c-claims', C.claims, 'CLAIMS · last answer', `${s.claims.total} checked`, w, [
285 s.claims.total === 0
286 ? <Text key="none" dimColor wrap="wrap">Not assessed: no completion claims audited yet.</Text>
287 : (
288 <Text key="verdicts" wrap="truncate-end">
289 <Text {...fg(C.ok)}>✓</Text>{` ${s.claims.supported} supported · ◐ ${s.claims.partial} partial · `}
290 <Text {...fg(C.warn)}>?</Text>{` ${s.claims.unsupported} unsupported · `}
291 <Text {...fg(C.bad)}>✗</Text>{` ${s.claims.contradicted} contradicted · · ${s.claims.notAssessable} n/a`}
292 </Text>
293 ),
294 ...worst.map(c => (
295 <Text key={`wc:${c.id}`} wrap="truncate-end">
296 <Text {...fg(c.verdict === 'contradicted' ? C.bad : C.warn)}>{E.GLYPH[c.verdict]}</Text>
297 {` "${clip(c.text, 40)}": `}
298 <Text dimColor>{c.reason}</Text>
299 </Text>
300 )),
301 ])
302 }
303
304 // Flightdeck's receipt: one line in its own card, here the top concern.
305 const receipt = (w: number): RenderElement => {
306 const line = s.concern
307 ? [
308 <Text key="concern" wrap="truncate-end" {...fg(C.warn)}>▸ {clip(s.concern.text, Math.max(10, inner(w) - 12))} </Text>,
309 <Button key="concern-go" plain label="inspect" hotkey="i" onPress={go(s.concern.tab)} />,
310 ]
311 : [<Text key="concern" {...fg(C.ok)}>✓ No outstanding concern.</Text>]
312 return isMini
313 ? <Box key="receipt" flexDirection="row">{line}</Box>
314 : <Box key="receipt" flexDirection="row" borderStyle="round" {...(prefs.color ? { borderColor: s.concern ? C.warn : C.ok } : {})} borderDimColor={s.concern === undefined} paddingX={1} width={w}>{line}</Box>
315 }
316
317 const actionRow = (
318 <Box key="actions" flexDirection="row" columnGap={2} {...(isMini ? {} : { paddingX: 1 })}>
319 {s.approval !== 'approved' && <Button key="approve" plain hotkey="a" label="approve contract" onPress={() => flash(actions.approve())} />}
320 <Button key="recheck" plain hotkey="c" label="re-check" onPress={() => flash(actions.recheck())} />
321 {ui.tab === 'overview' && !isMini && <Button key="ov-report" plain hotkey="r" label="report" dimColor onPress={go('report')} />}
322 </Box>
323 )
324
325 /** Flightdeck's session log: dim time, colored kind, the text. */
326 const log = (key: string, title: string, w: number, rows: number): RenderElement => {
327 const lines = task.activity.slice(-rows).reverse()
328 const body = lines.length === 0 ? [<Text key="none" {...faint}>nothing yet</Text>] : lines.map((a, i) => (
329 <Box key={`a${i}`} flexDirection="row">
330 <Box key="t" width={6} flexShrink={0}><Text {...faint}>{time(a.at)}</Text></Box>
331 <Box key="k" width={12} flexShrink={0}><Text bold wrap="truncate-end" {...fg(KIND_COLOR[a.kind])}>{a.kind}</Text></Box>
332 <Text key="x" wrap="truncate-end">{a.text}</Text>
333 </Box>
334 ))
335 return isMini
336 ? <Box key={key} flexDirection="column">{body}</Box>
337 : (
338 <Box key={key} flexDirection="column" borderStyle="round" {...(prefs.color ? { borderColor: C.faint } : {})} paddingX={1} width={w}>
339 <Text key="title" dimColor>{title}</Text>
340 {body}
341 </Box>
342 )
343 }
344
345 const overview = (): RenderElement => {
346 if (isMini) {
347 const g = gauge(s.progress, 6)
348 return (
349 <Box flexDirection="column" width={W}>
350 <Text key="mini1" wrap="truncate-end">
351 <Text bold {...fg(C.main)}>Integrity </Text>
352 <Text {...fg(C.ok)}>{g.on}</Text>
353 <Text {...faint}>{g.off}</Text>
354 <Text bold>{` ${s.requirements.verified}/${inForce} verified`}</Text>
355 <Text dimColor>{` · ${approval} · ${s.mode}`}</Text>
356 </Text>
357 {counts}
358 {receipt(W)}
359 {actionRow}
360 </Box>
361 )
362 }
363 if (layout === 'wide') {
364 const colW = Math.floor((W - 2) / 2)
365 return (
366 <Box flexDirection="column" width={W}>
367 <Box key="cols" flexDirection="row" columnGap={2}>
368 <Box key="left" flexDirection="column" width={colW}>
369 {contractCard(colW)}
370 {rail('rail-l', colW)}
371 {evidenceCard(colW)}
372 </Box>
373 <Box key="right" flexDirection="column" width={colW}>
374 {findingsCard(colW)}
375 {rail('rail-r', colW)}
376 {claimsCard(colW)}
377 </Box>
378 </Box>
379 {receipt(W)}
380 {actionRow}
381 {log('c-log', 'activity log', W, 6)}
382 </Box>
383 )
384 }
385 return (
386 <Box flexDirection="column" width={W}>
387 {contractCard(W)}
388 {rail('rail-1', W)}
389 {evidenceCard(W)}
390 {rail('rail-2', W)}
391 {findingsCard(W)}
392 {rail('rail-3', W)}
393 {claimsCard(W)}
394 {receipt(W)}
395 {actionRow}
396 {log('c-log', 'activity log', W, 4)}
397 </Box>
398 )
399 }
400
401 // ---- list views with a drill-down -------------------------------------------------------
402
403 const hasDetail = layout === 'wide' && ui.selected !== undefined
404 const listW = hasDetail ? W - Math.floor(W / 2) - 2 : W
405 const rowWidth = inner(listW)
406
407 const detail = (color: ThemeKey, lines: (string | RenderElement)[]): RenderElement => (
408 <Box key="detail" flexDirection="column" borderStyle="double" {...(prefs.color ? { borderColor: color } : {})} paddingX={1}
409 {...(layout === 'wide' ? { width: Math.floor(W / 2) } : { marginTop: 1 })}>
410 {lines.map((l, i) => (typeof l === 'string' ? <Text key={`d${i}`} wrap="wrap">{l}</Text> : l))}
411 <Button key="detail-close" plain hotkey="x" label="close detail" dimColor onPress={() => void actions.setUi(u => ({ ...u, selected: undefined, waiving: undefined }))} />
412 </Box>
413 )
414
415 const waiveField = (id: string, kind: 'finding' | 'requirement'): RenderElement => (
416 ui.waiving === id
417 ? Input !== undefined
418 ? <Input key="waive-reason" label="Waiver reason" placeholder="why this is acceptable" autoFocus
419 onSubmit={reason => { if (reason.trim() !== '') flash(kind === 'finding' ? actions.setFinding(id, 'waived', reason) : actions.waiveRequirement(id, reason)) }} />
420 : <Text key="waive-reason" dimColor>Type /integrity-waive {id} <reason> to waive.</Text>
421 : <Button key="waive" plain hotkey="w" label="waive (with reason)" onPress={() => void actions.setUi(u => ({ ...u, waiving: id }))} />
422 )
423
424 const split = (list: RenderElement, side: RenderElement | null): RenderElement =>
425 layout === 'wide' && side !== null
426 ? <Box flexDirection="row" columnGap={2}><Box flexDirection="column" width={listW}>{list}</Box>{side}</Box>
427 : <Box flexDirection="column">{list}{side}</Box>
428
429 const requirements = (): RenderElement => {
430 const reqs = task.contract.requirements
431 const list = card('v-req', C.main, 'REQUIREMENTS', `${s.requirements.verified}/${inForce} verified`, listW, [
432 ...(reqs.length === 0 ? [<Text key="none" dimColor>No requirements. /integrity-contract add <requirement></Text>] : []),
433 ...reqs.map(r => {
434 const st = E.requirementState(r, task).state
435 return (
436 <Button key={`req:${r.id}`} plain label={clip(`${REQ_GLYPH[st]} ${r.id} [${r.criticality}] ${r.description}`, rowWidth)}
437 dimColor={st === 'waived'} onPress={go('requirements', r.id)} />
438 )
439 }),
440 task.contract.exclusions.length > 0 && <Text key="excl" dimColor wrap="truncate-end">Exclusions: {task.contract.exclusions.join(', ')}</Text>,
441 ])
442 const r = reqs.find(x => x.id === ui.selected)
443 if (r === undefined) return list
444 const st = E.requirementState(r, task)
445 return split(list, detail(REQ_COLOR[st.state], [
446 `${REQ_GLYPH[st.state]} ${r.id}: ${REQ_LABEL[st.state]}`,
447 r.description,
448 `${r.criticality} · ${r.category} · verified by ${r.verification} evidence · ${r.origin}`,
449 `Why: ${st.reason}`,
450 ...(r.checks.length ? [`Checks: ${r.checks.join(', ')}`] : ['No declared check: /integrity-contract check ' + r.id + ' <command>']),
451 ...(st.evidence.length ? st.evidence.slice(-5).map(ev => ` ${ev.id} ${ev.check} ${ev.outcome} · ${E.freshness(ev, task)} · ${ev.basis}`) : ['No linked evidence. /integrity-verify ' + r.id + ' <what you checked>']),
452 ...(r.waiver ? [`Waived: ${r.waiver.reason}`] : [waiveField(r.id, 'requirement')]),
453 ]))
454 }
455
456 const evidence = (): RenderElement => {
457 const shown = task.evidence.slice(-(isMini ? 6 : 20)).reverse()
458 const list = card('v-ev', C.evidence, 'EVIDENCE', `${s.evidence.fresh} fresh · ${s.evidence.stale} stale`, listW, [
459 ...(shown.length === 0 ? [<Text key="none" dimColor wrap="wrap">Not observed: no checks yet. Test, build, lint and typecheck runs Claude makes appear here. Nothing is verified without one.</Text>] : []),
460 ...shown.map(ev => {
461 const fresh = E.freshness(ev, task)
462 const mark = ev.synthetic ? '⊘' : ev.outcome === 'pass' ? '✓' : ev.outcome === 'fail' ? '✗' : '·'
463 return (
464 <Button key={`ev:${ev.id}`} plain onPress={go('evidence', ev.id)} dimColor={fresh === 'stale' || ev.outcome === 'inconclusive'}
465 label={clip(`${mark} ${ev.id} ${ev.check} ${ev.outcome} · ${fresh === 'stale' ? '⧗ stale' : fresh} · ${ev.command ?? ev.source}`, rowWidth)} />
466 )
467 }),
468 ])
469 const ev = task.evidence.find(x => x.id === ui.selected)
470 if (ev === undefined) return list
471 const fresh = E.freshness(ev, task)
472 return split(list, detail(ev.outcome === 'fail' ? C.bad : fresh === 'stale' ? C.warn : C.evidence, [
473 `${ev.id} · ${ev.kind} · ${ev.check} ${ev.outcome}${ev.synthetic ? ' · SYNTHETIC (never proof)' : ''}`,
474 `Basis: ${ev.basis}`,
475 `Freshness: ${fresh}${fresh === 'stale' ? ` (changed since: ${E.staleBecause(ev, task).slice(0, 4).join(', ')})` : ''}`,
476 ...(ev.command ? [`Command: ${ev.command}`] : []),
477 ...(ev.counts ? [`Counts: ${ev.counts.passed} passed, ${ev.counts.failed} failed, ${ev.counts.skipped} skipped`] : []),
478 `Scope: ${ev.scope.length ? ev.scope.join(', ') : 'whole project'}${ev.truncated ? ' · output truncated' : ''}`,
479 `Source: ${ev.source}${ev.agentId ? ` (subagent ${ev.agentId})` : ''} · ${new Date(ev.timestamp).toISOString()}`,
480 ...(ev.revision ? [`Revision: ${ev.revision}`] : []),
481 ...(ev.cwd ? [`Working dir: ${ev.cwd}`] : []),
482 `Requirements: ${ev.requirementIds.join(', ') || 'none linked (/integrity-evidence link ' + ev.id + ' R#)'}`,
483 ]))
484 }
485
486 const findings = (): RenderElement => {
487 const per = isMini ? Math.min(4, prefs.maxFindings) : prefs.maxFindings
488 const pages = Math.max(1, Math.ceil(sorted.length / per))
489 const page = Math.min(ui.page, pages - 1)
490 const list = card('v-find', C.warn, 'FINDINGS', `${s.findings.open} open`, listW, [
491 ...(sorted.length === 0 ? [<Text key="none" dimColor wrap="wrap">No open findings. Mirage and drift checks run on every observed edit.</Text>] : []),
492 ...sorted.slice(page * per, page * per + per).map(f => (
493 <Button key={`f:${f.id}`} plain onPress={go('findings', f.id)} dimColor={f.status !== 'open' || f.confidence === 'expected'}
494 label={clip(findingLabel(f), rowWidth)} />
495 )),
496 pages > 1 && (
497 <Box key="pager" flexDirection="row" columnGap={1}>
498 <Text dimColor>page {page + 1}/{pages}</Text>
499 {page > 0 && <Button key="prev" plain hotkey="p" label="prev" onPress={() => void actions.setUi(u => ({ ...u, page: page - 1 }))} />}
500 {page < pages - 1 && <Button key="next" plain hotkey="n" label="next" onPress={() => void actions.setUi(u => ({ ...u, page: page + 1 }))} />}
501 </Box>
502 ),
503 ])
504 const f = task.findings.find(x => x.id === ui.selected)
505 if (f === undefined) return list
506 return split(list, detail(f.severity === 'high' ? C.bad : C.warn, [
507 `${SEV_GLYPH[f.severity]} ${f.severity} · ${f.confidence} · ${f.status} · ${f.ruleId} (${f.system})`,
508 `${f.file}:${f.line}`,
509 <Text key="snippet" dimColor wrap="truncate-end"> {f.snippet}</Text>,
510 f.explanation,
511 `Evidence: ${f.evidence}`,
512 `Verify: ${f.verify}`,
513 ...(f.requirementIds.length ? [`Requirements: ${f.requirementIds.join(', ')}`] : []),
514 ...(f.waiver ? [`Waived: ${f.waiver.reason}`] : []),
515 <Box key="f-actions" flexDirection="row" columnGap={1}>
516 {f.status === 'open' && <Button key="ack" plain hotkey="k" label="acknowledge" onPress={() => flash(actions.setFinding(f.id, 'acknowledged'))} />}
517 {f.status !== 'open' && <Button key="reopen" plain hotkey="o" label="reopen" onPress={() => flash(actions.setFinding(f.id, 'open'))} />}
518 </Box>,
519 ...(f.status === 'waived' ? [] : [waiveField(f.id, 'finding')]),
520 ]))
521 }
522
523 const claims = (): RenderElement => {
524 const shown = task.claims.slice(-(isMini ? 4 : 12)).reverse()
525 const list = card('v-claims', C.claims, 'CLAIMS', `${task.claims.length} audited`, listW, [
526 ...(shown.length === 0 ? [<Text key="none" dimColor wrap="wrap">Not assessed: no completion claims audited yet. They are checked when Claude finishes an answer.</Text>] : []),
527 ...shown.map(c => (
528 <Button key={`c:${c.id}`} plain onPress={go('claims', c.id)} dimColor={c.verdict === 'not-assessable'}
529 label={clip(`${E.GLYPH[c.verdict]} ${c.verdict}: "${c.text}"`, rowWidth)} />
530 )),
531 ])
532 const c = task.claims.find(x => x.id === ui.selected)
533 if (c === undefined) return list
534 return split(list, detail(c.verdict === 'contradicted' ? C.bad : c.verdict === 'supported' ? C.ok : C.claims, [
535 `${E.GLYPH[c.verdict]} ${c.verdict} · ${c.kind} · ${time(c.at)}`,
536 `"${c.text}"`,
537 `Why: ${c.reason}`,
538 `Evidence: ${c.evidenceIds.join(', ') || (c.verdict === 'unsupported' ? 'none observed (absence of evidence, not proof of failure)' : 'none')}`,
539 ]))
540 }
541
542 const activity = (): RenderElement => log('v-log', 'ACTIVITY', W, isMini ? 6 : 30)
543
544 const report = (): RenderElement =>
545 card('v-report', C.main, 'REPORT', `task ${task.id}`, W, [
546 <Text key="same" wrap="wrap">Same counts as /integrity-report: {s.requirements.verified}/{inForce} verified, {s.requirements.stale} stale, {s.requirements.failed} failed, {s.findings.open} open findings, {s.evidence.fresh} fresh passing checks.</Text>,
547 <Box key="exports" flexDirection="row" columnGap={2}>
548 <Button key="export-md" plain hotkey="e" label="export Markdown" onPress={() => flash(actions.exportReport('md'))} />
549 <Button key="export-json" plain hotkey="j" label="export JSON" onPress={() => flash(actions.exportReport('json'))} />
550 </Box>,
551 <Text key="where" dimColor wrap="wrap">Files go to .claude/integrity/{task.id}.md|.json in the project. Headless: /integrity-report [md|json] [export].</Text>,
552 ])
553
554 const views: Record<IntegrityTab, () => RenderElement> = { overview, requirements, evidence, findings, claims, activity, report }
555
556 return (
557 <Box flexDirection="column" width={W}>
558 {header}
559 {tabBar}
560 <Box key="body" flexDirection="column" marginTop={isMini ? 0 : 1}>{views[ui.tab]()}</Box>
561 {footer}
562 </Box>
563 )
564}
565src/engine/types.ts 206 lines1// Claude Integrity engine: the data model. Plain JSON-safe types, no runtime
2// dependencies, so the engine runs in a Mod, a CLI or a CI job alike.
3
4export type Mode = 'observe' | 'review' | 'strict'
5export type NotifyLevel = 'off' | 'important' | 'all'
6export type Layout = 'auto' | 'mini' | 'compact' | 'wide'
7
8export type RequirementCategory = 'functional' | 'constraint' | 'non-functional' | 'test' | 'preservation'
9export type Criticality = 'must' | 'should' | 'optional'
10export type VerificationKind = 'static' | 'test' | 'runtime' | 'manual' | 'unknown'
11/** Stored status: what a person or the engine decided. Display state adds freshness. */
12export type RequirementStatus = 'unverified' | 'implemented' | 'verified' | 'failed' | 'waived'
13
14export type Requirement = {
15 id: string
16 description: string
17 category: RequirementCategory
18 criticality: Criticality
19 verification: VerificationKind
20 /** `explicit`: quoted from the prompt; `manual`: typed by the person; `proposed`: a model suggestion awaiting approval. */
21 origin: 'explicit' | 'manual' | 'proposed'
22 /** Command substrings whose passing run verifies this requirement (`npm test -- events`). */
23 checks: string[]
24 /** Set by a person: the code for it exists. Never by a model. */
25 implemented?: boolean
26 /** Evidence a person linked by hand (`/integrity-link`). */
27 evidenceIds: string[]
28 waiver?: { reason: string; at: number }
29}
30
31export type ContractIntent = 'production' | 'prototype' | 'unknown'
32
33export type ContractSnapshot = {
34 summary: string
35 requirements: Requirement[]
36 exclusions: string[]
37 scope: string[]
38 assumptions: string[]
39 intent: ContractIntent
40}
41
42export type Revision = { version: number; at: number; change: string; snapshot: ContractSnapshot }
43
44export type Contract = ContractSnapshot & {
45 taskId: string
46 prompt: { hash: string; excerpt: string; capturedAt: number; source: 'prompt' | 'manual' }
47 createdAt: number
48 updatedAt: number
49 version: number
50 approval: { status: 'draft' | 'approved'; at?: number; version?: number }
51 /** Frozen at first approval; amendments become revisions, never overwrite it. */
52 baseline?: ContractSnapshot
53 revisions: Revision[]
54}
55
56export type CheckKind = 'test' | 'build' | 'typecheck' | 'lint' | 'migration' | 'browser' | 'manual' | 'static' | 'other'
57export type EvidenceKind = 'tool-result' | 'static-analysis' | 'test-result' | 'runtime-check' | 'manual-confirmation'
58export type Outcome = 'pass' | 'fail' | 'inconclusive'
59
60export type Evidence = {
61 id: string
62 requirementIds: string[]
63 kind: EvidenceKind
64 check: CheckKind
65 outcome: Outcome
66 timestamp: number
67 /** The tool that ran (`Bash`, `mcp__playwright__browser_snapshot`, `person`). */
68 source: string
69 summary: string
70 /** Redacted, length-bounded command line. */
71 command?: string
72 /** Why the outcome is what it is: `exit 0`, `exit non-zero`, `denied`, `synthetic`, ... */
73 basis: string
74 /** Paths the command targeted; empty means the whole project. */
75 scope: string[]
76 counts?: { passed: number; failed: number; skipped: number }
77 truncated?: boolean
78 /** Answered by a plugin, not by the tool: never proof of execution. */
79 synthetic?: boolean
80 cwd?: string
81 revision?: string
82 agentId?: string
83 /** The task's change sequence when this evidence was recorded. */
84 seq: number
85 artifactPath?: string
86}
87
88export type Freshness = 'fresh' | 'stale' | 'unrelated'
89
90export type FileState = {
91 path: string
92 created: boolean
93 deleted: boolean
94 firstSeq: number
95 lastSeq: number
96 changes: number
97 hash?: string
98 /** Literal `/api/...` endpoints the file calls, for the cross-file endpoint check. */
99 calls?: { path: string; line: number }[]
100 /** The endpoint a route file serves (`app/api/events/route.ts` → `/api/events`). */
101 route?: string
102}
103
104export type Severity = 'high' | 'medium' | 'low' | 'info'
105/** `confirmed`: deterministic proof; `potential`: suspicious, needs a look; `expected`: matches the contract (a prototype's mock). */
106export type Confidence = 'confirmed' | 'potential' | 'expected'
107export type FindingStatus = 'open' | 'acknowledged' | 'waived' | 'resolved'
108
109export type Finding = {
110 id: string
111 ruleId: string
112 system: 'mirage' | 'drift' | 'policy'
113 category: string
114 file: string
115 line: number
116 snippet: string
117 explanation: string
118 evidence: string
119 severity: Severity
120 confidence: Confidence
121 requirementIds: string[]
122 verify: string
123 status: FindingStatus
124 waiver?: { reason: string; at: number }
125 firstSeq: number
126 lastSeq: number
127}
128
129export type ClaimKind =
130 | 'tests-pass' | 'build' | 'typecheck' | 'lint' | 'migration' | 'implemented'
131 | 'production-ready' | 'responsive' | 'browser' | 'live-data' | 'works'
132export type ClaimVerdict = 'supported' | 'partial' | 'unsupported' | 'contradicted' | 'not-assessable'
133
134export type Claim = {
135 id: string
136 turnId: string
137 text: string
138 kind: ClaimKind
139 verdict: ClaimVerdict
140 reason: string
141 evidenceIds: string[]
142 at: number
143}
144
145export type ActivityKind =
146 | 'task' | 'contract' | 'approval' | 'requirement' | 'evidence' | 'invalidated'
147 | 'finding' | 'resolved' | 'waiver' | 'claims' | 'mode' | 'turn'
148export type Activity = { at: number; kind: ActivityKind; text: string }
149
150export type Task = {
151 id: string
152 createdAt: number
153 updatedAt: number
154 contract: Contract
155 evidence: Evidence[]
156 files: Record<string, FileState>
157 findings: Finding[]
158 claims: Claim[]
159 activity: Activity[]
160 /** Monotonic count of observed source changes; evidence freshness is measured against it. */
161 seq: number
162 /** Keys of hints already shown, so a repeat never toasts twice. */
163 notified: string[]
164 nextEvidence: number
165}
166
167export type ProjectRecord = {
168 schema: 1
169 key: string
170 label: string
171 /** Bumped on every write; a writer that finds another rev merges instead of overwriting. */
172 rev: number
173 writer: string
174 updatedAt: number
175 activeTaskId?: string
176 mode: Mode
177 tasks: Record<string, Task>
178}
179
180/** A hint the adapter may show, already deduplicated by the engine. */
181export type Hint = { key: string; level: 'important' | 'info'; text: string }
182
183/** One observed tool call, normalised by the adapter from `tool.call`. */
184export type ToolObservation = {
185 tool: string
186 input: Record<string, unknown>
187 /** Core's answer had `ref`: the tool really ran. */
188 ranInCore: boolean
189 denied?: string
190 isError: boolean
191 text: string
192 result: unknown
193 cwd?: string
194 revision?: string
195 agentId?: string
196}
197
198export type FileChange = {
199 path: string
200 kind: 'create' | 'update' | 'delete'
201 /** The file's text after the change, when known (Write content, a read after Edit). */
202 after?: string
203 /** The text before, when the tool reported it. */
204 before?: string | null
205}
206src/engine/contract.ts 139 lines1// Intent Fingerprint: requirement contracts captured from the developer's own
2// words. Nothing here invents acceptance criteria: only sentences the prompt
3// states become requirements; everything else is typed by a person or marked
4// `proposed` and waits for approval.
5
6import type {
7 Contract, ContractIntent, ContractSnapshot, Criticality, Requirement, RequirementCategory,
8} from './types'
9import { clip, hash, isPathPattern, redact, sentences } from './util'
10
11const ACTION = /\b(implement|add|build|create|make|fix|refactor|enrich|update|change|replace|remove|delete|migrate|integrate|wire|connect|write|extend|improve|support|rename|port|convert|set up|hook up|ship)\b/i
12const QUESTION = /^(what|why|how|when|where|who|which|can|could|does|do|is|are|should|would|explain|tell me|show me)\b/i
13
14/** A prompt worth a contract: an instruction to change something, not a question or a one-liner. */
15export function isMeaningfulTask(prompt: string): boolean {
16 const text = prompt.trim()
17 if (text.length < 25 || text.startsWith('/')) return false
18 if (QUESTION.test(text) && text.endsWith('?')) return false
19 return ACTION.test(text)
20}
21
22const NEGATIVE = /\b(do not|don't|dont|never|must not|mustn't|should not|shouldn't|avoid|without)\b/i
23const PRESERVE = /\b(preserve|keep|retain|maintain|leave)\b.*\b(existing|current|unchanged|intact|as is|as-is|navigation|behaviou?r|api|route|layout)\b|\b(existing|current)\b.*\b(must|should)\b.*\b(stay|remain|work)\b/i
24const MUST = /\b(must|need to|needs to|has to|have to|required|require|ensure|make sure)\b/i
25const SHOULD = /\b(should|ideally|prefer)\b/i
26const TESTS = /\b(tests?|spec|coverage|e2e|end-to-end)\b/i
27const PERF = /\b(fast|performance|latency|accessib|a11y|responsive|secure|security)\b/i
28
29function categorize(sentence: string): { category: RequirementCategory; criticality: Criticality } | undefined {
30 if (PRESERVE.test(sentence)) return { category: 'preservation', criticality: 'must' }
31 if (NEGATIVE.test(sentence)) {
32 return { category: /\b(existing|current|navigation|route)\b/i.test(sentence) ? 'preservation' : 'constraint', criticality: 'must' }
33 }
34 if (TESTS.test(sentence) && MUST.test(sentence)) return { category: 'test', criticality: 'must' }
35 if (MUST.test(sentence)) return { category: PERF.test(sentence) ? 'non-functional' : 'functional', criticality: 'must' }
36 if (SHOULD.test(sentence)) return { category: PERF.test(sentence) ? 'non-functional' : 'functional', criticality: 'should' }
37 return undefined
38}
39
40export function detectIntent(text: string): ContractIntent {
41 const prototype = /\b(prototype|mock(ed|up)?s?|mockup|demo|placeholder|sample data|fake data|fixtures?|wireframe|stub(bed)?|poc|proof of concept|static)\b/i.test(text)
42 const production = /\b(live|real|production|persist(ed|ence)?|database|db|backend|api|endpoint|server|save[ds]?|store[ds]?|fetch(ed)?)\b/i.test(text)
43 if (prototype && !/\b(not|no|instead of|replace)\b[^.]*\b(mock|placeholder|fake|static|fixture)/i.test(text)) return 'prototype'
44 return production ? 'production' : 'unknown'
45}
46
47/** Path-like tokens a "do not touch/modify/create" sentence names become exclusion globs. */
48function exclusionsFrom(sentence: string): string[] {
49 if (!/\b(do not|don't|never|must not)\b.*\b(touch|modify|change|edit|create|delete|remove)\b/i.test(sentence)) return []
50 return (sentence.match(/[`'"]?[\w.*/-]+[`'"]?/g) ?? [])
51 .map(t => t.replace(/[`'"]/g, '').replace(/[.,;:]+$/, ''))
52 .filter(t => isPathPattern(t) && /[/*]/.test(t))
53}
54
55function nextRequirementId(requirements: readonly Requirement[]): string {
56 const n = requirements.reduce((max, r) => Math.max(max, Number(r.id.slice(1)) || 0), 0)
57 return `R${n + 1}`
58}
59
60export function makeRequirement(
61 existing: readonly Requirement[],
62 description: string,
63 fields: Partial<Omit<Requirement, 'id' | 'description'>> = {},
64): Requirement {
65 const category = fields.category ?? 'functional'
66 return {
67 id: nextRequirementId(existing),
68 description: clip(redact(description), 300),
69 category,
70 criticality: fields.criticality ?? 'must',
71 verification: fields.verification ?? (category === 'test' ? 'test' : category === 'preservation' || category === 'constraint' ? 'static' : 'unknown'),
72 origin: fields.origin ?? 'manual',
73 checks: fields.checks ?? [],
74 evidenceIds: [],
75 ...(fields.implemented === undefined ? {} : { implemented: fields.implemented }),
76 }
77}
78
79/**
80 * Captures a contract from a prompt without a model: the request is kept by
81 * hash and a short redacted excerpt; explicit constraints become requirements.
82 */
83export function createContract(taskId: string, prompt: string, now: number, source: 'prompt' | 'manual'): Contract {
84 const requirements: Requirement[] = []
85 const exclusions: string[] = []
86 for (const sentence of sentences(prompt)) {
87 const kind = categorize(sentence)
88 if (kind !== undefined) requirements.push(makeRequirement(requirements, sentence, { ...kind, origin: 'explicit' }))
89 exclusions.push(...exclusionsFrom(sentence))
90 }
91 const first = sentences(prompt)[0] ?? prompt
92 return {
93 taskId,
94 summary: clip(redact(first), 140),
95 requirements,
96 exclusions: [...new Set(exclusions)],
97 scope: [],
98 assumptions: [],
99 intent: detectIntent(prompt),
100 prompt: { hash: hash(prompt), excerpt: clip(redact(prompt), 280), capturedAt: now, source },
101 createdAt: now,
102 updatedAt: now,
103 version: 1,
104 approval: { status: 'draft' },
105 revisions: [],
106 }
107}
108
109export const snapshot = (c: Contract): ContractSnapshot =>
110 JSON.parse(JSON.stringify({
111 summary: c.summary, requirements: c.requirements, exclusions: c.exclusions,
112 scope: c.scope, assumptions: c.assumptions, intent: c.intent,
113 })) as ContractSnapshot
114
115export function approve(c: Contract, now: number): void {
116 c.approval = { status: 'approved', at: now, version: c.version }
117 c.baseline ??= snapshot(c)
118 c.updatedAt = now
119}
120
121/**
122 * Every change after the first approval is a versioned revision: the baseline
123 * stays as approved, and a change to what is required asks for re-approval.
124 */
125export function amend(c: Contract, now: number, change: string, mutate: (c: Contract) => void, needsReapproval = true): void {
126 mutate(c)
127 c.updatedAt = now
128 if (c.baseline === undefined) return
129 c.version += 1
130 c.revisions.push({ version: c.version, at: now, change: clip(change, 200), snapshot: snapshot(c) })
131 if (c.revisions.length > 50) c.revisions.splice(0, c.revisions.length - 50)
132 if (needsReapproval) c.approval = { ...c.approval, status: 'draft' }
133}
134
135export const findRequirement = (c: Contract, id: string): Requirement | undefined =>
136 c.requirements.find(r => r.id.toLowerCase() === id.trim().toLowerCase())
137
138export const isApproved = (c: Contract): boolean => c.approval.status === 'approved'
139src/engine/evidence.ts 216 lines1// The Witness: evidence from observed tool calls, its provenance and freshness.
2// Outcomes come from the tool's completion state (core's isError, interrupt,
3// background, deny, synthetic answer) first; output text only refines it, and
4// never turns a failed or unobserved run into a pass.
5
6import type { CheckKind, Evidence, EvidenceKind, Freshness, Outcome, Task, ToolObservation } from './types'
7import { clip, isCodePath, isDocPath, redact } from './util'
8
9type Classified = { check: CheckKind; scope: string[] }
10
11const CHECKS: readonly [CheckKind, RegExp][] = [
12 ['browser', /\b(playwright\s+test|cypress\s+run|wdio\s+run|testcafe)\b/i],
13 ['migration', /\b(prisma\s+migrate|drizzle-kit\s+(push|migrate)|knex\s+migrate|sequelize\s+db:migrate|alembic\s+upgrade|rails\s+db:migrate|manage\.py\s+migrate|typeorm\s+migration:run|supabase\s+db\s+push)\b/i],
14 ['test', /\b((npm|pnpm|yarn|bun)\s+(run\s+)?test(:\w+)?|jest|vitest|mocha|ava|pytest|python3?\s+-m\s+(pytest|unittest)|go\s+test|cargo\s+test|dotnet\s+test|node\s+--test|deno\s+test|claude\s+plugin\s+test|rspec|phpunit|(mvn|mvnw)(\s+-\S+)*\s+test|gradlew?\s+test)\b/i],
15 ['typecheck', /\b(tsc|vue-tsc|mypy|pyright|(npm|pnpm|yarn|bun)\s+(run\s+)?(typecheck|type-check|check-types))\b|typescript\.js/i],
16 ['lint', /\b(eslint|biome\s+(check|lint)|ruff|flake8|clippy|golangci-lint|prettier\s+--check|(npm|pnpm|yarn|bun)\s+(run\s+)?lint|claude\s+plugin\s+validate)\b/i],
17 ['build', /\b((npm|pnpm|yarn|bun)\s+(run\s+)?build|next\s+build|vite\s+build|cargo\s+build|go\s+build|dotnet\s+build|webpack|(mvn|mvnw)(\s+-\S+)*\s+(package|install|compile|verify)|gradlew?\s+(build|assemble))\b/i],
18]
19
20const SUBCOMMANDS = new Set(['run', 'test', 'exec', 'watch', 'npx', 'npm', 'pnpm', 'yarn', 'bun', '--'])
21
22/** Paths or filters a check names: `vitest run src/events` scopes to `src/events`. Empty: the whole project. */
23function scopeOf(segment: string, match: RegExpMatchArray): string[] {
24 const rest = segment.slice((match.index ?? 0) + match[0].length)
25 return rest
26 .split(/\s+/)
27 .map(t => t.replace(/^['"]|['"]$/g, '').split('::')[0] ?? '')
28 .filter(t => t !== '' && !t.startsWith('-') && !SUBCOMMANDS.has(t) && (/[/\\]/.test(t) || /\.\w{1,5}$/.test(t)))
29 .map(t => t.replace(/\\/g, '/').replace(/^\.\//, ''))
30}
31
32/** Splits a shell line into its `&&`/`;`/`||` segments and classifies each that is a known check. */
33export function classifyCommand(command: string): Classified[] {
34 const found: Classified[] = []
35 for (const segment of command.split(/&&|\|\||;|\n/)) {
36 const head = segment.split('|')[0] ?? ''
37 for (const [check, pattern] of CHECKS) {
38 const match = head.match(pattern)
39 if (match) {
40 found.push({ check, scope: scopeOf(head, match) })
41 break
42 }
43 }
44 }
45 return found
46}
47
48/** A pipe after the check, `|| true` or `; exit 0` hide the check's own exit status. */
49export function isExitMasked(command: string): boolean {
50 return /\|\|\s*(true|:|exit\s+0)\b|;\s*(true|exit\s+0)\s*$/.test(command) ||
51 command.split(/&&|;|\n/).some(seg => {
52 const pipe = seg.indexOf('|')
53 return pipe >= 0 && seg[pipe + 1] !== '|' && classifyCommand(seg.slice(0, pipe)).length > 0 && !/set\s+-o\s+pipefail/.test(command)
54 })
55}
56
57export type Counts = { passed: number; failed: number; skipped: number }
58
59/** Reads the summary line of common runners (jest, vitest, mocha, pytest, cargo, go, node:test). */
60export function parseCounts(output: string): Counts | undefined {
61 const lines = output.split(/\r?\n/)
62 const summary =
63 lines.filter(l => /^\s*Tests?:?\s/i.test(l) && /\d/.test(l)).at(-1) ??
64 lines.filter(l => /test result:|^=+ .*\d+ (passed|failed)|\d+ (passing|failing)/i.test(l)).join(' ') ??
65 ''
66 const text = summary !== '' ? summary : lines.filter(l => /^#\s*(pass|fail|skipped|todo)\s+\d+/i.test(l) || /\b\d+\s+(passed|failed|skipped)\b/i.test(l)).join(' ')
67 if (text === '') {
68 const goFails = lines.filter(l => /^--- FAIL:/.test(l)).length
69 const goOk = lines.filter(l => /^ok\s+\S+/.test(l)).length
70 return goFails + goOk > 0 ? { passed: goOk, failed: goFails, skipped: 0 } : undefined
71 }
72 const sum = (re: RegExp): number => [...text.matchAll(re)].reduce((n, m) => n + Number(m[1] ?? m[2] ?? 0), 0)
73 return {
74 passed: sum(/(\d+)\s+(?:passed|passing)\b|#\s*pass\s+(\d+)/gi),
75 failed: sum(/(\d+)\s+(?:failed|failing)\b|#\s*fail\s+(\d+)/gi),
76 skipped: sum(/(\d+)\s+(?:skipped|pending|ignored|todo)\b|#\s*(?:skipped|todo)\s+(\d+)/gi),
77 }
78}
79
80const EVIDENCE_KIND: Record<CheckKind, EvidenceKind> = {
81 test: 'test-result', browser: 'runtime-check', migration: 'runtime-check', build: 'tool-result',
82 typecheck: 'tool-result', lint: 'tool-result', manual: 'manual-confirmation', static: 'static-analysis', other: 'tool-result',
83}
84
85type BashRecord = {
86 interrupted?: boolean
87 backgroundTaskId?: string
88 timedOutAfterMs?: number
89 persistedOutputPath?: string
90 stdout?: string
91 stderr?: string
92}
93
94/**
95 * Evidence from one observed Bash call, or none when the command runs no
96 * known check. One record per check the line runs.
97 */
98export function evidenceFromBash(obs: ToolObservation, seq: number, now: number, nextId: () => string): Evidence[] {
99 const command = typeof obs.input['command'] === 'string' ? obs.input['command'] : ''
100 const checks = classifyCommand(command)
101 if (checks.length === 0) return []
102 const record = (obs.result ?? {}) as BashRecord
103 const output = `${record.stdout ?? ''}\n${record.stderr ?? ''}\n${obs.text}`
104 const counts = parseCounts(output)
105 const truncated = record.persistedOutputPath !== undefined || /output too large|\[truncated\]|… \d+ (more )?lines/i.test(obs.text)
106 const masked = isExitMasked(command)
107
108 let outcome: Outcome
109 let basis: string
110 if (obs.denied !== undefined) [outcome, basis] = ['inconclusive', 'permission denied: not executed']
111 else if (!obs.ranInCore) [outcome, basis] = ['inconclusive', 'synthetic: answered by a plugin, not executed']
112 else if (record.backgroundTaskId !== undefined) [outcome, basis] = ['inconclusive', 'moved to background: completion not observed']
113 else if (record.timedOutAfterMs !== undefined) [outcome, basis] = ['inconclusive', `timed out after ${record.timedOutAfterMs} ms`]
114 else if (record.interrupted === true) [outcome, basis] = ['inconclusive', 'interrupted']
115 else if (counts !== undefined && counts.failed > 0) [outcome, basis] = ['fail', masked ? 'output reports failures (exit status masked)' : 'output reports failures']
116 else if (masked) {
117 ;[outcome, basis] = counts !== undefined && counts.passed > 0
118 ? ['pass', 'output summary; exit status masked by a pipe or `|| true`']
119 : ['inconclusive', 'exit status masked by a pipe or `|| true`']
120 } else if (obs.isError) [outcome, basis] = ['fail', exitCode(obs.text) ?? 'exit non-zero']
121 else if ((counts === undefined || counts.passed + counts.failed === 0) && /no tests? (found|ran)|no test files/i.test(output)) {
122 ;[outcome, basis] = ['inconclusive', 'exit 0 but no tests ran']
123 } else [outcome, basis] = ['pass', 'exit 0']
124
125 return checks.map((c, i) => {
126 // In a chain that failed, only a check whose own output says so is known to have failed.
127 const chainFailed = checks.length > 1 && outcome === 'fail' && !(c.check === 'test' && (counts?.failed ?? 0) > 0)
128 const own = chainFailed ? 'inconclusive' : outcome
129 const ownBasis = chainFailed ? `${basis}; chained with other commands, failing step unknown` : basis
130 const tally = c.check === 'test' || c.check === 'browser' ? counts : undefined
131 return {
132 id: nextId(),
133 requirementIds: [],
134 kind: EVIDENCE_KIND[c.check],
135 check: c.check,
136 outcome: own,
137 timestamp: now + i,
138 source: obs.tool,
139 summary: summarize(c.check, own, tally, truncated),
140 command: clip(redact(command), 240),
141 basis: ownBasis,
142 scope: c.scope,
143 ...(tally === undefined ? {} : { counts: tally }),
144 ...(truncated ? { truncated } : {}),
145 ...(obs.ranInCore ? {} : { synthetic: true }),
146 ...(obs.cwd === undefined ? {} : { cwd: redact(obs.cwd) }),
147 ...(obs.revision === undefined ? {} : { revision: obs.revision }),
148 ...(obs.agentId === undefined ? {} : { agentId: obs.agentId }),
149 seq,
150 }
151 })
152}
153
154const exitCode = (text: string): string | undefined => {
155 const m = text.match(/exit (?:code|status)[:\s]+(\d+)/i)
156 return m ? `exit ${m[1]}` : undefined
157}
158
159function summarize(check: CheckKind, outcome: Outcome, counts: Counts | undefined, truncated: boolean): string {
160 const tally = counts === undefined ? '' : ` (${counts.passed} passed, ${counts.failed} failed${counts.skipped ? `, ${counts.skipped} skipped` : ''})`
161 return `${check} ${outcome}${tally}${truncated ? ', output truncated' : ''}`
162}
163
164/** A browser automation call is an observed session, not an assertion: inconclusive until a person confirms. */
165export function evidenceFromBrowserTool(obs: ToolObservation, seq: number, now: number, id: string): Evidence | undefined {
166 if (!/^mcp__.*(playwright|chrome|browser|puppeteer)/i.test(obs.tool)) return undefined
167 return {
168 id, requirementIds: [], kind: 'runtime-check', check: 'browser', outcome: 'inconclusive', timestamp: now,
169 source: obs.tool, summary: 'browser session observed (no assertion)', basis: obs.ranInCore ? 'interaction only' : 'synthetic',
170 scope: [], seq, ...(obs.ranInCore ? {} : { synthetic: true }),
171 }
172}
173
174const CONFIG = /(^|\/)(package\.json|tsconfig[^/]*\.json|pnpm-lock\.yaml|package-lock\.json|yarn\.lock|vite\.config\.\w+|jest\.config\.\w+|vitest\.config\.\w+|next\.config\.\w+|babel\.config\.\w+)$/i
175
176const stem = (path: string): string => (path.split('/').pop() ?? path).replace(/\.(test|spec)(?=\.)/i, '').replace(/\.\w+$/, '').toLowerCase()
177const dir = (path: string): string => path.includes('/') ? path.slice(0, path.lastIndexOf('/')) : ''
178
179/** Whether a change at `path` can affect evidence scoped to `scope`. */
180export function relates(path: string, scope: string): boolean {
181 const p = path.toLowerCase()
182 const s = scope.toLowerCase().replace(/\/$/, '')
183 if (p === s || p.startsWith(`${s}/`)) return true
184 if (stem(p) === stem(s)) return true
185 const sd = /\.\w{1,5}$/.test(s) ? dir(s) : s
186 return sd !== '' && dir(p) === sd
187}
188
189/** Changes that invalidate `e`: code or config, never docs, made after it was recorded. */
190function invalidators(e: Evidence, task: Task): string[] {
191 return Object.values(task.files)
192 .filter(f => f.lastSeq > e.seq && !isDocPath(f.path) && (isCodePath(f.path) || CONFIG.test(f.path)))
193 .filter(f => e.scope.length === 0 || CONFIG.test(f.path) || e.scope.some(s => relates(f.path, s)))
194 .map(f => f.path)
195}
196
197export function freshness(e: Evidence, task: Task): Freshness {
198 if (invalidators(e, task).length > 0) return 'stale'
199 if (e.scope.length > 0 && Object.keys(task.files).length > 0 &&
200 !Object.keys(task.files).some(p => e.scope.some(s => relates(p, s)))) return 'unrelated'
201 return 'fresh'
202}
203
204export const staleBecause = (e: Evidence, task: Task): string[] => invalidators(e, task)
205
206/** Evidence that counts as proof: a real, passing, current run. */
207export const isProof = (e: Evidence, task: Task): boolean =>
208 e.outcome === 'pass' && e.synthetic !== true && freshness(e, task) !== 'stale'
209
210/** The latest record of each check kind, the one the summary shows. */
211export function latestByCheck(task: Task): Map<CheckKind, Evidence> {
212 const out = new Map<CheckKind, Evidence>()
213 for (const e of task.evidence) if (e.check !== 'manual' && e.check !== 'static') out.set(e.check, e)
214 return out
215}
216src/engine/mirage.ts 340 lines1// Mirage Detector: deterministic rules for fake completeness in JS/TS/React
2// sources. Each rule reads one file's text against the contract's intent and
3// says how sure it is. Nothing here is a verdict on its own: `confirmed` is
4// kept for code that literally says it is not implemented; everything else
5// is `potential` (look at it) or `expected` (the contract asked for it).
6//
7// Token-level, not an AST: the Mod runtime has no TypeScript compiler, and a
8// regex over comment-stripped text with brace matching covers these patterns
9// with known, documented blind spots (docs/DETECTION-RULES.md).
10
11import type { Confidence, ContractIntent, Severity } from './types'
12import { clip, hash, isTestPath } from './util'
13
14export type RuleContext = {
15 path: string
16 source: string
17 /** The source with comments blanked, offsets and newlines kept. */
18 code: string
19 intent: ContractIntent
20 /** The contract's own words, lower-cased: rules read what was asked for. */
21 asked: string
22}
23
24export type RawFinding = {
25 ruleId: string
26 category: string
27 line: number
28 snippet: string
29 explanation: string
30 evidence: string
31 severity: Severity
32 confidence: Confidence
33 verify: string
34}
35
36export type MirageRule = {
37 id: string
38 category: string
39 title: string
40 appliesTo: (path: string) => boolean
41 detect: (ctx: RuleContext) => RawFinding[]
42}
43
44/** Blanks `//` and `/* *\/` comments outside strings, keeping every offset. */
45export function stripComments(src: string): string {
46 let out = ''
47 let quote: string | undefined
48 for (let i = 0; i < src.length; i++) {
49 const c = src[i] as string
50 if (quote !== undefined) {
51 out += c
52 if (c === '\\') out += src[++i] ?? ''
53 else if (c === quote) quote = undefined
54 } else if (c === '"' || c === "'" || c === '`') {
55 quote = c
56 out += c
57 } else if (c === '/' && src[i + 1] === '/') {
58 while (i < src.length && src[i] !== '\n') { out += ' '; i++ }
59 if (i < src.length) out += '\n'
60 } else if (c === '/' && src[i + 1] === '*') {
61 const end = src.indexOf('*/', i + 2)
62 const stop = end < 0 ? src.length : end + 2
63 out += src.slice(i, stop).replace(/[^\n]/g, ' ')
64 i = stop - 1
65 } else out += c
66 }
67 return out
68}
69
70export const lineOf = (src: string, index: number): number => src.slice(0, index).split('\n').length
71
72const lineText = (src: string, index: number): string => {
73 const start = src.lastIndexOf('\n', index - 1) + 1
74 const end = src.indexOf('\n', index)
75 return src.slice(start, end < 0 ? src.length : end)
76}
77
78/** Index of the `}` closing the `{` at `open`, or the end of the text. */
79function closeOf(code: string, open: number): number {
80 let depth = 0
81 for (let i = open; i < code.length; i++) {
82 if (code[i] === '{') depth++
83 else if (code[i] === '}' && --depth === 0) return i
84 }
85 return code.length
86}
87
88/** The body of the innermost function around `index`: `function ... {` or `=> {`. */
89export function enclosingFunction(code: string, index: number): { start: number; body: string } | undefined {
90 const starts = [...code.slice(0, index).matchAll(/(\bfunction\b[^{;]*|=>\s*)\{/g)]
91 for (let i = starts.length - 1; i >= 0; i--) {
92 const m = starts[i] as RegExpMatchArray
93 const open = (m.index ?? 0) + m[0].length - 1
94 const close = closeOf(code, open)
95 if (close >= index) return { start: m.index ?? 0, body: code.slice(open, close + 1) }
96 }
97 return undefined
98}
99
100const byIntent = (intent: ContractIntent, production: Severity, unknown: Severity): Severity =>
101 intent === 'production' ? production : unknown
102
103const find = (ctx: RuleContext, re: RegExp, make: (m: RegExpMatchArray) => Omit<RawFinding, 'line' | 'snippet'> | undefined): RawFinding[] =>
104 [...ctx.code.matchAll(re)].flatMap(m => {
105 const made = make(m)
106 const at = m.index ?? 0
107 return made === undefined ? [] : [{ ...made, line: lineOf(ctx.code, at), snippet: clip(lineText(ctx.source, at), 140) }]
108 })
109
110const SOURCE = (path: string): boolean => /\.(c|m)?(t|j)sx?$/i.test(path) && !isTestPath(path) && !/\.d\.ts$/i.test(path)
111const UI = (path: string): boolean => SOURCE(path) && (/\.(j|t)sx$/i.test(path) || /(^|\/)(components?|pages|app|views|screens|routes)\//i.test(path))
112const FIXTURE_PATH = /(^|\/)(mocks?|__mocks__|fixtures?|stories|storybook|seed|demo)(\/|\.)|\.(stories|mock|fixture)\.\w+$/i
113
114/** Route handlers: Next.js app/pages API, Express-style `app.post(...)`. */
115const HANDLER = /export\s+(?:async\s+)?function\s+(POST|PUT|PATCH|DELETE)\b|export\s+const\s+(POST|PUT|PATCH|DELETE)\s*=|\b(?:app|router|server)\.(post|put|patch|delete)\s*\(/g
116const IO = /\b(await\s+(?!(?:req|request)\.(?:json|text|formData)\(\))[\w.]+\(|prisma|db\.|sql`|\.query\(|\.insert|\.update\(|\.upsert|\.save\(|\.create\(|\.delete\(|writeFile|fetch\(|axios|supabase|firestore|mongo|redis|knex|drizzle|kv\.|repository\.|\w+Service\.)/
117const SUCCESS = /\b(success|ok)\s*:\s*true|status\(\s*20[01]\s*\)|status:\s*20[01]|Response\.json\(|NextResponse\.json\(|res\.(json|send)\(|return\s+new\s+Response\(/
118
119export const RULES: readonly MirageRule[] = [
120 {
121 id: 'MIR-E1', category: 'E. Placeholder implementation', title: 'Throws "not implemented"', appliesTo: SOURCE,
122 detect: ctx => find(ctx, /throw\s+new\s+\w*Error\(\s*['"`](not\s+(yet\s+)?implemented|unimplemented|todo)\b/gi, () => ({
123 ruleId: 'MIR-E1', category: 'E. Placeholder implementation', severity: 'high', confidence: 'confirmed',
124 explanation: 'This code path throws "not implemented": the behaviour does not exist yet.',
125 evidence: 'Literal throw of a not-implemented error in a non-test source file.',
126 verify: 'Implement the branch, or record it as an explicit exclusion in the contract.',
127 })),
128 },
129 {
130 id: 'MIR-E2', category: 'E. Placeholder implementation', title: 'TODO that defers the real work', appliesTo: SOURCE,
131 // TODOs live in comments, so this rule reads the raw source; plain TODOs are not flagged.
132 detect: ctx => [...ctx.source.matchAll(/(\/\/|\/\*|\{\s*\/\*)\s*(TODO|FIXME|XXX)\b[:\s-]*(.{0,120})/g)]
133 .filter(m => /\b(implement|wire|hook up|connect|persist|save|real|call (the )?api|backend|replace|fetch|auth)/i.test(m[3] ?? ''))
134 .map(m => ({
135 ruleId: 'MIR-E2', category: 'E. Placeholder implementation', line: lineOf(ctx.source, m.index ?? 0),
136 snippet: clip(lineText(ctx.source, m.index ?? 0), 140), severity: byIntent(ctx.intent, 'medium', 'low'), confidence: 'potential' as Confidence,
137 explanation: 'A TODO says the real behaviour is still to be wired up here.',
138 evidence: `Comment: "${clip(m[3] ?? '', 80)}"`,
139 verify: 'Confirm the deferred work is done elsewhere, or track it as unfinished.',
140 })),
141 },
142 {
143 id: 'MIR-E3', category: 'E. Placeholder implementation', title: 'Placeholder endpoint or key', appliesTo: SOURCE,
144 detect: ctx => find(ctx, /['"`](https?:\/\/(api\.)?example\.(com|org)[^'"`]*|\/api\/(todo|placeholder|xxx|example)[^'"`]*|YOUR_[A-Z_]+|<your[^>]*>|REPLACE_ME)['"`]/gi, m => ({
145 ruleId: 'MIR-E3', category: 'E. Placeholder implementation', severity: byIntent(ctx.intent, 'high', 'medium'), confidence: 'potential',
146 explanation: 'A placeholder URL or credential stands where a real one is needed.',
147 evidence: `Literal ${clip(m[1] ?? '', 60)}`,
148 verify: 'Point it at the real endpoint or configuration value.',
149 })),
150 },
151 {
152 id: 'MIR-B1', category: 'B. Simulated success', title: 'Handler returns success without doing the work',
153 appliesTo: SOURCE,
154 detect: ctx => [...ctx.code.matchAll(HANDLER)].flatMap(m => {
155 const at = m.index ?? 0
156 const open = ctx.code.indexOf('{', at)
157 if (open < 0) return []
158 const body = ctx.code.slice(open, closeOf(ctx.code, open) + 1)
159 if (!SUCCESS.test(body) || IO.test(body)) return []
160 const sev: Severity = ctx.intent === 'prototype' ? 'info' : byIntent(ctx.intent, 'high', 'medium')
161 return [{
162 ruleId: 'MIR-B1', category: 'B. Simulated success', line: lineOf(ctx.code, at), snippet: clip(lineText(ctx.source, at), 140),
163 severity: sev, confidence: (ctx.intent === 'prototype' ? 'expected' : 'potential') as Confidence,
164 explanation: `The ${m[1] ?? m[2] ?? m[3]?.toUpperCase()} handler answers success but performs no write, query or outbound call.`,
165 evidence: 'Success response in the handler body; no database, fetch, service or file call found in it.',
166 verify: 'Run a request against the endpoint and confirm the change is persisted.',
167 }]
168 }),
169 },
170 {
171 id: 'MIR-B2', category: 'B. Simulated success', title: 'Artificial delay standing in for a request', appliesTo: SOURCE,
172 detect: ctx => find(ctx, /await\s+new\s+Promise\s*\(\s*\(?\s*\w+\s*\)?\s*=>\s*setTimeout\s*\(|await\s+(sleep|delay|wait)\s*\(\s*\d+/g, m => {
173 const fn = enclosingFunction(ctx.code, m.index ?? 0)
174 const body = fn?.body ?? ''
175 const fakesSuccess = /\b(setSuccess|setSaved|setDone|toast\.success|setStatus\(\s*['"](success|saved|done))|success\s*:\s*true/i.test(body)
176 if (!fakesSuccess || /\bfetch\(|axios|mutate|api\.\w+\(/.test(body)) return undefined
177 return {
178 ruleId: 'MIR-B2', category: 'B. Simulated success', severity: ctx.intent === 'prototype' ? 'info' : 'high',
179 confidence: ctx.intent === 'prototype' ? 'expected' : 'potential',
180 explanation: 'A timer delay is followed by a success state, with no request in between: it imitates a backend call.',
181 evidence: 'setTimeout/sleep plus a success state in the same function; no fetch or mutation call.',
182 verify: 'Replace the delay with the real request and check its failure path.',
183 }
184 }),
185 },
186 {
187 id: 'MIR-C1', category: 'C. Non-functional control', title: 'Control with an empty or log-only action', appliesTo: UI,
188 detect: ctx => find(ctx, /\b(on(?:Click|Press|Submit|Change|Select))=\{\s*(\(\s*\w*\s*\)\s*=>\s*(\{\s*\}|null|undefined|console\.\w+\([^)]*\)|alert\([^)]*\)|\w+\.preventDefault\(\))|undefined|null)\s*\}|\bhref=["'](#|javascript:void\(0\);?)?["']/g, m => ({
189 ruleId: 'MIR-C1', category: 'C. Non-functional control', severity: 'medium', confidence: 'potential',
190 explanation: 'This control is drawn but its action does nothing (empty, log-only or placeholder link).',
191 evidence: `Handler: ${clip(m[0], 80)}`,
192 verify: 'Press the control and confirm it does what the requirement says.',
193 })),
194 },
195 {
196 id: 'MIR-A1', category: 'A. Mock data as production data', title: 'Hardcoded or fixture data in a shipped view',
197 appliesTo: path => UI(path) && !FIXTURE_PATH.test(path),
198 detect: ctx => {
199 const named = /\b(?:const|let|var)\s+((?:mock|fake|dummy|sample|placeholder|hardcoded|static|test)\w*)\s*(?::[^=]+)?=\s*[[{]/gi
200 const imported = /\bimport\s+[^;]*?from\s+['"]([^'"]*(?:mocks?|fixtures?|sample|fake|dummy)[^'"]*)['"]/gi
201 const confidence: Confidence = ctx.intent === 'prototype' ? 'expected' : 'potential'
202 const severity: Severity = ctx.intent === 'prototype' ? 'info' : byIntent(ctx.intent, 'high', 'medium')
203 const why = ctx.intent === 'prototype'
204 ? 'Fixture data, as the contract asks for a prototype.'
205 : 'A view renders fixture data; nothing shows it is replaced by real data.'
206 return [
207 ...find(ctx, named, m => ({
208 ruleId: 'MIR-A1', category: 'A. Mock data as production data', severity, confidence, explanation: why,
209 evidence: `Data declared as "${m[1]}" in a UI source file.`,
210 verify: 'Check where the view gets its data at runtime; trace it to the API or store.',
211 })),
212 ...find(ctx, imported, m => ({
213 ruleId: 'MIR-A1', category: 'A. Mock data as production data', severity, confidence, explanation: why,
214 evidence: `Imports from "${m[1]}".`,
215 verify: 'Check where the view gets its data at runtime; trace it to the API or store.',
216 })),
217 ]
218 },
219 },
220 {
221 id: 'MIR-D1', category: 'D. Incomplete data integration', title: 'Fetched response is discarded', appliesTo: SOURCE,
222 detect: ctx => find(ctx, /\b(?:const|let)\s+(\w+)\s*=\s*await\s+(fetch|axios(?:\.\w+)?|api\.\w+|\w+Client\.\w+)\s*\(/g, m => {
223 const name = m[1] as string
224 const uses = ctx.code.match(new RegExp(`\\b${name}\\b`, 'g'))?.length ?? 0
225 if (uses > 1) return undefined
226 return {
227 ruleId: 'MIR-D1', category: 'D. Incomplete data integration', severity: byIntent(ctx.intent, 'high', 'medium'), confidence: 'potential',
228 explanation: `The response "${name}" is requested and never read: the request runs but its data is not used.`,
229 evidence: `"${name}" is assigned from ${m[2]}(...) and referenced nowhere else in the file.`,
230 verify: 'Confirm the UI renders the response, not static data.',
231 }
232 }),
233 },
234 {
235 id: 'MIR-F1', category: 'F. Missing persistence', title: 'Submit updates local state only',
236 appliesTo: UI,
237 detect: ctx => {
238 if (!/\b(persist|save|saved|store|database|db|backend|server|api)\b/.test(ctx.asked) && ctx.intent !== 'production') return []
239 return find(ctx, /\b(?:const|function)\s+(handleSubmit|onSubmit|submit\w*|save\w*|handleSave\w*|create\w*)\b/g, m => {
240 const fn = enclosingFunction(ctx.code, (m.index ?? 0) + m[0].length + 40)
241 if (fn === undefined || fn.start < (m.index ?? 0)) return undefined
242 const body = fn.body
243 if (!/\bset[A-Z]\w*\(|toast|alert\(/.test(body)) return undefined
244 if (/\bfetch\(|axios|mutate|mutation|\bapi\.|\bpost\(|\bput\(|\bdb\.|prisma|supabase|firestore|localStorage|indexedDB|action\(|Action\(|startTransition|\bsave\w*\(|\bcreate\w*\(|\bupdate\w*\(/.test(body.replace(/\bset[A-Z]\w*\(/g, ''))) return undefined
245 return {
246 ruleId: 'MIR-F1', category: 'F. Missing persistence', severity: 'high', confidence: 'potential',
247 explanation: `${m[1]} changes component state but sends nothing anywhere, while the task asks for saved data.`,
248 evidence: 'State setter or toast in the handler; no request, mutation or storage call.',
249 verify: 'Submit, reload the page, and confirm the data is still there.',
250 }
251 })
252 },
253 },
254 {
255 id: 'MIR-G1', category: 'G. Authorization omission', title: 'Permission enforced in the UI only', appliesTo: SOURCE,
256 detect: ctx => {
257 const out: RawFinding[] = []
258 const mutatesClient = /method:\s*['"](POST|PUT|PATCH|DELETE)['"]/i.test(ctx.code)
259 if (UI(ctx.path) && mutatesClient) {
260 out.push(...find(ctx, /\b(?:user|session|currentUser|me)\??\.(?:role|isAdmin|permissions?)\b|\bisAdmin\s*&&|\bcan\w*\s*&&\s*</g, () => ({
261 ruleId: 'MIR-G1', category: 'G. Authorization omission', severity: 'medium', confidence: 'potential',
262 explanation: 'The UI hides or shows an action by role; a hidden button is not authorization.',
263 evidence: 'Client-side role check in a component that sends a mutating request.',
264 verify: 'Call the endpoint as an unprivileged user and confirm it is refused.',
265 })).slice(0, 1))
266 }
267 if (/\b(auth|permission|admin|role|owner|only)\b/.test(ctx.asked)) {
268 out.push(...[...ctx.code.matchAll(HANDLER)].flatMap(m => {
269 const open = ctx.code.indexOf('{', m.index ?? 0)
270 const body = ctx.code.slice(open, closeOf(ctx.code, open) + 1)
271 if (/\b(auth|session|getServerSession|currentUser|requireUser|verify\w*Token|role|permission|forbidden|401|403)\b/i.test(body)) return []
272 return [{
273 ruleId: 'MIR-G1', category: 'G. Authorization omission', line: lineOf(ctx.code, m.index ?? 0),
274 snippet: clip(lineText(ctx.source, m.index ?? 0), 140), severity: 'high' as Severity, confidence: 'potential' as Confidence,
275 explanation: 'The task mentions permissions, and this mutating handler checks no session, role or token.',
276 evidence: 'No auth, session, role or 401/403 reference in the handler body.',
277 verify: 'Call it without credentials and confirm it is refused.',
278 }]
279 }))
280 }
281 return out
282 },
283 },
284 {
285 id: 'MIR-H1', category: 'H. Missing failure handling', title: 'Request with no failure path', appliesTo: SOURCE,
286 detect: ctx => find(ctx, /\bawait\s+fetch\s*\(/g, m => {
287 const fn = enclosingFunction(ctx.code, m.index ?? 0)
288 const body = fn?.body ?? ctx.code
289 if (/\btry\s*\{|\.catch\(|\.ok\b|\.status\b|throwOnError|onError|catch\s*\(/.test(body)) return undefined
290 return {
291 ruleId: 'MIR-H1', category: 'H. Missing failure handling', severity: byIntent(ctx.intent, 'medium', 'low'), confidence: 'potential',
292 explanation: 'A request is awaited with no try/catch, .catch or status check: a failure has no defined behaviour.',
293 evidence: 'await fetch(...) in a function with no error or status handling.',
294 verify: 'Make the request fail (offline, 500) and check what the user sees.',
295 }
296 }),
297 },
298]
299
300/** Runs every rule that applies to `path`. Findings carry a fingerprint id stable across line moves. */
301export function scanSource(path: string, source: string, intent: ContractIntent, asked: string): (RawFinding & { id: string })[] {
302 if (source.length > 400_000) return []
303 const ctx: RuleContext = { path, source, code: stripComments(source), intent, asked: asked.toLowerCase() }
304 const seen = new Map<string, number>()
305 return RULES.filter(r => r.appliesTo(path)).flatMap(r => r.detect(ctx)).map(f => {
306 const base = hash(`${f.ruleId}|${path}|${f.snippet.replace(/\s+/g, '')}`)
307 const n = seen.get(base) ?? 0
308 seen.set(base, n + 1)
309 return { ...f, id: n === 0 ? base : `${base}-${n}` }
310 })
311}
312
313/** Literal `/api/...` URLs a file requests (template literals with substitutions are skipped). */
314export function apiCalls(source: string): { path: string; line: number }[] {
315 const code = stripComments(source)
316 return [...code.matchAll(/\b(?:fetch|axios(?:\.(?:get|post|put|patch|delete))?|\w+\.(?:get|post|put|patch|delete))\(\s*['"`](\/api\/[^'"`?#$]*)['"`?#]/g)]
317 .map(m => ({ path: (m[1] as string).replace(/\/$/, ''), line: lineOf(code, m.index ?? 0) }))
318}
319
320/** The endpoint a Next.js route file serves, or undefined. Route groups `(x)` are dropped. */
321export function routeOf(path: string): string | undefined {
322 const app = path.match(/(?:^|\/)app\/((?:.+\/)?api(?:\/.+)?)\/route\.(?:t|j)sx?$/)
323 if (app) return '/' + (app[1] as string).split('/').filter(s => !/^\(.*\)$/.test(s)).join('/')
324 const pages = path.match(/(?:^|\/)pages\/(api(?:\/.+?)?)(?:\/index)?\.(?:t|j)sx?$/)
325 return pages ? `/${pages[1]}` : undefined
326}
327
328/** `/api/events/[id]` serves `/api/events/42`; `[...slug]` serves the rest. */
329export function routeMatches(route: string, call: string): boolean {
330 const r = route.split('/')
331 const c = call.split('/')
332 for (let i = 0; i < r.length; i++) {
333 const seg = r[i] as string
334 if (/^\[\[?\.\.\./.test(seg)) return true
335 if (c[i] === undefined) return false
336 if (!/^\[.+\]$/.test(seg) && seg !== c[i]) return false
337 }
338 return r.length === c.length
339}
340src/engine/drift.ts 121 lines1// Intent drift: observed changes compared with the contract. A file name alone
2// never proves intent, so only a match against an exclusion glob the person
3// approved is `confirmed`; every other rule reports a `potential` drift and
4// says what to check.
5
6import type { Confidence, Contract, FileChange, Severity } from './types'
7import type { RawFinding } from './mirage'
8import { clip, hash, isCodePath, isPathPattern, matchesGlob } from './util'
9
10export type DriftFinding = RawFinding & { id: string; requirementIds: string[] }
11
12const ROUTE_FILE = /(^|\/)app\/(.+\/)?(page|route)\.(t|j)sx?$|(^|\/)pages\/(?!_app|_document|api\/).+\.(t|j)sx?$|(^|\/)routes\/.+\.(t|j)sx?$/i
13const NO_NEW = /\b(?:do not|don't|never|must not|mustn't|without|avoid)\s+(?:creat\w*|add\w*|introduc\w*|mak\w*|build\w*)\s+(?:another|a new|new|a second|second|separate|an additional|additional|a duplicate|duplicate)\s+([\w-]+)(?:\s+(page|route|screen|view|component|endpoint|module|file|service|table))?/i
14
15function make(change: FileChange, ruleId: string, fields: Omit<RawFinding, 'ruleId' | 'line' | 'snippet'> & { requirementIds?: string[] }): DriftFinding {
16 const { requirementIds = [], ...rest } = fields
17 return {
18 ruleId, line: 1, snippet: `${change.kind} ${change.path}`, ...rest, requirementIds,
19 id: hash(`${ruleId}|${change.path}|${requirementIds.join(',')}`),
20 }
21}
22
23function dependencies(json: string | null | undefined): Set<string> {
24 if (json === null || json === undefined) return new Set()
25 try {
26 const parsed = JSON.parse(json) as Record<string, Record<string, string> | undefined>
27 return new Set([...Object.keys(parsed['dependencies'] ?? {}), ...Object.keys(parsed['devDependencies'] ?? {})])
28 } catch {
29 return new Set()
30 }
31}
32
33export function driftFindings(change: FileChange, contract: Contract): DriftFinding[] {
34 const out: DriftFinding[] = []
35 const approved = contract.approval.status === 'approved' || contract.baseline !== undefined
36 const asked = [contract.summary, ...contract.requirements.map(r => r.description), ...contract.assumptions].join(' ').toLowerCase()
37 const preservation = contract.requirements.filter(r => r.category === 'preservation' && r.waiver === undefined)
38
39 for (const glob of contract.exclusions.filter(isPathPattern)) {
40 if (!matchesGlob(change.path, glob)) continue
41 out.push(make(change, 'DRIFT-X1', {
42 category: 'Excluded path changed', severity: 'high', confidence: approved ? 'confirmed' : 'potential',
43 requirementIds: contract.requirements.filter(r => r.description.includes(glob)).map(r => r.id),
44 explanation: `The contract excludes "${glob}", and ${change.path} was ${change.kind === 'create' ? 'created' : change.kind === 'delete' ? 'deleted' : 'modified'}.`,
45 evidence: `Observed ${change.kind} of ${change.path}; exclusion "${glob}" ${approved ? 'is in the approved contract' : 'is in an unapproved draft'}.`,
46 verify: 'Revert the change, or amend the contract if the exclusion no longer holds.',
47 }))
48 }
49
50 if (contract.scope.length > 0 && isCodePath(change.path) && !contract.scope.some(g => matchesGlob(change.path, g))) {
51 out.push(make(change, 'DRIFT-S1', {
52 category: 'Change outside scope', severity: 'medium', confidence: 'potential',
53 explanation: `${change.path} is outside the contract's scope (${contract.scope.join(', ')}).`,
54 evidence: `Observed ${change.kind} of ${change.path}.`,
55 verify: 'Confirm the change is needed for the task, or widen the scope explicitly.',
56 }))
57 }
58
59 if (change.kind === 'create') {
60 for (const r of contract.requirements) {
61 const m = r.waiver === undefined ? r.description.match(NO_NEW) : null
62 const noun = m?.[1]?.toLowerCase()
63 if (noun === undefined || noun.length < 3 || !change.path.toLowerCase().includes(noun)) continue
64 out.push(make(change, 'DRIFT-N1', {
65 category: 'Possible duplicate of existing feature', severity: 'high', confidence: 'potential', requirementIds: [r.id],
66 explanation: `${r.id} says "${clip(r.description, 90)}", and a new file ${change.path} names "${noun}".`,
67 evidence: `New file whose path contains "${noun}"${ROUTE_FILE.test(change.path) ? ' and that defines a route' : ''}. The name alone does not prove a duplicate.`,
68 verify: `Check whether ${change.path} replaces or duplicates the existing ${noun}; extend the existing one instead if so.`,
69 }))
70 }
71 if (ROUTE_FILE.test(change.path) && preservation.some(r => /\b(navigation|route|routing|url|page)\b/i.test(r.description))) {
72 const ids = preservation.filter(r => /\b(navigation|route|routing|url|page)\b/i.test(r.description)).map(r => r.id)
73 out.push(make(change, 'DRIFT-R1', {
74 category: 'New route while routes are preserved', severity: 'medium', confidence: 'potential', requirementIds: ids,
75 explanation: `${change.path} adds a route, and ${ids.join(', ')} asks to preserve the current navigation.`,
76 evidence: `New route file ${change.path}.`,
77 verify: 'Open the app and confirm the existing navigation still leads where it did.',
78 }))
79 }
80 }
81
82 if (change.kind === 'delete' && isCodePath(change.path)) {
83 out.push(make(change, 'DRIFT-D1', {
84 category: 'Existing behaviour deleted', severity: preservation.length > 0 ? 'high' : 'medium', confidence: 'potential',
85 requirementIds: preservation.map(r => r.id),
86 explanation: `${change.path} was deleted; whatever it did is gone unless moved elsewhere.`,
87 evidence: `Observed deletion of ${change.path}.`,
88 verify: 'Confirm its behaviour moved elsewhere or is no longer needed.',
89 }))
90 }
91
92 if (change.kind === 'update' && typeof change.before === 'string' && change.after !== undefined && isCodePath(change.path)) {
93 const before = change.before.split('\n').map(l => l.trim()).filter(Boolean)
94 const after = new Set(change.after.split('\n').map(l => l.trim()))
95 const kept = before.filter(l => after.has(l)).length
96 if (before.length >= 30 && kept / before.length < 0.4) {
97 out.push(make(change, 'DRIFT-W1', {
98 category: 'Rewrite instead of extension', severity: preservation.length > 0 ? 'medium' : 'low', confidence: 'potential',
99 requirementIds: preservation.map(r => r.id),
100 explanation: `${change.path} kept ${Math.round((kept / before.length) * 100)}% of its previous lines: it was largely rewritten.`,
101 evidence: `${kept} of ${before.length} non-empty lines survive the change.`,
102 verify: 'Check that the existing behaviour of this file still works.',
103 }))
104 }
105 }
106
107 if (/(^|\/)package\.json$/.test(change.path) && change.after !== undefined) {
108 const before = dependencies(change.before)
109 const added = [...dependencies(change.after)].filter(d => !before.has(d) && !asked.includes(d.toLowerCase().replace(/^@[^/]+\//, '')))
110 if (added.length > 0 && change.before !== undefined) {
111 out.push(make(change, 'DRIFT-P1', {
112 category: 'Dependency added without a requirement', severity: 'medium', confidence: 'potential' as Confidence,
113 explanation: `New dependencies not mentioned by the contract: ${added.slice(0, 6).join(', ')}.`,
114 evidence: `package.json gained ${added.length} dependenc${added.length === 1 ? 'y' : 'ies'}.`,
115 verify: 'Confirm each dependency is needed, or record it as an assumption.',
116 }))
117 }
118 }
119 return out.map(f => ({ ...f, severity: f.severity as Severity }))
120}
121src/engine/claims.ts 127 lines1// Completion Claim Auditor: finds concrete claims of finished work in a final
2// answer and weighs each against recorded evidence. It tells "no evidence"
3// (unsupported) apart from "evidence says otherwise" (contradicted), and never
4// treats a missing observation as proof that something did not happen.
5
6import type { CheckKind, Claim, ClaimKind, ClaimVerdict, Evidence, Task } from './types'
7import { freshness, isProof } from './evidence'
8import { requirementState } from './status'
9import { clip, hash, redact, sentences } from './util'
10
11const PATTERNS: readonly [ClaimKind, RegExp][] = [
12 ['tests-pass', /\b(all\s+)?(the\s+)?(\w+\s+)?tests?\s+(now\s+)?(all\s+)?(pass(es|ed|ing)?|succeed(s|ed)?|are\s+(passing|green)|(are\s+|is\s+)?green)\b|\b\d+\s*(\/\s*\d+\s+)?tests?\s+pass(ed|ing)?\b|\btest suite\s+(passes|passed|is green)\b/i],
13 ['build', /\bbuild\s+(now\s+)?(succeeds|succeeded|passes|passed|works|is\s+green)\b|\bbuilds?\s+(cleanly|successfully)\b|\bcompiles?\s+(cleanly|successfully|without\s+errors)\b/i],
14 ['typecheck', /\b(type[- ]?check(s|ing)?|tsc|types?)\s+(now\s+)?(pass(es|ed)?|is\s+clean|are\s+clean|succeeds?)\b|\bno\s+type\s+errors\b/i],
15 ['lint', /\blint(ing|er)?\s+(now\s+)?(pass(es|ed)?|is\s+clean|succeeds?)\b|\bno\s+lint(ing)?\s+(errors|warnings|issues)\b/i],
16 ['migration', /\bmigrations?\s+(ran|run|applied|succeeded|completed)\b|\bmigrations?\s+(was|were|has been|have been)\s+(applied|run)\b/i],
17 ['production-ready', /\bproduction[- ]ready\b|\bready\s+for\s+production\b/i],
18 ['responsive', /\b(fully\s+)?responsive\b/i],
19 ['browser', /\b(works|verified|tested|checked)\s+in\s+(the\s+)?browser\b|\bend[- ]to[- ]end\b.*\b(works|pass)|\b(user\s+(flow|journey)|ui)\s+(works|is working)\b/i],
20 ['live-data', /\b(fetch(es|ed)?|load(s|ed)?|pull(s|ed)?|comes?|retrieved?)\b[^.]{0,40}\bfrom\s+the\s+(live\s+|real\s+)?(api|backend|server|database|db)\b|\b(live|real)\s+(api\s+)?data\b/i],
21 ['implemented', /\b[\w-]+(\s+[\w-]+){0,2}\s+(is|are|has\s+been|have\s+been)\s+(now\s+)?(fully\s+)?(implemented|complete(d)?|done|finished)\b|\b(i\s+(have\s+)?|i've\s+)(implemented|completed|finished)\b/i],
22 ['works', /\b(verified|confirmed)\s+(that\s+)?(it|this|everything)\s+works\b|\beverything\s+works\b|\b(it|this)\s+(now\s+)?works\s+(correctly|as expected)\b/i],
23]
24
25/** Hedges, negations, instructions and plans are not claims of completed work. */
26const NOT_A_CLAIM = /\b(not|n't|never|nothing|none|no longer|unable|cannot|could not|didn't|did not|haven't|have not|hasn't|without running|unverified|untested|should|will|would|may|might|once|if|when you|you can|you could|to verify|try|please|run `|need to|needs to|todo|next step|failing|fails|failed)\b|\?$/i
27
28export function extractClaims(answer: string): { kind: ClaimKind; text: string }[] {
29 const out: { kind: ClaimKind; text: string }[] = []
30 for (const s of sentences(answer)) {
31 if (NOT_A_CLAIM.test(s)) continue
32 for (const [kind, re] of PATTERNS) {
33 if (re.test(s) && !out.some(c => c.kind === kind && c.text === s)) out.push({ kind, text: s })
34 }
35 }
36 return out
37}
38
39type Judged = { verdict: ClaimVerdict; reason: string; evidenceIds: string[] }
40
41const CHECK_OF: Partial<Record<ClaimKind, CheckKind>> = {
42 'tests-pass': 'test', build: 'build', typecheck: 'typecheck', lint: 'lint', migration: 'migration', browser: 'browser',
43}
44
45function judgeCheck(kind: ClaimKind, text: string, task: Task): Judged {
46 const check = CHECK_OF[kind] as CheckKind
47 const runs = task.evidence.filter(e => e.check === check || (check === 'test' && e.check === 'browser' && /\b(e2e|end[- ]to[- ]end|browser)\b/i.test(text)))
48 const real = runs.filter(e => e.synthetic !== true)
49 const current = real.filter(e => freshness(e, task) !== 'stale')
50 const ids = (list: Evidence[]): string[] => list.slice(-3).map(e => e.id)
51 const failed = current.filter(e => e.outcome === 'fail').at(-1)
52 const passed = current.filter(e => e.outcome === 'pass')
53 // Observation order decides which run is latest; clocks can tie.
54 const order = (e: Evidence | undefined): number => (e === undefined ? -1 : task.evidence.indexOf(e))
55
56 if (failed !== undefined && order(passed.at(-1)) < order(failed)) {
57 return { verdict: 'contradicted', reason: `latest ${check} run failed: ${failed.summary}`, evidenceIds: [failed.id] }
58 }
59 if (passed.length === 0) {
60 if (kind === 'browser' && task.evidence.some(e => e.check === 'test' && e.outcome === 'pass')) {
61 return { verdict: 'unsupported', reason: 'unit tests passed, but no browser test ran: unit tests do not prove a browser journey', evidenceIds: [] }
62 }
63 if (real.some(e => e.outcome === 'pass')) return { verdict: 'unsupported', reason: `only stale ${check} evidence: code changed after the last passing run`, evidenceIds: ids(real) }
64 if (runs.length > real.length) return { verdict: 'unsupported', reason: `only synthetic ${check} results: nothing was executed`, evidenceIds: ids(runs) }
65 if (runs.length > 0) return { verdict: 'unsupported', reason: `${check} runs were inconclusive (${runs.at(-1)?.basis})`, evidenceIds: ids(runs) }
66 return { verdict: 'unsupported', reason: `no ${check} run was observed for this task`, evidenceIds: [] }
67 }
68 const last = passed.at(-1) as Evidence
69 const broad = /\ball\b|\bevery\b|\bentire\b|\bfull\b/i.test(text)
70 const namesOther = kind === 'tests-pass' && /\b(integration|e2e|end[- ]to[- ]end|browser)\b/i.test(text) && !passed.some(e => /integration|e2e|playwright|cypress/i.test(e.command ?? ''))
71 if (namesOther) return { verdict: 'partial', reason: `${last.summary}, but no integration/e2e run was observed`, evidenceIds: ids(passed) }
72 if (broad && (last.scope.length > 0 || (last.counts?.skipped ?? 0) > 0)) {
73 const why = last.scope.length > 0 ? `the run was scoped to ${last.scope.join(', ')}` : `${last.counts?.skipped} tests were skipped`
74 return { verdict: 'partial', reason: `${last.summary}; ${why}`, evidenceIds: ids(passed) }
75 }
76 if (last.truncated && last.counts === undefined) return { verdict: 'partial', reason: `${last.summary}; output truncated, summary not seen`, evidenceIds: ids(passed) }
77 return { verdict: 'supported', reason: `${last.summary} via ${last.basis} (${last.id})`, evidenceIds: ids(passed) }
78}
79
80function judge(kind: ClaimKind, text: string, task: Task): Judged {
81 if (CHECK_OF[kind] !== undefined) return judgeCheck(kind, text, task)
82 const openMirage = task.findings.filter(f => f.system === 'mirage' && f.status === 'open')
83 switch (kind) {
84 case 'live-data': {
85 const mock = openMirage.find(f => /^MIR-(A1|D1|B1|B2)/.test(f.ruleId) && f.confidence !== 'expected')
86 if (mock) return { verdict: 'contradicted', reason: `potentially contradicted: ${mock.ruleId} in ${mock.file}:${mock.line} (${mock.explanation})`, evidenceIds: [] }
87 const runtime = task.evidence.filter(e => (e.kind === 'runtime-check' || e.kind === 'manual-confirmation') && isProof(e, task))
88 return runtime.length > 0
89 ? { verdict: 'partial', reason: 'a runtime check passed; it does not show where the data came from', evidenceIds: runtime.map(e => e.id) }
90 : { verdict: 'not-assessable', reason: 'static analysis cannot prove live data; no runtime check observed', evidenceIds: [] }
91 }
92 case 'implemented': {
93 const placeholder = openMirage.find(f => f.confidence === 'confirmed')
94 if (placeholder) return { verdict: 'contradicted', reason: `${placeholder.ruleId} in ${placeholder.file}:${placeholder.line}: ${placeholder.explanation}`, evidenceIds: [] }
95 const must = task.contract.requirements.filter(r => r.criticality === 'must' && r.waiver === undefined)
96 const states = must.map(r => requirementState(r, task).state)
97 const verified = states.filter(s => s === 'verified').length
98 if (must.length > 0 && verified === must.length) return { verdict: 'supported', reason: `all ${must.length} must requirements verified`, evidenceIds: [] }
99 if (states.includes('failed')) return { verdict: 'contradicted', reason: 'a must requirement has failing evidence', evidenceIds: [] }
100 if (verified > 0 || Object.keys(task.files).length > 0) {
101 return { verdict: 'partial', reason: `code changes observed; ${verified}/${must.length} must requirements verified`, evidenceIds: [] }
102 }
103 return { verdict: 'unsupported', reason: 'no code changes or verification observed', evidenceIds: [] }
104 }
105 case 'production-ready':
106 case 'responsive':
107 return { verdict: 'not-assessable', reason: `"${kind}" cannot be judged from source and terminal output; needs a manual or runtime check`, evidenceIds: [] }
108 default: {
109 const proof = task.evidence.filter(e => isProof(e, task))
110 return proof.length > 0
111 ? { verdict: 'partial', reason: `${proof.length} passing check(s); "works" is broader than any one of them`, evidenceIds: proof.slice(-3).map(e => e.id) }
112 : { verdict: 'unsupported', reason: 'no passing check observed', evidenceIds: [] }
113 }
114 }
115}
116
117export function auditAnswer(answer: string, task: Task, turnId: string, now: number): Claim[] {
118 return extractClaims(answer).map(({ kind, text }) => ({
119 id: hash(`${turnId}|${kind}|${text}`),
120 turnId,
121 text: clip(redact(text), 200),
122 kind,
123 ...judge(kind, text, task),
124 at: now,
125 }))
126}
127src/engine/status.ts 77 lines1// Requirement status from evidence, and the explicit policies Strict mode
2// enforces. "Verified" needs a real, passing, current run of a kind that can
3// prove the requirement: a model's opinion never counts, and one passing suite
4// verifies only the requirements linked to it.
5
6import type { Evidence, EvidenceKind, Mode, Requirement, Task, VerificationKind } from './types'
7import { freshness } from './evidence'
8import { matchesGlob } from './util'
9
10export type RequirementState = 'verified' | 'stale' | 'failed' | 'waived' | 'implemented' | 'unverified'
11
12const PROVES: Record<VerificationKind, readonly EvidenceKind[]> = {
13 test: ['test-result', 'runtime-check', 'manual-confirmation'],
14 runtime: ['runtime-check', 'manual-confirmation'],
15 manual: ['manual-confirmation'],
16 static: ['tool-result', 'static-analysis', 'test-result', 'runtime-check', 'manual-confirmation'],
17 unknown: ['tool-result', 'static-analysis', 'test-result', 'runtime-check', 'manual-confirmation'],
18}
19
20/** Evidence tied to a requirement: linked by id, matched by one of its `checks`, or any test run for a test requirement. */
21export function linkedEvidence(r: Requirement, task: Task): Evidence[] {
22 const checks = r.checks.map(c => c.toLowerCase())
23 return task.evidence.filter(e =>
24 e.requirementIds.includes(r.id) ||
25 r.evidenceIds.includes(e.id) ||
26 (checks.length > 0 && e.command !== undefined && checks.some(c => e.command!.toLowerCase().includes(c))) ||
27 (r.category === 'test' && checks.length === 0 && (e.check === 'test' || e.check === 'browser')))
28}
29
30export function requirementState(r: Requirement, task: Task): { state: RequirementState; reason: string; evidence: Evidence[] } {
31 const linked = linkedEvidence(r, task)
32 if (r.waiver !== undefined) return { state: 'waived', reason: `waived: ${r.waiver.reason}`, evidence: linked }
33 const violated = task.findings.find(f => f.requirementIds.includes(r.id) && f.confidence === 'confirmed' && f.status === 'open')
34 if (violated) return { state: 'failed', reason: `${violated.ruleId}: ${violated.explanation}`, evidence: linked }
35
36 const usable = linked.filter(e => e.synthetic !== true && e.outcome !== 'inconclusive' && PROVES[r.verification].includes(e.kind))
37 const current = usable.filter(e => freshness(e, task) !== 'stale')
38 const latest = current.at(-1)
39 if (latest?.outcome === 'fail') return { state: 'failed', reason: `${latest.summary} (${latest.id})`, evidence: linked }
40 const pass = current.filter(e => e.outcome === 'pass').at(-1)
41 if (pass) return { state: 'verified', reason: `${pass.summary} (${pass.id})`, evidence: linked }
42 if (usable.some(e => e.outcome === 'pass')) return { state: 'stale', reason: 'passing evidence predates later changes', evidence: linked }
43 if (linked.length > 0 && usable.length === 0) {
44 return { state: r.implemented ? 'implemented' : 'unverified', reason: 'only inconclusive, synthetic or non-qualifying evidence', evidence: linked }
45 }
46 if (r.implemented) return { state: 'implemented', reason: 'marked implemented; no qualifying evidence', evidence: linked }
47 return { state: 'unverified', reason: linked.length === 0 ? 'no evidence observed' : 'no qualifying evidence', evidence: linked }
48}
49
50export type PolicyViolation = { id: string; text: string }
51
52/**
53 * Strict mode's deterministic policies, each one the person enabled by writing
54 * it into the approved contract: an excluded path, a must requirement's
55 * declared check. Uncertain judgments are never here.
56 */
57export function policyViolations(task: Task): PolicyViolation[] {
58 const out: PolicyViolation[] = []
59 if (task.contract.approval.status !== 'approved' && task.contract.baseline === undefined) return out
60 for (const f of task.findings) {
61 if (f.ruleId === 'DRIFT-X1' && f.status === 'open') out.push({ id: f.id, text: `excluded path modified: ${f.file}` })
62 }
63 for (const r of task.contract.requirements) {
64 if (r.criticality !== 'must' || r.waiver !== undefined || (r.checks.length === 0 && r.category !== 'test')) continue
65 const { state } = requirementState(r, task)
66 if (state === 'failed') out.push({ id: r.id, text: `${r.id} required check failed` })
67 else if (state !== 'verified') out.push({ id: r.id, text: `${r.id} required check has not passed on the current changes (${r.checks.join(', ') || 'tests'})` })
68 }
69 return out
70}
71
72/** Strict mode's one hard gate the API supports: refuse an edit to a path the approved contract excludes. */
73export function deniedPath(task: Task | undefined, mode: Mode, path: string): string | undefined {
74 if (mode !== 'strict' || task === undefined || task.contract.baseline === undefined) return undefined
75 return task.contract.exclusions.find(g => /[/*]|\.\w{1,5}$/.test(g) && !/\s/.test(g) && matchesGlob(path, g))
76}
77src/engine/util.ts 95 lines1// Small pure helpers: hashing, paths, globs and redaction.
2
3/** cyrb53: a fast, deterministic 53-bit string hash, hex. Not cryptographic; identity only. */
4export function hash(text: string, seed = 0): string {
5 let h1 = 0xdeadbeef ^ seed
6 let h2 = 0x41c6ce57 ^ seed
7 for (let i = 0; i < text.length; i++) {
8 const ch = text.charCodeAt(i)
9 h1 = Math.imul(h1 ^ ch, 2654435761)
10 h2 = Math.imul(h2 ^ ch, 1597334677)
11 }
12 h1 = Math.imul(h1 ^ (h1 >>> 16), 2246822507) ^ Math.imul(h2 ^ (h2 >>> 13), 3266489909)
13 h2 = Math.imul(h2 ^ (h2 >>> 16), 2246822507) ^ Math.imul(h1 ^ (h1 >>> 13), 3266489909)
14 return (4294967296 * (2097151 & h2) + (h1 >>> 0)).toString(16).padStart(14, '0')
15}
16
17export function newId(prefix: string, now: number): string {
18 return `${prefix}-${now.toString(36)}-${Math.floor(Math.random() * 1679616).toString(36).padStart(4, '0')}`
19}
20
21const CODE_EXT = /\.(c|m)?(t|j)sx?$|\.(py|go|rs|java|kt|rb|php|cs|swift|vue|svelte|sql|prisma)$/i
22const DOC_EXT = /\.(md|mdx|txt|rst|adoc)$|(^|\/)(LICENSE|CHANGELOG)[^/]*$/i
23const TEST_PATH = /(^|\/)(__tests__|tests?|spec|e2e|__mocks__)\/|\.(test|spec)\.[cm]?[jt]sx?$|_test\.(go|py)$/i
24
25export const isCodePath = (path: string): boolean => CODE_EXT.test(path)
26export const isDocPath = (path: string): boolean => DOC_EXT.test(path)
27export const isTestPath = (path: string): boolean => TEST_PATH.test(path)
28export const isScannable = (path: string): boolean => /\.(c|m)?(t|j)sx?$/i.test(path)
29
30/** Forward slashes, and relative to `root` when inside it (compared case-insensitively for Windows drives). */
31export function normalizePath(path: string, root?: string): string {
32 const p = path.replace(/\\/g, '/')
33 if (root === undefined || root === '') return p
34 const r = root.replace(/\\/g, '/').replace(/\/$/, '')
35 return p.toLowerCase().startsWith(r.toLowerCase() + '/') ? p.slice(r.length + 1) : p
36}
37
38/** `**` any depth, `*` within a segment, `?` one character. A pattern without `/` matches the base name too. */
39export function globToRegExp(glob: string): RegExp {
40 const g = glob.replace(/\\/g, '/').replace(/^\.\//, '')
41 let out = ''
42 for (let i = 0; i < g.length; i++) {
43 const c = g[i] as string
44 if (c === '*' && g[i + 1] === '*') {
45 out += g[i + 2] === '/' ? '(?:.*/)?' : '.*'
46 i += g[i + 2] === '/' ? 2 : 1
47 } else if (c === '*') out += '[^/]*'
48 else if (c === '?') out += '[^/]'
49 else out += c.replace(/[.+^${}()|[\]\\]/g, '\\$&')
50 }
51 return new RegExp(g.includes('/') ? `^${out}$` : `(^|/)${out}$`, 'i')
52}
53
54export const matchesGlob = (path: string, glob: string): boolean => globToRegExp(glob).test(path.replace(/\\/g, '/'))
55
56/** Looks like a path glob rather than prose: has a slash, a star or a file extension, and no spaces. */
57export const isPathPattern = (text: string): boolean => !/\s/.test(text.trim()) && /[/*]|\.\w{1,5}$/.test(text.trim())
58
59const SECRET_PATTERNS: readonly [RegExp, string][] = [
60 [/-----BEGIN [A-Z ]*PRIVATE KEY-----[\s\S]*?-----END [A-Z ]*PRIVATE KEY-----/g, '[private key]'],
61 [/\b(AKIA|ASIA)[0-9A-Z]{16}\b/g, '[aws key]'],
62 [/\bgh[pousr]_[A-Za-z0-9]{30,}\b/g, '[github token]'],
63 [/\bxox[abprs]-[A-Za-z0-9-]{10,}\b/g, '[slack token]'],
64 [/\bsk-[A-Za-z0-9_-]{20,}\b/g, '[api key]'],
65 [/\beyJ[A-Za-z0-9_-]{8,}\.[A-Za-z0-9_-]{8,}\.[A-Za-z0-9_-]{8,}\b/g, '[jwt]'],
66 [/\b(Bearer|Basic)\s+[A-Za-z0-9._~+/=-]{8,}/gi, '$1 ***'],
67 [/(\/\/)[^/\s:@]+:[^/\s@]+@/g, '$1***:***@'],
68 [/\b([A-Za-z0-9_]*(?:api[_-]?key|token|secret|passw(?:or)?d|pwd|credential|auth)[A-Za-z0-9_]*)(\s*[:=]\s*|\s+)("[^"]*"|'[^']*'|[^\s'"&;|]+)/gi, '$1$2***'],
69]
70
71/** Removes secrets and personal home paths. Applied to everything stored or exported. */
72export function redact(text: string): string {
73 let out = text
74 for (const [pattern, replacement] of SECRET_PATTERNS) out = out.replace(pattern, replacement)
75 return out
76 .replace(/\b[A-Za-z]:[\\/]Users[\\/][^\\/\s]+/gi, '~')
77 .replace(/\/(home|Users)\/[^/\s]+/g, '~')
78}
79
80export function clip(text: string, max: number): string {
81 const one = text.replace(/\s+/g, ' ').trim()
82 return one.length <= max ? one : `${one.slice(0, Math.max(0, max - 1))}…`
83}
84
85export const sentences = (text: string): string[] =>
86 text
87 .replace(/```[\s\S]*?```/g, ' ')
88 .split(/(?<=[.!?])\s+|\n+/)
89 .map(s => s.replace(/^[\s>*\-•\d.)#]+/, '').trim())
90 .filter(s => s.length > 3)
91
92/** The store key of a project: its repository root and remote, so two repositories never share a contract. */
93export const projectKey = (root: string, remote: string | null | undefined): string =>
94 `p:${hash(`${normalizePath(root).replace(/\/$/, '').toLowerCase()}|${remote ?? ''}`)}`
95src/engine/summary.ts 140 lines1// The one canonical summary: the status line, every pane view and the report
2// all read their counts from `summarize`, so they cannot disagree.
3
4import type { CheckKind, Claim, Evidence, Freshness, Mode, Task } from './types'
5import { freshness, latestByCheck } from './evidence'
6import { requirementState, type RequirementState } from './status'
7
8export type Tab = 'overview' | 'requirements' | 'evidence' | 'findings' | 'claims' | 'activity' | 'report'
9
10export type Summary = {
11 hasTask: boolean
12 taskId?: string
13 title: string
14 approval: 'none' | 'draft' | 'approved' | 'amended'
15 version: number
16 intent: string
17 mode: Mode
18 requirements: Record<RequirementState, number> & { total: number }
19 must: { total: number; verified: number }
20 findings: { open: number; high: number; medium: number; low: number; info: number; expected: number; waived: number; acknowledged: number; confirmed: number }
21 evidence: { total: number; fresh: number; stale: number; latest: { check: CheckKind; outcome: string; freshness: Freshness; summary: string; id: string; at: number }[] }
22 claims: { total: number; supported: number; partial: number; unsupported: number; contradicted: number; notAssessable: number; latest: Claim[] }
23 concern?: { text: string; tab: Tab }
24 /** Verified share of the requirements still in force (waived ones left out), 0..1. */
25 progress: number
26}
27
28const EMPTY_REQ = { verified: 0, stale: 0, failed: 0, waived: 0, implemented: 0, unverified: 0, total: 0 }
29
30export function summarize(task: Task | undefined, mode: Mode): Summary {
31 if (task === undefined) {
32 return {
33 hasTask: false, title: 'No task contract', approval: 'none', version: 0, intent: 'unknown', mode,
34 requirements: { ...EMPTY_REQ }, must: { total: 0, verified: 0 },
35 findings: { open: 0, high: 0, medium: 0, low: 0, info: 0, expected: 0, waived: 0, acknowledged: 0, confirmed: 0 },
36 evidence: { total: 0, fresh: 0, stale: 0, latest: [] },
37 claims: { total: 0, supported: 0, partial: 0, unsupported: 0, contradicted: 0, notAssessable: 0, latest: [] },
38 progress: 0,
39 }
40 }
41 const c = task.contract
42 const requirements = { ...EMPTY_REQ }
43 const must = { total: 0, verified: 0 }
44 for (const r of c.requirements) {
45 const { state } = requirementState(r, task)
46 requirements[state]++
47 requirements.total++
48 if (r.criticality === 'must' && state !== 'waived') {
49 must.total++
50 if (state === 'verified') must.verified++
51 }
52 }
53
54 const findings = { open: 0, high: 0, medium: 0, low: 0, info: 0, expected: 0, waived: 0, acknowledged: 0, confirmed: 0 }
55 for (const f of task.findings) {
56 if (f.status === 'waived') findings.waived++
57 else if (f.status === 'acknowledged') findings.acknowledged++
58 else if (f.status === 'open' && f.confidence === 'expected') findings.expected++
59 else if (f.status === 'open') {
60 findings.open++
61 findings[f.severity]++
62 if (f.confidence === 'confirmed') findings.confirmed++
63 }
64 }
65
66 const fresh = (e: Evidence): Freshness => freshness(e, task)
67 const real = task.evidence.filter(e => e.synthetic !== true)
68 const latest = [...latestByCheck(task).values()].map(e => ({ check: e.check, outcome: e.outcome, freshness: fresh(e), summary: e.summary, id: e.id, at: e.timestamp }))
69 const lastTurn = task.claims.at(-1)?.turnId
70 const latestClaims = task.claims.filter(cl => cl.turnId === lastTurn)
71 const count = (v: Claim['verdict']): number => latestClaims.filter(cl => cl.verdict === v).length
72
73 const summary: Summary = {
74 hasTask: true,
75 taskId: task.id,
76 title: c.summary,
77 approval: c.approval.status === 'approved' ? 'approved' : c.baseline !== undefined ? 'amended' : 'draft',
78 version: c.version,
79 intent: c.intent,
80 mode,
81 requirements,
82 must,
83 findings,
84 evidence: {
85 total: task.evidence.length,
86 fresh: real.filter(e => e.outcome === 'pass' && fresh(e) === 'fresh').length,
87 stale: real.filter(e => e.outcome === 'pass' && fresh(e) === 'stale').length,
88 latest,
89 },
90 claims: {
91 total: latestClaims.length, supported: count('supported'), partial: count('partial'), unsupported: count('unsupported'),
92 contradicted: count('contradicted'), notAssessable: count('not-assessable'), latest: latestClaims,
93 },
94 progress: requirements.total - requirements.waived > 0 ? requirements.verified / (requirements.total - requirements.waived) : 0,
95 }
96 summary.concern = concernOf(summary, task)
97 return summary
98}
99
100function concernOf(s: Summary, task: Task): Summary['concern'] {
101 const contradicted = s.claims.latest.find(cl => cl.verdict === 'contradicted')
102 if (contradicted) return { text: `Claim contradicted: "${contradicted.text}"`, tab: 'claims' }
103 const confirmed = task.findings.find(f => f.status === 'open' && f.confidence === 'confirmed')
104 if (confirmed) return { text: `${confirmed.ruleId} ${confirmed.file}:${confirmed.line}: ${confirmed.category}`, tab: 'findings' }
105 if (s.requirements.failed > 0) return { text: `${s.requirements.failed} requirement(s) have failing evidence`, tab: 'requirements' }
106 const high = task.findings.find(f => f.status === 'open' && f.severity === 'high' && f.confidence === 'potential')
107 if (high) return { text: `Possible ${high.category.replace(/^\w\.\s*/, '').toLowerCase()} in ${high.file}`, tab: 'findings' }
108 const stale = s.evidence.latest.find(e => e.freshness === 'stale' && e.outcome === 'pass')
109 if (stale) return { text: `${stale.check} evidence is stale after later edits`, tab: 'evidence' }
110 const unsupported = s.claims.latest.find(cl => cl.verdict === 'unsupported')
111 if (unsupported) return { text: `Unsupported claim: "${unsupported.text}"`, tab: 'claims' }
112 if (s.approval !== 'approved') return { text: s.approval === 'amended' ? 'Contract amended: re-approve with /integrity-approve' : 'Draft contract: review and /integrity-approve', tab: 'requirements' }
113 if (s.must.total > s.must.verified) return { text: `${s.must.total - s.must.verified} must requirement(s) not verified yet`, tab: 'requirements' }
114 return undefined
115}
116
117/** The glanceable one-line status; undefined when there is nothing to say. Glyphs carry the meaning, not color. */
118export function statusLine(s: Summary, columns = 120): string | undefined {
119 if (!s.hasTask) return undefined
120 if (s.approval === 'draft' && s.evidence.total === 0 && s.findings.open === 0) return 'Integrity ◇ draft contract · /integrity'
121 const parts: string[] = []
122 const inForce = s.requirements.total - s.requirements.waived
123 if (inForce > 0) parts.push(`✓ ${s.requirements.verified}/${inForce} verified`)
124 const unverified = s.requirements.unverified + s.requirements.implemented
125 if (unverified > 0) parts.push(`? ${unverified} unverified`)
126 if (s.requirements.failed > 0) parts.push(`✗ ${s.requirements.failed} failed`)
127 if (s.findings.open > 0) parts.push(`! ${s.findings.open} finding${s.findings.open === 1 ? '' : 's'}`)
128 if (s.evidence.stale > 0) parts.push(`⧗ ${s.evidence.stale} stale`)
129 if (s.claims.contradicted + s.claims.unsupported > 0) parts.push(`✗ ${s.claims.contradicted + s.claims.unsupported} claim${s.claims.contradicted + s.claims.unsupported === 1 ? '' : 's'}`)
130 if (s.approval !== 'approved') parts.push(s.approval === 'amended' ? '◇ amended' : '◇ draft')
131 if (parts.length === 0) parts.push('no evidence yet')
132 // Narrow terminals keep the leading (most important) parts.
133 let line = 'Integrity ' + parts.join(' · ')
134 while (line.length > columns && parts.length > 1) {
135 parts.pop()
136 line = 'Integrity ' + parts.join(' · ')
137 }
138 return line
139}
140