SLOPSHOPPER

Claude Integrity

Tells implemented apart from verified: a requirement contract from your own words, a Mirage Detector for fake completeness, evidence from real test/build/lint…

newpaneguardcommandtoaststatus
v0.1.0MITupdated 2026-10-09NMenzel/claude-integrity-mod
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · integrity
│ ┃ Integrity ✕ › fix the failing auth╭────────────────────────────────────────────╮ │ ┃ INTEGRITY · NO CONTRACT │ integrity │ │ ┃ ╭─────────────────────────────────────────── ⏺ Read(src/auth.ts) │ Integrity: test failed (exit non-zero) · │ │ ┃ │ CONTRACT · none yet ⎿ Read 6 lines │ /integrity evidence │ │ ┃ │ ◇ No task contract. Nothing is tracked as ⏺ Update(src/auth.ts) ╰────────────────────────────────────────────╯ │ ┃ │ yet. ⎿ Added 2 lines, removed 1 line │ ┃ │ Describe what you want built (an implement ⏺ Bash(bun test) │ ┃ │ prompt drafts a contract, no model call), ⎿ 3 pass, 1 fail │ ┃ │ /integrity-start <task>. │ ┃ │ /integrity-help explains the workflow · /i ● Done. refresh now rejects expired claims and logs an audit event. │ ┃ │ for /integrity │ ┃ ╰─────────────────────────────────────────── ✻ Worked for 42s · done 4:20 PM │ │ › /integrity │ ⎿ integrity: Integrity ! 2 findings · ◇ draft │ ⎿ integrity: Top concern: Draft contract: review and /integrity-ap │ │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Pane · Integrity
INTEGRITY · NO CONTRACT ╭──────────────────────────────────────────────────────╮ │ CONTRACT · none yet │ │ ◇ No task contract. Nothing is tracked as verified │ │ yet. │ │ Describe what you want built (an implementation │ │ prompt drafts a contract, no model call), or run │ │ /integrity-start <task>. │ │ /integrity-help explains the workflow · /ig is short │ │ for /integrity │ ╰──────────────────────────────────────────────────────╯
README

<img src="docs/media/banner.svg" alt="Claude Integrity: a verified requirement, an unverified one, a fake save flagged by the Mirage Detector, and a contradicted 'All tests pass' claim" width="720">

<h1 align="center">Claude Integrity</h1>

<strong>"All tests pass." Did they run? Claude Integrity checks what your coding agent claims against what it actually did.</strong>

<sub>A Claude Code mod that keeps a requirement contract from your own words, flags code that only looks finished, counts a requirement as verified only after a real passing check, and audits completion claims. Everything stays on your machine.</sub>

<a href="https://github.com/NMenzel/claude-integrity-mod/stargazers"><img src="https://img.shields.io/github/stars/NMenzel/claude-integrity-mod?style=flat-square&color=yellow&label=stars" alt="GitHub stars"></a>&nbsp; <a href="https://github.com/NMenzel/claude-integrity-mod/releases/latest"><img src="https://img.shields.io/github/v/release/NMenzel/claude-integrity-mod?style=flat-square&label=version&color=blue" alt="Latest release"></a>&nbsp; <a href="LICENSE"><img src="https://img.shields.io/badge/license-MIT-green?style=flat-square" alt="License: MIT"></a>&nbsp; <a href="https://claude.com/blog/claude-code-mods"><img src="https://img.shields.io/badge/Claude%20Code-2.1.294%2B%20mod-d97757?style=flat-square" alt="Claude Code 2.1.294+ mod"></a>

<a href="#install">Install</a> · <a href="#use-it">Use it</a> · <a href="#the-dashboard">Dashboard</a> · <a href="#what-it-checks">What it checks</a> · <a href="#commands">Commands</a> · <a href="#options">Options</a> · <a href="docs/SECURITY.md">Security</a>


The problem

Claude Code ends a task with a confident summary. "Implemented the save flow. All tests pass." Often that is true. Sometimes:

  • The tests never ran. Or they ran before the last three edits, or failed behind a | tail.
  • The code only looks finished. A save handler returns { ok: true } and writes nothing. A button's onClick is () => {}. The "live" list still reads mockEvents. A service throws new Error("Not implemented").
  • The task drifted. You said "do not create another Discover page", and there is now app/discover-v2/page.tsx. You said to leave app/legacy/** alone, and it was edited.

Checking all of that by hand, after every turn, is the work you wanted to hand off.

The solution

Claude Integrity compares four things while Claude works, built on Claude Code's native Mods API:

  1. What you asked for: a requirement contract drafted from your own words (Intent Fingerprint).
  2. What changed: every edit, scanned for fake completeness and scope drift (Mirage Detector).
  3. What proves it: real test, build, lint and typecheck runs, tied to the code they ran against (The Witness).
  4. What was claimed: "all tests pass", "it's implemented" and similar lines in Claude's final answer, checked against that evidence (Claim Auditor).
In Claude Code todayWith Claude Integrity
"All tests pass." in the final answerA note beneath the answer: ? "All tests pass.": no test run was observed for this task
A test run that scrolled past three edits agoEvidence tied to the source state. Edits after the run mark it ⧗ stale
return Response.json({ ok: true }) and nothing savedMIR-B1: a POST handler returns success with no write, query, fetch or service call
"Do not create another Discover page" in your promptDRIFT-N1 when a new file is named after a second Discover page
Re-reading the diff to see what is really doneA status line, Integrity ✓ 2/5 verified · ? 3 unverified · ! 1 finding · ⧗ 1 stale, and a dashboard per requirement
Hoping review catches itStrict mode refuses edits to excluded paths, and claude -p "/integrity-check" exits 1 on violations

Nothing is verified because a model said so. A requirement is verified only by a real, passing, current check of a kind that can prove it, or by your own confirmation.

Native mod. No wrappers, no API keys, no network calls. No model calls unless you turn on /integrity-suggest. Two commands to install.

It sees what the Mods API exposes: your prompts, Claude's tool calls and their results, and the files Claude edits. It does not see tests you run in your own terminal or in CI (confirm those with /integrity-verify). Its rules read source text, not a syntax tree, so every finding that is not literal is marked potential. It never approves a tool call and never rewrites Claude's answer.

[!TIP] If Claude Integrity catches a "done" that wasn't, a ⭐ on the repo helps other developers find it.

Claude Integrity is a community project. It is not made or endorsed by Anthropic.

Requirements

  • Claude Code with Mods support. Built and tested on 2.1.294 (Windows 10). Mods need 2.1.287 or later, and builds older than 2.1.294 are untested.
  • The dashboard pane is built for the terminal and the Code tab of Claude Desktop. In the VS Code chat panel and in claude -p the hooks still run, and every /integrity-* command answers as text.

Install

Inside Claude Code:

/plugin marketplace add NMenzel/claude-integrity-mod
/plugin install integrity@claude-integrity-mod
/reload-plugins
/integrity-help

Or from a terminal, in one command (it adds the marketplace too):

claude plugin install integrity --marketplace NMenzel/claude-integrity-mod

Then run /reload-plugins in an open session, or start a new one.

The installer may say config options aren't set. The defaults are fine, and /config changes them. Mods are an early-access Claude Code feature, and their API can change between releases.

Check that it loaded: run /plugin and look for integrity in the mods active line. You can also run /integrity-help.

To update: claude plugin update integrity@claude-integrity-mod. To remove: claude plugin uninstall integrity@claude-integrity-mod.

Install with an AI agent

Paste this into Claude Code, or any coding agent that can run shell commands:

Install the Claude Integrity mod for Claude Code from
https://github.com/NMenzel/claude-integrity-mod. Follow the
"Instructions for agents" in its README, then tell me what to run next.
Instructions for agents

If you are an AI agent installing Claude Integrity for a user, follow these steps.

  1. Check the version. Run claude --version. Claude Code 2.1.287 or later supports mods, and Claude Integrity is tested on 2.1.294. If the version is older, stop and tell the user to update Claude Code (claude update).
  2. Install, without prompts: ``sh claude plugin install integrity --marketplace NMenzel/claude-integrity-mod --json ` This adds the marketplace to the user's settings and installs at user scope. Exit code 0 means it is installed. To share it with everyone working in the current repository, add --scope project (it is then recorded in .claude/settings.json`). Ask the user before using a scope other than the default.
  3. Options are optional. The defaults suit most users. Set an option only if the user asked for it, with --config <option>=<value> on the install command. Examples: --config notify=all, --config autoCapture=false. All options are listed under Options.
  4. Verify: claude plugin list --json includes integrity@claude-integrity-mod.
  5. Hand over to the user. An agent cannot activate a mod inside the user's running session. Tell the user to run /reload-plugins, then /integrity-help. A newly started session loads the mod by itself.

Notes for agents:

  • Nothing needs to be built, and it needs no npm packages, global installs or sudo. Do not clone the repository to install it.
  • The mod makes no network calls and no model calls by default. It never approves a tool call and does not change the user's permission settings.
  • To test it headlessly: claude -p "/integrity-help" prints the help without a model call. In Git Bash on Windows, prefix MSYS_NO_PATHCONV=1 so the shell does not turn /integrity-help into a path.

Use it

  1. Describe a task the way you normally would: "Enrich the existing Discover page with live events. Do not create another Discover page. Tests must pass." Integrity drafts a contract from the explicit requirements and constraints. No model call is made, and nothing is approved for you. To start one by hand: /integrity-start <task>.
  2. Review it with /integrity-contract (short: /igc). Edit it in place: ``text /integrity-contract add Events are fetched from /api/events /integrity-contract edit R2 Keep the existing /discover route /integrity-contract crit R2 should # must | should | optional /integrity-contract check R1 npm test # this command proves R1 /integrity-contract exclude app/legacy/** # never touch these paths /integrity-contract intent prototype # fixtures are expected, not findings ``
  3. Approve it with /integrity-approve (/iga). That freezes the baseline. Later changes become versioned revisions.
  4. Work normally. The status line under the prompt keeps count: Integrity ✓ 2/5 verified · ? 3 unverified · ! 1 finding · ⧗ 1 stale.
  5. Open the dashboard with /integrity, or /ig. Keys 1-7 switch between Overview, Requirements, Evidence, Findings, Claims, Activity and Report. Press a row for details. k acknowledges a finding, w waives it (a reason is required), c re-checks, e/j export.
  6. Hand it over with /integrity-report export (/igr export). It writes .claude/integrity/<task>.md and .json.

The dashboard

/ig opens it. It is drawn in the style of Claude Flightdeck, and docked beside a fullscreen transcript it asks for the same 66 columns:

  • A centered title, INTEGRITY · STRICT MODE · APPROVED V2, over a color legend.
  • CONTRACT: the task, one colored cell per requirement (verified, implemented, not verified, stale, failed), a ▰▰▰▱▱ 3/5 verified · must 2/3 gauge, and the open problems.
  • EVIDENCE: the latest run of each kind of check, fresh or ⧗ stale.
  • FINDINGS: counts by severity and the top three open findings. Press one to open its details.
  • CLAIMS: the verdicts on Claude's last answer, and the worst two claims with their reasons.
  • The top concern in a card of its own (i inspects it), a approve, c re-check, and the activity log.

From 110 columns the cards sit side by side in two columns. Inline above the prompt it is a three-line summary. With the color option off, the glyphs (✓ ◐ ? ⧗ ✗) carry every state on their own.

What it checks

Intent Fingerprint: the contract

Each requirement is your own sentence, with a criticality (must, should, optional), a category (functional, constraint, non-functional, test, preservation) and how it can be proved (static, test, runtime, manual). Exclusions (exclude <glob>), scope (scope <glob>) and intent (production, prototype) shape how strict the checks are.

The contract lives in the mod's own store, keyed per repository, outside the conversation. Compaction cannot drop it, and a new session picks it up.

Mirage Detector: code that only looks finished

Twelve rules for JavaScript, TypeScript, React and Next.js, each with its confidence:

  • Placeholders: throw new Error("Not implemented") (confirmed), TODOs that defer the real work, example.com endpoints and YOUR_API_KEY.
  • Simulated success: a mutating handler that returns success with no IO, a setTimeout delay followed by a success state.
  • Inert controls: onClick={() => {}}, log-only handlers, href="#".
  • Mock data shipped as real: mockEvents in a production view (expected in a prototype).
  • Incomplete integration: a fetch whose result is never read, a client calling /api/x that no observed route serves, a save that only sets state, a fetch with no error handling, role-gated UI over unguarded handlers.

And seven intent drift rules: edits to excluded paths, edits outside the scope, a new file named after something you said not to recreate, a new route while routes must be preserved, deleted code files, near-total rewrites, and new dependencies the contract never mentions.

Every finding is confirmed, potential (it says what to check) or expected (allowed by the contract, shown but never counted). Acknowledged and waived findings stay quiet across later scans. The full list, with each rule's false-positive guard, is in docs/DETECTION-RULES.md.

The Witness: evidence

Test, build, lint, typecheck, migration and browser-test runs that Claude makes become evidence: the command, the outcome, the counts the runner printed (jest, vitest, mocha, pytest, cargo, go, node:test), and the git HEAD they ran against.

  • A run that was denied, synthetic, sent to the background, interrupted or had its exit status masked by a pipe or || true is never a pass.
  • An edit to a related file after the run makes it ⧗ stale.
  • A browser session without an assertion is inconclusive. Clicking around is not a test.
  • What Claude cannot run, you confirm: /integrity-verify R3 submitted the form and reloaded: saved.

Claim Auditor

When Claude finishes, Integrity reads the final answer for completion claims ("all tests pass", "it's implemented", "uses live data", "production-ready") and checks each against the evidence: supported, partial, unsupported, contradicted or not assessable. When one is unsupported or contradicted, a short note appears beneath the answer:

Integrity · 2 completion claims checked: 1 partial · 1 unsupported
  ? "All tests pass.": no test run was observed for this task
  details: /integrity claims

The answer itself is never changed. Hedged ("should pass") and negated sentences are skipped on purpose: a missed claim is better than a false alarm.

Modes

  • observe (default): records, flags, and adds a note beneath an answer when a claim is unsupported or contradicted. Never blocks anything.
  • review: also lists open must-requirements and high findings after each answer, and opens the pane.
  • strict: refuses edits to paths the approved contract excludes (a real tool-call refusal), and makes /integrity-check exit 1 on policy violations. It cannot stop a turn from completing, because the Mods API has no such hook. It says so instead of pretending.

Set it per project with /integrity-mode observe|review|strict.

To use the check as a gate in a script or a git hook on the machine where the contract lives:

claude -p "/integrity-check"     # exit 1 under strict mode when a policy is violated

Commands

| Command | Does | | - | - | | /integrity [view] | Open the pane (in a headless session: print the status). /integrity <subcommand> runs /integrity-<subcommand> | | /integrity-start [text] | Start a contract from the text, or from your last prompt | | /integrity-contract [action] | View or edit the contract (add, edit, remove, crit, category, verify-by, check, exclude, scope, assume, intent, summary) | | /integrity-approve | Approve the baseline or a revision | | /integrity-check | Re-read tracked files, re-scan and evaluate policies. Exits 1 under strict mode on violations | | /integrity-findings [all\|ack <id>\|reopen <id>] | List or triage findings | | /integrity-evidence [link E# R#] | List evidence, or link a check to a requirement | | /integrity-verify R# <what you checked> | Record your own confirmation | | /integrity-report [md\|json] [export] | Print or export the delivery report | | /integrity-mode observe\|review\|strict | Set the mode for this project | | /integrity-waive <id\|R#> <reason> | Waive a finding or requirement, with a reason | | /integrity-reset [all] [confirm] | Delete the active task (or all of this project's data) after confirmation | | /integrity-ui [setting value ...] | Status hint, notifications, color, findings per page, open or close the pane | | /integrity-layout auto\|mini\|compact\|wide | Pane layout | | /integrity-suggest | Opt-in AI requirement suggestions: one model call, results marked proposed | | /integrity-help | The workflow in short | | /ig [view\|subcommand] | Short for /integrity: /ig check, /ig start <task> and every other subcommand work too | | /igc · /iga · /igv · /igf · /igr | Short for /integrity-contract, /integrity-approve, /integrity-verify, /integrity-findings and /integrity-report |

Options

Set them in /config, or at install time with claude plugin install ... --config <option>=<value>:

| Option | Default | Meaning | | - | - | - | | notify | important | Hints and notes beneath answers: off, important (only unsupported or contradicted claims) or all | | statusHint | true | The one-line status under the prompt | | paneAutoOpen | false | Open the pane when a session starts (only where it docks as a sidebar) | | layout | auto | Pane layout: auto, mini, compact or wide | | maxFindings | 8 | Findings per pane page, 1 to 50 | | color | true | Color the status glyphs (the glyphs always carry the meaning) | | autoCapture | true | Draft a contract from an implementation prompt when no task is active. Never approves it | | retentionDays | 30 | Idle tasks older than this are deleted, 1 to 3650 | | gitProvenance | true | Record the git HEAD with each piece of evidence (git rev-parse, read-only) | | aiAssist | false | Allow /integrity-suggest to make one model call per request |

/integrity-ui and /integrity-layout change the display options without leaving the session.

Privacy

Local only: no telemetry and no network calls. The only processes it starts are two read-only git commands for provenance (off with gitProvenance). It stores a hash of your prompt and a short redacted excerpt, never the full prompt and never raw tool output. Credentials, tokens and home-directory paths are redacted. Details: docs/SECURITY.md.

Verified vs. limited

Covered by 66 automated tests (claude plugin test) and a real headless load (claude -p --plugin-dir):

  • Contracts: capture, approval, versioned amendments, persistence across processes, per-repository isolation.
  • Evidence: test, build, lint, typecheck, migration and browser runs; denied, synthetic, background, interrupted and exit-masked runs are never a pass; freshness and staleness.
  • The twelve Mirage rules and seven drift rules on fixture projects, with expected-mock handling and persistent waivers.
  • Claim auditing (supported, partial, unsupported, contradicted, not assessable) and the note beneath the answer.
  • The pane on the terminal, desktop and mobile element tables at narrow, typical and wide widths, docked and inline, with and without color, with working presses, inputs and paging.
  • The short aliases answering exactly as the commands they stand for.
  • The strict gate exiting 1 in a real headless process.

Not covered by tests: how the terminal and Desktop paint the pane (the test kit checks the drawn trees, not the pixels). See docs/LIMITATIONS.md.

Develop and test

npm install                      # local TypeScript only; nothing global
claude plugin validate .         # manifest, hooks, calls, state contract
claude plugin test .             # 66 tests: pure engine + real hooks and pane through claude-code/testing
npx tsc -p .                     # type-check (after one load has laid .claude-plugin/types)

The engine writes .claude-plugin/types/ the first time it loads the folder. To lay it without starting an interactive session, run once: claude -p --plugin-dir . "/integrity-help".

See docs/ARCHITECTURE.md, docs/DETECTION-RULES.md, docs/API-COMPATIBILITY.md, docs/SECURITY.md and docs/LIMITATIONS.md. CONTRIBUTING.md has the ground rules; changes are listed in CHANGELOG.md.

License

MIT

Source 15 files
hooks/register.tsx 780 lines
1// Claude Integrity: the Claude Code Mod adapter. It observes prompts, tool
2// calls and turn completions, keeps the engine's state in `$.store`, and draws
3// the status line and the /integrity pane. All judgement lives in the pure
4// engine (../src/engine); this file does IO and nothing else.
5
6import { atom, read, update } from 'claude-code'
7import type { EngineInterface, PluginOptions, Register } from 'claude-code'
8
9import type { IntegrityUi } from '../types'
10import * as E from '../src/engine/index'
11import { renderPane, type PaneActions } from './pane'
12
13const PANE = 'integrity'
14const rev = atom({ plugin: 'integrity', key: 'rev' } as const, 0)
15const uiState = atom({ plugin: 'integrity', key: 'ui' } as const, { tab: 'overview', page: 0 } as IntegrityUi)
16
17export type Prefs = {
18  notify: E.NotifyLevel
19  statusHint: boolean
20  paneAutoOpen: boolean
21  layout: E.Layout
22  maxFindings: number
23  color: boolean
24  autoCapture: boolean
25  retentionDays: number
26  gitProvenance: boolean
27  aiAssist: boolean
28}
29
30function prefsFrom(options: PluginOptions, stored: Partial<Prefs> | undefined): Prefs {
31  const pick = <T,>(key: keyof Prefs, fallback: T, ok: (v: unknown) => boolean): T => {
32    const s = stored?.[key]
33    if (s !== undefined && ok(s)) return s as T
34    const o = options[key]
35    return o !== undefined && ok(o) ? (o as T) : fallback
36  }
37  const isBool = (v: unknown) => typeof v === 'boolean'
38  const isNum = (v: unknown) => typeof v === 'number' && Number.isFinite(v) && v > 0
39  return {
40    notify: pick('notify', 'important', v => v === 'off' || v === 'important' || v === 'all'),
41    statusHint: pick('statusHint', true, isBool),
42    paneAutoOpen: pick('paneAutoOpen', false, isBool),
43    layout: pick('layout', 'auto', v => v === 'auto' || v === 'mini' || v === 'compact' || v === 'wide'),
44    maxFindings: Math.min(50, Math.round(pick('maxFindings', 8, isNum))),
45    color: pick('color', true, isBool),
46    autoCapture: pick('autoCapture', true, isBool),
47    retentionDays: pick('retentionDays', 30, isNum),
48    gitProvenance: pick('gitProvenance', true, isBool),
49    aiAssist: pick('aiAssist', false, isBool),
50  }
51}
52
53const COMMANDS: readonly { name: string; description: string; argumentHint?: string }[] = [
54  { name: 'integrity', description: 'Claude Integrity: open the dashboard (or /integrity <subcommand>)', argumentHint: '[help|status|...]' },
55  { name: 'integrity-start', description: 'Integrity: start a task contract from text or your last prompt', argumentHint: '[task description]' },
56  { name: 'integrity-contract', description: 'Integrity: view or edit the contract (add, edit, remove, crit, check, exclude, scope, intent, ...)', argumentHint: '[add|edit|remove|crit|check|exclude|scope|assume|intent|implemented] ...' },
57  { name: 'integrity-approve', description: 'Integrity: approve the contract as the baseline' },
58  { name: 'integrity-check', description: 'Integrity: re-check tracked files and evaluate policies (exit 1 in strict mode on violations)' },
59  { name: 'integrity-findings', description: 'Integrity: list findings, or ack/reopen one', argumentHint: '[all | ack <id> | reopen <id>]' },
60  { name: 'integrity-evidence', description: 'Integrity: list evidence, or link one to a requirement', argumentHint: '[link <E#> <R#>]' },
61  { name: 'integrity-verify', description: 'Integrity: record your own manual confirmation of a requirement', argumentHint: '<R#> <what you checked>' },
62  { name: 'integrity-report', description: 'Integrity: print the delivery report, or export it to files', argumentHint: '[md|json] [export]' },
63  { name: 'integrity-mode', description: 'Integrity: set the mode (observe, review, strict)', argumentHint: '[observe|review|strict]' },
64  { name: 'integrity-waive', description: 'Integrity: waive a finding or requirement, with a reason', argumentHint: '<finding-id|R#> <reason>' },
65  { name: 'integrity-reset', description: 'Integrity: delete the active task (or all stored data) after confirmation', argumentHint: '[all] [confirm]' },
66  { name: 'integrity-ui', description: 'Integrity: configure the status hint, pane and notifications', argumentHint: '[status on|off] [pane open|close] [autoopen on|off] [notify off|important|all] [color on|off] [max N]' },
67  { name: 'integrity-layout', description: 'Integrity: set the pane layout', argumentHint: '[auto|mini|compact|wide]' },
68  { name: 'integrity-suggest', description: 'Integrity: suggest requirements with one opt-in model call (proposed, needs approval)' },
69  { name: 'integrity-help', description: 'Integrity: how it works and every command' },
70]
71
72// Short aliases, like Claude DevTools' /bp family: /ig is /integrity (so /ig check, /ig approve ... work too).
73const SHORTCUTS: readonly { name: string; to: string; description: string; argumentHint?: string }[] = [
74  { name: 'ig', to: 'integrity', description: 'Integrity: open the dashboard (= /integrity); /ig check, /ig start ... run a subcommand', argumentHint: '[view|subcommand] ...' },
75  { name: 'igc', to: 'integrity-contract', description: 'Integrity: view or edit the contract (= /integrity-contract)', argumentHint: '[add|edit|remove|crit|check|exclude|scope|intent] ...' },
76  { name: 'iga', to: 'integrity-approve', description: 'Integrity: approve the contract (= /integrity-approve)' },
77  { name: 'igv', to: 'integrity-verify', description: 'Integrity: record your own confirmation (= /integrity-verify)', argumentHint: '<R#> <what you checked>' },
78  { name: 'igf', to: 'integrity-findings', description: 'Integrity: list or triage findings (= /integrity-findings)', argumentHint: '[all | ack <id> | reopen <id>]' },
79  { name: 'igr', to: 'integrity-report', description: 'Integrity: print or export the report (= /integrity-report)', argumentHint: '[md|json] [export]' },
80]
81const ALIASES: Readonly<Record<string, string>> = Object.fromEntries(SHORTCUTS.map(s => [s.name, s.to]))
82
83const HELP = `**Claude Integrity**: never confuse AI-generated implementation with verified completion.
84
85It keeps a *contract* (what you asked for), watches *changes* and *checks* Claude runs, flags *mirages* (code that looks finished but isn't) and audits *completion claims* against evidence. Nothing is "verified" without a real, passing, current check, or your own confirmation.
86
871. Describe a task, or \`/integrity-start <task>\`: a draft contract is captured from your words (no model call).
882. \`/integrity-contract\` to review; \`add\`, \`edit R2 ...\`, \`crit R2 should\`, \`check R1 npm test\`, \`exclude app/legacy/**\`, \`intent prototype\`.
893. \`/integrity-approve\` freezes the baseline; later edits become versioned revisions.
904. Work normally. Test/build/lint/typecheck runs become evidence; later edits make it stale.
915. \`/integrity\` opens the dashboard (keys 1-7 switch views). \`/integrity-report export\` writes Markdown + JSON.
92
93Other: \`/integrity-verify R3 <what you checked>\` · \`/integrity-waive <id> <reason>\` · \`/integrity-findings ack <id>\` · \`/integrity-evidence link E4 R2\` · \`/integrity-mode observe|review|strict\` · \`/integrity-check\` · \`/integrity-ui\` · \`/integrity-layout\` · \`/integrity-suggest\` (opt-in AI) · \`/integrity-reset\`.
94
95Short aliases: \`/ig\` (= \`/integrity\`; \`/ig check\`, \`/ig start ...\`) · \`/igc\` contract · \`/iga\` approve · \`/igv\` verify · \`/igf\` findings · \`/igr\` report.
96
97Modes: **observe** (default) records and notes; **review** adds a summary of open must-requirements and high findings after each answer; **strict** also refuses edits to paths the approved contract excludes and makes \`/integrity-check\` exit 1 on policy violations. Nothing blocks a turn from completing; the Mods API cannot, so Integrity says so instead of pretending.`
98
99type Runtime = {
100  options: PluginOptions
101  project?: E.ProjectRecord
102  loading?: Promise<E.ProjectRecord>
103  key: string
104  root: string
105  session: string
106  prefs: Prefs
107  storedPrefs?: Partial<Prefs>
108  lastStatus?: string
109  dirty: boolean
110  writing?: Promise<void>
111  deleted: Set<string>
112  lastPrompt?: string
113  suggested: Map<string, string[]>
114}
115
116
117function freshRuntime(options: PluginOptions): Runtime {
118  return { options, key: '', root: '', session: 'session', prefs: prefsFrom(options, undefined), dirty: false, deleted: new Set(), suggested: new Map() }
119}
120
121/** Module state, renewed each time the engine loads the module (a hot reload is a fresh load). */
122let rt: Runtime = freshRuntime({})
123
124function now($: EngineInterface): Promise<number> {
125  return $.clock.now()
126}
127
128function rel(path: string): string {
129  return E.normalizePath(path, rt.root)
130}
131
132function ensure($: EngineInterface): Promise<E.ProjectRecord> {
133  if (rt.project !== undefined) return Promise.resolve(rt.project)
134  rt.loading ??= (async () => {
135    const repo = await $.session.repo().catch(() => null)
136    const root = repo?.root ?? (await $.session.root().catch(() => $.session.cwd()))
137    rt.root = root
138    rt.key = E.projectKey(root, repo?.remote)
139    rt.session = await $.session.id().catch(() => 'session')
140    rt.storedPrefs = ((await $.store.get('prefs')) ?? undefined) as Partial<Prefs> | undefined
141    rt.prefs = prefsFrom(rt.options, rt.storedPrefs)
142    const stored = (await $.store.get(rt.key)) as E.ProjectRecord | undefined
143    const label = repo?.remote ?? root.replace(/\\/g, '/').split('/').slice(-2).join('/')
144    rt.project = stored?.schema === 1 ? stored : E.newProject(rt.key, label, rt.session, await now($))
145    return rt.project
146  })()
147  return rt.loading
148}
149
150/** Re-reads, merges and writes; writes coalesce, so a burst of changes is one or two writes. */
151async function writeOnce($: EngineInterface): Promise<void> {
152  const mine = rt.project
153  if (mine === undefined) return
154  const t = await now($)
155  const stored = (await $.store.get(rt.key)) as E.ProjectRecord | undefined
156  const merged = E.mergeProjects(mine, stored, t)
157  for (const id of rt.deleted) delete merged.tasks[id]
158  if (merged.activeTaskId !== undefined && merged.tasks[merged.activeTaskId] === undefined) delete merged.activeTaskId
159  merged.writer = rt.session
160  E.trimProject(merged, t, rt.prefs.retentionDays, 600_000)
161  await $.store.set(rt.key, merged)
162  rt.project = merged
163}
164
165function save($: EngineInterface): Promise<void> {
166  rt.dirty = true
167  rt.writing ??= (async () => {
168    try {
169      while (rt.dirty) {
170        rt.dirty = false
171        await writeOnce($)
172      }
173    } catch (err) {
174      $.ui.log(`integrity: could not save state (${String(err)})`, { to: 'debug' })
175    } finally {
176      rt.writing = undefined
177    }
178  })()
179  return rt.writing
180}
181
182async function refreshStatus($: EngineInterface): Promise<void> {
183  const p = await ensure($)
184  const line = rt.prefs.statusHint ? E.statusLine(E.summarize(E.activeTask(p), p.mode)) : undefined
185  if (line === rt.lastStatus) return
186  rt.lastStatus = line
187  $.ui.status(line)
188}
189
190function showHints($: EngineInterface, hints: readonly E.Hint[]): void {
191  for (const h of hints) {
192    if (rt.prefs.notify === 'off' || (rt.prefs.notify === 'important' && h.level !== 'important')) continue
193    $.ui.toast(h.text, { timeoutMs: 6000 })
194  }
195}
196
197/** After any state change: redraw subscribers, refresh the status line, show hints, persist. */
198async function changed($: EngineInterface, hints: readonly E.Hint[] = []): Promise<void> {
199  await update($, rev, n => n + 1)
200  await refreshStatus($)
201  showHints($, hints)
202  await save($)
203}
204
205/** Docked it asks Flightdeck's width beside the transcript; inline above the prompt, a short block. */
206function openPane($: EngineInterface) {
207  return $.ui.open({ id: PANE, title: 'Integrity', columns: 66, rows: 10 })
208}
209
210async function gitHead($: EngineInterface): Promise<string | undefined> {
211  if (!rt.prefs.gitProvenance) return undefined
212  const run = await $.process.run(['git', 'rev-parse', '--short', 'HEAD'], { timeoutMs: 3000 }).catch(() => undefined)
213  if (run === undefined || run.exitCode !== 0) return undefined
214  const dirty = await $.process.run(['git', 'status', '--porcelain', '--untracked-files=no'], { timeoutMs: 3000 }).catch(() => undefined)
215  return `${run.stdout.trim()}${dirty !== undefined && dirty.stdout.trim() !== '' ? `+dirty:${E.hash(dirty.stdout).slice(0, 6)}` : ''}`
216}
217
218async function readText($: EngineInterface, path: string): Promise<string | undefined> {
219  const text = await $.fs.read(path).catch(() => undefined)
220  return typeof text === 'string' ? text : undefined
221}
222
223function filePath(e: { tool: string } & Record<string, unknown>): string | undefined {
224  const path = e['file_path'] ?? e['notebook_path']
225  return (e.tool === 'Edit' || e.tool === 'Write' || e.tool === 'NotebookEdit' || e.tool === 'MultiEdit') && typeof path === 'string' ? path : undefined
226}
227
228// ---- Commands -------------------------------------------------------------------
229
230function words(args: string): string[] {
231  return args.trim().split(/\s+/).filter(Boolean)
232}
233
234async function requireTask($: EngineInterface): Promise<E.Task | string> {
235  const task = E.activeTask(await ensure($))
236  return task ?? 'No active task. Describe what you want built, or run `/integrity-start <task>`.'
237}
238
239function contractText(task: E.Task, mode: E.Mode): string {
240  const c = task.contract
241  const s = E.summarize(task, mode)
242  const lines = [
243    `**Contract** ${c.summary} · ${s.approval} v${c.version} · intent ${c.intent} · task \`${task.id}\``,
244    `Prompt: "${c.prompt.excerpt}" (hash ${c.prompt.hash.slice(0, 10)})`,
245    '',
246    ...(c.requirements.length === 0 ? ['_No requirements yet: `/integrity-contract add <requirement>`_'] : c.requirements.map(r => {
247      const st = E.requirementState(r, task)
248      return `- **${r.id}** [${r.criticality}/${r.category}/${r.verification}${r.origin === 'proposed' ? '/proposed' : ''}] ${r.description}  \n  → ${st.state}: ${st.reason}${r.checks.length ? ` · checks: ${r.checks.join(', ')}` : ''}`
249    })),
250  ]
251  if (c.exclusions.length) lines.push('', `Exclusions: ${c.exclusions.map(x => `\`${x}\``).join(', ')}`)
252  if (c.scope.length) lines.push(`Scope: ${c.scope.map(x => `\`${x}\``).join(', ')}`)
253  if (c.assumptions.length) lines.push(`Assumptions: ${c.assumptions.join('; ')}`)
254  if (c.revisions.length) lines.push(`Revisions: ${c.revisions.slice(-5).map(r => `v${r.version} ${r.change}`).join(' · ')}`)
255  lines.push('', s.approval === 'approved' ? 'Approved baseline kept; edits create revisions.' : 'Not approved yet: `/integrity-approve` when it reads right.')
256  return lines.join('\n')
257}
258
259async function editContract($: EngineInterface, args: string): Promise<string> {
260  const task = await requireTask($)
261  if (typeof task === 'string') return task
262  const p = rt.project as E.ProjectRecord
263  const [verb = '', ...rest] = words(args)
264  const t = await now($)
265  const c = task.contract
266  if (verb === '') return contractText(task, p.mode)
267  const id = rest[0] ?? ''
268  const text = rest.slice(1).join(' ')
269  const req = E.findRequirement(c, id)
270  const needReq = (): string | undefined => (req === undefined ? `No requirement "${id}". Requirements: ${c.requirements.map(r => r.id).join(', ') || 'none'}.` : undefined)
271  let done: string
272  switch (verb) {
273    case 'add': {
274      const description = rest.join(' ')
275      if (description === '') return 'Usage: `/integrity-contract add <requirement>`'
276      let added: E.Requirement | undefined
277      E.amend(c, t, `added requirement: ${description}`, x => { added = E.makeRequirement(x.requirements, description); x.requirements.push(added) })
278      done = `Added ${added?.id}.`
279      break
280    }
281    case 'edit':
282      if (needReq() || text === '') return needReq() ?? 'Usage: `/integrity-contract edit R2 <new text>`'
283      E.amend(c, t, `${req!.id} reworded`, () => { req!.description = E.redact(text).slice(0, 300); if (req!.origin === 'proposed') req!.origin = 'manual' })
284      done = `Updated ${req!.id}.`
285      break
286    case 'remove':
287      if (needReq()) return needReq() as string
288      E.amend(c, t, `${req!.id} removed: ${req!.description}`, x => { x.requirements = x.requirements.filter(r => r !== req) })
289      done = `Removed ${req!.id}.`
290      break
291    case 'crit':
292    case 'criticality':
293      if (needReq() || !['must', 'should', 'optional'].includes(text)) return needReq() ?? 'Usage: `/integrity-contract crit R2 must|should|optional`'
294      E.amend(c, t, `${req!.id} criticality ${req!.criticality} → ${text}`, () => { req!.criticality = text as E.Criticality })
295      done = `${req!.id} is now ${text}.`
296      break
297    case 'category':
298      if (needReq() || !['functional', 'constraint', 'non-functional', 'test', 'preservation'].includes(text)) return needReq() ?? 'Usage: `/integrity-contract category R2 functional|constraint|non-functional|test|preservation`'
299      E.amend(c, t, `${req!.id} category → ${text}`, () => { req!.category = text as E.RequirementCategory })
300      done = `${req!.id} category is now ${text}.`
301      break
302    case 'verify-by':
303      if (needReq() || !['static', 'test', 'runtime', 'manual', 'unknown'].includes(text)) return needReq() ?? 'Usage: `/integrity-contract verify-by R2 static|test|runtime|manual`'
304      E.amend(c, t, `${req!.id} verification → ${text}`, () => { req!.verification = text as E.VerificationKind })
305      done = `${req!.id} is verified by ${text} evidence.`
306      break
307    case 'check':
308      if (needReq() || text === '') return needReq() ?? 'Usage: `/integrity-contract check R2 <command substring, e.g. npm test -- events>`'
309      E.amend(c, t, `${req!.id} check: ${text}`, () => { req!.checks = [...new Set([...req!.checks, text])] })
310      done = `${req!.id} is verified by a passing run of a command containing "${text}".`
311      break
312    case 'implemented':
313      if (needReq()) return needReq() as string
314      E.amend(c, t, `${req!.id} marked implemented by the developer`, () => { req!.implemented = true }, false)
315      done = `${req!.id} marked implemented (not verified: that still needs evidence).`
316      break
317    case 'exclude':
318    case 'scope':
319    case 'assume': {
320      const value = rest.join(' ')
321      if (value === '') return `Usage: \`/integrity-contract ${verb} <${verb === 'assume' ? 'text' : 'glob, e.g. app/legacy/**'}>\``
322      const list = verb === 'exclude' ? 'exclusions' : verb === 'scope' ? 'scope' : 'assumptions'
323      E.amend(c, t, `${verb}: ${value}`, x => { x[list] = [...new Set([...x[list], E.redact(value)])] }, verb !== 'assume')
324      done = `Added to ${list}: ${value}`
325      break
326    }
327    case 'intent':
328      if (!['production', 'prototype', 'unknown'].includes(id)) return 'Usage: `/integrity-contract intent production|prototype|unknown`'
329      E.amend(c, t, `intent ${c.intent} → ${id}`, x => { x.intent = id as E.ContractIntent })
330      done = `Intent is now ${id}. Mirage rules read mock data as ${id === 'prototype' ? 'expected' : 'suspicious'}.`
331      break
332    case 'summary':
333      if (rest.length === 0) return 'Usage: `/integrity-contract summary <one line>`'
334      E.amend(c, t, 'summary changed', x => { x.summary = E.redact(rest.join(' ')).slice(0, 140) }, false)
335      done = 'Summary updated.'
336      break
337    default:
338      return `Unknown contract action "${verb}". Try: add, edit, remove, crit, category, verify-by, check, implemented, exclude, scope, assume, intent, summary.`
339  }
340  E.log(task, 'contract', done, t)
341  await changed($)
342  const note = c.baseline !== undefined && c.approval.status !== 'approved' ? ' Contract amended to v' + c.version + ': re-approve with `/integrity-approve`.' : ''
343  return done + note
344}
345
346async function findingsText($: EngineInterface, args: string): Promise<string> {
347  const task = await requireTask($)
348  if (typeof task === 'string') return task
349  const [verb = '', id = ''] = words(args)
350  if (verb === 'ack' || verb === 'reopen') {
351    const f = E.setFindingStatus(task, id, verb === 'ack' ? 'acknowledged' : 'open', await now($))
352    if (f === undefined) return `No finding "${id}".`
353    await changed($)
354    return `${f.ruleId} ${f.file}:${f.line} ${verb === 'ack' ? 'acknowledged' : 'reopened'}.`
355  }
356  const order: Record<E.Severity, number> = { high: 0, medium: 1, low: 2, info: 3 }
357  const list = task.findings
358    .filter(f => verb === 'all' || f.status === 'open')
359    .sort((a, b) => order[a.severity] - order[b.severity])
360  if (list.length === 0) return verb === 'all' ? 'No findings recorded.' : 'No open findings. `/integrity-findings all` lists resolved, waived and acknowledged ones.'
361  return list.slice(0, 40).map(f =>
362    `- \`${f.id.slice(0, 8)}\` **${f.severity}/${f.confidence}** ${f.ruleId} \`${f.file}:${f.line}\` (${f.status}): ${f.explanation}\n  Verify: ${f.verify}`,
363  ).join('\n') + (list.length > 40 ? `\n…and ${list.length - 40} more` : '')
364}
365
366async function evidenceText($: EngineInterface, args: string): Promise<string> {
367  const task = await requireTask($)
368  if (typeof task === 'string') return task
369  const [verb = '', eid = '', rid = ''] = words(args)
370  if (verb === 'link') {
371    const ev = task.evidence.find(x => x.id.toLowerCase() === eid.toLowerCase())
372    const req = E.findRequirement(task.contract, rid)
373    if (ev === undefined || req === undefined) return 'Usage: `/integrity-evidence link E4 R2` (both must exist).'
374    req.evidenceIds = [...new Set([...req.evidenceIds, ev.id])]
375    E.log(task, 'requirement', `${ev.id} linked to ${req.id} by the developer`, await now($))
376    await changed($)
377    return `${ev.id} linked to ${req.id}: now ${E.requirementState(req, task).state}.`
378  }
379  if (task.evidence.length === 0) return 'No checks observed yet. Test, build, lint and typecheck runs Claude makes are recorded automatically.'
380  return task.evidence.slice(-30).map(ev => {
381    const fresh = E.freshness(ev, task)
382    const why = fresh === 'stale' ? ` (stale: ${E.staleBecause(ev, task).slice(0, 3).join(', ')} changed)` : ''
383    return `- **${ev.id}** ${ev.check} ${ev.outcome} · ${fresh}${why} · ${ev.command ? `\`${ev.command}\`` : ev.source} · ${ev.basis}${ev.revision ? ` @ ${ev.revision}` : ''}`
384  }).join('\n')
385}
386
387async function check($: EngineInterface): Promise<{ text: string; exitCode?: number }> {
388  const task = await requireTask($)
389  if (typeof task === 'string') return { text: task }
390  const p = rt.project as E.ProjectRecord
391  const current: Record<string, string | null> = {}
392  for (const path of Object.keys(task.files).slice(0, 200)) {
393    const abs = /^([A-Za-z]:)?[\\/]/.test(path) ? path : `${rt.root.replace(/[\\/]$/, '')}/${path}`
394    current[path] = (await $.fs.exists(abs).catch(() => false)) ? (await readText($, abs)) ?? null : null
395  }
396  const t = await now($)
397  const hints = E.reconcileFiles(task, current, t)
398  // Re-scan with the current contract (intent or wording may have changed since).
399  for (const [path, text] of Object.entries(current)) {
400    if (text !== null && /\.(c|m)?(t|j)sx?$/i.test(path)) {
401      const asked = [task.contract.summary, task.contract.prompt.excerpt, ...task.contract.requirements.map(r => r.description)].join(' ')
402      hints.push(...E.applyFindings(task, path, 'mirage', E.scanSource(path, text, task.contract.intent, asked), t))
403    }
404  }
405  const violations = E.policyViolations(task)
406  E.log(task, 'evidence', `re-check: ${Object.keys(current).length} tracked file(s), ${violations.length} policy violation(s)`, t)
407  await changed($, hints)
408  const s = E.summarize(task, p.mode)
409  const lines = [
410    E.statusLine(s, 200) ?? 'Integrity: no status',
411    `Re-read ${Object.keys(current).length} tracked file(s); ${hints.length} new hint(s).`,
412    ...(violations.length ? ['', `**Policy (${p.mode}):**`, ...violations.map(v => `- ✗ ${v.text}`)] : ['Policies: none violated.']),
413  ]
414  return { text: lines.join('\n'), ...(p.mode === 'strict' && violations.length > 0 ? { exitCode: 1 } : {}) }
415}
416
417async function exportReport($: EngineInterface, format: 'md' | 'json' | 'both'): Promise<string> {
418  const task = await requireTask($)
419  if (typeof task === 'string') return task
420  const p = rt.project as E.ProjectRecord
421  const t = await now($)
422  const dir = `${rt.root.replace(/[\\/]$/, '')}/.claude/integrity`
423  const written: string[] = []
424  if (format !== 'json') {
425    await $.fs.write(`${dir}/${task.id}.md`, E.toMarkdown(task, p.mode, p.label, t))
426    written.push(`.claude/integrity/${task.id}.md`)
427  }
428  if (format !== 'md') {
429    await $.fs.write(`${dir}/${task.id}.json`, JSON.stringify(E.toJson(task, p.mode, p.label, t), null, 2))
430    written.push(`.claude/integrity/${task.id}.json`)
431  }
432  E.log(task, 'task', `report exported: ${written.join(', ')}`, t)
433  await changed($)
434  return `Exported ${written.join(' and ')}.`
435}
436
437async function suggest($: EngineInterface): Promise<string> {
438  const task = await requireTask($)
439  if (typeof task === 'string') return task
440  if (!rt.prefs.aiAssist) {
441    return 'AI suggestions are off. Turn on "AI requirement suggestions" for integrity in /config, then run `/integrity-suggest` again. Each run makes one model call (about 1-2k tokens) on this session\'s own model provider; nothing goes to any other service.'
442  }
443  const c = task.contract
444  const source = rt.lastPrompt !== undefined && E.hash(rt.lastPrompt) === c.prompt.hash ? rt.lastPrompt : c.prompt.excerpt
445  const cacheKey = E.hash(source)
446  let lines = rt.suggested.get(cacheKey)
447  let cost = 'cached, no model call'
448  if (lines === undefined) {
449    const r = await $.model.complete({
450      model: 'haiku',
451      system: 'You list testable software requirements. Output 3 to 7 lines, each starting with "- ". No commentary. Do not invent features the request does not imply.',
452      prompt: `Developer request (secrets redacted):\n${E.redact(source).slice(0, 4000)}\n\nAlready recorded:\n${c.requirements.map(r => `- ${r.description}`).join('\n') || '(none)'}\n\nList missing acceptance criteria.`,
453      maxTokens: 500,
454      effort: 'low',
455      timeoutMs: 30_000,
456    })
457    if (!r.isAnswered) return `No suggestions: the model call ended with ${r.reason}. Nothing was changed.`
458    lines = r.text.split('\n').map(l => l.replace(/^\s*[-*•]\s*/, '').trim()).filter(l => l.length > 8).slice(0, 7)
459    rt.suggested.set(cacheKey, lines)
460    cost = `${r.usage.input_tokens} input + ${r.usage.output_tokens} output tokens`
461  }
462  if (lines.length === 0) return 'The model proposed nothing new.'
463  const t = await now($)
464  E.amend(c, t, `${lines.length} proposed requirement(s) added (model suggestion)`, x => {
465    for (const l of lines!) x.requirements.push(E.makeRequirement(x.requirements, l, { origin: 'proposed', criticality: 'should' }))
466  })
467  E.log(task, 'contract', `${lines.length} model-proposed requirement(s) added; awaiting approval`, t)
468  await changed($)
469  return `Added ${lines.length} **proposed** requirement(s) (${cost}). They count only once you approve the contract; remove any with \`/integrity-contract remove R#\`.\n\n${lines.map(l => `- ${l}`).join('\n')}`
470}
471
472async function savePrefs($: EngineInterface, change: Partial<Prefs>): Promise<void> {
473  rt.storedPrefs = { ...rt.storedPrefs, ...change }
474  await $.store.set('prefs', rt.storedPrefs)
475  rt.prefs = prefsFrom(rt.options, rt.storedPrefs)
476  rt.lastStatus = '\u0000'
477  await changed($)
478}
479
480async function run($: EngineInterface, command: string, args: string, presentation: { isFullscreen: boolean }): Promise<{ text: string; exitCode?: number }> {
481  const p = await ensure($)
482  const t = await now($)
483  switch (command) {
484    case 'integrity': {
485      const [sub = '', ...rest] = words(args)
486      if (sub !== '' && COMMANDS.some(c => c.name === `integrity-${sub}`)) return run($, `integrity-${sub}`, rest.join(' '), presentation)
487      if (sub === 'status' || sub === 'claims' || sub === '' || ['overview', 'requirements', 'evidence', 'findings', 'activity', 'report'].includes(sub)) {
488        const tab = (['overview', 'requirements', 'evidence', 'findings', 'claims', 'activity', 'report'].includes(sub) ? sub : 'overview') as E.Tab
489        if (sub !== 'status') {
490          await update($, uiState, u => ({ ...u, tab, page: 0 }))
491          void openPane($)
492        }
493        const s = E.summarize(E.activeTask(p), p.mode)
494        if (!s.hasTask) return { text: 'Claude Integrity: no task contract yet. Describe what you want built, or `/integrity-start <task>`. `/integrity-help` explains.' }
495        const claims = s.claims.latest.map(c => `- ${E.GLYPH[c.verdict]} ${c.verdict}: "${c.text}" (${c.reason})`)
496        return { text: [E.statusLine(s, 200), s.concern ? `Top concern: ${s.concern.text}` : 'No outstanding concern.', ...(sub === 'claims' ? ['', ...claims] : [])].join('\n') }
497      }
498      return { text: `Unknown subcommand "${sub}". \`/integrity-help\` lists them.` }
499    }
500    case 'integrity-start': {
501      const text = args.trim() !== '' ? args : rt.lastPrompt ?? ''
502      if (text === '') return { text: 'Usage: `/integrity-start <what you want built>` (or run it right after describing the task).' }
503      const old = E.activeTask(p)
504      const task = E.startTask(p, text, args.trim() !== '' ? 'manual' : 'prompt', t)
505      if (old) E.log(task, 'task', `previous task ${old.id} kept in history`, t)
506      await changed($)
507      return { text: `Started task \`${task.id}\`.\n\n${contractText(task, p.mode)}` }
508    }
509    case 'integrity-contract':
510      return { text: await editContract($, args) }
511    case 'integrity-approve': {
512      const task = await requireTask($)
513      if (typeof task === 'string') return { text: task }
514      if (task.contract.approval.status === 'approved') return { text: `Already approved (v${task.contract.version}).` }
515      if (task.contract.requirements.length === 0) return { text: 'Nothing to approve: add at least one requirement (`/integrity-contract add ...`).' }
516      const first = task.contract.baseline === undefined
517      E.approve(task.contract, t)
518      E.log(task, 'approval', `contract ${first ? 'approved as baseline' : 're-approved'} at v${task.contract.version}`, t)
519      await changed($)
520      const proposed = task.contract.requirements.filter(r => r.origin === 'proposed').length
521      return { text: `Contract ${first ? 'approved: baseline frozen' : 're-approved'} at v${task.contract.version} (${task.contract.requirements.length} requirements${proposed ? `, ${proposed} model-proposed` : ''}).` }
522    }
523    case 'integrity-check':
524      return check($)
525    case 'integrity-findings':
526      return { text: await findingsText($, args) }
527    case 'integrity-evidence':
528      return { text: await evidenceText($, args) }
529    case 'integrity-verify': {
530      const task = await requireTask($)
531      if (typeof task === 'string') return { text: task }
532      const [id = '', ...note] = words(args)
533      const req = E.findRequirement(task.contract, id)
534      if (req === undefined || note.length === 0) return { text: 'Usage: `/integrity-verify R3 <what you checked, e.g. "submitted the form and reloaded: saved">`' }
535      const ev = E.recordManual(task, req.id, note.join(' '), t)
536      await changed($)
537      return { text: `${ev.id}: your confirmation of ${req.id} recorded (${E.requirementState(req, task).state}). It goes stale if related code changes later.` }
538    }
539    case 'integrity-report': {
540      const task = await requireTask($)
541      if (typeof task === 'string') return { text: task }
542      const w = words(args)
543      const format = w.includes('json') ? 'json' : w.includes('md') ? 'md' : 'both'
544      if (w.includes('export')) return { text: await exportReport($, format) }
545      return { text: format === 'json' ? '```json\n' + JSON.stringify(E.toJson(task, p.mode, p.label, t), null, 2).slice(0, 60_000) + '\n```' : E.toMarkdown(task, p.mode, p.label, t) }
546    }
547    case 'integrity-mode': {
548      const mode = words(args)[0]
549      if (mode === undefined) return { text: `Mode: **${p.mode}**. Options: observe (default, never blocks), review (summary after answers), strict (refuses edits to excluded paths; /integrity-check exits 1 on violations).` }
550      if (mode !== 'observe' && mode !== 'review' && mode !== 'strict') return { text: 'Usage: `/integrity-mode observe|review|strict`' }
551      p.mode = mode
552      const task = E.activeTask(p)
553      if (task) E.log(task, 'mode', `mode → ${mode}`, t)
554      await changed($)
555      const strictNote = mode === 'strict' && (task?.contract.baseline === undefined) ? ' Strict policies apply once a contract is approved.' : ''
556      return { text: `Mode set to **${mode}** for this project.${strictNote}` }
557    }
558    case 'integrity-waive': {
559      const task = await requireTask($)
560      if (typeof task === 'string') return { text: task }
561      const [id = '', ...reason] = words(args)
562      if (id === '' || reason.length === 0) return { text: 'Usage: `/integrity-waive <finding-id|R#> <reason>`. A reason is required.' }
563      const req = E.findRequirement(task.contract, id)
564      if (req !== undefined) {
565        E.amend(task.contract, t, `${req.id} waived: ${reason.join(' ')}`, () => { req.waiver = { reason: E.redact(reason.join(' ')), at: t } }, false)
566        E.log(task, 'waiver', `${req.id} waived: ${reason.join(' ')}`, t)
567        await changed($)
568        return { text: `${req.id} waived: ${reason.join(' ')}` }
569      }
570      const f = E.setFindingStatus(task, id, 'waived', t, reason.join(' '))
571      if (f === undefined) return { text: `No finding or requirement "${id}". \`/integrity-findings\` lists ids.` }
572      await changed($)
573      return { text: `Waived ${f.ruleId} at ${f.file}:${f.line}: ${reason.join(' ')}. The same finding stays waived on later scans.` }
574    }
575    case 'integrity-reset': {
576      const w = words(args)
577      const all = w.includes('all')
578      const task = E.activeTask(p)
579      if (!all && task === undefined) return { text: 'No active task to reset.' }
580      const what = all ? `all Integrity data for this project (${Object.keys(p.tasks).length} task(s))` : `task \`${task!.id}\` (${task!.contract.summary})`
581      let confirmed = w.includes('confirm')
582      if (!confirmed) {
583        const answer = await $.ui.ask(`Delete ${what}? This cannot be undone.`, ['Delete', 'Keep']).catch(() => undefined)
584        if (answer === undefined) return { text: `This deletes ${what}. Run \`/integrity-reset ${all ? 'all ' : ''}confirm\` to proceed.` }
585        confirmed = answer === 'Delete'
586      }
587      if (!confirmed) return { text: 'Kept. Nothing was deleted.' }
588      const ids = all ? Object.keys(p.tasks) : [task!.id]
589      for (const id of ids) { rt.deleted.add(id); delete p.tasks[id] }
590      delete p.activeTaskId
591      await changed($)
592      if (all) { await $.store.delete(rt.key); rt.project = E.newProject(rt.key, p.label, rt.session, t); rt.deleted.clear() }
593      return { text: `Deleted ${what}.` }
594    }
595    case 'integrity-ui': {
596      const w = words(args)
597      if (w.length === 0) return { text: `Status hint ${rt.prefs.statusHint ? 'on' : 'off'} · pane auto-open ${rt.prefs.paneAutoOpen ? 'on' : 'off'} · notify ${rt.prefs.notify} · color ${rt.prefs.color ? 'on' : 'off'} · max findings ${rt.prefs.maxFindings} · layout ${rt.prefs.layout}` }
598      const change: Partial<Prefs> = {}
599      for (let i = 0; i < w.length; i += 2) {
600        const [k, v] = [w[i], w[i + 1]]
601        if (k === 'status') change.statusHint = v === 'on'
602        else if (k === 'autoopen') change.paneAutoOpen = v === 'on'
603        else if (k === 'color') change.color = v === 'on'
604        else if (k === 'notify' && (v === 'off' || v === 'important' || v === 'all')) change.notify = v
605        else if (k === 'max' && Number(v) > 0) change.maxFindings = Number(v)
606        else if (k === 'pane' && v === 'open') void openPane($)
607        else if (k === 'pane' && v === 'close') await $.ui.close({ id: PANE }).catch(() => undefined)
608        else return { text: `Unknown setting "${k} ${v ?? ''}". \`/integrity-ui\` with no arguments shows the current ones.` }
609      }
610      if (Object.keys(change).length > 0) await savePrefs($, change)
611      return { text: `Saved: ${Object.entries(change).map(([k, v]) => `${k}=${v}`).join(', ') || 'pane toggled'}.` }
612    }
613    case 'integrity-layout': {
614      const l = words(args)[0]
615      if (l !== 'auto' && l !== 'mini' && l !== 'compact' && l !== 'wide') return { text: `Layout: ${rt.prefs.layout}. Usage: \`/integrity-layout auto|mini|compact|wide\`` }
616      await savePrefs($, { layout: l })
617      return { text: `Layout set to ${l}.` }
618    }
619    case 'integrity-suggest':
620      return { text: await suggest($) }
621    case 'integrity-help':
622      return { text: HELP }
623  }
624  return { text: `Unknown command ${command}` }
625}
626
627// ---- The pane -----------------------------------------------------------------
628
629function actions($: EngineInterface): PaneActions {
630  return {
631    setUi: fn => update($, uiState, fn),
632    approve: async () => (await run($, 'integrity-approve', '', { isFullscreen: false })).text,
633    recheck: async () => (await check($)).text.split('\n')[1] ?? 'Re-checked.',
634    exportReport: format => exportReport($, format),
635    setFinding: async (id, status, reason) => {
636      const task = E.activeTask(await ensure($))
637      if (task === undefined) return 'No active task.'
638      const f = E.setFindingStatus(task, id, status, await now($), reason)
639      await changed($)
640      return f ? `${f.ruleId} ${status}.` : 'Finding not found.'
641    },
642    waiveRequirement: async (id, reason) => (await run($, 'integrity-waive', `${id} ${reason}`, { isFullscreen: false })).text,
643    close: () => $.ui.close({ id: PANE }),
644  }
645}
646
647export const register: Register = (on, options) => {
648  rt = freshRuntime(options)
649
650  // ---- lifecycle -------------------------------------------------------------
651
652  on('session.start', async ($, e, next) => {
653    await Promise.all(COMMANDS.map(c => $.command.register(c)))
654    await Promise.all(SHORTCUTS.map(({ to: _, ...c }) => $.command.register(c)))
655    await ensure($)
656    await refreshStatus($)
657    if (rt.prefs.paneAutoOpen && e.isInteractive && E.activeTask(rt.project as E.ProjectRecord) !== undefined) {
658      void openPane($)
659    }
660    return next(e)
661  })
662
663  on('session.end', async ($, e, next) => {
664    if (rt.dirty || rt.writing !== undefined) await save($)
665    return next(e)
666  })
667
668  // ---- Intent Fingerprint: capture --------------------------------------------
669
670  on('prompt.submit', async ($, e, next) => {
671    const human = e.origin.kind === 'composer' || e.origin.kind === 'bridge' || e.origin.kind === 'sdk'
672    if (human && !e.text.trim().startsWith('/')) {
673      rt.lastPrompt = e.text
674      const p = await ensure($)
675      if (rt.prefs.autoCapture && E.activeTask(p) === undefined && E.isMeaningfulTask(e.text)) {
676        const task = E.startTask(p, e.text, 'prompt', await now($))
677        await changed($, rt.prefs.notify === 'all' ? [{ key: `draft:${task.id}`, level: 'info', text: `Integrity: drafted a contract (${task.contract.requirements.length} explicit requirement(s)) · /integrity-contract` }] : [])
678      }
679    }
680    return next(e)
681  })
682
683  // ---- The Witness: observe tool calls ----------------------------------------
684
685  on('tool.call', async ($, e, next) => {
686    const args = e as unknown as { tool: string } & Record<string, unknown>
687    const p = await ensure($)
688    const task = E.activeTask(p)
689    const target = filePath(args)
690    if (task !== undefined && target !== undefined) {
691      const glob = E.deniedPath(task, p.mode, rel(target))
692      if (glob !== undefined) {
693        E.log(task, 'finding', `strict: refused ${args.tool} of ${rel(target)} (excluded by "${glob}")`, await now($))
694        await changed($)
695        return { deny: `Claude Integrity (strict mode): ${rel(target)} is excluded by the approved contract ("${glob}"). Amend the contract (/integrity-contract) or change mode (/integrity-mode) to edit it.` }
696      }
697    }
698
699    const res = await next(e)
700    if (task === undefined) return res
701    try {
702      const t = await now($)
703      const ranInCore = res.ref !== undefined
704      const result = res.result as Record<string, unknown> | undefined
705      const hints: E.Hint[] = []
706
707      if (target !== undefined) {
708        // A refused, failed, staged or plugin-answered write changed nothing core can vouch for.
709        if (res.deny === undefined && res.isError !== true && ranInCore && result?.['staged'] !== true) {
710          const path = rel(target)
711          const after = args.tool === 'Write' && typeof args['content'] === 'string' ? args['content'] : await readText($, target)
712          const before = typeof result?.['originalFile'] === 'string' ? (result['originalFile'] as string) : result?.['originalFile'] === null ? null : undefined
713          const kind = args.tool === 'Write' && result?.['type'] === 'create' ? 'create' : 'update'
714          hints.push(...E.recordChange(task, { path, kind, ...(after === undefined ? {} : { after }), ...(before === undefined ? {} : { before }) }, t))
715        }
716      } else if (args.tool === 'Bash' || /^mcp__.*(playwright|chrome|browser|puppeteer)/i.test(args.tool)) {
717        // File changes a shell command made come first: a check in the same line ran after them.
718        const diff = result?.['bashEditDiff'] as { files?: { filePath: string; created?: true; deleted?: true }[] } | undefined
719        for (const f of (ranInCore ? diff?.files ?? [] : []).slice(0, 30)) {
720          const after = f.deleted ? undefined : await readText($, f.filePath)
721          hints.push(...E.recordChange(task, { path: rel(f.filePath), kind: f.deleted ? 'delete' : f.created ? 'create' : 'update', ...(after === undefined ? {} : { after }) }, t))
722        }
723        const obs: E.ToolObservation = {
724          tool: args.tool, input: args, ranInCore, isError: res.isError === true, text: res.text ?? '', result: res.result,
725          ...(res.deny === undefined ? {} : { denied: res.deny }),
726          ...(e.agentId === undefined ? {} : { agentId: e.agentId }),
727          cwd: await $.session.cwd().catch(() => rt.root),
728        }
729        const recorded = E.recordTool(task, obs, t)
730        if (recorded.evidence.length > 0) {
731          const head = ranInCore ? await gitHead($) : undefined
732          if (head !== undefined) for (const ev of recorded.evidence) ev.revision = head
733        }
734        hints.push(...recorded.hints)
735        if (recorded.evidence.length === 0 && hints.length === 0 && (diff?.files ?? []).length === 0) return res
736      } else return res
737      await changed($, hints)
738    } catch (err) {
739      $.ui.log(`integrity: could not record ${args.tool} (${String(err)})`, { to: 'debug' })
740    }
741    return res
742  })
743
744  // ---- Completion Claim Auditor -------------------------------------------------
745
746  on('turn.complete', async ($, e, next) => {
747    const res = await next(e)
748    if (e.agentId !== undefined) return res
749    const p = await ensure($)
750    const task = E.activeTask(p)
751    if (task === undefined) return res
752    const t = await now($)
753    if (e.reason === 'aborted' || e.isAborted) {
754      E.log(task, 'turn', 'turn interrupted: its claims were not audited', t)
755      await changed($)
756      return res
757    }
758    if (e.reason !== 'answer') return res
759    const { notice } = E.recordAudit(task, e.answer, e.turnId, t, p.mode, rt.prefs.notify)
760    await changed($)
761    if (p.mode === 'review' && notice !== undefined) void openPane($)
762    if (notice === undefined) return res
763    // Another plugin's line beneath the answer stays; ours follows it. The answer itself is never rewritten.
764    const theirs = res.text !== e.answer ? `${res.text}\n` : ''
765    return { ...res, text: `${theirs}${notice}` }
766  })
767
768  // ---- Commands and pane ---------------------------------------------------------
769
770  on('command.run', { command: [...COMMANDS, ...SHORTCUTS].map(c => c.name) }, async ($, e) => run($, ALIASES[e.command] ?? e.command, e.args, e.presentation))
771
772  on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
773    await read($, rev)
774    const ui = await read($, uiState)
775    const p = await ensure($)
776    const view = { surface: e.surface, bodyColumns: e.props.bodyColumns, columns: e.viewport?.columns, placement: e.props.placement }
777    return renderPane($.ui.resolve(e), view, { project: p, task: E.activeTask(p), prefs: rt.prefs, ui, actions: actions($) })
778  })
779}
780
src/engine/index.ts 283 lines
1// The engine's state transitions. Pure: every function takes the state and the
2// time, mutates the task in place, and returns the hints to show. The Mod
3// adapter does all IO; a CLI or CI job can drive the same functions.
4
5import type {
6  Activity, ActivityKind, Evidence, FileChange, FileState, Finding, Hint, Mode, ProjectRecord, Task, ToolObservation,
7} from './types'
8import { createContract } from './contract'
9import { evidenceFromBash, evidenceFromBrowserTool, freshness, isProof, staleBecause } from './evidence'
10import { apiCalls, routeMatches, routeOf, scanSource, type RawFinding } from './mirage'
11import { driftFindings } from './drift'
12import { auditAnswer } from './claims'
13import { requirementState } from './status'
14import { clip, hash, isScannable, newId, redact } from './util'
15
16export * from './types'
17export { summarize, statusLine, type Summary, type Tab } from './summary'
18export { requirementState, policyViolations, deniedPath, linkedEvidence } from './status'
19export { freshness, staleBecause, isProof, classifyCommand, parseCounts } from './evidence'
20export { createContract, approve, amend, makeRequirement, findRequirement, isMeaningfulTask, detectIntent } from './contract'
21export { scanSource, RULES, apiCalls, routeOf, routeMatches } from './mirage'
22export { driftFindings } from './drift'
23export { auditAnswer, extractClaims } from './claims'
24export { toMarkdown, toJson } from './report'
25export { mergeProjects, trimProject } from './persist'
26export { redact, normalizePath, hash, projectKey } from './util'
27
28const LIMITS = { activity: 200, notified: 300 }
29
30export function newProject(key: string, label: string, writer: string, now: number): ProjectRecord {
31  return { schema: 1, key, label: redact(label), rev: 0, writer, updatedAt: now, mode: 'observe', tasks: {} }
32}
33
34export const activeTask = (p: ProjectRecord): Task | undefined =>
35  p.activeTaskId === undefined ? undefined : p.tasks[p.activeTaskId]
36
37export function startTask(p: ProjectRecord, prompt: string, source: 'prompt' | 'manual', now: number): Task {
38  const id = newId('t', now)
39  const task: Task = {
40    id, createdAt: now, updatedAt: now, contract: createContract(id, prompt, now, source),
41    evidence: [], files: {}, findings: [], claims: [], activity: [], seq: 0, notified: [], nextEvidence: 1,
42  }
43  p.tasks[id] = task
44  p.activeTaskId = id
45  log(task, 'task', `task started (${source === 'prompt' ? 'captured from prompt' : 'manual'}): ${task.contract.summary}`, now)
46  return task
47}
48
49export function log(task: Task, kind: ActivityKind, text: string, now: number): void {
50  const entry: Activity = { at: now, kind, text: clip(redact(text), 200) }
51  task.activity.push(entry)
52  if (task.activity.length > LIMITS.activity) task.activity.splice(0, task.activity.length - LIMITS.activity)
53  task.updatedAt = now
54}
55
56/** A hint at most once per key, ever, for this task. */
57function hint(task: Task, key: string, level: Hint['level'], text: string): Hint[] {
58  if (task.notified.includes(key)) return []
59  task.notified.push(key)
60  if (task.notified.length > LIMITS.notified) task.notified.splice(0, task.notified.length - LIMITS.notified)
61  return [{ key, level, text }]
62}
63
64const nextEvidenceId = (task: Task) => (): string => `E${task.nextEvidence++}`
65
66function contractText(task: Task): string {
67  const c = task.contract
68  return [c.summary, c.prompt.excerpt, ...c.requirements.map(r => r.description), ...c.assumptions].join(' ')
69}
70
71/**
72 * Merges a fresh scan of one file into the findings: an id seen before keeps
73 * its waiver or acknowledgement; one no longer found is resolved.
74 */
75export function applyFindings(
76  task: Task, file: string, system: Finding['system'], fresh: (RawFinding & { id: string; requirementIds?: string[] })[], now: number,
77  replaceRules?: (ruleId: string) => boolean,
78): Hint[] {
79  const hints: Hint[] = []
80  const ids = new Set(fresh.map(f => f.id))
81  for (const old of task.findings) {
82    if (old.file !== file || old.system !== system || ids.has(old.id)) continue
83    if (replaceRules !== undefined && !replaceRules(old.ruleId)) continue
84    if (old.status === 'open' || old.status === 'acknowledged') {
85      old.status = 'resolved'
86      log(task, 'resolved', `${old.ruleId} no longer found in ${file}`, now)
87    }
88  }
89  for (const f of fresh) {
90    const existing = task.findings.find(x => x.id === f.id)
91    if (existing) {
92      Object.assign(existing, { line: f.line, snippet: redact(f.snippet), lastSeq: task.seq, severity: f.severity, confidence: f.confidence })
93      if (existing.status === 'resolved') existing.status = 'open'
94      continue
95    }
96    task.findings.push({
97      id: f.id, ruleId: f.ruleId, system, category: f.category, file, line: f.line, snippet: redact(f.snippet),
98      explanation: f.explanation, evidence: f.evidence, severity: f.severity, confidence: f.confidence,
99      requirementIds: f.requirementIds ?? [], verify: f.verify, status: 'open', firstSeq: task.seq, lastSeq: task.seq,
100    })
101    log(task, 'finding', `${f.ruleId} ${f.confidence} in ${file}:${f.line}`, now)
102    if (f.confidence === 'expected' || f.severity === 'low' || f.severity === 'info') continue
103    const text = system === 'drift'
104      ? `Integrity: possible scope drift in ${file} (${f.category.toLowerCase()}) · /integrity findings`
105      : `Integrity: ${f.confidence === 'confirmed' ? '' : 'possible '}${f.category.replace(/^\w\.\s*/, '').toLowerCase()} in ${file}:${f.line} · /integrity findings`
106    hints.push(...hint(task, `finding:${f.id}`, f.severity === 'high' ? 'important' : 'info', text))
107  }
108  if (task.findings.length > 300) {
109    const drop = task.findings.filter(f => f.status === 'resolved').slice(0, task.findings.length - 300)
110    task.findings = task.findings.filter(f => !drop.includes(f))
111  }
112  return hints
113}
114
115/** One observed source change: bumps the change sequence, then drift and mirage checks, then staleness. */
116export function recordChange(task: Task, change: FileChange, now: number): Hint[] {
117  const proofBefore = task.evidence.filter(e => isProof(e, task))
118  task.seq += 1
119  const prev = task.files[change.path]
120  task.files[change.path] = {
121    path: change.path,
122    created: prev?.created ?? change.kind === 'create',
123    deleted: change.kind === 'delete',
124    firstSeq: prev?.firstSeq ?? task.seq,
125    lastSeq: task.seq,
126    changes: (prev?.changes ?? 0) + 1,
127    ...(change.after === undefined ? {} : { hash: hash(change.after) }),
128  }
129  if (Object.keys(task.files).length > 500) {
130    const oldest = Object.values(task.files).sort((a, b) => a.lastSeq - b.lastSeq)[0]
131    if (oldest) delete task.files[oldest.path]
132  }
133  log(task, 'evidence', `${change.kind} ${change.path}`, now)
134
135  const hints: Hint[] = []
136  // Drift is history: a later edit does not undo a creation or deletion, so nothing auto-resolves.
137  hints.push(...applyFindings(task, change.path, 'drift', driftFindings(change, task.contract), now, () => false))
138  if (change.kind === 'delete') hints.push(...applyFindings(task, change.path, 'mirage', [], now))
139  else if (change.after !== undefined && isScannable(change.path)) {
140    const file = task.files[change.path] as FileState
141    file.calls = apiCalls(change.after)
142    const route = routeOf(change.path)
143    if (route !== undefined) file.route = route
144    hints.push(...applyFindings(task, change.path, 'mirage', scanSource(change.path, change.after, task.contract.intent, contractText(task)), now, id => id !== 'MIR-D2'))
145  }
146  hints.push(...endpointCheck(task, now))
147
148  const invalidated = proofBefore.filter(e => freshness(e, task) === 'stale')
149  for (const e of invalidated) log(task, 'invalidated', `${e.id} ${e.check} evidence stale after ${change.path} changed`, now)
150  if (invalidated.length > 0) {
151    const checks = [...new Set(invalidated.map(e => e.check))].join('/')
152    hints.push(...hint(task, `stale:${invalidated.map(e => e.id).join(',')}`, 'important', `Integrity: ${checks} evidence is stale after ${change.path} changed · /integrity evidence`))
153  }
154  task.updatedAt = now
155  return hints
156}
157
158/**
159 * MIR-D2, across files: a client calls an `/api/...` endpoint that no observed
160 * route file serves. Runs only once the task has seen at least one route file.
161 */
162function endpointCheck(task: Task, now: number): Hint[] {
163  const files = Object.values(task.files).filter(f => !f.deleted)
164  const routes = files.map(f => f.route).filter((r): r is string => r !== undefined)
165  if (routes.length === 0) return []
166  const hints: Hint[] = []
167  for (const f of files) {
168    if (f.calls === undefined) continue
169    const missing = f.calls.filter(c => !routes.some(r => routeMatches(r, c.path)))
170    hints.push(...applyFindings(task, f.path, 'mirage', missing.map(c => ({
171      id: hash(`MIR-D2|${f.path}|${c.path}`), ruleId: 'MIR-D2', category: 'D. Incomplete data integration', line: c.line,
172      snippet: `request to ${c.path}`, severity: 'high' as const, confidence: 'potential' as const,
173      explanation: `${f.path} calls ${c.path}, but no route file seen in this task serves it (seen: ${routes.slice(0, 4).join(', ')}).`,
174      evidence: 'Literal endpoint compared with the route files observed in this task; routes outside the task are not seen.',
175      verify: `Confirm ${c.path} exists, or point the client at the route that does.`,
176    })), now, id => id === 'MIR-D2'))
177  }
178  return hints
179}
180
181/** One observed tool call that may be evidence: a check run, a browser session. */
182export function recordTool(task: Task, obs: ToolObservation, now: number): { evidence: Evidence[]; hints: Hint[] } {
183  const verifiedBefore = new Set(task.contract.requirements.filter(r => requirementState(r, task).state === 'verified').map(r => r.id))
184  let evidence: Evidence[] = []
185  if (obs.tool === 'Bash') evidence = evidenceFromBash(obs, task.seq, now, nextEvidenceId(task))
186  else {
187    const browser = evidenceFromBrowserTool(obs, task.seq, now, `E${task.nextEvidence}`)
188    if (browser) {
189      task.nextEvidence++
190      evidence = [browser]
191    }
192  }
193  if (evidence.length === 0) return { evidence, hints: [] }
194
195  for (const e of evidence) {
196    e.requirementIds = task.contract.requirements
197      .filter(r => r.checks.length > 0 && e.command !== undefined && r.checks.some(c => e.command!.toLowerCase().includes(c.toLowerCase())))
198      .map(r => r.id)
199    task.evidence.push(e)
200    log(task, 'evidence', `${e.id} ${e.summary} [${e.basis}]`, now)
201  }
202  if (task.evidence.length > 200) task.evidence.splice(0, task.evidence.length - 200)
203
204  const hints: Hint[] = []
205  const failed = evidence.find(e => e.outcome === 'fail')
206  if (failed) hints.push(...hint(task, `fail:${failed.id}`, 'important', `Integrity: ${failed.check} failed (${failed.basis}) · /integrity evidence`))
207  const newlyVerified = task.contract.requirements.filter(r => !verifiedBefore.has(r.id) && requirementState(r, task).state === 'verified')
208  if (newlyVerified.length > 0) {
209    hints.push(...hint(task, `verified:${newlyVerified.map(r => r.id).join(',')}:${task.seq}`, 'info', `Integrity: ${newlyVerified.map(r => r.id).join(', ')} verified by fresh evidence`))
210  }
211  task.updatedAt = now
212  return { evidence, hints }
213}
214
215/** A person's own confirmation: the only way a manual requirement is verified. */
216export function recordManual(task: Task, requirementId: string, note: string, now: number): Evidence {
217  const e: Evidence = {
218    id: `E${task.nextEvidence++}`, requirementIds: [requirementId], kind: 'manual-confirmation', check: 'manual', outcome: 'pass',
219    timestamp: now, source: 'person', summary: `confirmed by the developer: ${clip(redact(note), 120)}`, basis: 'explicit user verification',
220    scope: [], seq: task.seq,
221  }
222  task.evidence.push(e)
223  log(task, 'evidence', `${e.id} manual confirmation of ${requirementId}`, now)
224  return e
225}
226
227export type AuditResult = { notice?: string; hints: Hint[] }
228
229/** Audits a final answer and composes the notice shown beneath it, or none per the notification level. */
230export function recordAudit(task: Task, answer: string, turnId: string, now: number, mode: Mode, level: 'off' | 'important' | 'all'): AuditResult {
231  const claims = auditAnswer(answer, task, turnId, now)
232  const known = new Set(task.claims.map(c => c.id))
233  task.claims.push(...claims.filter(c => !known.has(c.id)))
234  if (task.claims.length > 100) task.claims.splice(0, task.claims.length - 100)
235  if (claims.length > 0) log(task, 'claims', `${claims.length} claim(s) audited: ${claims.map(c => c.verdict).join(', ')}`, now)
236
237  const bad = claims.filter(c => c.verdict === 'contradicted' || c.verdict === 'unsupported')
238  const lines: string[] = []
239  if (level !== 'off' && (bad.length > 0 || (level === 'all' && claims.length > 0))) {
240    const counts = (['supported', 'partial', 'unsupported', 'contradicted', 'not-assessable'] as const)
241      .map(v => [v, claims.filter(c => c.verdict === v).length] as const).filter(([, n]) => n > 0).map(([v, n]) => `${n} ${v}`)
242    lines.push(`Integrity · ${claims.length} completion claim${claims.length === 1 ? '' : 's'} checked: ${counts.join(' · ')}`)
243    for (const c of (level === 'all' ? claims : bad).slice(0, 4)) lines.push(`  ${GLYPH[c.verdict]} "${clip(c.text, 70)}": ${c.reason}`)
244  }
245  if (mode !== 'observe' && level !== 'off') {
246    const high = task.findings.filter(f => f.status === 'open' && f.severity === 'high' && f.confidence !== 'expected')
247    const mustOpen = task.contract.requirements.filter(r => r.criticality === 'must' && r.waiver === undefined && requirementState(r, task).state !== 'verified')
248    if (high.length > 0 || mustOpen.length > 0) {
249      lines.push(`Integrity ${mode}: ${mustOpen.length} must requirement(s) not verified · ${high.length} high finding(s) open`)
250    }
251  }
252  if (lines.length > 0) lines.push('  details: /integrity claims')
253  return { ...(lines.length > 0 ? { notice: lines.join('\n') } : {}), hints: [] }
254}
255
256export const GLYPH = { supported: '✓', partial: '◐', unsupported: '?', contradicted: '✗', 'not-assessable': '·' } as const
257
258export function setFindingStatus(task: Task, id: string, status: 'acknowledged' | 'waived' | 'open', now: number, reason?: string): Finding | undefined {
259  const f = task.findings.find(x => x.id === id || x.id.startsWith(id))
260  if (f === undefined) return undefined
261  f.status = status
262  if (status === 'waived') f.waiver = { reason: clip(redact(reason ?? ''), 200), at: now }
263  else delete f.waiver
264  log(task, status === 'waived' ? 'waiver' : 'finding', `${f.ruleId} ${f.file}:${f.line} ${status}${reason ? `: ${reason}` : ''}`, now)
265  return f
266}
267
268/** Re-reads tracked files (passed in by the adapter) and records out-of-band edits as changes. */
269export function reconcileFiles(task: Task, current: Record<string, string | null>, now: number): Hint[] {
270  const hints: Hint[] = []
271  for (const [path, text] of Object.entries(current)) {
272    const known = task.files[path]
273    if (known === undefined) continue
274    if (text === null && !known.deleted) hints.push(...recordChange(task, { path, kind: 'delete' }, now))
275    else if (text !== null && known.hash !== undefined && hash(text) !== known.hash) {
276      hints.push(...recordChange(task, { path, kind: 'update', after: text }, now))
277      log(task, 'invalidated', `${path} changed outside observed tools`, now)
278    }
279  }
280  return hints
281}
282
283
hooks/pane.tsx 565 lines
1// The /integrity pane, drawn the way Claude Flightdeck draws its own: a centered
2// title over a color legend, rounded cards per subsystem (title left, count
3// right), a strip of requirement states, and an activity log. Everything comes
4// from the engine's canonical summary, so its counts always match
5// /integrity-report. Keys 1-7 switch views while the pane holds the keyboard;
6// every action is also a command for headless use.
7
8import type { ElementTable, RenderElement, ThemeKey } from 'claude-code'
9
10import type { IntegrityTab, IntegrityUi } from '../types'
11import * as E from '../src/engine/index'
12
13export type PaneActions = {
14  setUi: (fn: (u: IntegrityUi) => IntegrityUi) => Promise<unknown>
15  approve: () => Promise<string>
16  recheck: () => Promise<string>
17  exportReport: (format: 'md' | 'json' | 'both') => Promise<string>
18  setFinding: (id: string, status: 'acknowledged' | 'waived' | 'open', reason?: string) => Promise<string>
19  waiveRequirement: (id: string, reason: string) => Promise<string>
20  close: () => Promise<void>
21}
22
23type Prefs = { layout: E.Layout; maxFindings: number; color: boolean }
24
25export type PaneContext = {
26  project: E.ProjectRecord
27  task: E.Task | undefined
28  prefs: Prefs
29  ui: IntegrityUi
30  actions: PaneActions
31}
32
33const TABS: readonly [IntegrityTab, string, string][] = [
34  ['overview', 'Overview', 'Ovw'], ['requirements', 'Requirements', 'Reqs'], ['evidence', 'Evidence', 'Evid'],
35  ['findings', 'Findings', 'Find'], ['claims', 'Claims', 'Clms'], ['activity', 'Activity', 'Log'], ['report', 'Report', 'Rpt'],
36]
37
38// Theme keys, like Flightdeck's default palette: they follow light, dark and color-blind themes.
39const C = {
40  main: 'claude', ok: 'success', evidence: 'suggestion', warn: 'warning', bad: 'error', claims: 'merged', dim: 'inactive', faint: 'subtle',
41} as const satisfies Record<string, ThemeKey>
42
43type ReqState = Exclude<keyof E.Summary['requirements'], 'total'>
44const REQ_GLYPH: Record<ReqState, string> = { verified: '✓', stale: '⧗', failed: '✗', waived: '–', implemented: '◐', unverified: '?' }
45const REQ_COLOR: Record<ReqState, ThemeKey> = { verified: C.ok, implemented: C.evidence, unverified: C.dim, stale: C.warn, failed: C.bad, waived: C.faint }
46const REQ_LABEL: Record<ReqState, string> = {
47  verified: 'Verified', stale: 'Stale', failed: 'Failed', waived: 'Waived', implemented: 'Implemented, not verified', unverified: 'Not verified',
48}
49const LEGEND: readonly ReqState[] = ['verified', 'implemented', 'unverified', 'stale', 'failed']
50const SEV_GLYPH: Record<E.Severity, string> = { high: '‼', medium: '!', low: '·', info: '○' }
51const KIND_COLOR: Record<E.Activity['kind'], ThemeKey> = {
52  task: C.main, contract: C.main, approval: C.main, requirement: C.main, mode: C.main, evidence: C.evidence,
53  invalidated: C.warn, finding: C.warn, resolved: C.ok, waiver: C.dim, claims: C.claims, turn: 'text',
54}
55
56const clip = (text: string, max: number): string => (text.length <= max ? text : `${text.slice(0, Math.max(1, max - 1))}…`)
57const time = (at: number): string => new Date(at).toTimeString().slice(0, 5)
58const plural = (n: number, word: string): string => `${n} ${word}${n === 1 ? '' : 's'}`
59/** Flightdeck's gauge: ▰ filled, ▱ empty. */
60const gauge = (ratio: number, cells: number): { on: string; off: string } => {
61  const n = Math.max(0, Math.min(cells, Math.round(ratio * cells)))
62  return { on: '▰'.repeat(n), off: '▱'.repeat(cells - n) }
63}
64
65/** Inline above the prompt the pane is a short summary, as Flightdeck's mini; docked it follows the width. */
66export function layoutFor(columns: number, pref: E.Layout, placement: 'dock' | 'inline' = 'dock'): Exclude<E.Layout, 'auto'> {
67  if (pref !== 'auto') return pref
68  if (placement === 'inline') return 'mini'
69  return columns < 56 ? 'mini' : columns < 110 ? 'compact' : 'wide'
70}
71
72/** What the pane is drawn into: the Pane site's body width and placement, and the viewport's width when measured. */
73export type PaneView = {
74  surface: 'terminal' | 'desktop' | 'mobile' | 'vscode'
75  bodyColumns: number
76  columns?: number | undefined
77  placement?: 'dock' | 'inline' | undefined
78}
79
80export function renderPane(els: ElementTable, view: PaneView, ctx: PaneContext): RenderElement {
81  const { Box, Text, Button } = els
82  // The mobile app draws no field yet: waivers there go through the command.
83  const Input = view.surface !== 'mobile' && 'Input' in els ? els.Input : undefined
84  const { task, prefs, ui, actions, project } = ctx
85  const W = Math.max(24, view.bodyColumns || view.columns || 80)
86  const layout = layoutFor(W, prefs.layout, view.placement)
87  const isMini = layout === 'mini'
88  // Color supplements the glyphs and words, never replaces them, so it can be switched off.
89  const fg = (c: ThemeKey): { color?: ThemeKey } => (prefs.color ? { color: c } : {})
90  const faint = prefs.color ? { color: C.faint } : { dimColor: true }
91  const s = E.summarize(task, project.mode)
92
93  const go = (tab: IntegrityTab, selected?: string) => (): void => {
94    void actions.setUi(u => ({ tab, page: 0, ...(selected === undefined ? {} : { selected: u.selected === selected && u.tab === tab ? undefined : selected }) }))
95  }
96  const flash = (p: Promise<string>): void => {
97    void p.then(msg => actions.setUi(u => ({ ...u, flash: msg, waiving: undefined })))
98  }
99
100  /** A card: a rounded border in the section's color, its title left and a count right. Mini drops the border. */
101  const card = (key: string, color: ThemeKey, title: string, right: string, w: number, body: (RenderElement | null | false)[], double = false): RenderElement =>
102    isMini ? (
103      <Box key={key} flexDirection="column" width={w}>
104        <Text key={`${key}:title`} wrap="truncate-end">
105          <Text bold {...fg(color)}>{title}</Text>
106          {right !== '' ? <Text dimColor>{`  ${right}`}</Text> : null}
107        </Text>
108        {body}
109      </Box>
110    ) : (
111      <Box key={key} flexDirection="column" borderStyle={double ? 'double' : 'round'} {...(prefs.color ? { borderColor: color } : {})} paddingX={1} width={w}>
112        <Box key={`${key}:title`} justifyContent="space-between">
113          <Text bold wrap="truncate-end" {...fg(color)}>{title}</Text>
114          {right !== '' ? <Text dimColor>{right}</Text> : null}
115        </Box>
116        {body}
117      </Box>
118    )
119  /** Inside a card's border and padding. */
120  const inner = (w: number): number => (isMini ? w : w - 4)
121  const rail = (key: string, w: number): RenderElement => <Text key={key} {...faint}>{'─'.repeat(Math.max(1, w))}</Text>
122
123  const tabBar = (
124    <Box key="tabs" flexDirection="row" flexWrap="wrap" columnGap={1} {...(isMini ? {} : { justifyContent: 'center' as const })}>
125      {TABS.map(([id, long, short], i) => (
126        <Button key={`tab:${id}`} plain hotkey={String(i + 1)} label={layout === 'wide' ? long : short}
127          variant={ui.tab === id ? 'primary' : 'secondary'} dimColor={ui.tab !== id} onPress={go(id)} />
128      ))}
129    </Box>
130  )
131
132  if (!s.hasTask || task === undefined) {
133    return (
134      <Box flexDirection="column" width={W}>
135        {!isMini && (
136          <Box key="head" justifyContent="center">
137            <Text bold wrap="truncate-end">INTEGRITY<Text {...faint}> · </Text><Text dimColor>NO CONTRACT</Text></Text>
138          </Box>
139        )}
140        {card('c-empty', C.faint, 'CONTRACT · none yet', '', W, [
141          <Text key="none" dimColor wrap="wrap">◇ No task contract. Nothing is tracked as verified yet.</Text>,
142          <Text key="how" wrap="wrap">Describe what you want built (an implementation prompt drafts a contract, no model call), or run /integrity-start &lt;task&gt;.</Text>,
143          <Text key="help" dimColor wrap="wrap">/integrity-help explains the workflow · /ig is short for /integrity</Text>,
144        ])}
145      </Box>
146    )
147  }
148
149  const inForce = s.requirements.total - s.requirements.waived
150  const approval = s.approval === 'approved' ? `approved v${s.version}` : s.approval === 'amended' ? 'amended · re-approve' : `draft v${s.version}`
151
152  // ---- header: Flightdeck's centered title and legend ------------------------------------
153
154  const legend = (() => {
155    const out: ReqState[] = []
156    let used = 0
157    for (const st of LEGEND) {
158      const cell = st.length + 4
159      if (used + cell > W) break
160      out.push(st)
161      used += cell
162    }
163    return out
164  })()
165  const header = isMini ? null : (
166    <Box key="header" flexDirection="column">
167      <Box key="head" justifyContent="center">
168        <Text bold wrap="truncate-end">
169          <Text>INTEGRITY</Text>
170          <Text {...faint}> · </Text>
171          <Text {...fg(C.main)}>{s.mode.toUpperCase()}</Text>
172          <Text> MODE</Text>
173          <Text {...faint}> · </Text>
174          <Text {...fg(s.approval === 'approved' ? C.ok : C.warn)}>{approval.toUpperCase()}</Text>
175        </Text>
176      </Box>
177      <Box key="legend" justifyContent="center" columnGap={2}>
178        {legend.map(st => (
179          <Text key={`lg:${st}`}>
180            <Text {...fg(REQ_COLOR[st])}>{prefs.color ? '■' : REQ_GLYPH[st]}</Text>
181            <Text dimColor>{` ${st}`}</Text>
182          </Text>
183        ))}
184      </Box>
185    </Box>
186  )
187  const footer = ui.flash ? <Text key="flash" dimColor wrap="truncate-end">{ui.flash}</Text> : null
188
189  // ---- overview pieces --------------------------------------------------------------------
190
191  /** One cell per requirement, colored by state (its glyph when color is off), like Flightdeck's gate strip. */
192  const strip = (w: number): RenderElement => {
193    const reqs = task.contract.requirements
194    const shown = reqs.slice(0, Math.max(4, w - 10))
195    return (
196      <Text key="strip" wrap="truncate-end">
197        <Text dimColor>reqs </Text>
198        {shown.map(r => {
199          const st = E.requirementState(r, task).state
200          return <Text key={`st:${r.id}`} {...fg(REQ_COLOR[st])}>{prefs.color ? '■' : REQ_GLYPH[st]}</Text>
201        })}
202        {reqs.length > shown.length ? <Text dimColor>{` +${reqs.length - shown.length}`}</Text> : null}
203        {reqs.length === 0 ? <Text dimColor>none yet</Text> : null}
204      </Text>
205    )
206  }
207  const progress = (w: number): RenderElement => {
208    const g = gauge(s.progress, Math.max(5, Math.min(14, w - 34)))
209    return (
210      <Text key="progress" wrap="truncate-end">
211        <Text {...fg(C.ok)}>{g.on}</Text>
212        <Text {...faint}>{g.off}</Text>
213        <Text bold>{` ${s.requirements.verified}/${inForce} verified`}</Text>
214        <Text dimColor>{` · must ${s.must.verified}/${s.must.total}`}</Text>
215      </Text>
216    )
217  }
218  const notVerified = s.requirements.unverified + s.requirements.implemented
219  const badClaims = s.claims.contradicted + s.claims.unsupported
220  const parts: [string, ThemeKey][] = [
221    ...(notVerified > 0 ? [[`? ${notVerified} not verified`, C.dim] as [string, ThemeKey]] : []),
222    ...(s.requirements.stale > 0 ? [[`⧗ ${s.requirements.stale} stale`, C.warn] as [string, ThemeKey]] : []),
223    ...(s.requirements.failed > 0 ? [[`✗ ${s.requirements.failed} failed`, C.bad] as [string, ThemeKey]] : []),
224    ...(s.findings.open > 0 ? [[`! ${plural(s.findings.open, 'finding')}${s.findings.high ? ` (${s.findings.high} high)` : ''}`, C.warn] as [string, ThemeKey]] : []),
225    ...(badClaims > 0 ? [[`✗ ${badClaims} claim(s)`, C.bad] as [string, ThemeKey]] : []),
226  ]
227  const counts = (
228    <Text key="counts" wrap="truncate-end">
229      {parts.length === 0 ? <Text {...fg(C.ok)}>no open problems recorded</Text> : parts.map(([text, color], i) => (
230        <Text key={`n${i}`}>
231          {i > 0 ? <Text {...faint}> · </Text> : null}
232          <Text {...fg(color)}>{text}</Text>
233        </Text>
234      ))}
235    </Text>
236  )
237
238  const contractCard = (w: number): RenderElement =>
239    card('c-contract', C.main, `CONTRACT · ${approval}`, `intent ${s.intent}`, w, [
240      <Text key="title" wrap="truncate-end">◆ {s.title}</Text>,
241      strip(inner(w)),
242      progress(inner(w)),
243      counts,
244    ])
245
246  const evidenceCard = (w: number): RenderElement =>
247    card('c-evidence', C.evidence, 'EVIDENCE · witness', plural(s.evidence.total, 'check'), w,
248      s.evidence.latest.length === 0
249        ? [<Text key="none" dimColor wrap="wrap">Not observed: no test, build, lint or typecheck run yet.</Text>]
250        : s.evidence.latest.map(ev => {
251          const stale = ev.freshness === 'stale'
252          const color = ev.outcome === 'fail' ? C.bad : ev.outcome !== 'pass' ? C.dim : stale ? C.warn : C.ok
253          return (
254            <Text key={`latest:${ev.check}`} wrap="truncate-end">
255              <Text {...fg(color)}>{ev.outcome === 'pass' ? '✓' : ev.outcome === 'fail' ? '✗' : '·'}</Text>
256              {` ${ev.check} ${ev.outcome} · `}
257              <Text {...(stale ? fg(C.warn) : {})}>{stale ? '⧗ stale' : ev.freshness}</Text>
258              <Text dimColor>{` · ${time(ev.at)} ${ev.id}`}</Text>
259            </Text>
260          )
261        }))
262
263  const order: Record<E.Severity, number> = { high: 0, medium: 1, low: 2, info: 3 }
264  const sorted = task.findings.filter(f => f.status !== 'resolved')
265    .sort((a, b) => (a.status === 'open' ? 0 : 1) - (b.status === 'open' ? 0 : 1) || order[a.severity] - order[b.severity])
266  const findingLabel = (f: E.Finding): string =>
267    `${f.confidence === 'expected' ? '○' : SEV_GLYPH[f.severity]} ${f.ruleId} ${f.file}:${f.line} ${f.category.replace(/^\w\.\s*/, '')}${f.status !== 'open' ? ` (${f.status})` : f.confidence === 'confirmed' ? ' [confirmed]' : ''}`
268
269  const findingsCard = (w: number): RenderElement => {
270    const top = sorted.filter(f => f.status === 'open' && f.confidence !== 'expected').slice(0, 3)
271    return card('c-findings', C.warn, 'FINDINGS · mirage + drift', `${s.findings.open} open`, w, [
272      <Text key="sev" wrap="truncate-end">
273        <Text {...fg(C.bad)}>‼</Text>{` ${s.findings.high} high · `}
274        <Text {...fg(C.warn)}>!</Text>{` ${s.findings.medium} medium · · ${s.findings.low + s.findings.info} low · ○ ${s.findings.expected} expected · ${s.findings.waived} waived`}
275      </Text>,
276      ...(top.length === 0
277        ? [<Text key="none" dimColor wrap="truncate-end">none open · checked on every observed edit</Text>]
278        : top.map(f => <Button key={`ovf:${f.id}`} plain label={clip(findingLabel(f), inner(w))} onPress={go('findings', f.id)} />)),
279    ])
280  }
281
282  const claimsCard = (w: number): RenderElement => {
283    const worst = s.claims.latest.filter(c => c.verdict === 'contradicted' || c.verdict === 'unsupported').slice(0, 2)
284    return card('c-claims', C.claims, 'CLAIMS · last answer', `${s.claims.total} checked`, w, [
285      s.claims.total === 0
286        ? <Text key="none" dimColor wrap="wrap">Not assessed: no completion claims audited yet.</Text>
287        : (
288          <Text key="verdicts" wrap="truncate-end">
289            <Text {...fg(C.ok)}>✓</Text>{` ${s.claims.supported} supported · ◐ ${s.claims.partial} partial · `}
290            <Text {...fg(C.warn)}>?</Text>{` ${s.claims.unsupported} unsupported · `}
291            <Text {...fg(C.bad)}>✗</Text>{` ${s.claims.contradicted} contradicted · · ${s.claims.notAssessable} n/a`}
292          </Text>
293        ),
294      ...worst.map(c => (
295        <Text key={`wc:${c.id}`} wrap="truncate-end">
296          <Text {...fg(c.verdict === 'contradicted' ? C.bad : C.warn)}>{E.GLYPH[c.verdict]}</Text>
297          {` "${clip(c.text, 40)}": `}
298          <Text dimColor>{c.reason}</Text>
299        </Text>
300      )),
301    ])
302  }
303
304  // Flightdeck's receipt: one line in its own card, here the top concern.
305  const receipt = (w: number): RenderElement => {
306    const line = s.concern
307      ? [
308        <Text key="concern" wrap="truncate-end" {...fg(C.warn)}>▸ {clip(s.concern.text, Math.max(10, inner(w) - 12))} </Text>,
309        <Button key="concern-go" plain label="inspect" hotkey="i" onPress={go(s.concern.tab)} />,
310      ]
311      : [<Text key="concern" {...fg(C.ok)}>✓ No outstanding concern.</Text>]
312    return isMini
313      ? <Box key="receipt" flexDirection="row">{line}</Box>
314      : <Box key="receipt" flexDirection="row" borderStyle="round" {...(prefs.color ? { borderColor: s.concern ? C.warn : C.ok } : {})} borderDimColor={s.concern === undefined} paddingX={1} width={w}>{line}</Box>
315  }
316
317  const actionRow = (
318    <Box key="actions" flexDirection="row" columnGap={2} {...(isMini ? {} : { paddingX: 1 })}>
319      {s.approval !== 'approved' && <Button key="approve" plain hotkey="a" label="approve contract" onPress={() => flash(actions.approve())} />}
320      <Button key="recheck" plain hotkey="c" label="re-check" onPress={() => flash(actions.recheck())} />
321      {ui.tab === 'overview' && !isMini && <Button key="ov-report" plain hotkey="r" label="report" dimColor onPress={go('report')} />}
322    </Box>
323  )
324
325  /** Flightdeck's session log: dim time, colored kind, the text. */
326  const log = (key: string, title: string, w: number, rows: number): RenderElement => {
327    const lines = task.activity.slice(-rows).reverse()
328    const body = lines.length === 0 ? [<Text key="none" {...faint}>nothing yet</Text>] : lines.map((a, i) => (
329      <Box key={`a${i}`} flexDirection="row">
330        <Box key="t" width={6} flexShrink={0}><Text {...faint}>{time(a.at)}</Text></Box>
331        <Box key="k" width={12} flexShrink={0}><Text bold wrap="truncate-end" {...fg(KIND_COLOR[a.kind])}>{a.kind}</Text></Box>
332        <Text key="x" wrap="truncate-end">{a.text}</Text>
333      </Box>
334    ))
335    return isMini
336      ? <Box key={key} flexDirection="column">{body}</Box>
337      : (
338        <Box key={key} flexDirection="column" borderStyle="round" {...(prefs.color ? { borderColor: C.faint } : {})} paddingX={1} width={w}>
339          <Text key="title" dimColor>{title}</Text>
340          {body}
341        </Box>
342      )
343  }
344
345  const overview = (): RenderElement => {
346    if (isMini) {
347      const g = gauge(s.progress, 6)
348      return (
349        <Box flexDirection="column" width={W}>
350          <Text key="mini1" wrap="truncate-end">
351            <Text bold {...fg(C.main)}>Integrity </Text>
352            <Text {...fg(C.ok)}>{g.on}</Text>
353            <Text {...faint}>{g.off}</Text>
354            <Text bold>{` ${s.requirements.verified}/${inForce} verified`}</Text>
355            <Text dimColor>{` · ${approval} · ${s.mode}`}</Text>
356          </Text>
357          {counts}
358          {receipt(W)}
359          {actionRow}
360        </Box>
361      )
362    }
363    if (layout === 'wide') {
364      const colW = Math.floor((W - 2) / 2)
365      return (
366        <Box flexDirection="column" width={W}>
367          <Box key="cols" flexDirection="row" columnGap={2}>
368            <Box key="left" flexDirection="column" width={colW}>
369              {contractCard(colW)}
370              {rail('rail-l', colW)}
371              {evidenceCard(colW)}
372            </Box>
373            <Box key="right" flexDirection="column" width={colW}>
374              {findingsCard(colW)}
375              {rail('rail-r', colW)}
376              {claimsCard(colW)}
377            </Box>
378          </Box>
379          {receipt(W)}
380          {actionRow}
381          {log('c-log', 'activity log', W, 6)}
382        </Box>
383      )
384    }
385    return (
386      <Box flexDirection="column" width={W}>
387        {contractCard(W)}
388        {rail('rail-1', W)}
389        {evidenceCard(W)}
390        {rail('rail-2', W)}
391        {findingsCard(W)}
392        {rail('rail-3', W)}
393        {claimsCard(W)}
394        {receipt(W)}
395        {actionRow}
396        {log('c-log', 'activity log', W, 4)}
397      </Box>
398    )
399  }
400
401  // ---- list views with a drill-down -------------------------------------------------------
402
403  const hasDetail = layout === 'wide' && ui.selected !== undefined
404  const listW = hasDetail ? W - Math.floor(W / 2) - 2 : W
405  const rowWidth = inner(listW)
406
407  const detail = (color: ThemeKey, lines: (string | RenderElement)[]): RenderElement => (
408    <Box key="detail" flexDirection="column" borderStyle="double" {...(prefs.color ? { borderColor: color } : {})} paddingX={1}
409      {...(layout === 'wide' ? { width: Math.floor(W / 2) } : { marginTop: 1 })}>
410      {lines.map((l, i) => (typeof l === 'string' ? <Text key={`d${i}`} wrap="wrap">{l}</Text> : l))}
411      <Button key="detail-close" plain hotkey="x" label="close detail" dimColor onPress={() => void actions.setUi(u => ({ ...u, selected: undefined, waiving: undefined }))} />
412    </Box>
413  )
414
415  const waiveField = (id: string, kind: 'finding' | 'requirement'): RenderElement => (
416    ui.waiving === id
417      ? Input !== undefined
418        ? <Input key="waive-reason" label="Waiver reason" placeholder="why this is acceptable" autoFocus
419          onSubmit={reason => { if (reason.trim() !== '') flash(kind === 'finding' ? actions.setFinding(id, 'waived', reason) : actions.waiveRequirement(id, reason)) }} />
420        : <Text key="waive-reason" dimColor>Type /integrity-waive {id} &lt;reason&gt; to waive.</Text>
421      : <Button key="waive" plain hotkey="w" label="waive (with reason)" onPress={() => void actions.setUi(u => ({ ...u, waiving: id }))} />
422  )
423
424  const split = (list: RenderElement, side: RenderElement | null): RenderElement =>
425    layout === 'wide' && side !== null
426      ? <Box flexDirection="row" columnGap={2}><Box flexDirection="column" width={listW}>{list}</Box>{side}</Box>
427      : <Box flexDirection="column">{list}{side}</Box>
428
429  const requirements = (): RenderElement => {
430    const reqs = task.contract.requirements
431    const list = card('v-req', C.main, 'REQUIREMENTS', `${s.requirements.verified}/${inForce} verified`, listW, [
432      ...(reqs.length === 0 ? [<Text key="none" dimColor>No requirements. /integrity-contract add &lt;requirement&gt;</Text>] : []),
433      ...reqs.map(r => {
434        const st = E.requirementState(r, task).state
435        return (
436          <Button key={`req:${r.id}`} plain label={clip(`${REQ_GLYPH[st]} ${r.id} [${r.criticality}] ${r.description}`, rowWidth)}
437            dimColor={st === 'waived'} onPress={go('requirements', r.id)} />
438        )
439      }),
440      task.contract.exclusions.length > 0 && <Text key="excl" dimColor wrap="truncate-end">Exclusions: {task.contract.exclusions.join(', ')}</Text>,
441    ])
442    const r = reqs.find(x => x.id === ui.selected)
443    if (r === undefined) return list
444    const st = E.requirementState(r, task)
445    return split(list, detail(REQ_COLOR[st.state], [
446      `${REQ_GLYPH[st.state]} ${r.id}: ${REQ_LABEL[st.state]}`,
447      r.description,
448      `${r.criticality} · ${r.category} · verified by ${r.verification} evidence · ${r.origin}`,
449      `Why: ${st.reason}`,
450      ...(r.checks.length ? [`Checks: ${r.checks.join(', ')}`] : ['No declared check: /integrity-contract check ' + r.id + ' <command>']),
451      ...(st.evidence.length ? st.evidence.slice(-5).map(ev => `  ${ev.id} ${ev.check} ${ev.outcome} · ${E.freshness(ev, task)} · ${ev.basis}`) : ['No linked evidence. /integrity-verify ' + r.id + ' <what you checked>']),
452      ...(r.waiver ? [`Waived: ${r.waiver.reason}`] : [waiveField(r.id, 'requirement')]),
453    ]))
454  }
455
456  const evidence = (): RenderElement => {
457    const shown = task.evidence.slice(-(isMini ? 6 : 20)).reverse()
458    const list = card('v-ev', C.evidence, 'EVIDENCE', `${s.evidence.fresh} fresh · ${s.evidence.stale} stale`, listW, [
459      ...(shown.length === 0 ? [<Text key="none" dimColor wrap="wrap">Not observed: no checks yet. Test, build, lint and typecheck runs Claude makes appear here. Nothing is verified without one.</Text>] : []),
460      ...shown.map(ev => {
461        const fresh = E.freshness(ev, task)
462        const mark = ev.synthetic ? '⊘' : ev.outcome === 'pass' ? '✓' : ev.outcome === 'fail' ? '✗' : '·'
463        return (
464          <Button key={`ev:${ev.id}`} plain onPress={go('evidence', ev.id)} dimColor={fresh === 'stale' || ev.outcome === 'inconclusive'}
465            label={clip(`${mark} ${ev.id} ${ev.check} ${ev.outcome} · ${fresh === 'stale' ? '⧗ stale' : fresh} · ${ev.command ?? ev.source}`, rowWidth)} />
466        )
467      }),
468    ])
469    const ev = task.evidence.find(x => x.id === ui.selected)
470    if (ev === undefined) return list
471    const fresh = E.freshness(ev, task)
472    return split(list, detail(ev.outcome === 'fail' ? C.bad : fresh === 'stale' ? C.warn : C.evidence, [
473      `${ev.id} · ${ev.kind} · ${ev.check} ${ev.outcome}${ev.synthetic ? ' · SYNTHETIC (never proof)' : ''}`,
474      `Basis: ${ev.basis}`,
475      `Freshness: ${fresh}${fresh === 'stale' ? ` (changed since: ${E.staleBecause(ev, task).slice(0, 4).join(', ')})` : ''}`,
476      ...(ev.command ? [`Command: ${ev.command}`] : []),
477      ...(ev.counts ? [`Counts: ${ev.counts.passed} passed, ${ev.counts.failed} failed, ${ev.counts.skipped} skipped`] : []),
478      `Scope: ${ev.scope.length ? ev.scope.join(', ') : 'whole project'}${ev.truncated ? ' · output truncated' : ''}`,
479      `Source: ${ev.source}${ev.agentId ? ` (subagent ${ev.agentId})` : ''} · ${new Date(ev.timestamp).toISOString()}`,
480      ...(ev.revision ? [`Revision: ${ev.revision}`] : []),
481      ...(ev.cwd ? [`Working dir: ${ev.cwd}`] : []),
482      `Requirements: ${ev.requirementIds.join(', ') || 'none linked (/integrity-evidence link ' + ev.id + ' R#)'}`,
483    ]))
484  }
485
486  const findings = (): RenderElement => {
487    const per = isMini ? Math.min(4, prefs.maxFindings) : prefs.maxFindings
488    const pages = Math.max(1, Math.ceil(sorted.length / per))
489    const page = Math.min(ui.page, pages - 1)
490    const list = card('v-find', C.warn, 'FINDINGS', `${s.findings.open} open`, listW, [
491      ...(sorted.length === 0 ? [<Text key="none" dimColor wrap="wrap">No open findings. Mirage and drift checks run on every observed edit.</Text>] : []),
492      ...sorted.slice(page * per, page * per + per).map(f => (
493        <Button key={`f:${f.id}`} plain onPress={go('findings', f.id)} dimColor={f.status !== 'open' || f.confidence === 'expected'}
494          label={clip(findingLabel(f), rowWidth)} />
495      )),
496      pages > 1 && (
497        <Box key="pager" flexDirection="row" columnGap={1}>
498          <Text dimColor>page {page + 1}/{pages}</Text>
499          {page > 0 && <Button key="prev" plain hotkey="p" label="prev" onPress={() => void actions.setUi(u => ({ ...u, page: page - 1 }))} />}
500          {page < pages - 1 && <Button key="next" plain hotkey="n" label="next" onPress={() => void actions.setUi(u => ({ ...u, page: page + 1 }))} />}
501        </Box>
502      ),
503    ])
504    const f = task.findings.find(x => x.id === ui.selected)
505    if (f === undefined) return list
506    return split(list, detail(f.severity === 'high' ? C.bad : C.warn, [
507      `${SEV_GLYPH[f.severity]} ${f.severity} · ${f.confidence} · ${f.status} · ${f.ruleId} (${f.system})`,
508      `${f.file}:${f.line}`,
509      <Text key="snippet" dimColor wrap="truncate-end">  {f.snippet}</Text>,
510      f.explanation,
511      `Evidence: ${f.evidence}`,
512      `Verify: ${f.verify}`,
513      ...(f.requirementIds.length ? [`Requirements: ${f.requirementIds.join(', ')}`] : []),
514      ...(f.waiver ? [`Waived: ${f.waiver.reason}`] : []),
515      <Box key="f-actions" flexDirection="row" columnGap={1}>
516        {f.status === 'open' && <Button key="ack" plain hotkey="k" label="acknowledge" onPress={() => flash(actions.setFinding(f.id, 'acknowledged'))} />}
517        {f.status !== 'open' && <Button key="reopen" plain hotkey="o" label="reopen" onPress={() => flash(actions.setFinding(f.id, 'open'))} />}
518      </Box>,
519      ...(f.status === 'waived' ? [] : [waiveField(f.id, 'finding')]),
520    ]))
521  }
522
523  const claims = (): RenderElement => {
524    const shown = task.claims.slice(-(isMini ? 4 : 12)).reverse()
525    const list = card('v-claims', C.claims, 'CLAIMS', `${task.claims.length} audited`, listW, [
526      ...(shown.length === 0 ? [<Text key="none" dimColor wrap="wrap">Not assessed: no completion claims audited yet. They are checked when Claude finishes an answer.</Text>] : []),
527      ...shown.map(c => (
528        <Button key={`c:${c.id}`} plain onPress={go('claims', c.id)} dimColor={c.verdict === 'not-assessable'}
529          label={clip(`${E.GLYPH[c.verdict]} ${c.verdict}: "${c.text}"`, rowWidth)} />
530      )),
531    ])
532    const c = task.claims.find(x => x.id === ui.selected)
533    if (c === undefined) return list
534    return split(list, detail(c.verdict === 'contradicted' ? C.bad : c.verdict === 'supported' ? C.ok : C.claims, [
535      `${E.GLYPH[c.verdict]} ${c.verdict} · ${c.kind} · ${time(c.at)}`,
536      `"${c.text}"`,
537      `Why: ${c.reason}`,
538      `Evidence: ${c.evidenceIds.join(', ') || (c.verdict === 'unsupported' ? 'none observed (absence of evidence, not proof of failure)' : 'none')}`,
539    ]))
540  }
541
542  const activity = (): RenderElement => log('v-log', 'ACTIVITY', W, isMini ? 6 : 30)
543
544  const report = (): RenderElement =>
545    card('v-report', C.main, 'REPORT', `task ${task.id}`, W, [
546      <Text key="same" wrap="wrap">Same counts as /integrity-report: {s.requirements.verified}/{inForce} verified, {s.requirements.stale} stale, {s.requirements.failed} failed, {s.findings.open} open findings, {s.evidence.fresh} fresh passing checks.</Text>,
547      <Box key="exports" flexDirection="row" columnGap={2}>
548        <Button key="export-md" plain hotkey="e" label="export Markdown" onPress={() => flash(actions.exportReport('md'))} />
549        <Button key="export-json" plain hotkey="j" label="export JSON" onPress={() => flash(actions.exportReport('json'))} />
550      </Box>,
551      <Text key="where" dimColor wrap="wrap">Files go to .claude/integrity/{task.id}.md|.json in the project. Headless: /integrity-report [md|json] [export].</Text>,
552    ])
553
554  const views: Record<IntegrityTab, () => RenderElement> = { overview, requirements, evidence, findings, claims, activity, report }
555
556  return (
557    <Box flexDirection="column" width={W}>
558      {header}
559      {tabBar}
560      <Box key="body" flexDirection="column" marginTop={isMini ? 0 : 1}>{views[ui.tab]()}</Box>
561      {footer}
562    </Box>
563  )
564}
565
src/engine/types.ts 206 lines
1// Claude Integrity engine: the data model. Plain JSON-safe types, no runtime
2// dependencies, so the engine runs in a Mod, a CLI or a CI job alike.
3
4export type Mode = 'observe' | 'review' | 'strict'
5export type NotifyLevel = 'off' | 'important' | 'all'
6export type Layout = 'auto' | 'mini' | 'compact' | 'wide'
7
8export type RequirementCategory = 'functional' | 'constraint' | 'non-functional' | 'test' | 'preservation'
9export type Criticality = 'must' | 'should' | 'optional'
10export type VerificationKind = 'static' | 'test' | 'runtime' | 'manual' | 'unknown'
11/** Stored status: what a person or the engine decided. Display state adds freshness. */
12export type RequirementStatus = 'unverified' | 'implemented' | 'verified' | 'failed' | 'waived'
13
14export type Requirement = {
15  id: string
16  description: string
17  category: RequirementCategory
18  criticality: Criticality
19  verification: VerificationKind
20  /** `explicit`: quoted from the prompt; `manual`: typed by the person; `proposed`: a model suggestion awaiting approval. */
21  origin: 'explicit' | 'manual' | 'proposed'
22  /** Command substrings whose passing run verifies this requirement (`npm test -- events`). */
23  checks: string[]
24  /** Set by a person: the code for it exists. Never by a model. */
25  implemented?: boolean
26  /** Evidence a person linked by hand (`/integrity-link`). */
27  evidenceIds: string[]
28  waiver?: { reason: string; at: number }
29}
30
31export type ContractIntent = 'production' | 'prototype' | 'unknown'
32
33export type ContractSnapshot = {
34  summary: string
35  requirements: Requirement[]
36  exclusions: string[]
37  scope: string[]
38  assumptions: string[]
39  intent: ContractIntent
40}
41
42export type Revision = { version: number; at: number; change: string; snapshot: ContractSnapshot }
43
44export type Contract = ContractSnapshot & {
45  taskId: string
46  prompt: { hash: string; excerpt: string; capturedAt: number; source: 'prompt' | 'manual' }
47  createdAt: number
48  updatedAt: number
49  version: number
50  approval: { status: 'draft' | 'approved'; at?: number; version?: number }
51  /** Frozen at first approval; amendments become revisions, never overwrite it. */
52  baseline?: ContractSnapshot
53  revisions: Revision[]
54}
55
56export type CheckKind = 'test' | 'build' | 'typecheck' | 'lint' | 'migration' | 'browser' | 'manual' | 'static' | 'other'
57export type EvidenceKind = 'tool-result' | 'static-analysis' | 'test-result' | 'runtime-check' | 'manual-confirmation'
58export type Outcome = 'pass' | 'fail' | 'inconclusive'
59
60export type Evidence = {
61  id: string
62  requirementIds: string[]
63  kind: EvidenceKind
64  check: CheckKind
65  outcome: Outcome
66  timestamp: number
67  /** The tool that ran (`Bash`, `mcp__playwright__browser_snapshot`, `person`). */
68  source: string
69  summary: string
70  /** Redacted, length-bounded command line. */
71  command?: string
72  /** Why the outcome is what it is: `exit 0`, `exit non-zero`, `denied`, `synthetic`, ... */
73  basis: string
74  /** Paths the command targeted; empty means the whole project. */
75  scope: string[]
76  counts?: { passed: number; failed: number; skipped: number }
77  truncated?: boolean
78  /** Answered by a plugin, not by the tool: never proof of execution. */
79  synthetic?: boolean
80  cwd?: string
81  revision?: string
82  agentId?: string
83  /** The task's change sequence when this evidence was recorded. */
84  seq: number
85  artifactPath?: string
86}
87
88export type Freshness = 'fresh' | 'stale' | 'unrelated'
89
90export type FileState = {
91  path: string
92  created: boolean
93  deleted: boolean
94  firstSeq: number
95  lastSeq: number
96  changes: number
97  hash?: string
98  /** Literal `/api/...` endpoints the file calls, for the cross-file endpoint check. */
99  calls?: { path: string; line: number }[]
100  /** The endpoint a route file serves (`app/api/events/route.ts` → `/api/events`). */
101  route?: string
102}
103
104export type Severity = 'high' | 'medium' | 'low' | 'info'
105/** `confirmed`: deterministic proof; `potential`: suspicious, needs a look; `expected`: matches the contract (a prototype's mock). */
106export type Confidence = 'confirmed' | 'potential' | 'expected'
107export type FindingStatus = 'open' | 'acknowledged' | 'waived' | 'resolved'
108
109export type Finding = {
110  id: string
111  ruleId: string
112  system: 'mirage' | 'drift' | 'policy'
113  category: string
114  file: string
115  line: number
116  snippet: string
117  explanation: string
118  evidence: string
119  severity: Severity
120  confidence: Confidence
121  requirementIds: string[]
122  verify: string
123  status: FindingStatus
124  waiver?: { reason: string; at: number }
125  firstSeq: number
126  lastSeq: number
127}
128
129export type ClaimKind =
130  | 'tests-pass' | 'build' | 'typecheck' | 'lint' | 'migration' | 'implemented'
131  | 'production-ready' | 'responsive' | 'browser' | 'live-data' | 'works'
132export type ClaimVerdict = 'supported' | 'partial' | 'unsupported' | 'contradicted' | 'not-assessable'
133
134export type Claim = {
135  id: string
136  turnId: string
137  text: string
138  kind: ClaimKind
139  verdict: ClaimVerdict
140  reason: string
141  evidenceIds: string[]
142  at: number
143}
144
145export type ActivityKind =
146  | 'task' | 'contract' | 'approval' | 'requirement' | 'evidence' | 'invalidated'
147  | 'finding' | 'resolved' | 'waiver' | 'claims' | 'mode' | 'turn'
148export type Activity = { at: number; kind: ActivityKind; text: string }
149
150export type Task = {
151  id: string
152  createdAt: number
153  updatedAt: number
154  contract: Contract
155  evidence: Evidence[]
156  files: Record<string, FileState>
157  findings: Finding[]
158  claims: Claim[]
159  activity: Activity[]
160  /** Monotonic count of observed source changes; evidence freshness is measured against it. */
161  seq: number
162  /** Keys of hints already shown, so a repeat never toasts twice. */
163  notified: string[]
164  nextEvidence: number
165}
166
167export type ProjectRecord = {
168  schema: 1
169  key: string
170  label: string
171  /** Bumped on every write; a writer that finds another rev merges instead of overwriting. */
172  rev: number
173  writer: string
174  updatedAt: number
175  activeTaskId?: string
176  mode: Mode
177  tasks: Record<string, Task>
178}
179
180/** A hint the adapter may show, already deduplicated by the engine. */
181export type Hint = { key: string; level: 'important' | 'info'; text: string }
182
183/** One observed tool call, normalised by the adapter from `tool.call`. */
184export type ToolObservation = {
185  tool: string
186  input: Record<string, unknown>
187  /** Core's answer had `ref`: the tool really ran. */
188  ranInCore: boolean
189  denied?: string
190  isError: boolean
191  text: string
192  result: unknown
193  cwd?: string
194  revision?: string
195  agentId?: string
196}
197
198export type FileChange = {
199  path: string
200  kind: 'create' | 'update' | 'delete'
201  /** The file's text after the change, when known (Write content, a read after Edit). */
202  after?: string
203  /** The text before, when the tool reported it. */
204  before?: string | null
205}
206
src/engine/contract.ts 139 lines
1// Intent Fingerprint: requirement contracts captured from the developer's own
2// words. Nothing here invents acceptance criteria: only sentences the prompt
3// states become requirements; everything else is typed by a person or marked
4// `proposed` and waits for approval.
5
6import type {
7  Contract, ContractIntent, ContractSnapshot, Criticality, Requirement, RequirementCategory,
8} from './types'
9import { clip, hash, isPathPattern, redact, sentences } from './util'
10
11const ACTION = /\b(implement|add|build|create|make|fix|refactor|enrich|update|change|replace|remove|delete|migrate|integrate|wire|connect|write|extend|improve|support|rename|port|convert|set up|hook up|ship)\b/i
12const QUESTION = /^(what|why|how|when|where|who|which|can|could|does|do|is|are|should|would|explain|tell me|show me)\b/i
13
14/** A prompt worth a contract: an instruction to change something, not a question or a one-liner. */
15export function isMeaningfulTask(prompt: string): boolean {
16  const text = prompt.trim()
17  if (text.length < 25 || text.startsWith('/')) return false
18  if (QUESTION.test(text) && text.endsWith('?')) return false
19  return ACTION.test(text)
20}
21
22const NEGATIVE = /\b(do not|don't|dont|never|must not|mustn't|should not|shouldn't|avoid|without)\b/i
23const PRESERVE = /\b(preserve|keep|retain|maintain|leave)\b.*\b(existing|current|unchanged|intact|as is|as-is|navigation|behaviou?r|api|route|layout)\b|\b(existing|current)\b.*\b(must|should)\b.*\b(stay|remain|work)\b/i
24const MUST = /\b(must|need to|needs to|has to|have to|required|require|ensure|make sure)\b/i
25const SHOULD = /\b(should|ideally|prefer)\b/i
26const TESTS = /\b(tests?|spec|coverage|e2e|end-to-end)\b/i
27const PERF = /\b(fast|performance|latency|accessib|a11y|responsive|secure|security)\b/i
28
29function categorize(sentence: string): { category: RequirementCategory; criticality: Criticality } | undefined {
30  if (PRESERVE.test(sentence)) return { category: 'preservation', criticality: 'must' }
31  if (NEGATIVE.test(sentence)) {
32    return { category: /\b(existing|current|navigation|route)\b/i.test(sentence) ? 'preservation' : 'constraint', criticality: 'must' }
33  }
34  if (TESTS.test(sentence) && MUST.test(sentence)) return { category: 'test', criticality: 'must' }
35  if (MUST.test(sentence)) return { category: PERF.test(sentence) ? 'non-functional' : 'functional', criticality: 'must' }
36  if (SHOULD.test(sentence)) return { category: PERF.test(sentence) ? 'non-functional' : 'functional', criticality: 'should' }
37  return undefined
38}
39
40export function detectIntent(text: string): ContractIntent {
41  const prototype = /\b(prototype|mock(ed|up)?s?|mockup|demo|placeholder|sample data|fake data|fixtures?|wireframe|stub(bed)?|poc|proof of concept|static)\b/i.test(text)
42  const production = /\b(live|real|production|persist(ed|ence)?|database|db|backend|api|endpoint|server|save[ds]?|store[ds]?|fetch(ed)?)\b/i.test(text)
43  if (prototype && !/\b(not|no|instead of|replace)\b[^.]*\b(mock|placeholder|fake|static|fixture)/i.test(text)) return 'prototype'
44  return production ? 'production' : 'unknown'
45}
46
47/** Path-like tokens a "do not touch/modify/create" sentence names become exclusion globs. */
48function exclusionsFrom(sentence: string): string[] {
49  if (!/\b(do not|don't|never|must not)\b.*\b(touch|modify|change|edit|create|delete|remove)\b/i.test(sentence)) return []
50  return (sentence.match(/[`'"]?[\w.*/-]+[`'"]?/g) ?? [])
51    .map(t => t.replace(/[`'"]/g, '').replace(/[.,;:]+$/, ''))
52    .filter(t => isPathPattern(t) && /[/*]/.test(t))
53}
54
55function nextRequirementId(requirements: readonly Requirement[]): string {
56  const n = requirements.reduce((max, r) => Math.max(max, Number(r.id.slice(1)) || 0), 0)
57  return `R${n + 1}`
58}
59
60export function makeRequirement(
61  existing: readonly Requirement[],
62  description: string,
63  fields: Partial<Omit<Requirement, 'id' | 'description'>> = {},
64): Requirement {
65  const category = fields.category ?? 'functional'
66  return {
67    id: nextRequirementId(existing),
68    description: clip(redact(description), 300),
69    category,
70    criticality: fields.criticality ?? 'must',
71    verification: fields.verification ?? (category === 'test' ? 'test' : category === 'preservation' || category === 'constraint' ? 'static' : 'unknown'),
72    origin: fields.origin ?? 'manual',
73    checks: fields.checks ?? [],
74    evidenceIds: [],
75    ...(fields.implemented === undefined ? {} : { implemented: fields.implemented }),
76  }
77}
78
79/**
80 * Captures a contract from a prompt without a model: the request is kept by
81 * hash and a short redacted excerpt; explicit constraints become requirements.
82 */
83export function createContract(taskId: string, prompt: string, now: number, source: 'prompt' | 'manual'): Contract {
84  const requirements: Requirement[] = []
85  const exclusions: string[] = []
86  for (const sentence of sentences(prompt)) {
87    const kind = categorize(sentence)
88    if (kind !== undefined) requirements.push(makeRequirement(requirements, sentence, { ...kind, origin: 'explicit' }))
89    exclusions.push(...exclusionsFrom(sentence))
90  }
91  const first = sentences(prompt)[0] ?? prompt
92  return {
93    taskId,
94    summary: clip(redact(first), 140),
95    requirements,
96    exclusions: [...new Set(exclusions)],
97    scope: [],
98    assumptions: [],
99    intent: detectIntent(prompt),
100    prompt: { hash: hash(prompt), excerpt: clip(redact(prompt), 280), capturedAt: now, source },
101    createdAt: now,
102    updatedAt: now,
103    version: 1,
104    approval: { status: 'draft' },
105    revisions: [],
106  }
107}
108
109export const snapshot = (c: Contract): ContractSnapshot =>
110  JSON.parse(JSON.stringify({
111    summary: c.summary, requirements: c.requirements, exclusions: c.exclusions,
112    scope: c.scope, assumptions: c.assumptions, intent: c.intent,
113  })) as ContractSnapshot
114
115export function approve(c: Contract, now: number): void {
116  c.approval = { status: 'approved', at: now, version: c.version }
117  c.baseline ??= snapshot(c)
118  c.updatedAt = now
119}
120
121/**
122 * Every change after the first approval is a versioned revision: the baseline
123 * stays as approved, and a change to what is required asks for re-approval.
124 */
125export function amend(c: Contract, now: number, change: string, mutate: (c: Contract) => void, needsReapproval = true): void {
126  mutate(c)
127  c.updatedAt = now
128  if (c.baseline === undefined) return
129  c.version += 1
130  c.revisions.push({ version: c.version, at: now, change: clip(change, 200), snapshot: snapshot(c) })
131  if (c.revisions.length > 50) c.revisions.splice(0, c.revisions.length - 50)
132  if (needsReapproval) c.approval = { ...c.approval, status: 'draft' }
133}
134
135export const findRequirement = (c: Contract, id: string): Requirement | undefined =>
136  c.requirements.find(r => r.id.toLowerCase() === id.trim().toLowerCase())
137
138export const isApproved = (c: Contract): boolean => c.approval.status === 'approved'
139
src/engine/evidence.ts 216 lines
1// The Witness: evidence from observed tool calls, its provenance and freshness.
2// Outcomes come from the tool's completion state (core's isError, interrupt,
3// background, deny, synthetic answer) first; output text only refines it, and
4// never turns a failed or unobserved run into a pass.
5
6import type { CheckKind, Evidence, EvidenceKind, Freshness, Outcome, Task, ToolObservation } from './types'
7import { clip, isCodePath, isDocPath, redact } from './util'
8
9type Classified = { check: CheckKind; scope: string[] }
10
11const CHECKS: readonly [CheckKind, RegExp][] = [
12  ['browser', /\b(playwright\s+test|cypress\s+run|wdio\s+run|testcafe)\b/i],
13  ['migration', /\b(prisma\s+migrate|drizzle-kit\s+(push|migrate)|knex\s+migrate|sequelize\s+db:migrate|alembic\s+upgrade|rails\s+db:migrate|manage\.py\s+migrate|typeorm\s+migration:run|supabase\s+db\s+push)\b/i],
14  ['test', /\b((npm|pnpm|yarn|bun)\s+(run\s+)?test(:\w+)?|jest|vitest|mocha|ava|pytest|python3?\s+-m\s+(pytest|unittest)|go\s+test|cargo\s+test|dotnet\s+test|node\s+--test|deno\s+test|claude\s+plugin\s+test|rspec|phpunit|(mvn|mvnw)(\s+-\S+)*\s+test|gradlew?\s+test)\b/i],
15  ['typecheck', /\b(tsc|vue-tsc|mypy|pyright|(npm|pnpm|yarn|bun)\s+(run\s+)?(typecheck|type-check|check-types))\b|typescript\.js/i],
16  ['lint', /\b(eslint|biome\s+(check|lint)|ruff|flake8|clippy|golangci-lint|prettier\s+--check|(npm|pnpm|yarn|bun)\s+(run\s+)?lint|claude\s+plugin\s+validate)\b/i],
17  ['build', /\b((npm|pnpm|yarn|bun)\s+(run\s+)?build|next\s+build|vite\s+build|cargo\s+build|go\s+build|dotnet\s+build|webpack|(mvn|mvnw)(\s+-\S+)*\s+(package|install|compile|verify)|gradlew?\s+(build|assemble))\b/i],
18]
19
20const SUBCOMMANDS = new Set(['run', 'test', 'exec', 'watch', 'npx', 'npm', 'pnpm', 'yarn', 'bun', '--'])
21
22/** Paths or filters a check names: `vitest run src/events` scopes to `src/events`. Empty: the whole project. */
23function scopeOf(segment: string, match: RegExpMatchArray): string[] {
24  const rest = segment.slice((match.index ?? 0) + match[0].length)
25  return rest
26    .split(/\s+/)
27    .map(t => t.replace(/^['"]|['"]$/g, '').split('::')[0] ?? '')
28    .filter(t => t !== '' && !t.startsWith('-') && !SUBCOMMANDS.has(t) && (/[/\\]/.test(t) || /\.\w{1,5}$/.test(t)))
29    .map(t => t.replace(/\\/g, '/').replace(/^\.\//, ''))
30}
31
32/** Splits a shell line into its `&&`/`;`/`||` segments and classifies each that is a known check. */
33export function classifyCommand(command: string): Classified[] {
34  const found: Classified[] = []
35  for (const segment of command.split(/&&|\|\||;|\n/)) {
36    const head = segment.split('|')[0] ?? ''
37    for (const [check, pattern] of CHECKS) {
38      const match = head.match(pattern)
39      if (match) {
40        found.push({ check, scope: scopeOf(head, match) })
41        break
42      }
43    }
44  }
45  return found
46}
47
48/** A pipe after the check, `|| true` or `; exit 0` hide the check's own exit status. */
49export function isExitMasked(command: string): boolean {
50  return /\|\|\s*(true|:|exit\s+0)\b|;\s*(true|exit\s+0)\s*$/.test(command) ||
51    command.split(/&&|;|\n/).some(seg => {
52      const pipe = seg.indexOf('|')
53      return pipe >= 0 && seg[pipe + 1] !== '|' && classifyCommand(seg.slice(0, pipe)).length > 0 && !/set\s+-o\s+pipefail/.test(command)
54    })
55}
56
57export type Counts = { passed: number; failed: number; skipped: number }
58
59/** Reads the summary line of common runners (jest, vitest, mocha, pytest, cargo, go, node:test). */
60export function parseCounts(output: string): Counts | undefined {
61  const lines = output.split(/\r?\n/)
62  const summary =
63    lines.filter(l => /^\s*Tests?:?\s/i.test(l) && /\d/.test(l)).at(-1) ??
64    lines.filter(l => /test result:|^=+ .*\d+ (passed|failed)|\d+ (passing|failing)/i.test(l)).join(' ') ??
65    ''
66  const text = summary !== '' ? summary : lines.filter(l => /^#\s*(pass|fail|skipped|todo)\s+\d+/i.test(l) || /\b\d+\s+(passed|failed|skipped)\b/i.test(l)).join(' ')
67  if (text === '') {
68    const goFails = lines.filter(l => /^--- FAIL:/.test(l)).length
69    const goOk = lines.filter(l => /^ok\s+\S+/.test(l)).length
70    return goFails + goOk > 0 ? { passed: goOk, failed: goFails, skipped: 0 } : undefined
71  }
72  const sum = (re: RegExp): number => [...text.matchAll(re)].reduce((n, m) => n + Number(m[1] ?? m[2] ?? 0), 0)
73  return {
74    passed: sum(/(\d+)\s+(?:passed|passing)\b|#\s*pass\s+(\d+)/gi),
75    failed: sum(/(\d+)\s+(?:failed|failing)\b|#\s*fail\s+(\d+)/gi),
76    skipped: sum(/(\d+)\s+(?:skipped|pending|ignored|todo)\b|#\s*(?:skipped|todo)\s+(\d+)/gi),
77  }
78}
79
80const EVIDENCE_KIND: Record<CheckKind, EvidenceKind> = {
81  test: 'test-result', browser: 'runtime-check', migration: 'runtime-check', build: 'tool-result',
82  typecheck: 'tool-result', lint: 'tool-result', manual: 'manual-confirmation', static: 'static-analysis', other: 'tool-result',
83}
84
85type BashRecord = {
86  interrupted?: boolean
87  backgroundTaskId?: string
88  timedOutAfterMs?: number
89  persistedOutputPath?: string
90  stdout?: string
91  stderr?: string
92}
93
94/**
95 * Evidence from one observed Bash call, or none when the command runs no
96 * known check. One record per check the line runs.
97 */
98export function evidenceFromBash(obs: ToolObservation, seq: number, now: number, nextId: () => string): Evidence[] {
99  const command = typeof obs.input['command'] === 'string' ? obs.input['command'] : ''
100  const checks = classifyCommand(command)
101  if (checks.length === 0) return []
102  const record = (obs.result ?? {}) as BashRecord
103  const output = `${record.stdout ?? ''}\n${record.stderr ?? ''}\n${obs.text}`
104  const counts = parseCounts(output)
105  const truncated = record.persistedOutputPath !== undefined || /output too large|\[truncated\]|… \d+ (more )?lines/i.test(obs.text)
106  const masked = isExitMasked(command)
107
108  let outcome: Outcome
109  let basis: string
110  if (obs.denied !== undefined) [outcome, basis] = ['inconclusive', 'permission denied: not executed']
111  else if (!obs.ranInCore) [outcome, basis] = ['inconclusive', 'synthetic: answered by a plugin, not executed']
112  else if (record.backgroundTaskId !== undefined) [outcome, basis] = ['inconclusive', 'moved to background: completion not observed']
113  else if (record.timedOutAfterMs !== undefined) [outcome, basis] = ['inconclusive', `timed out after ${record.timedOutAfterMs} ms`]
114  else if (record.interrupted === true) [outcome, basis] = ['inconclusive', 'interrupted']
115  else if (counts !== undefined && counts.failed > 0) [outcome, basis] = ['fail', masked ? 'output reports failures (exit status masked)' : 'output reports failures']
116  else if (masked) {
117    ;[outcome, basis] = counts !== undefined && counts.passed > 0
118      ? ['pass', 'output summary; exit status masked by a pipe or `|| true`']
119      : ['inconclusive', 'exit status masked by a pipe or `|| true`']
120  } else if (obs.isError) [outcome, basis] = ['fail', exitCode(obs.text) ?? 'exit non-zero']
121  else if ((counts === undefined || counts.passed + counts.failed === 0) && /no tests? (found|ran)|no test files/i.test(output)) {
122    ;[outcome, basis] = ['inconclusive', 'exit 0 but no tests ran']
123  } else [outcome, basis] = ['pass', 'exit 0']
124
125  return checks.map((c, i) => {
126    // In a chain that failed, only a check whose own output says so is known to have failed.
127    const chainFailed = checks.length > 1 && outcome === 'fail' && !(c.check === 'test' && (counts?.failed ?? 0) > 0)
128    const own = chainFailed ? 'inconclusive' : outcome
129    const ownBasis = chainFailed ? `${basis}; chained with other commands, failing step unknown` : basis
130    const tally = c.check === 'test' || c.check === 'browser' ? counts : undefined
131    return {
132      id: nextId(),
133      requirementIds: [],
134      kind: EVIDENCE_KIND[c.check],
135      check: c.check,
136      outcome: own,
137      timestamp: now + i,
138      source: obs.tool,
139      summary: summarize(c.check, own, tally, truncated),
140      command: clip(redact(command), 240),
141      basis: ownBasis,
142      scope: c.scope,
143      ...(tally === undefined ? {} : { counts: tally }),
144      ...(truncated ? { truncated } : {}),
145      ...(obs.ranInCore ? {} : { synthetic: true }),
146      ...(obs.cwd === undefined ? {} : { cwd: redact(obs.cwd) }),
147      ...(obs.revision === undefined ? {} : { revision: obs.revision }),
148      ...(obs.agentId === undefined ? {} : { agentId: obs.agentId }),
149      seq,
150    }
151  })
152}
153
154const exitCode = (text: string): string | undefined => {
155  const m = text.match(/exit (?:code|status)[:\s]+(\d+)/i)
156  return m ? `exit ${m[1]}` : undefined
157}
158
159function summarize(check: CheckKind, outcome: Outcome, counts: Counts | undefined, truncated: boolean): string {
160  const tally = counts === undefined ? '' : ` (${counts.passed} passed, ${counts.failed} failed${counts.skipped ? `, ${counts.skipped} skipped` : ''})`
161  return `${check} ${outcome}${tally}${truncated ? ', output truncated' : ''}`
162}
163
164/** A browser automation call is an observed session, not an assertion: inconclusive until a person confirms. */
165export function evidenceFromBrowserTool(obs: ToolObservation, seq: number, now: number, id: string): Evidence | undefined {
166  if (!/^mcp__.*(playwright|chrome|browser|puppeteer)/i.test(obs.tool)) return undefined
167  return {
168    id, requirementIds: [], kind: 'runtime-check', check: 'browser', outcome: 'inconclusive', timestamp: now,
169    source: obs.tool, summary: 'browser session observed (no assertion)', basis: obs.ranInCore ? 'interaction only' : 'synthetic',
170    scope: [], seq, ...(obs.ranInCore ? {} : { synthetic: true }),
171  }
172}
173
174const CONFIG = /(^|\/)(package\.json|tsconfig[^/]*\.json|pnpm-lock\.yaml|package-lock\.json|yarn\.lock|vite\.config\.\w+|jest\.config\.\w+|vitest\.config\.\w+|next\.config\.\w+|babel\.config\.\w+)$/i
175
176const stem = (path: string): string => (path.split('/').pop() ?? path).replace(/\.(test|spec)(?=\.)/i, '').replace(/\.\w+$/, '').toLowerCase()
177const dir = (path: string): string => path.includes('/') ? path.slice(0, path.lastIndexOf('/')) : ''
178
179/** Whether a change at `path` can affect evidence scoped to `scope`. */
180export function relates(path: string, scope: string): boolean {
181  const p = path.toLowerCase()
182  const s = scope.toLowerCase().replace(/\/$/, '')
183  if (p === s || p.startsWith(`${s}/`)) return true
184  if (stem(p) === stem(s)) return true
185  const sd = /\.\w{1,5}$/.test(s) ? dir(s) : s
186  return sd !== '' && dir(p) === sd
187}
188
189/** Changes that invalidate `e`: code or config, never docs, made after it was recorded. */
190function invalidators(e: Evidence, task: Task): string[] {
191  return Object.values(task.files)
192    .filter(f => f.lastSeq > e.seq && !isDocPath(f.path) && (isCodePath(f.path) || CONFIG.test(f.path)))
193    .filter(f => e.scope.length === 0 || CONFIG.test(f.path) || e.scope.some(s => relates(f.path, s)))
194    .map(f => f.path)
195}
196
197export function freshness(e: Evidence, task: Task): Freshness {
198  if (invalidators(e, task).length > 0) return 'stale'
199  if (e.scope.length > 0 && Object.keys(task.files).length > 0 &&
200    !Object.keys(task.files).some(p => e.scope.some(s => relates(p, s)))) return 'unrelated'
201  return 'fresh'
202}
203
204export const staleBecause = (e: Evidence, task: Task): string[] => invalidators(e, task)
205
206/** Evidence that counts as proof: a real, passing, current run. */
207export const isProof = (e: Evidence, task: Task): boolean =>
208  e.outcome === 'pass' && e.synthetic !== true && freshness(e, task) !== 'stale'
209
210/** The latest record of each check kind, the one the summary shows. */
211export function latestByCheck(task: Task): Map<CheckKind, Evidence> {
212  const out = new Map<CheckKind, Evidence>()
213  for (const e of task.evidence) if (e.check !== 'manual' && e.check !== 'static') out.set(e.check, e)
214  return out
215}
216
src/engine/mirage.ts 340 lines
1// Mirage Detector: deterministic rules for fake completeness in JS/TS/React
2// sources. Each rule reads one file's text against the contract's intent and
3// says how sure it is. Nothing here is a verdict on its own: `confirmed` is
4// kept for code that literally says it is not implemented; everything else
5// is `potential` (look at it) or `expected` (the contract asked for it).
6//
7// Token-level, not an AST: the Mod runtime has no TypeScript compiler, and a
8// regex over comment-stripped text with brace matching covers these patterns
9// with known, documented blind spots (docs/DETECTION-RULES.md).
10
11import type { Confidence, ContractIntent, Severity } from './types'
12import { clip, hash, isTestPath } from './util'
13
14export type RuleContext = {
15  path: string
16  source: string
17  /** The source with comments blanked, offsets and newlines kept. */
18  code: string
19  intent: ContractIntent
20  /** The contract's own words, lower-cased: rules read what was asked for. */
21  asked: string
22}
23
24export type RawFinding = {
25  ruleId: string
26  category: string
27  line: number
28  snippet: string
29  explanation: string
30  evidence: string
31  severity: Severity
32  confidence: Confidence
33  verify: string
34}
35
36export type MirageRule = {
37  id: string
38  category: string
39  title: string
40  appliesTo: (path: string) => boolean
41  detect: (ctx: RuleContext) => RawFinding[]
42}
43
44/** Blanks `//` and `/* *\/` comments outside strings, keeping every offset. */
45export function stripComments(src: string): string {
46  let out = ''
47  let quote: string | undefined
48  for (let i = 0; i < src.length; i++) {
49    const c = src[i] as string
50    if (quote !== undefined) {
51      out += c
52      if (c === '\\') out += src[++i] ?? ''
53      else if (c === quote) quote = undefined
54    } else if (c === '"' || c === "'" || c === '`') {
55      quote = c
56      out += c
57    } else if (c === '/' && src[i + 1] === '/') {
58      while (i < src.length && src[i] !== '\n') { out += ' '; i++ }
59      if (i < src.length) out += '\n'
60    } else if (c === '/' && src[i + 1] === '*') {
61      const end = src.indexOf('*/', i + 2)
62      const stop = end < 0 ? src.length : end + 2
63      out += src.slice(i, stop).replace(/[^\n]/g, ' ')
64      i = stop - 1
65    } else out += c
66  }
67  return out
68}
69
70export const lineOf = (src: string, index: number): number => src.slice(0, index).split('\n').length
71
72const lineText = (src: string, index: number): string => {
73  const start = src.lastIndexOf('\n', index - 1) + 1
74  const end = src.indexOf('\n', index)
75  return src.slice(start, end < 0 ? src.length : end)
76}
77
78/** Index of the `}` closing the `{` at `open`, or the end of the text. */
79function closeOf(code: string, open: number): number {
80  let depth = 0
81  for (let i = open; i < code.length; i++) {
82    if (code[i] === '{') depth++
83    else if (code[i] === '}' && --depth === 0) return i
84  }
85  return code.length
86}
87
88/** The body of the innermost function around `index`: `function ... {` or `=> {`. */
89export function enclosingFunction(code: string, index: number): { start: number; body: string } | undefined {
90  const starts = [...code.slice(0, index).matchAll(/(\bfunction\b[^{;]*|=>\s*)\{/g)]
91  for (let i = starts.length - 1; i >= 0; i--) {
92    const m = starts[i] as RegExpMatchArray
93    const open = (m.index ?? 0) + m[0].length - 1
94    const close = closeOf(code, open)
95    if (close >= index) return { start: m.index ?? 0, body: code.slice(open, close + 1) }
96  }
97  return undefined
98}
99
100const byIntent = (intent: ContractIntent, production: Severity, unknown: Severity): Severity =>
101  intent === 'production' ? production : unknown
102
103const find = (ctx: RuleContext, re: RegExp, make: (m: RegExpMatchArray) => Omit<RawFinding, 'line' | 'snippet'> | undefined): RawFinding[] =>
104  [...ctx.code.matchAll(re)].flatMap(m => {
105    const made = make(m)
106    const at = m.index ?? 0
107    return made === undefined ? [] : [{ ...made, line: lineOf(ctx.code, at), snippet: clip(lineText(ctx.source, at), 140) }]
108  })
109
110const SOURCE = (path: string): boolean => /\.(c|m)?(t|j)sx?$/i.test(path) && !isTestPath(path) && !/\.d\.ts$/i.test(path)
111const UI = (path: string): boolean => SOURCE(path) && (/\.(j|t)sx$/i.test(path) || /(^|\/)(components?|pages|app|views|screens|routes)\//i.test(path))
112const FIXTURE_PATH = /(^|\/)(mocks?|__mocks__|fixtures?|stories|storybook|seed|demo)(\/|\.)|\.(stories|mock|fixture)\.\w+$/i
113
114/** Route handlers: Next.js app/pages API, Express-style `app.post(...)`. */
115const HANDLER = /export\s+(?:async\s+)?function\s+(POST|PUT|PATCH|DELETE)\b|export\s+const\s+(POST|PUT|PATCH|DELETE)\s*=|\b(?:app|router|server)\.(post|put|patch|delete)\s*\(/g
116const IO = /\b(await\s+(?!(?:req|request)\.(?:json|text|formData)\(\))[\w.]+\(|prisma|db\.|sql`|\.query\(|\.insert|\.update\(|\.upsert|\.save\(|\.create\(|\.delete\(|writeFile|fetch\(|axios|supabase|firestore|mongo|redis|knex|drizzle|kv\.|repository\.|\w+Service\.)/
117const SUCCESS = /\b(success|ok)\s*:\s*true|status\(\s*20[01]\s*\)|status:\s*20[01]|Response\.json\(|NextResponse\.json\(|res\.(json|send)\(|return\s+new\s+Response\(/
118
119export const RULES: readonly MirageRule[] = [
120  {
121    id: 'MIR-E1', category: 'E. Placeholder implementation', title: 'Throws "not implemented"', appliesTo: SOURCE,
122    detect: ctx => find(ctx, /throw\s+new\s+\w*Error\(\s*['"`](not\s+(yet\s+)?implemented|unimplemented|todo)\b/gi, () => ({
123      ruleId: 'MIR-E1', category: 'E. Placeholder implementation', severity: 'high', confidence: 'confirmed',
124      explanation: 'This code path throws "not implemented": the behaviour does not exist yet.',
125      evidence: 'Literal throw of a not-implemented error in a non-test source file.',
126      verify: 'Implement the branch, or record it as an explicit exclusion in the contract.',
127    })),
128  },
129  {
130    id: 'MIR-E2', category: 'E. Placeholder implementation', title: 'TODO that defers the real work', appliesTo: SOURCE,
131    // TODOs live in comments, so this rule reads the raw source; plain TODOs are not flagged.
132    detect: ctx => [...ctx.source.matchAll(/(\/\/|\/\*|\{\s*\/\*)\s*(TODO|FIXME|XXX)\b[:\s-]*(.{0,120})/g)]
133      .filter(m => /\b(implement|wire|hook up|connect|persist|save|real|call (the )?api|backend|replace|fetch|auth)/i.test(m[3] ?? ''))
134      .map(m => ({
135        ruleId: 'MIR-E2', category: 'E. Placeholder implementation', line: lineOf(ctx.source, m.index ?? 0),
136        snippet: clip(lineText(ctx.source, m.index ?? 0), 140), severity: byIntent(ctx.intent, 'medium', 'low'), confidence: 'potential' as Confidence,
137        explanation: 'A TODO says the real behaviour is still to be wired up here.',
138        evidence: `Comment: "${clip(m[3] ?? '', 80)}"`,
139        verify: 'Confirm the deferred work is done elsewhere, or track it as unfinished.',
140      })),
141  },
142  {
143    id: 'MIR-E3', category: 'E. Placeholder implementation', title: 'Placeholder endpoint or key', appliesTo: SOURCE,
144    detect: ctx => find(ctx, /['"`](https?:\/\/(api\.)?example\.(com|org)[^'"`]*|\/api\/(todo|placeholder|xxx|example)[^'"`]*|YOUR_[A-Z_]+|<your[^>]*>|REPLACE_ME)['"`]/gi, m => ({
145      ruleId: 'MIR-E3', category: 'E. Placeholder implementation', severity: byIntent(ctx.intent, 'high', 'medium'), confidence: 'potential',
146      explanation: 'A placeholder URL or credential stands where a real one is needed.',
147      evidence: `Literal ${clip(m[1] ?? '', 60)}`,
148      verify: 'Point it at the real endpoint or configuration value.',
149    })),
150  },
151  {
152    id: 'MIR-B1', category: 'B. Simulated success', title: 'Handler returns success without doing the work',
153    appliesTo: SOURCE,
154    detect: ctx => [...ctx.code.matchAll(HANDLER)].flatMap(m => {
155      const at = m.index ?? 0
156      const open = ctx.code.indexOf('{', at)
157      if (open < 0) return []
158      const body = ctx.code.slice(open, closeOf(ctx.code, open) + 1)
159      if (!SUCCESS.test(body) || IO.test(body)) return []
160      const sev: Severity = ctx.intent === 'prototype' ? 'info' : byIntent(ctx.intent, 'high', 'medium')
161      return [{
162        ruleId: 'MIR-B1', category: 'B. Simulated success', line: lineOf(ctx.code, at), snippet: clip(lineText(ctx.source, at), 140),
163        severity: sev, confidence: (ctx.intent === 'prototype' ? 'expected' : 'potential') as Confidence,
164        explanation: `The ${m[1] ?? m[2] ?? m[3]?.toUpperCase()} handler answers success but performs no write, query or outbound call.`,
165        evidence: 'Success response in the handler body; no database, fetch, service or file call found in it.',
166        verify: 'Run a request against the endpoint and confirm the change is persisted.',
167      }]
168    }),
169  },
170  {
171    id: 'MIR-B2', category: 'B. Simulated success', title: 'Artificial delay standing in for a request', appliesTo: SOURCE,
172    detect: ctx => find(ctx, /await\s+new\s+Promise\s*\(\s*\(?\s*\w+\s*\)?\s*=>\s*setTimeout\s*\(|await\s+(sleep|delay|wait)\s*\(\s*\d+/g, m => {
173      const fn = enclosingFunction(ctx.code, m.index ?? 0)
174      const body = fn?.body ?? ''
175      const fakesSuccess = /\b(setSuccess|setSaved|setDone|toast\.success|setStatus\(\s*['"](success|saved|done))|success\s*:\s*true/i.test(body)
176      if (!fakesSuccess || /\bfetch\(|axios|mutate|api\.\w+\(/.test(body)) return undefined
177      return {
178        ruleId: 'MIR-B2', category: 'B. Simulated success', severity: ctx.intent === 'prototype' ? 'info' : 'high',
179        confidence: ctx.intent === 'prototype' ? 'expected' : 'potential',
180        explanation: 'A timer delay is followed by a success state, with no request in between: it imitates a backend call.',
181        evidence: 'setTimeout/sleep plus a success state in the same function; no fetch or mutation call.',
182        verify: 'Replace the delay with the real request and check its failure path.',
183      }
184    }),
185  },
186  {
187    id: 'MIR-C1', category: 'C. Non-functional control', title: 'Control with an empty or log-only action', appliesTo: UI,
188    detect: ctx => find(ctx, /\b(on(?:Click|Press|Submit|Change|Select))=\{\s*(\(\s*\w*\s*\)\s*=>\s*(\{\s*\}|null|undefined|console\.\w+\([^)]*\)|alert\([^)]*\)|\w+\.preventDefault\(\))|undefined|null)\s*\}|\bhref=["'](#|javascript:void\(0\);?)?["']/g, m => ({
189      ruleId: 'MIR-C1', category: 'C. Non-functional control', severity: 'medium', confidence: 'potential',
190      explanation: 'This control is drawn but its action does nothing (empty, log-only or placeholder link).',
191      evidence: `Handler: ${clip(m[0], 80)}`,
192      verify: 'Press the control and confirm it does what the requirement says.',
193    })),
194  },
195  {
196    id: 'MIR-A1', category: 'A. Mock data as production data', title: 'Hardcoded or fixture data in a shipped view',
197    appliesTo: path => UI(path) && !FIXTURE_PATH.test(path),
198    detect: ctx => {
199      const named = /\b(?:const|let|var)\s+((?:mock|fake|dummy|sample|placeholder|hardcoded|static|test)\w*)\s*(?::[^=]+)?=\s*[[{]/gi
200      const imported = /\bimport\s+[^;]*?from\s+['"]([^'"]*(?:mocks?|fixtures?|sample|fake|dummy)[^'"]*)['"]/gi
201      const confidence: Confidence = ctx.intent === 'prototype' ? 'expected' : 'potential'
202      const severity: Severity = ctx.intent === 'prototype' ? 'info' : byIntent(ctx.intent, 'high', 'medium')
203      const why = ctx.intent === 'prototype'
204        ? 'Fixture data, as the contract asks for a prototype.'
205        : 'A view renders fixture data; nothing shows it is replaced by real data.'
206      return [
207        ...find(ctx, named, m => ({
208          ruleId: 'MIR-A1', category: 'A. Mock data as production data', severity, confidence, explanation: why,
209          evidence: `Data declared as "${m[1]}" in a UI source file.`,
210          verify: 'Check where the view gets its data at runtime; trace it to the API or store.',
211        })),
212        ...find(ctx, imported, m => ({
213          ruleId: 'MIR-A1', category: 'A. Mock data as production data', severity, confidence, explanation: why,
214          evidence: `Imports from "${m[1]}".`,
215          verify: 'Check where the view gets its data at runtime; trace it to the API or store.',
216        })),
217      ]
218    },
219  },
220  {
221    id: 'MIR-D1', category: 'D. Incomplete data integration', title: 'Fetched response is discarded', appliesTo: SOURCE,
222    detect: ctx => find(ctx, /\b(?:const|let)\s+(\w+)\s*=\s*await\s+(fetch|axios(?:\.\w+)?|api\.\w+|\w+Client\.\w+)\s*\(/g, m => {
223      const name = m[1] as string
224      const uses = ctx.code.match(new RegExp(`\\b${name}\\b`, 'g'))?.length ?? 0
225      if (uses > 1) return undefined
226      return {
227        ruleId: 'MIR-D1', category: 'D. Incomplete data integration', severity: byIntent(ctx.intent, 'high', 'medium'), confidence: 'potential',
228        explanation: `The response "${name}" is requested and never read: the request runs but its data is not used.`,
229        evidence: `"${name}" is assigned from ${m[2]}(...) and referenced nowhere else in the file.`,
230        verify: 'Confirm the UI renders the response, not static data.',
231      }
232    }),
233  },
234  {
235    id: 'MIR-F1', category: 'F. Missing persistence', title: 'Submit updates local state only',
236    appliesTo: UI,
237    detect: ctx => {
238      if (!/\b(persist|save|saved|store|database|db|backend|server|api)\b/.test(ctx.asked) && ctx.intent !== 'production') return []
239      return find(ctx, /\b(?:const|function)\s+(handleSubmit|onSubmit|submit\w*|save\w*|handleSave\w*|create\w*)\b/g, m => {
240        const fn = enclosingFunction(ctx.code, (m.index ?? 0) + m[0].length + 40)
241        if (fn === undefined || fn.start < (m.index ?? 0)) return undefined
242        const body = fn.body
243        if (!/\bset[A-Z]\w*\(|toast|alert\(/.test(body)) return undefined
244        if (/\bfetch\(|axios|mutate|mutation|\bapi\.|\bpost\(|\bput\(|\bdb\.|prisma|supabase|firestore|localStorage|indexedDB|action\(|Action\(|startTransition|\bsave\w*\(|\bcreate\w*\(|\bupdate\w*\(/.test(body.replace(/\bset[A-Z]\w*\(/g, ''))) return undefined
245        return {
246          ruleId: 'MIR-F1', category: 'F. Missing persistence', severity: 'high', confidence: 'potential',
247          explanation: `${m[1]} changes component state but sends nothing anywhere, while the task asks for saved data.`,
248          evidence: 'State setter or toast in the handler; no request, mutation or storage call.',
249          verify: 'Submit, reload the page, and confirm the data is still there.',
250        }
251      })
252    },
253  },
254  {
255    id: 'MIR-G1', category: 'G. Authorization omission', title: 'Permission enforced in the UI only', appliesTo: SOURCE,
256    detect: ctx => {
257      const out: RawFinding[] = []
258      const mutatesClient = /method:\s*['"](POST|PUT|PATCH|DELETE)['"]/i.test(ctx.code)
259      if (UI(ctx.path) && mutatesClient) {
260        out.push(...find(ctx, /\b(?:user|session|currentUser|me)\??\.(?:role|isAdmin|permissions?)\b|\bisAdmin\s*&&|\bcan\w*\s*&&\s*</g, () => ({
261          ruleId: 'MIR-G1', category: 'G. Authorization omission', severity: 'medium', confidence: 'potential',
262          explanation: 'The UI hides or shows an action by role; a hidden button is not authorization.',
263          evidence: 'Client-side role check in a component that sends a mutating request.',
264          verify: 'Call the endpoint as an unprivileged user and confirm it is refused.',
265        })).slice(0, 1))
266      }
267      if (/\b(auth|permission|admin|role|owner|only)\b/.test(ctx.asked)) {
268        out.push(...[...ctx.code.matchAll(HANDLER)].flatMap(m => {
269          const open = ctx.code.indexOf('{', m.index ?? 0)
270          const body = ctx.code.slice(open, closeOf(ctx.code, open) + 1)
271          if (/\b(auth|session|getServerSession|currentUser|requireUser|verify\w*Token|role|permission|forbidden|401|403)\b/i.test(body)) return []
272          return [{
273            ruleId: 'MIR-G1', category: 'G. Authorization omission', line: lineOf(ctx.code, m.index ?? 0),
274            snippet: clip(lineText(ctx.source, m.index ?? 0), 140), severity: 'high' as Severity, confidence: 'potential' as Confidence,
275            explanation: 'The task mentions permissions, and this mutating handler checks no session, role or token.',
276            evidence: 'No auth, session, role or 401/403 reference in the handler body.',
277            verify: 'Call it without credentials and confirm it is refused.',
278          }]
279        }))
280      }
281      return out
282    },
283  },
284  {
285    id: 'MIR-H1', category: 'H. Missing failure handling', title: 'Request with no failure path', appliesTo: SOURCE,
286    detect: ctx => find(ctx, /\bawait\s+fetch\s*\(/g, m => {
287      const fn = enclosingFunction(ctx.code, m.index ?? 0)
288      const body = fn?.body ?? ctx.code
289      if (/\btry\s*\{|\.catch\(|\.ok\b|\.status\b|throwOnError|onError|catch\s*\(/.test(body)) return undefined
290      return {
291        ruleId: 'MIR-H1', category: 'H. Missing failure handling', severity: byIntent(ctx.intent, 'medium', 'low'), confidence: 'potential',
292        explanation: 'A request is awaited with no try/catch, .catch or status check: a failure has no defined behaviour.',
293        evidence: 'await fetch(...) in a function with no error or status handling.',
294        verify: 'Make the request fail (offline, 500) and check what the user sees.',
295      }
296    }),
297  },
298]
299
300/** Runs every rule that applies to `path`. Findings carry a fingerprint id stable across line moves. */
301export function scanSource(path: string, source: string, intent: ContractIntent, asked: string): (RawFinding & { id: string })[] {
302  if (source.length > 400_000) return []
303  const ctx: RuleContext = { path, source, code: stripComments(source), intent, asked: asked.toLowerCase() }
304  const seen = new Map<string, number>()
305  return RULES.filter(r => r.appliesTo(path)).flatMap(r => r.detect(ctx)).map(f => {
306    const base = hash(`${f.ruleId}|${path}|${f.snippet.replace(/\s+/g, '')}`)
307    const n = seen.get(base) ?? 0
308    seen.set(base, n + 1)
309    return { ...f, id: n === 0 ? base : `${base}-${n}` }
310  })
311}
312
313/** Literal `/api/...` URLs a file requests (template literals with substitutions are skipped). */
314export function apiCalls(source: string): { path: string; line: number }[] {
315  const code = stripComments(source)
316  return [...code.matchAll(/\b(?:fetch|axios(?:\.(?:get|post|put|patch|delete))?|\w+\.(?:get|post|put|patch|delete))\(\s*['"`](\/api\/[^'"`?#$]*)['"`?#]/g)]
317    .map(m => ({ path: (m[1] as string).replace(/\/$/, ''), line: lineOf(code, m.index ?? 0) }))
318}
319
320/** The endpoint a Next.js route file serves, or undefined. Route groups `(x)` are dropped. */
321export function routeOf(path: string): string | undefined {
322  const app = path.match(/(?:^|\/)app\/((?:.+\/)?api(?:\/.+)?)\/route\.(?:t|j)sx?$/)
323  if (app) return '/' + (app[1] as string).split('/').filter(s => !/^\(.*\)$/.test(s)).join('/')
324  const pages = path.match(/(?:^|\/)pages\/(api(?:\/.+?)?)(?:\/index)?\.(?:t|j)sx?$/)
325  return pages ? `/${pages[1]}` : undefined
326}
327
328/** `/api/events/[id]` serves `/api/events/42`; `[...slug]` serves the rest. */
329export function routeMatches(route: string, call: string): boolean {
330  const r = route.split('/')
331  const c = call.split('/')
332  for (let i = 0; i < r.length; i++) {
333    const seg = r[i] as string
334    if (/^\[\[?\.\.\./.test(seg)) return true
335    if (c[i] === undefined) return false
336    if (!/^\[.+\]$/.test(seg) && seg !== c[i]) return false
337  }
338  return r.length === c.length
339}
340
src/engine/drift.ts 121 lines
1// Intent drift: observed changes compared with the contract. A file name alone
2// never proves intent, so only a match against an exclusion glob the person
3// approved is `confirmed`; every other rule reports a `potential` drift and
4// says what to check.
5
6import type { Confidence, Contract, FileChange, Severity } from './types'
7import type { RawFinding } from './mirage'
8import { clip, hash, isCodePath, isPathPattern, matchesGlob } from './util'
9
10export type DriftFinding = RawFinding & { id: string; requirementIds: string[] }
11
12const ROUTE_FILE = /(^|\/)app\/(.+\/)?(page|route)\.(t|j)sx?$|(^|\/)pages\/(?!_app|_document|api\/).+\.(t|j)sx?$|(^|\/)routes\/.+\.(t|j)sx?$/i
13const NO_NEW = /\b(?:do not|don't|never|must not|mustn't|without|avoid)\s+(?:creat\w*|add\w*|introduc\w*|mak\w*|build\w*)\s+(?:another|a new|new|a second|second|separate|an additional|additional|a duplicate|duplicate)\s+([\w-]+)(?:\s+(page|route|screen|view|component|endpoint|module|file|service|table))?/i
14
15function make(change: FileChange, ruleId: string, fields: Omit<RawFinding, 'ruleId' | 'line' | 'snippet'> & { requirementIds?: string[] }): DriftFinding {
16  const { requirementIds = [], ...rest } = fields
17  return {
18    ruleId, line: 1, snippet: `${change.kind} ${change.path}`, ...rest, requirementIds,
19    id: hash(`${ruleId}|${change.path}|${requirementIds.join(',')}`),
20  }
21}
22
23function dependencies(json: string | null | undefined): Set<string> {
24  if (json === null || json === undefined) return new Set()
25  try {
26    const parsed = JSON.parse(json) as Record<string, Record<string, string> | undefined>
27    return new Set([...Object.keys(parsed['dependencies'] ?? {}), ...Object.keys(parsed['devDependencies'] ?? {})])
28  } catch {
29    return new Set()
30  }
31}
32
33export function driftFindings(change: FileChange, contract: Contract): DriftFinding[] {
34  const out: DriftFinding[] = []
35  const approved = contract.approval.status === 'approved' || contract.baseline !== undefined
36  const asked = [contract.summary, ...contract.requirements.map(r => r.description), ...contract.assumptions].join(' ').toLowerCase()
37  const preservation = contract.requirements.filter(r => r.category === 'preservation' && r.waiver === undefined)
38
39  for (const glob of contract.exclusions.filter(isPathPattern)) {
40    if (!matchesGlob(change.path, glob)) continue
41    out.push(make(change, 'DRIFT-X1', {
42      category: 'Excluded path changed', severity: 'high', confidence: approved ? 'confirmed' : 'potential',
43      requirementIds: contract.requirements.filter(r => r.description.includes(glob)).map(r => r.id),
44      explanation: `The contract excludes "${glob}", and ${change.path} was ${change.kind === 'create' ? 'created' : change.kind === 'delete' ? 'deleted' : 'modified'}.`,
45      evidence: `Observed ${change.kind} of ${change.path}; exclusion "${glob}" ${approved ? 'is in the approved contract' : 'is in an unapproved draft'}.`,
46      verify: 'Revert the change, or amend the contract if the exclusion no longer holds.',
47    }))
48  }
49
50  if (contract.scope.length > 0 && isCodePath(change.path) && !contract.scope.some(g => matchesGlob(change.path, g))) {
51    out.push(make(change, 'DRIFT-S1', {
52      category: 'Change outside scope', severity: 'medium', confidence: 'potential',
53      explanation: `${change.path} is outside the contract's scope (${contract.scope.join(', ')}).`,
54      evidence: `Observed ${change.kind} of ${change.path}.`,
55      verify: 'Confirm the change is needed for the task, or widen the scope explicitly.',
56    }))
57  }
58
59  if (change.kind === 'create') {
60    for (const r of contract.requirements) {
61      const m = r.waiver === undefined ? r.description.match(NO_NEW) : null
62      const noun = m?.[1]?.toLowerCase()
63      if (noun === undefined || noun.length < 3 || !change.path.toLowerCase().includes(noun)) continue
64      out.push(make(change, 'DRIFT-N1', {
65        category: 'Possible duplicate of existing feature', severity: 'high', confidence: 'potential', requirementIds: [r.id],
66        explanation: `${r.id} says "${clip(r.description, 90)}", and a new file ${change.path} names "${noun}".`,
67        evidence: `New file whose path contains "${noun}"${ROUTE_FILE.test(change.path) ? ' and that defines a route' : ''}. The name alone does not prove a duplicate.`,
68        verify: `Check whether ${change.path} replaces or duplicates the existing ${noun}; extend the existing one instead if so.`,
69      }))
70    }
71    if (ROUTE_FILE.test(change.path) && preservation.some(r => /\b(navigation|route|routing|url|page)\b/i.test(r.description))) {
72      const ids = preservation.filter(r => /\b(navigation|route|routing|url|page)\b/i.test(r.description)).map(r => r.id)
73      out.push(make(change, 'DRIFT-R1', {
74        category: 'New route while routes are preserved', severity: 'medium', confidence: 'potential', requirementIds: ids,
75        explanation: `${change.path} adds a route, and ${ids.join(', ')} asks to preserve the current navigation.`,
76        evidence: `New route file ${change.path}.`,
77        verify: 'Open the app and confirm the existing navigation still leads where it did.',
78      }))
79    }
80  }
81
82  if (change.kind === 'delete' && isCodePath(change.path)) {
83    out.push(make(change, 'DRIFT-D1', {
84      category: 'Existing behaviour deleted', severity: preservation.length > 0 ? 'high' : 'medium', confidence: 'potential',
85      requirementIds: preservation.map(r => r.id),
86      explanation: `${change.path} was deleted; whatever it did is gone unless moved elsewhere.`,
87      evidence: `Observed deletion of ${change.path}.`,
88      verify: 'Confirm its behaviour moved elsewhere or is no longer needed.',
89    }))
90  }
91
92  if (change.kind === 'update' && typeof change.before === 'string' && change.after !== undefined && isCodePath(change.path)) {
93    const before = change.before.split('\n').map(l => l.trim()).filter(Boolean)
94    const after = new Set(change.after.split('\n').map(l => l.trim()))
95    const kept = before.filter(l => after.has(l)).length
96    if (before.length >= 30 && kept / before.length < 0.4) {
97      out.push(make(change, 'DRIFT-W1', {
98        category: 'Rewrite instead of extension', severity: preservation.length > 0 ? 'medium' : 'low', confidence: 'potential',
99        requirementIds: preservation.map(r => r.id),
100        explanation: `${change.path} kept ${Math.round((kept / before.length) * 100)}% of its previous lines: it was largely rewritten.`,
101        evidence: `${kept} of ${before.length} non-empty lines survive the change.`,
102        verify: 'Check that the existing behaviour of this file still works.',
103      }))
104    }
105  }
106
107  if (/(^|\/)package\.json$/.test(change.path) && change.after !== undefined) {
108    const before = dependencies(change.before)
109    const added = [...dependencies(change.after)].filter(d => !before.has(d) && !asked.includes(d.toLowerCase().replace(/^@[^/]+\//, '')))
110    if (added.length > 0 && change.before !== undefined) {
111      out.push(make(change, 'DRIFT-P1', {
112        category: 'Dependency added without a requirement', severity: 'medium', confidence: 'potential' as Confidence,
113        explanation: `New dependencies not mentioned by the contract: ${added.slice(0, 6).join(', ')}.`,
114        evidence: `package.json gained ${added.length} dependenc${added.length === 1 ? 'y' : 'ies'}.`,
115        verify: 'Confirm each dependency is needed, or record it as an assumption.',
116      }))
117    }
118  }
119  return out.map(f => ({ ...f, severity: f.severity as Severity }))
120}
121
src/engine/claims.ts 127 lines
1// Completion Claim Auditor: finds concrete claims of finished work in a final
2// answer and weighs each against recorded evidence. It tells "no evidence"
3// (unsupported) apart from "evidence says otherwise" (contradicted), and never
4// treats a missing observation as proof that something did not happen.
5
6import type { CheckKind, Claim, ClaimKind, ClaimVerdict, Evidence, Task } from './types'
7import { freshness, isProof } from './evidence'
8import { requirementState } from './status'
9import { clip, hash, redact, sentences } from './util'
10
11const PATTERNS: readonly [ClaimKind, RegExp][] = [
12  ['tests-pass', /\b(all\s+)?(the\s+)?(\w+\s+)?tests?\s+(now\s+)?(all\s+)?(pass(es|ed|ing)?|succeed(s|ed)?|are\s+(passing|green)|(are\s+|is\s+)?green)\b|\b\d+\s*(\/\s*\d+\s+)?tests?\s+pass(ed|ing)?\b|\btest suite\s+(passes|passed|is green)\b/i],
13  ['build', /\bbuild\s+(now\s+)?(succeeds|succeeded|passes|passed|works|is\s+green)\b|\bbuilds?\s+(cleanly|successfully)\b|\bcompiles?\s+(cleanly|successfully|without\s+errors)\b/i],
14  ['typecheck', /\b(type[- ]?check(s|ing)?|tsc|types?)\s+(now\s+)?(pass(es|ed)?|is\s+clean|are\s+clean|succeeds?)\b|\bno\s+type\s+errors\b/i],
15  ['lint', /\blint(ing|er)?\s+(now\s+)?(pass(es|ed)?|is\s+clean|succeeds?)\b|\bno\s+lint(ing)?\s+(errors|warnings|issues)\b/i],
16  ['migration', /\bmigrations?\s+(ran|run|applied|succeeded|completed)\b|\bmigrations?\s+(was|were|has been|have been)\s+(applied|run)\b/i],
17  ['production-ready', /\bproduction[- ]ready\b|\bready\s+for\s+production\b/i],
18  ['responsive', /\b(fully\s+)?responsive\b/i],
19  ['browser', /\b(works|verified|tested|checked)\s+in\s+(the\s+)?browser\b|\bend[- ]to[- ]end\b.*\b(works|pass)|\b(user\s+(flow|journey)|ui)\s+(works|is working)\b/i],
20  ['live-data', /\b(fetch(es|ed)?|load(s|ed)?|pull(s|ed)?|comes?|retrieved?)\b[^.]{0,40}\bfrom\s+the\s+(live\s+|real\s+)?(api|backend|server|database|db)\b|\b(live|real)\s+(api\s+)?data\b/i],
21  ['implemented', /\b[\w-]+(\s+[\w-]+){0,2}\s+(is|are|has\s+been|have\s+been)\s+(now\s+)?(fully\s+)?(implemented|complete(d)?|done|finished)\b|\b(i\s+(have\s+)?|i've\s+)(implemented|completed|finished)\b/i],
22  ['works', /\b(verified|confirmed)\s+(that\s+)?(it|this|everything)\s+works\b|\beverything\s+works\b|\b(it|this)\s+(now\s+)?works\s+(correctly|as expected)\b/i],
23]
24
25/** Hedges, negations, instructions and plans are not claims of completed work. */
26const NOT_A_CLAIM = /\b(not|n't|never|nothing|none|no longer|unable|cannot|could not|didn't|did not|haven't|have not|hasn't|without running|unverified|untested|should|will|would|may|might|once|if|when you|you can|you could|to verify|try|please|run `|need to|needs to|todo|next step|failing|fails|failed)\b|\?$/i
27
28export function extractClaims(answer: string): { kind: ClaimKind; text: string }[] {
29  const out: { kind: ClaimKind; text: string }[] = []
30  for (const s of sentences(answer)) {
31    if (NOT_A_CLAIM.test(s)) continue
32    for (const [kind, re] of PATTERNS) {
33      if (re.test(s) && !out.some(c => c.kind === kind && c.text === s)) out.push({ kind, text: s })
34    }
35  }
36  return out
37}
38
39type Judged = { verdict: ClaimVerdict; reason: string; evidenceIds: string[] }
40
41const CHECK_OF: Partial<Record<ClaimKind, CheckKind>> = {
42  'tests-pass': 'test', build: 'build', typecheck: 'typecheck', lint: 'lint', migration: 'migration', browser: 'browser',
43}
44
45function judgeCheck(kind: ClaimKind, text: string, task: Task): Judged {
46  const check = CHECK_OF[kind] as CheckKind
47  const runs = task.evidence.filter(e => e.check === check || (check === 'test' && e.check === 'browser' && /\b(e2e|end[- ]to[- ]end|browser)\b/i.test(text)))
48  const real = runs.filter(e => e.synthetic !== true)
49  const current = real.filter(e => freshness(e, task) !== 'stale')
50  const ids = (list: Evidence[]): string[] => list.slice(-3).map(e => e.id)
51  const failed = current.filter(e => e.outcome === 'fail').at(-1)
52  const passed = current.filter(e => e.outcome === 'pass')
53  // Observation order decides which run is latest; clocks can tie.
54  const order = (e: Evidence | undefined): number => (e === undefined ? -1 : task.evidence.indexOf(e))
55
56  if (failed !== undefined && order(passed.at(-1)) < order(failed)) {
57    return { verdict: 'contradicted', reason: `latest ${check} run failed: ${failed.summary}`, evidenceIds: [failed.id] }
58  }
59  if (passed.length === 0) {
60    if (kind === 'browser' && task.evidence.some(e => e.check === 'test' && e.outcome === 'pass')) {
61      return { verdict: 'unsupported', reason: 'unit tests passed, but no browser test ran: unit tests do not prove a browser journey', evidenceIds: [] }
62    }
63    if (real.some(e => e.outcome === 'pass')) return { verdict: 'unsupported', reason: `only stale ${check} evidence: code changed after the last passing run`, evidenceIds: ids(real) }
64    if (runs.length > real.length) return { verdict: 'unsupported', reason: `only synthetic ${check} results: nothing was executed`, evidenceIds: ids(runs) }
65    if (runs.length > 0) return { verdict: 'unsupported', reason: `${check} runs were inconclusive (${runs.at(-1)?.basis})`, evidenceIds: ids(runs) }
66    return { verdict: 'unsupported', reason: `no ${check} run was observed for this task`, evidenceIds: [] }
67  }
68  const last = passed.at(-1) as Evidence
69  const broad = /\ball\b|\bevery\b|\bentire\b|\bfull\b/i.test(text)
70  const namesOther = kind === 'tests-pass' && /\b(integration|e2e|end[- ]to[- ]end|browser)\b/i.test(text) && !passed.some(e => /integration|e2e|playwright|cypress/i.test(e.command ?? ''))
71  if (namesOther) return { verdict: 'partial', reason: `${last.summary}, but no integration/e2e run was observed`, evidenceIds: ids(passed) }
72  if (broad && (last.scope.length > 0 || (last.counts?.skipped ?? 0) > 0)) {
73    const why = last.scope.length > 0 ? `the run was scoped to ${last.scope.join(', ')}` : `${last.counts?.skipped} tests were skipped`
74    return { verdict: 'partial', reason: `${last.summary}; ${why}`, evidenceIds: ids(passed) }
75  }
76  if (last.truncated && last.counts === undefined) return { verdict: 'partial', reason: `${last.summary}; output truncated, summary not seen`, evidenceIds: ids(passed) }
77  return { verdict: 'supported', reason: `${last.summary} via ${last.basis} (${last.id})`, evidenceIds: ids(passed) }
78}
79
80function judge(kind: ClaimKind, text: string, task: Task): Judged {
81  if (CHECK_OF[kind] !== undefined) return judgeCheck(kind, text, task)
82  const openMirage = task.findings.filter(f => f.system === 'mirage' && f.status === 'open')
83  switch (kind) {
84    case 'live-data': {
85      const mock = openMirage.find(f => /^MIR-(A1|D1|B1|B2)/.test(f.ruleId) && f.confidence !== 'expected')
86      if (mock) return { verdict: 'contradicted', reason: `potentially contradicted: ${mock.ruleId} in ${mock.file}:${mock.line} (${mock.explanation})`, evidenceIds: [] }
87      const runtime = task.evidence.filter(e => (e.kind === 'runtime-check' || e.kind === 'manual-confirmation') && isProof(e, task))
88      return runtime.length > 0
89        ? { verdict: 'partial', reason: 'a runtime check passed; it does not show where the data came from', evidenceIds: runtime.map(e => e.id) }
90        : { verdict: 'not-assessable', reason: 'static analysis cannot prove live data; no runtime check observed', evidenceIds: [] }
91    }
92    case 'implemented': {
93      const placeholder = openMirage.find(f => f.confidence === 'confirmed')
94      if (placeholder) return { verdict: 'contradicted', reason: `${placeholder.ruleId} in ${placeholder.file}:${placeholder.line}: ${placeholder.explanation}`, evidenceIds: [] }
95      const must = task.contract.requirements.filter(r => r.criticality === 'must' && r.waiver === undefined)
96      const states = must.map(r => requirementState(r, task).state)
97      const verified = states.filter(s => s === 'verified').length
98      if (must.length > 0 && verified === must.length) return { verdict: 'supported', reason: `all ${must.length} must requirements verified`, evidenceIds: [] }
99      if (states.includes('failed')) return { verdict: 'contradicted', reason: 'a must requirement has failing evidence', evidenceIds: [] }
100      if (verified > 0 || Object.keys(task.files).length > 0) {
101        return { verdict: 'partial', reason: `code changes observed; ${verified}/${must.length} must requirements verified`, evidenceIds: [] }
102      }
103      return { verdict: 'unsupported', reason: 'no code changes or verification observed', evidenceIds: [] }
104    }
105    case 'production-ready':
106    case 'responsive':
107      return { verdict: 'not-assessable', reason: `"${kind}" cannot be judged from source and terminal output; needs a manual or runtime check`, evidenceIds: [] }
108    default: {
109      const proof = task.evidence.filter(e => isProof(e, task))
110      return proof.length > 0
111        ? { verdict: 'partial', reason: `${proof.length} passing check(s); "works" is broader than any one of them`, evidenceIds: proof.slice(-3).map(e => e.id) }
112        : { verdict: 'unsupported', reason: 'no passing check observed', evidenceIds: [] }
113    }
114  }
115}
116
117export function auditAnswer(answer: string, task: Task, turnId: string, now: number): Claim[] {
118  return extractClaims(answer).map(({ kind, text }) => ({
119    id: hash(`${turnId}|${kind}|${text}`),
120    turnId,
121    text: clip(redact(text), 200),
122    kind,
123    ...judge(kind, text, task),
124    at: now,
125  }))
126}
127
src/engine/status.ts 77 lines
1// Requirement status from evidence, and the explicit policies Strict mode
2// enforces. "Verified" needs a real, passing, current run of a kind that can
3// prove the requirement: a model's opinion never counts, and one passing suite
4// verifies only the requirements linked to it.
5
6import type { Evidence, EvidenceKind, Mode, Requirement, Task, VerificationKind } from './types'
7import { freshness } from './evidence'
8import { matchesGlob } from './util'
9
10export type RequirementState = 'verified' | 'stale' | 'failed' | 'waived' | 'implemented' | 'unverified'
11
12const PROVES: Record<VerificationKind, readonly EvidenceKind[]> = {
13  test: ['test-result', 'runtime-check', 'manual-confirmation'],
14  runtime: ['runtime-check', 'manual-confirmation'],
15  manual: ['manual-confirmation'],
16  static: ['tool-result', 'static-analysis', 'test-result', 'runtime-check', 'manual-confirmation'],
17  unknown: ['tool-result', 'static-analysis', 'test-result', 'runtime-check', 'manual-confirmation'],
18}
19
20/** Evidence tied to a requirement: linked by id, matched by one of its `checks`, or any test run for a test requirement. */
21export function linkedEvidence(r: Requirement, task: Task): Evidence[] {
22  const checks = r.checks.map(c => c.toLowerCase())
23  return task.evidence.filter(e =>
24    e.requirementIds.includes(r.id) ||
25    r.evidenceIds.includes(e.id) ||
26    (checks.length > 0 && e.command !== undefined && checks.some(c => e.command!.toLowerCase().includes(c))) ||
27    (r.category === 'test' && checks.length === 0 && (e.check === 'test' || e.check === 'browser')))
28}
29
30export function requirementState(r: Requirement, task: Task): { state: RequirementState; reason: string; evidence: Evidence[] } {
31  const linked = linkedEvidence(r, task)
32  if (r.waiver !== undefined) return { state: 'waived', reason: `waived: ${r.waiver.reason}`, evidence: linked }
33  const violated = task.findings.find(f => f.requirementIds.includes(r.id) && f.confidence === 'confirmed' && f.status === 'open')
34  if (violated) return { state: 'failed', reason: `${violated.ruleId}: ${violated.explanation}`, evidence: linked }
35
36  const usable = linked.filter(e => e.synthetic !== true && e.outcome !== 'inconclusive' && PROVES[r.verification].includes(e.kind))
37  const current = usable.filter(e => freshness(e, task) !== 'stale')
38  const latest = current.at(-1)
39  if (latest?.outcome === 'fail') return { state: 'failed', reason: `${latest.summary} (${latest.id})`, evidence: linked }
40  const pass = current.filter(e => e.outcome === 'pass').at(-1)
41  if (pass) return { state: 'verified', reason: `${pass.summary} (${pass.id})`, evidence: linked }
42  if (usable.some(e => e.outcome === 'pass')) return { state: 'stale', reason: 'passing evidence predates later changes', evidence: linked }
43  if (linked.length > 0 && usable.length === 0) {
44    return { state: r.implemented ? 'implemented' : 'unverified', reason: 'only inconclusive, synthetic or non-qualifying evidence', evidence: linked }
45  }
46  if (r.implemented) return { state: 'implemented', reason: 'marked implemented; no qualifying evidence', evidence: linked }
47  return { state: 'unverified', reason: linked.length === 0 ? 'no evidence observed' : 'no qualifying evidence', evidence: linked }
48}
49
50export type PolicyViolation = { id: string; text: string }
51
52/**
53 * Strict mode's deterministic policies, each one the person enabled by writing
54 * it into the approved contract: an excluded path, a must requirement's
55 * declared check. Uncertain judgments are never here.
56 */
57export function policyViolations(task: Task): PolicyViolation[] {
58  const out: PolicyViolation[] = []
59  if (task.contract.approval.status !== 'approved' && task.contract.baseline === undefined) return out
60  for (const f of task.findings) {
61    if (f.ruleId === 'DRIFT-X1' && f.status === 'open') out.push({ id: f.id, text: `excluded path modified: ${f.file}` })
62  }
63  for (const r of task.contract.requirements) {
64    if (r.criticality !== 'must' || r.waiver !== undefined || (r.checks.length === 0 && r.category !== 'test')) continue
65    const { state } = requirementState(r, task)
66    if (state === 'failed') out.push({ id: r.id, text: `${r.id} required check failed` })
67    else if (state !== 'verified') out.push({ id: r.id, text: `${r.id} required check has not passed on the current changes (${r.checks.join(', ') || 'tests'})` })
68  }
69  return out
70}
71
72/** Strict mode's one hard gate the API supports: refuse an edit to a path the approved contract excludes. */
73export function deniedPath(task: Task | undefined, mode: Mode, path: string): string | undefined {
74  if (mode !== 'strict' || task === undefined || task.contract.baseline === undefined) return undefined
75  return task.contract.exclusions.find(g => /[/*]|\.\w{1,5}$/.test(g) && !/\s/.test(g) && matchesGlob(path, g))
76}
77
src/engine/util.ts 95 lines
1// Small pure helpers: hashing, paths, globs and redaction.
2
3/** cyrb53: a fast, deterministic 53-bit string hash, hex. Not cryptographic; identity only. */
4export function hash(text: string, seed = 0): string {
5  let h1 = 0xdeadbeef ^ seed
6  let h2 = 0x41c6ce57 ^ seed
7  for (let i = 0; i < text.length; i++) {
8    const ch = text.charCodeAt(i)
9    h1 = Math.imul(h1 ^ ch, 2654435761)
10    h2 = Math.imul(h2 ^ ch, 1597334677)
11  }
12  h1 = Math.imul(h1 ^ (h1 >>> 16), 2246822507) ^ Math.imul(h2 ^ (h2 >>> 13), 3266489909)
13  h2 = Math.imul(h2 ^ (h2 >>> 16), 2246822507) ^ Math.imul(h1 ^ (h1 >>> 13), 3266489909)
14  return (4294967296 * (2097151 & h2) + (h1 >>> 0)).toString(16).padStart(14, '0')
15}
16
17export function newId(prefix: string, now: number): string {
18  return `${prefix}-${now.toString(36)}-${Math.floor(Math.random() * 1679616).toString(36).padStart(4, '0')}`
19}
20
21const CODE_EXT = /\.(c|m)?(t|j)sx?$|\.(py|go|rs|java|kt|rb|php|cs|swift|vue|svelte|sql|prisma)$/i
22const DOC_EXT = /\.(md|mdx|txt|rst|adoc)$|(^|\/)(LICENSE|CHANGELOG)[^/]*$/i
23const TEST_PATH = /(^|\/)(__tests__|tests?|spec|e2e|__mocks__)\/|\.(test|spec)\.[cm]?[jt]sx?$|_test\.(go|py)$/i
24
25export const isCodePath = (path: string): boolean => CODE_EXT.test(path)
26export const isDocPath = (path: string): boolean => DOC_EXT.test(path)
27export const isTestPath = (path: string): boolean => TEST_PATH.test(path)
28export const isScannable = (path: string): boolean => /\.(c|m)?(t|j)sx?$/i.test(path)
29
30/** Forward slashes, and relative to `root` when inside it (compared case-insensitively for Windows drives). */
31export function normalizePath(path: string, root?: string): string {
32  const p = path.replace(/\\/g, '/')
33  if (root === undefined || root === '') return p
34  const r = root.replace(/\\/g, '/').replace(/\/$/, '')
35  return p.toLowerCase().startsWith(r.toLowerCase() + '/') ? p.slice(r.length + 1) : p
36}
37
38/** `**` any depth, `*` within a segment, `?` one character. A pattern without `/` matches the base name too. */
39export function globToRegExp(glob: string): RegExp {
40  const g = glob.replace(/\\/g, '/').replace(/^\.\//, '')
41  let out = ''
42  for (let i = 0; i < g.length; i++) {
43    const c = g[i] as string
44    if (c === '*' && g[i + 1] === '*') {
45      out += g[i + 2] === '/' ? '(?:.*/)?' : '.*'
46      i += g[i + 2] === '/' ? 2 : 1
47    } else if (c === '*') out += '[^/]*'
48    else if (c === '?') out += '[^/]'
49    else out += c.replace(/[.+^${}()|[\]\\]/g, '\\$&')
50  }
51  return new RegExp(g.includes('/') ? `^${out}$` : `(^|/)${out}$`, 'i')
52}
53
54export const matchesGlob = (path: string, glob: string): boolean => globToRegExp(glob).test(path.replace(/\\/g, '/'))
55
56/** Looks like a path glob rather than prose: has a slash, a star or a file extension, and no spaces. */
57export const isPathPattern = (text: string): boolean => !/\s/.test(text.trim()) && /[/*]|\.\w{1,5}$/.test(text.trim())
58
59const SECRET_PATTERNS: readonly [RegExp, string][] = [
60  [/-----BEGIN [A-Z ]*PRIVATE KEY-----[\s\S]*?-----END [A-Z ]*PRIVATE KEY-----/g, '[private key]'],
61  [/\b(AKIA|ASIA)[0-9A-Z]{16}\b/g, '[aws key]'],
62  [/\bgh[pousr]_[A-Za-z0-9]{30,}\b/g, '[github token]'],
63  [/\bxox[abprs]-[A-Za-z0-9-]{10,}\b/g, '[slack token]'],
64  [/\bsk-[A-Za-z0-9_-]{20,}\b/g, '[api key]'],
65  [/\beyJ[A-Za-z0-9_-]{8,}\.[A-Za-z0-9_-]{8,}\.[A-Za-z0-9_-]{8,}\b/g, '[jwt]'],
66  [/\b(Bearer|Basic)\s+[A-Za-z0-9._~+/=-]{8,}/gi, '$1 ***'],
67  [/(\/\/)[^/\s:@]+:[^/\s@]+@/g, '$1***:***@'],
68  [/\b([A-Za-z0-9_]*(?:api[_-]?key|token|secret|passw(?:or)?d|pwd|credential|auth)[A-Za-z0-9_]*)(\s*[:=]\s*|\s+)("[^"]*"|'[^']*'|[^\s'"&;|]+)/gi, '$1$2***'],
69]
70
71/** Removes secrets and personal home paths. Applied to everything stored or exported. */
72export function redact(text: string): string {
73  let out = text
74  for (const [pattern, replacement] of SECRET_PATTERNS) out = out.replace(pattern, replacement)
75  return out
76    .replace(/\b[A-Za-z]:[\\/]Users[\\/][^\\/\s]+/gi, '~')
77    .replace(/\/(home|Users)\/[^/\s]+/g, '~')
78}
79
80export function clip(text: string, max: number): string {
81  const one = text.replace(/\s+/g, ' ').trim()
82  return one.length <= max ? one : `${one.slice(0, Math.max(0, max - 1))}…`
83}
84
85export const sentences = (text: string): string[] =>
86  text
87    .replace(/```[\s\S]*?```/g, ' ')
88    .split(/(?<=[.!?])\s+|\n+/)
89    .map(s => s.replace(/^[\s>*\-•\d.)#]+/, '').trim())
90    .filter(s => s.length > 3)
91
92/** The store key of a project: its repository root and remote, so two repositories never share a contract. */
93export const projectKey = (root: string, remote: string | null | undefined): string =>
94  `p:${hash(`${normalizePath(root).replace(/\/$/, '').toLowerCase()}|${remote ?? ''}`)}`
95
src/engine/summary.ts 140 lines
1// The one canonical summary: the status line, every pane view and the report
2// all read their counts from `summarize`, so they cannot disagree.
3
4import type { CheckKind, Claim, Evidence, Freshness, Mode, Task } from './types'
5import { freshness, latestByCheck } from './evidence'
6import { requirementState, type RequirementState } from './status'
7
8export type Tab = 'overview' | 'requirements' | 'evidence' | 'findings' | 'claims' | 'activity' | 'report'
9
10export type Summary = {
11  hasTask: boolean
12  taskId?: string
13  title: string
14  approval: 'none' | 'draft' | 'approved' | 'amended'
15  version: number
16  intent: string
17  mode: Mode
18  requirements: Record<RequirementState, number> & { total: number }
19  must: { total: number; verified: number }
20  findings: { open: number; high: number; medium: number; low: number; info: number; expected: number; waived: number; acknowledged: number; confirmed: number }
21  evidence: { total: number; fresh: number; stale: number; latest: { check: CheckKind; outcome: string; freshness: Freshness; summary: string; id: string; at: number }[] }
22  claims: { total: number; supported: number; partial: number; unsupported: number; contradicted: number; notAssessable: number; latest: Claim[] }
23  concern?: { text: string; tab: Tab }
24  /** Verified share of the requirements still in force (waived ones left out), 0..1. */
25  progress: number
26}
27
28const EMPTY_REQ = { verified: 0, stale: 0, failed: 0, waived: 0, implemented: 0, unverified: 0, total: 0 }
29
30export function summarize(task: Task | undefined, mode: Mode): Summary {
31  if (task === undefined) {
32    return {
33      hasTask: false, title: 'No task contract', approval: 'none', version: 0, intent: 'unknown', mode,
34      requirements: { ...EMPTY_REQ }, must: { total: 0, verified: 0 },
35      findings: { open: 0, high: 0, medium: 0, low: 0, info: 0, expected: 0, waived: 0, acknowledged: 0, confirmed: 0 },
36      evidence: { total: 0, fresh: 0, stale: 0, latest: [] },
37      claims: { total: 0, supported: 0, partial: 0, unsupported: 0, contradicted: 0, notAssessable: 0, latest: [] },
38      progress: 0,
39    }
40  }
41  const c = task.contract
42  const requirements = { ...EMPTY_REQ }
43  const must = { total: 0, verified: 0 }
44  for (const r of c.requirements) {
45    const { state } = requirementState(r, task)
46    requirements[state]++
47    requirements.total++
48    if (r.criticality === 'must' && state !== 'waived') {
49      must.total++
50      if (state === 'verified') must.verified++
51    }
52  }
53
54  const findings = { open: 0, high: 0, medium: 0, low: 0, info: 0, expected: 0, waived: 0, acknowledged: 0, confirmed: 0 }
55  for (const f of task.findings) {
56    if (f.status === 'waived') findings.waived++
57    else if (f.status === 'acknowledged') findings.acknowledged++
58    else if (f.status === 'open' && f.confidence === 'expected') findings.expected++
59    else if (f.status === 'open') {
60      findings.open++
61      findings[f.severity]++
62      if (f.confidence === 'confirmed') findings.confirmed++
63    }
64  }
65
66  const fresh = (e: Evidence): Freshness => freshness(e, task)
67  const real = task.evidence.filter(e => e.synthetic !== true)
68  const latest = [...latestByCheck(task).values()].map(e => ({ check: e.check, outcome: e.outcome, freshness: fresh(e), summary: e.summary, id: e.id, at: e.timestamp }))
69  const lastTurn = task.claims.at(-1)?.turnId
70  const latestClaims = task.claims.filter(cl => cl.turnId === lastTurn)
71  const count = (v: Claim['verdict']): number => latestClaims.filter(cl => cl.verdict === v).length
72
73  const summary: Summary = {
74    hasTask: true,
75    taskId: task.id,
76    title: c.summary,
77    approval: c.approval.status === 'approved' ? 'approved' : c.baseline !== undefined ? 'amended' : 'draft',
78    version: c.version,
79    intent: c.intent,
80    mode,
81    requirements,
82    must,
83    findings,
84    evidence: {
85      total: task.evidence.length,
86      fresh: real.filter(e => e.outcome === 'pass' && fresh(e) === 'fresh').length,
87      stale: real.filter(e => e.outcome === 'pass' && fresh(e) === 'stale').length,
88      latest,
89    },
90    claims: {
91      total: latestClaims.length, supported: count('supported'), partial: count('partial'), unsupported: count('unsupported'),
92      contradicted: count('contradicted'), notAssessable: count('not-assessable'), latest: latestClaims,
93    },
94    progress: requirements.total - requirements.waived > 0 ? requirements.verified / (requirements.total - requirements.waived) : 0,
95  }
96  summary.concern = concernOf(summary, task)
97  return summary
98}
99
100function concernOf(s: Summary, task: Task): Summary['concern'] {
101  const contradicted = s.claims.latest.find(cl => cl.verdict === 'contradicted')
102  if (contradicted) return { text: `Claim contradicted: "${contradicted.text}"`, tab: 'claims' }
103  const confirmed = task.findings.find(f => f.status === 'open' && f.confidence === 'confirmed')
104  if (confirmed) return { text: `${confirmed.ruleId} ${confirmed.file}:${confirmed.line}: ${confirmed.category}`, tab: 'findings' }
105  if (s.requirements.failed > 0) return { text: `${s.requirements.failed} requirement(s) have failing evidence`, tab: 'requirements' }
106  const high = task.findings.find(f => f.status === 'open' && f.severity === 'high' && f.confidence === 'potential')
107  if (high) return { text: `Possible ${high.category.replace(/^\w\.\s*/, '').toLowerCase()} in ${high.file}`, tab: 'findings' }
108  const stale = s.evidence.latest.find(e => e.freshness === 'stale' && e.outcome === 'pass')
109  if (stale) return { text: `${stale.check} evidence is stale after later edits`, tab: 'evidence' }
110  const unsupported = s.claims.latest.find(cl => cl.verdict === 'unsupported')
111  if (unsupported) return { text: `Unsupported claim: "${unsupported.text}"`, tab: 'claims' }
112  if (s.approval !== 'approved') return { text: s.approval === 'amended' ? 'Contract amended: re-approve with /integrity-approve' : 'Draft contract: review and /integrity-approve', tab: 'requirements' }
113  if (s.must.total > s.must.verified) return { text: `${s.must.total - s.must.verified} must requirement(s) not verified yet`, tab: 'requirements' }
114  return undefined
115}
116
117/** The glanceable one-line status; undefined when there is nothing to say. Glyphs carry the meaning, not color. */
118export function statusLine(s: Summary, columns = 120): string | undefined {
119  if (!s.hasTask) return undefined
120  if (s.approval === 'draft' && s.evidence.total === 0 && s.findings.open === 0) return 'Integrity ◇ draft contract · /integrity'
121  const parts: string[] = []
122  const inForce = s.requirements.total - s.requirements.waived
123  if (inForce > 0) parts.push(`✓ ${s.requirements.verified}/${inForce} verified`)
124  const unverified = s.requirements.unverified + s.requirements.implemented
125  if (unverified > 0) parts.push(`? ${unverified} unverified`)
126  if (s.requirements.failed > 0) parts.push(`✗ ${s.requirements.failed} failed`)
127  if (s.findings.open > 0) parts.push(`! ${s.findings.open} finding${s.findings.open === 1 ? '' : 's'}`)
128  if (s.evidence.stale > 0) parts.push(`⧗ ${s.evidence.stale} stale`)
129  if (s.claims.contradicted + s.claims.unsupported > 0) parts.push(`✗ ${s.claims.contradicted + s.claims.unsupported} claim${s.claims.contradicted + s.claims.unsupported === 1 ? '' : 's'}`)
130  if (s.approval !== 'approved') parts.push(s.approval === 'amended' ? '◇ amended' : '◇ draft')
131  if (parts.length === 0) parts.push('no evidence yet')
132  // Narrow terminals keep the leading (most important) parts.
133  let line = 'Integrity  ' + parts.join('  ·  ')
134  while (line.length > columns && parts.length > 1) {
135    parts.pop()
136    line = 'Integrity ' + parts.join(' · ')
137  }
138  return line
139}
140