SLOPSHOPPER

Temper

Claude cannot write code before you approve the intent. Gated plan, TDD, review and check for AI written code, with a phase bar in Claude Code.

newpanebandspinnerrowsguard
★ 17v9.6.7MITupdated 2026-10-07galando/temper
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · temper
│ ┃ temper-game ✕ › fix the failing auth test and add an audit log call │ ┃ ▣ client module ./ui/game-client.tsx │ ┃ [ w Jump ] [ s Duck ] [ r Run ] [ q Quit ⏺ Read(src/auth.ts) │ ⎿ Read 6 lines │ ⏺ Update(src/auth.ts) │ ⎿ Added 2 lines, removed 1 line │ ⏺ Bash(bun test) │ ⎿ 3 pass, 1 fail │ │ ● Done. refresh now rejects expired claims and logs an audit event. │ │ ✻ Worked for 42s · done 4:20 PM │ │ │ │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Pane · temper-game
▣ client module ./ui/game-client.tsx [ w Jump ] [ s Duck ] [ r Run ] [ q Quit ]
Pane · temper
◈ Temper No Temper run is active. Start one with /temper:temper and a feature description.
README

Temper

Claude cannot write code before you approve the intent.

An intent gated workflow for AI generated code. Every gate verdict is computed by a small CLI, never asserted by a model. With Claude Code 2.1.287 or later a mod refuses writes outside the current phase through Claude's editing tools (details in "Where enforcement works").

Plugin directory Version Claude Code License: MIT

Temper in a terminal: a refused write, then an approval with one key moves the phase bar from Intent to Plan

Website · Getting Started · Commands · Releases

Install

/plugin marketplace add galando/temper
/plugin install temper

You can also open the Plugins page in Claude, choose Discover and search for "temper".

Your first /temper:temper "describe the feature" sets the project up: the config, the .temper/ folder and a pre-commit hook that blocks git commit while any gate is red. The short form /temper is an interactive shortcut that may not resolve in every surface. Claude Code 2.1.287 or later adds the phase bar and the refusals below; older versions run every phase as prompts.

The problem

AI writes code fast, with predictable failures: happy paths without edge cases, features nobody asked for, calls to methods that do not exist, correct code that is never wired in. Most tools check that the code compiles. Temper checks that it solves the right problem, mechanically, not by asking the model to grade itself.

How it works

One loop with a human gate at every stage. The cheapest artifact is reviewed first. The order is Intent, Plan, Build, Review, Check, then Done. A failed check goes to Fix and back to Check.

flowchart LR
  I["Intent<br/>press 1 to approve"] --> P["Plan<br/>files it may touch"]
  P --> B["Build<br/>failing test first"]
  B --> R["Review<br/>fix or accept"]
  R --> C["Check<br/>run all checks"]
  C --> D(("Done<br/>commit allowed"))
  C -- "a check fails" --> F["Fix<br/>three loops at most"]
  F --> C
  • The intent gate comes first. You approve the problem and the success criteria before exploration or architecture spends tokens. Correcting a wrong intent costs words here and the whole plan later.
  • Every gate is computed. scripts/temper is auditable bash with no network. It reads an evidence ledger and prints PASS or FAIL per requirement. A red gate blocks git commit through a real hook. A person can override a gate (recorded with their identity). A confused model cannot.
  • The mod makes the phases real. Claude Code 2.1.287 or later runs a small mod that refuses a write that does not belong to the current phase, refuses git commit until Check passes, draws the phase bar, and keeps a report of the run (/temper:temper report).

One flow, two views

The Temper bar holds the same choices as the questions Temper asks at each gate, one digit away, plus Discuss, Play and Skip with a reason. With the mod loaded Temper does not ask twice: it prints the result of the stage and waits for the bar, a typed /temper:temper word or a message (table).

The three modes

You choose how much Temper draws with /temper:temper mode. Denials work in every mode. Full draws the bar with action buttons, the pane, toasts and suggestions. Minimal draws the phase bar only. Off draws nothing, and a write outside the phase is still refused. A toast confirms an enforcement change in every mode. The rows below are Full, Minimal and Off. The same mod runs in the desktop app (Code tab).

DarkLight
Full mode, dark: phase bar with action buttons and the paneFull mode, light: phase bar with action buttons and the pane
Minimal mode, dark: the phase bar onlyMinimal mode, light: the phase bar only
Off mode, dark: nothing drawn, denials still applyOff mode, light: nothing drawn, denials still apply

Each phase

Key 1 is the main action and changes when the phase is ready to move on. Key 9 is override everywhere and always asks for a reason. Key 0 shows every action. The keys of each phase are in Commands.

PhaseWrites allowed
Intentintent.md only
Planintent.md, plan.md, tasks.md, design.md and new decision records
BuildThe plan's files, test files and the spec folder. Other files raise scope drift.
ReviewThe spec folder only, unless a fix for that file is active
CheckThe spec folder only. git commit stays refused until Check passes.
FixThe failing files. After three failed loops Temper stops and offers Plan again, Override or Take over.

A game while you wait

While a phase works, the band, the pane and the prompt hint offer "Play while you wait". Press 8 at the empty prompt, or run /temper:temper play, to open Temper Run: Ember, a small dragon, runs in a forge hall. r runs, w jumps over anvils and buckets of cold water, s ducks under flying hammers, and q or Esc leaves. If no key reaches the game within 3 seconds, it says how to give it the keys. The game never opens by itself, shows a banner when a phase is ready, and refusals still apply while it is open. The plugin setting game is on (the default), command (the command only) or off. It runs on the terminal and the desktop app only, verified by hand on the terminal with the keyboard.

Temper Run, the optional game: Ember the dragon jumps over an anvil while Claude works

Where enforcement works

The mod needs Claude Code 2.1.287 or later. Here is where that holds, where not, and what is unverified.

Older versions, or without the mod. Before 2.1.287 the mod is inert and the skills say once that enforcement is off. Then, or whenever the mod does not load, you keep the full pipeline: every phase as a prompt, every CLI gate verdict, the native pre-commit hook and the evidence ledger. You lose the live refusals, bar and report. Checked on 2.1.200 and 2.1.259. The settings declare no picker options on purpose: a settings field with options stops the whole plugin loading on versions before 2.1.271.

SurfaceRefusals (hooks)Drawing
claude in a terminal, including editor terminalsyesyes
Desktop app, Code tabyesyes, except terminal only elements
Desktop app, WSL sessionno (plugins are unavailable)no
VS Code extension chat panelyesno
claude -p and the Agent SDKyes; with enforcement on a run stops at the first gate, because only a person in an interactive session can approve (with enforcement off it runs on)no
Remote Control (phone or web)yes, on your machineonly in your machine's terminal
Cloud sessions (claude.ai/code)yes, if server managed settings bring the plugin (most sessions get the prompt based phases)no
claude.ai chat, Coworknot documented, so unverifiednot documented
GitHub Actionsnot documented; it runs claude -p, so probably yes (unverified)no

Early access API. Claude Code's mods API is early access and may change. The adapter is thin, the rules are plain tested functions, and CI runs the suite on Claude Code 2.1.287.

Organization policy. An administrator can switch parts of this off:

  • allowManagedModsOnly, an option of the built in sec-default guard: with it on, Temper's mod does not load unless the organization ships Temper. Commands, skills, agents and hooks still load.
  • allowManagedHooksOnly also stops hooks from plugins, so the classic Temper hooks stop too unless the plugin is force enabled.
  • disableAllHooks in managed settings stops every mod and every settings hook.
  • A managed guard that runs first and denies a call wins, so two guards never conflict. This was tested with a simulated sec-default only, as the real one needs managed settings on the machine.

You can turn it off. Anyone can disable the plugin or run /temper:temper enforcement off. It guards a workflow for honest use; it is not a security boundary.

Bash is best effort. The hard guarantee covers the tool layer: Write, Edit, NotebookEdit, MultiEdit and git commit. For Bash the mod resolves variables in order, expands braces, follows cd, and refuses a write it cannot check when the command names Temper state, so common tricks fail closed. It cannot see a variable set in an earlier call or a profile, a Bash command can still write ordinary source files, and MCP file tools are not covered. The native pre-commit hook is the backstop. Button presses and the reason field carry no origin, so their authenticity rests on Claude Code.

Limits you should know. While a run is active, a Bash command that names the Temper script and hides what it runs ($(...), ${...}, $'...', a here-string, a script the same command writes and starts, a launcher such as env -S, make, awk or find -exec) is refused, even when the text shows no decision word. A shell, or a builtin that runs text as commands (such as source), given a program the text does not show (a pipe from an unknown command, a file on stdin, a word split by quotes, $ or braces) is refused too. A command that names a file of the run (gates.json, build-state.json, the evidence ledger, .claude/temper.config, git's and Temper's hooks) must be a plain read; chmod, find -delete, git clean and --no-verify are refused. A run whose build-state.json is hidden or removed stays enforced from the last known state until you turn enforcement off, except after the TRIVIAL exit (state clear at Intent before any intent is written), which ends the run. Shell tricks that a text reader cannot see are still possible: a link or a script made in an earlier call, a script already on disk and started later, a program that builds the script name or a path at run time, or the names inside a patch or an archive. What is staged is the session's own picture (a script that stages is not seen). MCP and PowerShell file tools are not evaluated. So the hard guarantees are the editing tools and the native pre-commit hook, not the Bash reader. A Temper enforcement: line can also appear in any file Claude can read. An injected copy can only hide a question, never advance a phase, because every advance still needs the decision of the person or a passed check. When the run is Done, a model git commit is allowed: the run is complete and the person pressed Continue. A later CLI could check a one time decision token.

What the mod reads and writes

Mods are not sandboxed, so this is the full list. The mod makes no network, process or env call and starts no agent. Its one tool call is Claude Code's question dialog ($.ui.ask, AskUserQuestion), to ask you for a mode, a drift choice or a reason. CI fails on a call outside scripts/check-mod-calls.sh.

  • Reads: .claude/temper.config; the run files in .temper/ (build-state.json, gates.json, status.json, overrides.json, feedback-loops.json, evidence/, and a report.md an older Temper wrote); the spec folder of the run (intent.md, plan.md, tasks.md, design.md, events/, config-suggestions.json); and .git/HEAD. To find the project it stats the session folder and looks for .temper/build-state.json there and in up to 11 folders above it. It reads its settings, store and state and the Claude Code version. When you change the mode or enforcement it reads the /config list (every row) to find its two rows and whether your organization locked them. When reviewerModel is set it reads the session's agent list and keeps only the id and type, to find the Temper review agent.
  • Writes no file itself. Decisions, run events, the report, the game's best score and whether you were asked for a mode stay in its plugin store, a JSON file in your Claude Code settings folder.
  • Session state: for its drawing it keeps the bar's view (run title, phase, criteria, findings), the mode, the project folder and the game's key counts in $.state, which other plugins can read.
  • Configuration and environment it sets: no environment variable. Only temper.uiMode (you type /temper:temper mode, or answer the mode question your first /temper:temper or a bare mode asks) and temper.enforcement (you type /temper:temper enforcement). A row your organization locked stays.
  • Slash commands it runs, and when: only /temper:temper and /temper:temper continue <stage> (intent, plan, design, build, review or check), each written as fixed text, and only when you press a button or Enter in the reason field. No command is built from data.
  • What it puts in the prompts it submits: only on that press, the fixed text of the action, with the phase, a finding number, your reason, and the plugin folder's path wherever the text names a Temper file (the scripts/temper command that records your choice, the Stop and Commit steps, the plan review files). Discuss and Change put a fixed draft in your prompt box. After an answer in full mode it may suggest the next fixed prompt; it never sends one. Each prompt is a turn of your session, marked as from the Temper plugin. Apart from these, the refusals below and its state, it sends no text out.
  • System prompt: prompt.compose adds one section, temper:phase, to each request: enforcement on or off, the phase (and whether paused), task, run title, passed criteria, stale phases, the next step, a warning when its state and the CLI's disagree, and one fixed line (answer a message at a gate; after a requested change, run the gate again).
  • Hooks:
  • tool.call sees every tool call, a subagent's too. It reads the path of Write, Edit, MultiEdit and NotebookEdit and the text of Bash, then refuses the call with a fixed reason and next step for Claude (a next step that runs scripts/temper gives the plugin folder's path), or passes it on and returns its result unchanged; it never answers for a tool. A write outside the plan asks you what to do (after Revert, the refusal asks Claude to restore the file).
  • command.run handles only /temper:temper. It answers status, timeline, help, report, mode, enforcement, pane, play, pause and resume. It records an accepted decision (approve, next, back, override, accept, drift) and passes it on; a refused one is answered with the reason. A word that changes state, and play, is refused unless you typed it in your prompt box; with enforcement off, a decision from any origin is accepted and its origin recorded. A bare /temper:temper you type toggles the pane in full mode during a phase. pr, discuss, continue and any other word or command pass on unchanged, even the two the mod runs.
  • session.start and classic.SessionStart find the project root and load the run; in full mode session.start also opens the pane during a phase. Both return what the engine gives them: no context, instruction or setting is added.
  • turn.step sets the model and effort from phaseModels and the Temper review agent's model from reviewerModel; empty options change nothing. attribution.text adds one Temper line to a pull request description during a run when prAttribution is on. turn.complete adds a one line status under an answer in full mode.
  • ui.render draws the bar, pane, game, spinner word, prompt hint and a line above Claude's question dialog, which stays unchanged. ui.message takes the game's score; ui.close notes a closed pane.
  • The game is the mod's one surface module (the game client, with its runner, art and palette files), named as fixed text and imported statically. It makes no engine call and posts only the score.
  • Other: key 9 moves the keys to the reason field. Toasts tell a phase change (full mode), an enforcement change, the default mode when you dismiss the mode question, and the result of a press. It waits 60 ms ($.clock.sleep, at most 4 times) to reread.
  • The tests never run in your session. The mod's test suite runs only under claude plugin test; the plugin never loads it. Every $ call there ($.session.start, $.ui.mount, $.tool.call, $.agent.spawn and the rest) goes to Claude Code's own test kit (claude-code/testing), not another plugin. Tests hand Bash, Write, Edit, NotebookEdit and Read calls to the guard. Temper's fake engine answers each tool call, question, prompt, config.set, fs.write and command.run with a stub and keeps what the mod sent in memory for the test to check, so nothing runs and nothing leaves the test. One test spawns a stub review agent (prompt review it, type temper:temper-review, no model); a stub answers it, so no agent runs, and the mod leaves the spawn unchanged.
  • Tests, lint, git and scripts/temper run as prompts to Claude with its normal permissions. Auto mode may refuse a skip as a gate bypass; the bar then says "Press 1 to record it". Allow that one scripts/temper override command, or run it yourself with !.

Commands

Three you will actually type. /temper:temper runs and routes the rest.

CommandPurpose
/temper:temper "..."The whole pipeline, intent gate to commit
/temper:fix "..."Root cause, a failing test that is write protected, a minimal fix
/temper:intent "..."Capture an idea as a committed draft, build it later

/temper:temper also takes subcommands such as status, approve, override <reason>, back, mode, pane and play. See Commands.

Granular control. Each stage on its own: /temper:plan, /temper:design, /temper:build, /temper:review, /temper:check. Utilities: /temper:status, /temper:pack, /temper:init.

Autonomy (opt in) runs stages after the plan gate unattended and never commits, pushes or merges. Packs: docs/packs.md. CI: examples/workflow/README.md.

Trust

Markdown, a mod written in TypeScript, about 4,400 lines of auditable bash (the CLI and the guard scripts) whose inline Python parses and writes JSON and computes the gate requirements, and four Python scripts (about 1,000 lines, standard library only). Temper itself makes no network calls, sends no telemetry and adds no packages. The committed artifacts (intent, plan, design, gate ledger and diff) are the audit trail, in the same commits as the code.

What Temper runs and changes

Temper's scripts run locally with bash, git and python3, and write only inside your project and its git folder. Apart from the mod's plugin store, Temper reads only one place in your home folder: its global pack folder ~/.claude/packs, if you made one. A pack's link targets come from the skills and commands your session lists.

  • Plugin hooks. The plugin's hooks file registers two classic hooks and the mod module. UserPromptSubmit runs scripts/guards/stage-marker.sh, which notes which gate a standalone stage command owes. Stop runs scripts/guards/verify-stage-gate.sh, which can ask Claude to keep working (at most twice per stage) until that gate has a verdict. Both fail open.
  • Git hook. On the first run scripts/guards/install.sh writes a pre-commit hook (a secret scan of the staged files, then temper gate commit) to temper-gate/pre-commit in the repository's git folder and points core.hooksPath at that folder, so every worktree runs it. It never writes into a folder named hooks. When git would stop running other hooks (in .git/hooks, Temper's older folder, or husky's or lefthook's folder), it leaves that setting and prints a path-free line for your hook, with a hint; 9.6.5's line still counts. After a move of the repository, the next /temper points the setting at the new place. Unset it before adding the pre-commit framework or lefthook. To remove, unset it, drop any Temper line, delete temper-gate and any 9.6.5 temper-pre-commit.
  • Your toolchain. Build and check run the test, lint and type check commands of your stack (detected, or set in check.commands.* in .claude/temper.config) and record their exit codes as evidence.
  • Optional tools already on your machine. OCR (open code review) is off by default. With tools.ocr.mode set to auto or require and ocr on your PATH, /temper:review uses it, and ocr sends the diff to the provider you set up. Temper adds no tool; the guardrails pack and autonomy run only when you ask.
  • Guardrails pack (opt in). /temper:pack enable guardrails asks for the project's .claude/settings.json or .claude/settings.local.json, shows the change and, once you confirm, adds the guard hooks with the plugin folder's absolute path written in; /temper:pack disable guardrails removes them. /temper:pack and /temper:init offer to fix a guard path that no longer exists. They never touch the settings in your home folder.
  • Sharing a plan. Share HTML review publishes the plan review only as a Claude artifact, after you confirm.

Documentation, contributing and license

Source 29 files
hooks/temper-mod/register.tsx 1365 lines
1import type { CommandRunResult, EngineInterface, PluginOptions, Register } from 'claude-code'
2
3import { apply, commitFacts, composeText, consumeDecision, idleSnapshot, keptReport, live, loadSnapshot, nothingToLose, publish, reportText, runFingerprint, settingsFrom, statusText, syncCheck, timelineText, writeReport } from './adapter'
4import type { Io, Snapshot } from './adapter'
5import { findingActions } from './core/actions'
6import type { Action } from './core/actions'
7import { classifyBash } from './core/bash'
8import { pluginCliFrom, stageOf } from './core/cli'
9import { DECISION_WORDS, HELP, NEEDS_INTERACTIVE, followUp, parseArgs, planCommand } from './core/commands'
10import type { Bare, Parsed } from './core/commands'
11import { parseGameMode, parseOnOff, parsePhaseModel, parsePhaseModels, parseUiMode, versionAtLeast } from './core/config'
12import type { GameMode, UiMode } from './core/config'
13import type { Draft } from './core/events'
14import { phaseLabel } from './core/machine'
15import type { Command } from './core/machine'
16import { normalizePath } from './core/paths'
17import { evaluate } from './core/rules'
18import type { RuleContext } from './core/rules'
19import { SECTION_ID } from './core/section'
20import { suggestion, transitionToast, turnLine } from './core/view'
21import type { View } from './core/view'
22import { hintProps } from './ui/hint'
23import { renderBand } from './ui/band'
24import { PANE_ID, renderPane } from './ui/pane'
25import { REASON_HINT } from './ui/band'
26import type { GameButton } from './ui/band'
27import { REASON_KEY } from './ui/kit'
28import { CARD_BG, FG } from './ui/palette'
29import GameClient from './ui/game-client'
30import { renderQuestion } from './ui/question'
31import { spinnerProps } from './ui/spinner'
32
33type Api = EngineInterface
34
35// Module state. A hot reload runs register() again and session.start fires again, so
36// nothing here has to outlive a reload.
37let options: PluginOptions = {}
38let current: Promise<Snapshot> | null = null
39let root = ''
40// The out of plan path a scope drift choice is waiting on (set when a write is denied).
41let pendingDrift: string | null = null
42// Whether a person is at the prompt (session.start says so); the first run question is
43// never asked without one.
44let interactive = false
45let paneOpen = false
46// The last phase announced, so a toast goes out once per transition and never at load.
47let lastPhase: string | null | undefined
48// The game: whether its pane is open, the best score, and a banner Temper hands to it.
49const GAME_ID = 'temper-game'
50let gameOpen = false
51let gameBest = 0
52let gameSeed = 1
53let gameBanner: { text: string; until: number } | null = null
54// The line the person reads when the game opens. It is always this one. Whether the keys reach the
55// game, the game finds out by itself (it draws a hint after 3 seconds with no key).
56const GAME_KEYS_TEXT = 'The game is open. Press r to run, w to jump, s to duck, q or Esc to leave.'
57// The type of the counters in $.state key game (types/index.d.ts has the same shape).
58type GameCtl = { jumpCount: number; duckCount: number; startCount: number }
59// The props of the game module, read off its function (the module is the static import above).
60type GameProps = Parameters<typeof GameClient>[0]
61// After a game over the Run Button says Run again.
62let gameOver = false
63// Where the session draws (session.start says so); the game needs the terminal or the desktop app.
64let drawSurface: string | null = null
65
66const MODE_LABELS: Array<[UiMode, string]> = [
67  ['full', 'Full: bar, buttons and pane'],
68  ['minimal', 'Minimal: phases only'],
69  ['off', 'Off: show nothing'],
70]
71
72// One toast. Every toast of the mod goes through here.
73function showToast($: Api, text: string): void {
74  $.ui.toast(text)
75}
76
77// The mod writes no file. What it records (the decision events and the report) is kept in the plugin
78// store, under the full path it stands for: `vf:<path>` holds the text, `vfdir:<folder>` the names in
79// a folder. Reads and lists see a kept text as a file, and still read real files (event files a run
80// wrote before 9.6.2), so the adapter and the core work on paths as before.
81const VF = 'vf:'
82const VF_DIR = 'vfdir:'
83// The folders with kept texts, oldest first. Past VF_MAX_DIRS the oldest folder is dropped from the store,
84// so the store (4 MiB for every project on the machine) never fills up. A run keeps its events in one folder.
85const VF_DIRS = 'vfdirs'
86const VF_MAX_DIRS = 40
87
88function splitPath(full: string): [string, string] {
89  const cut = full.lastIndexOf('/')
90  return cut < 0 ? ['', full] : [full.slice(0, cut), full.slice(cut + 1)]
91}
92
93async function keptNames($: Api, dir: string): Promise<string[]> {
94  const names = await $.store.get(VF_DIR + dir)
95  return Array.isArray(names) ? names.filter((n): n is string => typeof n === 'string') : []
96}
97
98async function keepText($: Api, full: string, text: string): Promise<void> {
99  const [dir, name] = splitPath(full)
100  await $.store.set(VF + full, text)
101  const names = await keptNames($, dir)
102  if (!names.includes(name)) await $.store.set(VF_DIR + dir, [...names, name])
103  const raw = await $.store.get(VF_DIRS)
104  const dirs = (Array.isArray(raw) ? raw.filter((d): d is string => typeof d === 'string') : []).filter(d => d !== dir)
105  dirs.push(dir)
106  while (dirs.length > VF_MAX_DIRS) {
107    const old = dirs.shift() as string
108    for (const n of await keptNames($, old)) await $.store.delete(VF + old + '/' + n)
109    await $.store.delete(VF_DIR + old)
110  }
111  await $.store.set(VF_DIRS, dirs)
112}
113
114// `$` is spelled only in this file, at each call site, so the adapter and the pure core
115// stay free of it.
116function makeIo($: Api): Io {
117  // A relative path means the project the session started in, even after Claude ran `cd` in a Bash
118  // call: without this the files of the run are not found and the band goes away.
119  const abs = (path: string): string => (root && !path.startsWith('/') ? `${root.replace(/\/$/, '')}/${path}` : path)
120  return {
121    read: async path => {
122      const full = abs(path)
123      const kept = await $.store.get(VF + full)
124      if (typeof kept === 'string') return kept
125      return $.fs.read(full).then(t => (typeof t === 'string' ? t : null))
126    },
127    list: async path => {
128      const full = abs(path).replace(/\/$/, '')
129      const kept = await keptNames($, full)
130      let onDisk: Array<{ name: string; kind: string }>
131      try {
132        onDisk = await $.fs.list(full)
133      } catch (err) {
134        if (kept.length === 0) throw err
135        onDisk = []
136      }
137      // A listing that is not a list is passed on as it came, so a broken load fails open as before.
138      if (!Array.isArray(onDisk)) return onDisk
139      const seen = new Set(onDisk.map(entry => entry.name))
140      return [...onDisk, ...kept.filter(n => !seen.has(n)).map(n => ({ name: n, kind: 'file' }))]
141    },
142    write: (path, text) => keepText($, abs(path), text),
143    pause: ms => $.clock.sleep(ms),
144    storeGet: key => $.store.get(key),
145    storeSet: (key, value) => $.store.set(key, value),
146    version: () => $.session.version().then(v => v.version),
147    setRun: async run => {
148      await $.state.set({ plugin: 'temper', key: 'run' } as const, run)
149    },
150    setMode: async mode => {
151      await $.state.set({ plugin: 'temper', key: 'mode' } as const, mode)
152    },
153    cli: pluginCli(),
154  }
155}
156
157// The last snapshot that held a run. Enforcement is sticky: once the mod has seen a run in progress, a
158// build-state.json that turns missing, unreadable or corrupt (chmod 000, a delete, a git clean) does not end it.
159// The last known state keeps applying until the file reads again, the person turns enforcement off, or the mod reloads.
160let lastRun: Snapshot | null = null
161
162const LOST_STATE =
163  'Temper state: .temper/build-state.json is missing or unreadable. The last known state still applies. Restore the file, or turn enforcement off with /temper:temper enforcement off.'
164
165// The last known run, when it was in progress; null when it was over (Done) or never seen.
166function heldRun(): Snapshot | null {
167  const r = lastRun
168  if (r === null || r.inert || r.sync.failOpen) return null
169  const phase = r.state.phase
170  if (phase === null || phase === 'done') return null
171  // The settings are the live ones (the person may have turned enforcement off since); the run is the last known.
172  return { ...r, ...settingsFrom(options), sync: { ...r.sync, line: LOST_STATE } }
173}
174
175// Whether the folder the mod reads the run from has been seen to hold one (see load).
176let rootVerified = false
177// Where the session began (session.start); used to look for the CLI's root again.
178let sessionCwd = ''
179
180// A failed load never throws. With no run it is a no-run snapshot (nothing enforced); a run that was in progress
181// stays enforced (see lastRun).
182async function load($: Api): Promise<Snapshot> {
183  // After a hot reload no session.start has run yet: the project root comes back from `$.state`.
184  if (!root) root = await rememberedRoot($)
185  let snap: Snapshot | null = await loadSnapshot(makeIo($), options).catch(() => null)
186  if (snap !== null && snap.slug === null && !snap.inert && lastRun === null && !rootVerified && sessionCwd) {
187    // No run in the folder the mod chose at the first session start. A run the CLI starts later may live in a folder
188    // above it (the CLI root): look again before deciding there is none.
189    const found = await locateRoot($, sessionCwd)
190    if (found !== null && found !== root) {
191      root = found
192      await $.state.set({ plugin: 'temper', key: 'root' } as const, root).catch(() => undefined)
193      snap = await loadSnapshot(makeIo($), options).catch(() => null)
194    }
195  }
196  if (snap !== null && snap.slug !== null) {
197    lastRun = snap
198    rootVerified = true
199    return snap
200  }
201  const held = heldRun()
202  if (held !== null) {
203    lastRun = held
204    return held
205  }
206  lastRun = null
207  return snap ?? idleSnapshot(options, false)
208}
209
210// ---- The project root ------------------------------------------------------------------------
211// The folder of the project is fixed once. A session folder reported later (after Claude ran `cd` into
212// the plugin folder, or another repo, and a hot reload started the session there) never replaces it:
213// the files of the run (build-state.json, gates.json, events, report) are always read and written under
214// the first root. `$.state` keeps it across a hot reload.
215
216async function rememberedRoot($: Api): Promise<string> {
217  const read = await $.state.get({ plugin: 'temper', key: 'root' } as const).catch(() => null)
218  return typeof read?.value === 'string' ? read.value : ''
219}
220
221// The nearest folder at or above `cwd` that holds .temper/build-state.json; null when none does.
222async function locateRoot($: Api, cwd: string): Promise<string | null> {
223  let dir = cwd.replace(/\/+$/, '')
224  for (let i = 0; i < 12 && dir !== ''; i++) {
225    if ((await $.fs.read(`${dir}/.temper/build-state.json`).catch(() => undefined)) !== undefined) return dir
226    dir = dir.slice(0, Math.max(0, dir.lastIndexOf('/')))
227  }
228  return null
229}
230
231// Session start: keep the remembered root; the first time, find it and remember it.
232async function settleRoot($: Api, cwd: string | undefined): Promise<void> {
233  const known = await rememberedRoot($)
234  if (known) {
235    root = known
236    return
237  }
238  if (!cwd) return
239  // An inert mod (a version below the minimum) touches no Temper file, not even to find the root.
240  const version = await $.session.version().then(v => v.version).catch(() => undefined)
241  if (!versionAtLeast(version)) {
242    root = cwd
243    return
244  }
245  const found = await locateRoot($, cwd)
246  root = found ?? cwd
247  rootVerified = found !== null
248  // No run was found: the root is the session folder for now, and a run that starts above it is looked for later (see load).
249  sessionCwd = found === null ? cwd : ''
250  await $.state.set({ plugin: 'temper', key: 'root' } as const, root).catch(() => undefined)
251}
252
253function ensure($: Api): Promise<Snapshot> {
254  current ??= load($)
255  return current
256}
257
258// One toast per phase transition, in full mode, never at the first load.
259function announce($: Api, snap: Snapshot): void {
260  const phase = snap.state.phase
261  if (lastPhase !== undefined && phase !== lastPhase && phase !== null && !snap.inert) {
262    // The game, if it is open, shows this line for a while. The toast is the normal single one.
263    const where = phase === 'done' ? 'the run is done' : `${phaseLabel(phase)} is open`
264    gameBanner = { text: `Temper: ${where}. Press Esc to go back.`, until: Date.now() + 20000 }
265    const toast = snap.mode === 'full' ? transitionToast(snap.state.history[snap.state.history.length - 1]) : null
266    if (toast) showToast($, toast)
267  }
268  lastPhase = phase
269}
270
271// Takes a snapshot an `apply` produced as the current one. A run it holds is the last known run too,
272// so a run that reached Done here is not held from memory as the phase before once its state goes.
273function adopt($: Api, snap: Snapshot): Snapshot {
274  current = Promise.resolve(snap)
275  if (snap.slug !== null) lastRun = snap
276  announce($, snap)
277  return snap
278}
279
280async function refresh($: Api): Promise<Snapshot> {
281  current = load($)
282  const snap = await current
283  await publish(makeIo($), snap).catch(() => undefined)
284  announce($, snap)
285  return snap
286}
287
288// The session's directory, to make absolute tool paths project relative: the cwd the
289// session reported, else where `.` resolves.
290async function rootOf($: Api): Promise<string> {
291  if (root) return root
292  const stat = await $.fs.stat('.', { resolve: true }).catch(() => undefined)
293  return stat?.realPath ?? ''
294}
295
296// What the drawing hooks read: the view and the live mode, from `$.state`, so a write
297// to either redraws the sites that read it. Null before the mod has published.
298async function readUi($: Api): Promise<{ view: View; mode: UiMode } | null> {
299  const run = await $.state.get({ plugin: 'temper', key: 'run' } as const)
300  const mode = await $.state.get({ plugin: 'temper', key: 'mode' } as const)
301  if (!run.value || !mode.value) return null
302  return { view: run.value.view, mode: mode.value }
303}
304
305// ---- The game --------------------------------------------------------------------
306
307// `game` is a plain string option (on or off), checked here and never as a picker.
308const gameMode = (): GameMode => parseGameMode(typeof options.game === 'string' ? options.game : undefined)
309// The command works in `on` and `command`. Offers (band, pane, hint) show in `on` only.
310const gameOn = (): boolean => gameMode() !== 'off'
311const gameOffered = (): boolean => gameMode() === 'on'
312// Whether Claude is working, as the band and the hint last saw it. The pane has no such prop.
313let working = false
314
315// The Client element exists on the terminal and the desktop app only.
316const GAME_SURFACES = ['terminal', 'desktop']
317
318// The line Temper hands to the game: a new phase for a short while, else a gate that waits for the
319// person (the first action is an approval or a move on).
320function bannerFor(view: View | null): string | null {
321  if (gameBanner && Date.now() < gameBanner.until) return gameBanner.text
322  if (view === null || view.phase === null || view.phase === 'done') return null
323  const first = view.actions?.primary[0]
324  const waits = first?.command === 'approve' || first?.command === 'next'
325  return waits ? `Temper: ${phaseLabel(view.phase)} is ready. Press Esc to go back.` : null
326}
327
328// Keeps the best score the game posted. This is the one place the game touches the plugin store.
329// Only a real number counts: a string, an array or a boolean is refused, not converted.
330async function saveBest($: Api, score: unknown): Promise<void> {
331  if (typeof score !== 'number' || !Number.isFinite(score)) return
332  const whole = Math.floor(score)
333  if (whole < 0 || whole > 99999 || whole <= gameBest) return
334  gameBest = whole
335  await $.store.set('gameBest', gameBest)
336}
337
338// The pane draws its game offer from `working`. When the band or the hint sees a change, redraw once.
339async function noteWorking($: Api, now: boolean): Promise<void> {
340  if (working === now) return
341  working = now
342  $.ui.invalidate('ui.render')
343}
344
345// Opens or closes the game pane. Only an action of the person calls this: the command or the band
346// button. Both ask for the keys (`focus`); the surface decides.
347async function toggleGame($: Api): Promise<string> {
348  if (!gameOn()) return 'The game is off. Set game to on in /config.'
349  if (drawSurface === null || !GAME_SURFACES.includes(drawSurface)) return 'The game needs the terminal or the desktop app.'
350  if (gameOpen) {
351    await $.ui.close({ id: GAME_ID })
352    gameOpen = false
353    return 'The game is closed.'
354  }
355  const stored = await $.store.get('gameBest')
356  gameBest = typeof stored === 'number' ? stored : 0
357  gameSeed = seedFor(0)
358  gameOver = false
359  // The pane asks for the keyboard and for Esc to close it. The surface may or may not grant the
360  // keys; the game draws a hint by itself when no key comes.
361  const placed = await $.ui.open({ id: GAME_ID, title: 'Temper Run', focus: true, closeOnEscape: true, rows: 14 })
362  gameOpen = placed.isPlaced
363  if (!placed.isPlaced) return 'The game needs a wider terminal.'
364  return GAME_KEYS_TEXT
365}
366
367// A seed for a run: the clock. A test sets the plugin option `gameSeed` (a number) to get a run it
368// knows; each new run then adds the count of the Run presses, so the seeds differ.
369function seedFor(salt: number): number {
370  const fixed = typeof options.gameSeed === 'string' || typeof options.gameSeed === 'number' ? Number(options.gameSeed) : NaN
371  if (Number.isFinite(fixed)) return (Math.floor(fixed) + salt * 7919) >>> 0
372  return ((Date.now() & 0x7fffffff) + salt) >>> 0
373}
374
375// The key presses that reached the pane's Buttons, as counters in $.state (key game). The drawing
376// reads them as props, compares them with the values it saw last, and applies each new press once.
377// A press is one write; the frame clock writes nothing here.
378async function readCtl($: Api): Promise<GameCtl> {
379  const read = await $.state.get({ plugin: 'temper', key: 'game' } as const).catch(() => null)
380  return read?.value ?? { jumpCount: 0, duckCount: 0, startCount: 0 }
381}
382
383async function pressGame($: Api, which: 'jumpCount' | 'duckCount' | 'startCount'): Promise<void> {
384  const cur = await readCtl($)
385  if (which === 'startCount') gameOver = false
386  await $.state.set({ plugin: 'temper', key: 'game' } as const, { ...cur, [which]: cur[which] + 1 })
387}
388
389// The Quit Button: close the pane, as Esc does.
390async function quitGame($: Api): Promise<void> {
391  await $.ui.close({ id: GAME_ID })
392  gameOpen = false
393}
394
395// ---- Pane ------------------------------------------------------------------------
396
397async function openPane($: Api): Promise<boolean> {
398  const placed = await $.ui.open({ id: PANE_ID, title: 'Temper' })
399  paneOpen = placed.isPlaced
400  live.paneOpen = placed.isPlaced
401  return placed.isPlaced
402}
403
404async function closePane($: Api): Promise<void> {
405  await $.ui.close({ id: PANE_ID })
406  paneOpen = false
407  live.paneOpen = false
408}
409
410async function togglePane($: Api): Promise<string> {
411  if (paneOpen) {
412    await closePane($)
413    return 'The Temper pane is closed.'
414  }
415  return (await openPane($)) ? 'The Temper pane is open.' : 'The Temper pane needs a wider terminal.'
416}
417
418// Opened unasked (a session start): only where it would dock. Where it would only wait
419// undrawn, it is closed again so it does not pop up later.
420async function autoOpenPane($: Api, snap: Snapshot): Promise<void> {
421  if (snap.mode !== 'full' || snap.state.phase === null || snap.state.phase === 'done') return
422  const placed = await $.ui.open({ id: PANE_ID, title: 'Temper' })
423  if (placed.isPlaced) {
424    paneOpen = true
425    live.paneOpen = true
426  } else {
427    await $.ui.close({ id: PANE_ID })
428    paneOpen = false
429    live.paneOpen = false
430  }
431}
432
433// ---- Decisions by the person: buttons -------------------------------------------
434
435async function askReason($: Api, what: string): Promise<string> {
436  try {
437    return (await $.ui.ask(`What is the reason for ${what}?`, { options: ['I accept the risk', 'It is not a real issue'], header: 'Reason' })).trim()
438  } catch {
439    return ''
440  }
441}
442
443// Records a decision as the given origin; a person's button press is the person's own.
444//
445// A button passes `drawnPhase`, the phase it was drawn for. When the run has moved since, the press
446// is stale and is ignored, and a move that follows another within a second is the same press twice.
447// Neither records an event. A typed command passes none.
448const STALE = 'That step is already done.'
449let lastMoveAt = 0
450const MOVES = new Set(['approve', 'advance', 'back', 'override'])
451// `from` is set for a decision from another origin that is accepted because enforcement is off: the origin's kind, which
452// the event keeps as its author.
453async function decideAs($: Api, command: Bare, origin: 'person' | 'model', drawnPhase?: string | null, from?: string): Promise<{ error?: string; events: Draft[] }> {
454  const fresh = await refresh($)
455  if (drawnPhase !== undefined && drawnPhase !== null) {
456    if (fresh.state.phase !== drawnPhase) return { error: STALE, events: [] }
457    // A test sets the plugin option `moveCooldownMs` (like `gameSeed`); nobody else needs to.
458    const cooldown = Number.isFinite(Number(options.moveCooldownMs)) && options.moveCooldownMs !== undefined ? Number(options.moveCooldownMs) : 1000
459    if (MOVES.has(command.type) && Date.now() - lastMoveAt < cooldown) return { error: STALE, events: [] }
460  }
461  const who = from === undefined ? { author: 'user' } : { author: `${from} origin`, anyOrigin: true }
462  const done = await apply(makeIo($), options, fresh, { ...command, origin, ...who } as Command)
463  adopt($, done.snap)
464  if (!done.error && MOVES.has(command.type)) lastMoveAt = Date.now()
465  return { error: done.error, events: done.events }
466}
467
468// One decision button at a time: a second press while the first runs is dropped at once, with no
469// toast and no event. The lock is taken before the first await.
470const pressing = new Set<string>()
471const LOCK: Record<string, string> = { approve: 'move', next: 'move', back: 'move', override: 'move', accept: 'accept', pause: 'pause', resume: 'pause', drift: 'drift' }
472async function withLock(word: string | undefined, run: () => Promise<void>): Promise<void> {
473  const key = word === undefined ? undefined : LOCK[word]
474  if (key === undefined) return run()
475  if (pressing.has(key)) return
476  pressing.add(key)
477  try {
478    await run()
479  } finally {
480    pressing.delete(key)
481  }
482}
483
484// prompt.submit is allowed here (a button press runs outside any held turn).
485// The Temper script in the plugin folder, as a full path, from this module's own URL (see pluginCliFrom). Anything
486// else (an unexpected location) gives the plain `scripts/temper`, and the prompt then says where to look. Exported so
487// a test reads the value the mod computes; the loader reads only `register`.
488export const MOD_CLI = pluginCliFrom((import.meta as { url?: string }).url)
489
490function pluginCli(): string {
491  return MOD_CLI
492}
493
494async function submitText($: Api, text: string | null): Promise<void> {
495  if (text) await $.prompt.submit({ text }).then(() => undefined, () => undefined)
496}
497
498// The band's 9: focus the reason field so the person types the reason and presses Enter. When the
499// field cannot take the focus (a different site), fall back to the question dialog.
500async function focusReason($: Api, requestId: string): Promise<boolean> {
501  const moved = await $.ui.focus({ requestId, key: REASON_KEY }).catch(() => undefined)
502  return moved !== undefined && !('deny' in moved && moved.deny)
503}
504
505// Enter in the band's reason field: an empty reason is refused, a real one records the override.
506function runReason($: Api, text: string, drawnPhase?: string | null): Promise<void> {
507  // The lock is taken here, before any await: two Enter presses at once record one skip.
508  return withLock('override', () => runReasonLocked($, text, drawnPhase))
509}
510
511async function runReasonLocked($: Api, text: string, drawnPhase?: string | null): Promise<void> {
512  const reason = text.trim()
513  if (!reason) {
514    showToast($, REASON_HINT)
515    return
516  }
517  const done = await decideAs($, { type: 'override', reason }, 'person', drawnPhase)
518  if (done.error) {
519    showToast($, done.error)
520    return
521  }
522  const launched = await handOver($, done.events)
523  // The step is skipped: the orchestrator runs the next one.
524  if (!launched) await resumeRun($)
525}
526
527// Tells Claude about decisions the person just made. A move forward goes to the orchestrator as
528// `/temper:temper continue <stage>`: the person approved <stage>, the decision event exists, and the
529// orchestrator does the "On Continue" steps of that stage exactly as written (state advance, status flip,
530// branch, commit of the artifacts) and launches the next stage. The guard lets its `state advance` through
531// once, because the matching decision exists. Anything else (back, override, accept) keeps the mirror
532// prompt. Returns whether the orchestrator was launched.
533async function handOver($: Api, drafts: readonly Draft[]): Promise<boolean> {
534  let launched = false
535  for (const draft of drafts) {
536    if (draft.type === 'advance' && draft.from !== 'fix') {
537      // Design is part of Plan in the mod. When the plan was approved and the CLI is at its design stage, this
538      // Continue approves the design: the orchestrator does the On Continue steps of Design.
539      const stage = draft.from === 'plan' && (await ensure($)).cliNext === 'design' ? 'design' : draft.from
540      await continueStage($, stage)
541      launched = true
542    } else {
543      const snap = await ensure($)
544      // The phase a step back left, from the run's own history: the CLI loop is from -> to.
545      const left = draft.type === 'back' ? ([...snap.state.history].reverse().find(step => step.kind === 'back' && step.to === draft.to)?.from ?? null) : null
546      await submitText($, followUp(draft, snap.complexity, pluginCli(), left))
547      // A skip is a move on, like a Continue: the CLI is still at the skipped stage, so the orchestrator does
548      // that stage's On Continue steps (`state advance`, which the guard lets through for a skip) and launches
549      // the next stage. Fix is the Check stage of the CLI.
550      if (draft.type === 'override') {
551        await continueStage($, stageOf(draft.phase))
552        launched = true
553      }
554    }
555  }
556  return launched
557}
558
559// Key 1 when the person's last move is not recorded in the CLI: the same pending decision, the same
560// mirror prompt, once more. No new event is written, and the decision stays single use. When the run
561// has moved by now, there is nothing to record.
562async function recordAgain($: Api): Promise<void> {
563  const snap = await refresh($)
564  const pending = snap.sync.pending
565  if (pending === null) {
566    showToast($, 'That step is already done.')
567    return
568  }
569  // A call that took the decision and never ran must not keep it away from the mirror call.
570  reservedDecisions.delete(pending.id)
571  if (!(await handOver($, [pending.draft]))) await resumeRun($)
572}
573
574// Key 0 (and the pane's More button): show or hide the menu of the other options. The band and the
575// pane both draw it, with the digits 1 to 9, in place of the main buttons. No pane is opened.
576async function showMore($: Api, expanded?: boolean): Promise<void> {
577  live.paneExpanded = expanded ?? !(live.paneExpanded ?? false)
578  await refresh($)
579  $.ui.invalidate('ui.render')
580}
581
582// A button that needs a stage to run ends with the orchestrator's own Resume: `/temper:temper` with no
583// arguments. Its brief for the stage (agents/*.md) is then the one that runs, not a prompt of ours.
584// `prompt.submit` refuses a text that starts with a slash, so the command is run as a command.
585async function resumeRun($: Api): Promise<void> {
586  await $.command.run({ command: 'temper:temper', args: '' }).then(ignore, ignore)
587}
588
589// `/temper:temper continue <stage>`: the orchestrator does the On Continue steps of a stage the person approved.
590// Each command is written out in full, so the plugin directory can read every command the mod runs. A stage
591// that is none of these runs nothing.
592async function continueStage($: Api, stage: string): Promise<void> {
593  switch (stage) {
594    case 'intent':
595      await $.command.run({ command: 'temper:temper', args: 'continue intent' }).then(ignore, ignore)
596      return
597    case 'plan':
598      await $.command.run({ command: 'temper:temper', args: 'continue plan' }).then(ignore, ignore)
599      return
600    case 'design':
601      await $.command.run({ command: 'temper:temper', args: 'continue design' }).then(ignore, ignore)
602      return
603    case 'build':
604      await $.command.run({ command: 'temper:temper', args: 'continue build' }).then(ignore, ignore)
605      return
606    case 'review':
607      await $.command.run({ command: 'temper:temper', args: 'continue review' }).then(ignore, ignore)
608      return
609    case 'check':
610      await $.command.run({ command: 'temper:temper', args: 'continue check' }).then(ignore, ignore)
611      return
612    default:
613      return
614  }
615}
616
617function ignore(): undefined {
618  return undefined
619}
620
621// Puts a draft in the prompt box so the person types the rest and presses Enter. The press itself
622// never moves the phase and writes no event.
623async function fillDraft($: Api, text: string): Promise<void> {
624  const filled = await $.prompt.fill({ text, mode: 'replace' }).catch(() => undefined)
625  showToast($, filled?.isFilled ? 'Type your message. Press Enter to send it.' : 'Close the pane, then type your message.')
626}
627
628function runAction($: Api, action: Action, requestId?: string, drawnPhase?: string | null): Promise<void> {
629  return withLock(action.command?.split(' ')[0], () => runActionLocked($, action, requestId, drawnPhase))
630}
631
632async function runActionLocked($: Api, action: Action, requestId?: string, drawnPhase?: string | null): Promise<void> {
633  if (action.id === 'play') {
634    showToast($, await toggleGame($))
635    return
636  }
637  if (action.id === 'more' || action.id === 'more-actions' || action.id === 'more-narrow') {
638    await showMore($)
639    return
640  }
641  // A choice from the menu leaves the menu: the main buttons come back.
642  if (live.paneExpanded) await showMore($, false)
643  if (action.record) {
644    await recordAgain($)
645    return
646  }
647  if (action.id === 'override' && requestId !== undefined && (await focusReason($, requestId))) return
648  if (action.fill !== undefined) {
649    await fillDraft($, action.fill)
650    return
651  }
652  // The original "Save for later" at the Commit question: nothing to record, the work stays as it is.
653  if (action.id === 'save-done') {
654    showToast($, 'Saved. Commit when you are ready.')
655    return
656  }
657  if (action.resume && !action.command && !action.prompt) {
658    if (working) {
659      showToast($, 'Claude is working. Wait for the answer, then press 1.')
660      return
661    }
662    await resumeRun($)
663    return
664  }
665  if (action.prompt && !action.command) {
666    await submitText($, action.prompt)
667    return
668  }
669  if (!action.command) return
670  let args = action.command
671  if (action.asksReason) {
672    const reason = await askReason($, action.label.toLowerCase())
673    if (!reason) {
674      showToast($, `${action.label} needs a reason.`)
675      return
676    }
677    args = `${action.command} ${reason}`
678  }
679  const parsed = parseArgs(args)
680  if (parsed === null) return
681  const plan = planCommand(parsed, pendingDrift)
682  if (plan.kind === 'error') {
683    showToast($, plan.text)
684    return
685  }
686  if (plan.kind === 'local') return
687  const done = await decideAs($, plan.command, 'person', drawnPhase)
688  if (done.error) {
689    showToast($, done.error)
690    return
691  }
692  const launched = await handOver($, done.events)
693  // The prompt of an action that both records and asks (Stop), then the stage the orchestrator runs.
694  if (action.prompt) await submitText($, action.prompt)
695  if (action.resume && !launched) await resumeRun($)
696}
697
698async function runFindingAction($: Api, kind: 'fix' | 'accept' | 'explain', id: string): Promise<void> {
699  const [fix, , explain] = findingActions(id)
700  if (kind === 'fix') return submitText($, fix?.prompt ?? null)
701  if (kind === 'explain') return submitText($, explain?.prompt ?? null)
702  const reason = await askReason($, `accepting finding ${id}`)
703  if (!reason) {
704    showToast($, 'Accept needs a reason.')
705    return
706  }
707  const done = await decideAs($, { type: 'acceptFinding', id, reason }, 'person')
708  if (done.error) {
709    showToast($, done.error)
710    return
711  }
712  await handOver($, done.events)
713}
714
715// ---- Modes (config rows) ------------------------------------------------------------
716
717type RowKey = 'temper.uiMode' | 'temper.enforcement'
718
719// Changes one of the mod's config rows the way the person changing it in /config does:
720// refused when an administrator locked it, else written (and the module reloads with the
721// new option). Returns the refusal text, or null when it took.
722async function setRow($: Api, key: RowKey, label: string, value: string): Promise<string | null> {
723  const rows = await $.config.list()
724  const row = rows.find(r => r.key === key)
725  if (row?.isLocked) return `Your organization set Temper's ${label} to ${String(row.value)}. Ask your admin to change it.`
726  // Each key is written as fixed text, so a reader (and the plugin directory) sees which setting changes.
727  const result = key === 'temper.uiMode' ? await $.config.set({ key: 'temper.uiMode', value: value }) : await $.config.set({ key: 'temper.enforcement', value: value })
728  return result.deny ?? null
729}
730
731async function switchMode($: Api, mode: UiMode): Promise<string> {
732  const refused = await setRow($, 'temper.uiMode', 'mode', mode)
733  if (refused) return refused
734  live.mode = mode
735  if (mode !== 'full' && paneOpen) await closePane($)
736  await refresh($)
737  $.ui.invalidate('ui.render')
738  return `Temper mode: ${mode}`
739}
740
741async function switchEnforcement($: Api, value: 'on' | 'off'): Promise<string> {
742  const refused = await setRow($, 'temper.enforcement', 'enforcement', value)
743  if (refused) return refused
744  live.enforcement = value
745  await refresh($)
746  $.ui.invalidate('ui.render')
747  showToast($, `Temper enforcement: ${value}`)
748  return `Temper enforcement: ${value}`
749}
750
751// Asks the person for a mode: the first interactive run (once, remembered in the store),
752// or `/temper:temper mode` with no argument. Dismissed means full.
753async function askMode($: Api): Promise<string> {
754  let picked: UiMode | null = null
755  try {
756    const answer = await $.ui.ask('How much do you want Temper to show?', { options: MODE_LABELS.map(([, label]) => label), header: 'Temper mode' })
757    picked = MODE_LABELS.find(([mode, label]) => answer === label || answer.trim().toLowerCase().startsWith(mode))?.[0] ?? null
758  } catch {
759    picked = null
760  }
761  if (picked === null) {
762    live.mode = 'full'
763    await refresh($)
764    showToast($, 'Temper mode is full. To change it, use /temper:temper mode <full|minimal|off>.')
765    return 'Temper mode: full'
766  }
767  return switchMode($, picked)
768}
769
770// Records that the mode question needs no first run ask. An explicit choice always counts; a
771// bare /temper:temper mode counts only where the question can be asked (never in a `-p` run, whose
772// store is the same machine wide one).
773async function markModeAsked($: Api, isExplicit: boolean): Promise<void> {
774  if (isExplicit || interactive) await $.store.set('modeAsked', 1)
775}
776
777async function firstRunAsk($: Api): Promise<void> {
778  if (!interactive) return
779  if (await $.store.get('modeAsked')) return
780  await $.store.set('modeAsked', 1)
781  await askMode($)
782}
783
784// ---- Scope drift ---------------------------------------------------------------------
785
786// Asks the person what to do about a write outside the plan, through the engine's own
787// dialog. Null when nobody can be asked (`claude -p`) or the dialog was dismissed.
788async function askDrift($: Api, path: string): Promise<{ choice: 'add' | 'revert' | 'allow-once'; reason: string } | null> {
789  try {
790    const answer = await $.ui.ask(`${path} is not in the plan. What do you want to do?`, {
791      options: ['Add to plan', 'Revert', 'Allow once'],
792      header: 'Scope drift',
793    })
794    if (answer === 'Add to plan') return { choice: 'add', reason: '' }
795    if (answer === 'Revert') return { choice: 'revert', reason: '' }
796    if (answer !== 'Allow once') return null
797    // Allow once needs a reason: ask again while the answer is empty, then give up.
798    for (let tries = 0; tries < 3; tries++) {
799      const reason = (
800        await $.ui.ask(`What is the reason to allow ${path} once?`, { options: ['Needed for this task', 'Short test'], header: 'Reason' })
801      ).trim()
802      if (reason) return { choice: 'allow-once', reason }
803    }
804    return null
805  } catch {
806    return null
807  }
808}
809
810// A write outside the plan: offer the three choices and log the decision as an event.
811// Returns the deny text, or null when the person let the write through.
812async function resolveDrift($: Api, snap: Snapshot, path: string, fallback: string): Promise<string | null> {
813  pendingDrift = path
814  const choice = await askDrift($, path)
815  if (choice === null) return fallback
816  const io = makeIo($)
817  const done = await apply(io, options, snap, { type: 'drift', path, choice: choice.choice, reason: choice.reason, origin: 'person', author: 'user' })
818  adopt($, done.snap)
819  if (done.error) return `Temper: scope drift on ${path} is not recorded. ${done.error} Next: ask the user to choose again (/temper:temper drift add|revert|allow <reason>).`
820  pendingDrift = null
821  if (choice.choice === 'revert') {
822    // prompt.submit cannot be called from a tool.call hook (the engine says it would wait
823    // on this very turn), so the instruction rides in the deny text Claude reads next.
824    return `Temper: scope drift. The user chose to revert ${path}. It stays out of the plan. Next: restore ${path} to its committed state. Then continue inside the plan files.`
825  }
826  // Add to plan, or allow once: the decision now lets this write through.
827  const again = evaluate(done.snap.state, ruleContext(done.snap, await rootOf($)), { tool: 'Edit', input: { file_path: path } })
828  if ('deny' in again) return again.deny
829  if (again.consume === 'drift' && again.driftPath) {
830    adopt($, (await apply(io, options, done.snap, { type: 'useDrift', path: again.driftPath, origin: 'system' })).snap)
831  }
832  return null
833}
834
835function ruleContext(snap: Snapshot, rootDir: string): RuleContext {
836  return { root: rootDir, specDir: snap.specDir, planFiles: snap.planFiles, humanDecisions: snap.humanDecisions, autonomyEnabled: snap.autonomyEnabled, designRequired: snap.designRequired, complexity: snap.complexity, failOpenWrites: snap.sync.looksReset, cli: pluginCli() }
837}
838
839// What the guard decided for one call: a deny text, or null to pass it on. `ids` are the human decisions
840// the call took. They are spent only when the call ran and succeeded (see the tool.call hook).
841// `fingerprint` is the CLI's files before a call that took decisions or a loop: after a call that ended in an error it
842// says whether the call took effect anyway. `loopIds` are back decisions a `state loop` call used.
843// `trivialExit` is a `state clear` of a run that never left Intent (see nothingToLose): once it ran, the run it
844// cleared is not held any more (see lastRun), so the work that follows is not judged against a run that is gone.
845type Guarded = { deny: string | null; ids: string[]; undoStaged?: { all: boolean; paths: string[] }; loopIds?: string[]; fingerprint?: string; trivialExit?: boolean }
846
847// The shell's folder, carried from one Bash call to the next (the engine keeps it). A folder outside the project, or one
848// that cannot be read, is taken as the project root: the engine puts the shell back there.
849let bashCwd = ''
850const loopedDecisions = new Set<string>()
851
852function carryCwd(after: string | null, rootDir: string): string {
853  if (after === null || after === '') return ''
854  if (!after.startsWith('/')) return after
855  const rel = normalizePath(after, rootDir)
856  return rel.startsWith('/') || rel.startsWith('..') ? '' : rel
857}
858
859const GUARD_ERROR =
860  'Temper: the guard hit an error while it checked this command, and a run is active. Temper does not pass a command it could not check. ' +
861  'Next: run the command again. If it keeps failing, ask the user (the user can turn enforcement off with /temper:temper enforcement off).'
862
863// A guard that threw: Bash is refused while a run is active (a command the guard cannot read is the one to refuse),
864// every other tool passes (the mod must not make Temper worse than without it).
865function failedGuard(tool: string): Guarded {
866  const r = lastRun
867  const active = r !== null && !r.inert && r.enforcement !== 'off' && !r.sync.failOpen && r.state.phase !== null && r.state.phase !== 'done'
868  return tool === 'Bash' && active ? { deny: GUARD_ERROR, ids: [] } : { deny: null, ids: [] }
869}
870
871// The one place this module denies.
872async function guard($: Api, tool: string, input: Record<string, unknown>): Promise<Guarded> {
873  let snap = await ensure($)
874  // Fail open: when the mod cannot tell where the run is, it never blocks Temper (the mod must not make
875  // Temper worse than without it).
876  if (snap.inert || snap.enforcement === 'off' || snap.sync.failOpen) return { deny: null, ids: [] }
877  const io = makeIo($)
878  const command = typeof input.command === 'string' ? input.command : ''
879  const cls = tool === 'Bash' ? classifyBash(command, bashCwd) : null
880  // What this session staged with `git add` so far: a commit of spec files only is the artifact chain.
881  if (cls) {
882    staged.all = staged.all || cls.staged.all
883    staged.paths.push(...cls.staged.paths)
884  }
885  if (cls?.commits) {
886    // The commit gate reads the CLI's latest verdict, so reload before deciding.
887    snap = adopt($, await syncCheck(io, options, await refresh($), false))
888  }
889  // A `state clear` at Intent (the TRIVIAL exit) is decided on the run as it is now.
890  let clearable = false
891  if (cls?.stateOps.some(op => op.op === 'clear') && snap.state.phase === 'intent') {
892    snap = adopt($, await refresh($))
893    clearable = snap.state.phase === 'intent' && (await nothingToLose(io, snap.specDir))
894  }
895  const root = await rootOf($)
896  const commit = cls?.commits ? await commitFacts(io, snap, root, staged) : undefined
897  // From here to the reservation there is no await: a parallel call cannot slip in between. A
898  // decision a running call has reserved is not offered to this one.
899  const ctx = { ...ruleContext(snap, root), cwd: bashCwd, loopedDecisions: [...loopedDecisions], ...(commit ? { commit } : {}), ...(clearable ? { nothingToLose: true } : {}) }
900  const r = evaluate(snap.state, { ...ctx, humanDecisions: (ctx.humanDecisions ?? []).filter(decision => !reservedDecisions.has(decision.id)) }, { tool, input })
901  // A commit that went through starts the next staging from nothing. If the commit then fails (an index lock,
902  // a hook), the files are still staged: tool.call gives the list back (found live: the retry was refused).
903  let undoStaged: Guarded['undoStaged']
904  if (cls?.commits && !('deny' in r)) {
905    undoStaged = { all: staged.all, paths: [...staged.paths] }
906    staged.all = false
907    staged.paths = []
908  }
909  if ('deny' in r) return { deny: r.drift ? await resolveDrift($, snap, r.drift, r.deny) : r.deny, ids: [] }
910  // The command will run: the shell's folder is where it leaves it.
911  if (cls) bashCwd = carryCwd(cls.cwdAfter, root)
912  // A `state loop` call used the person's back decision once: a second one needs another decision.
913  const loopIds = r.loopIds ?? []
914  for (const id of loopIds) loopedDecisions.add(id)
915  if (r.consume === 'drift' && r.driftPath) {
916    adopt($, (await apply(io, options, snap, { type: 'useDrift', path: r.driftPath, origin: 'system' })).snap)
917  } else if (r.consume && r.consume !== 'drift' && r.eventIds) {
918    // Every human event a chained command matched is reserved at once, so a parallel call cannot use
919    // it too. It is spent after the call ran, also when the call ended in an error but changed the CLI's files (see
920    // tookEffect). A call that failed and changed nothing, or never ran, gives the decision back, and the person's
921    // choice stays pending (key 1 records it again).
922    for (const id of r.eventIds) reservedDecisions.add(id)
923    return { deny: null, ids: r.eventIds, ...(undoStaged ? { undoStaged } : {}), ...(loopIds.length > 0 ? { loopIds } : {}), fingerprint: await runFingerprint(io).catch(() => ''), ...(clearable ? { trivialExit: true } : {}) }
924  }
925  return { deny: null, ids: [], ...(undoStaged ? { undoStaged } : {}), ...(loopIds.length > 0 ? { loopIds, fingerprint: await runFingerprint(io).catch(() => '') } : {}), ...(clearable ? { trivialExit: true } : {}) }
926}
927
928// After the TRIVIAL exit ran (or failed): the cleared run is no longer held, and the run is read again. When the clear
929// did not take effect, build-state.json is still there and the run reads back as it was.
930async function afterTrivialExit($: Api, g: Guarded): Promise<void> {
931  if (!g.trivialExit) return
932  lastRun = null
933  await refresh($)
934}
935
936// After a call that ended in an error: did it change the CLI's files all the same (`state advance ...; exit 1`)?
937// The decision it used is then spent, as for a call that ended well. The exit status says nothing about the effect.
938async function tookEffect($: Api, g: Guarded): Promise<boolean> {
939  if (g.fingerprint === undefined || g.fingerprint === '') return false
940  const after = await runFingerprint(makeIo($)).catch(() => g.fingerprint)
941  return after !== g.fingerprint
942}
943
944// A `state loop` call that failed with no effect gives the back decision's one use back.
945function releaseLoops(g: Guarded, failed: boolean): void {
946  if (failed) for (const id of g.loopIds ?? []) loopedDecisions.delete(id)
947}
948
949// A commit call that failed did not use the staged files: the list the guard cleared comes back.
950function restoreStaged(g: Guarded, failed: boolean): void {
951  if (!failed || !g.undoStaged) return
952  staged.all = staged.all || g.undoStaged.all
953  staged.paths = [...g.undoStaged.paths, ...staged.paths]
954}
955
956// After a call that took decisions: spend them when it succeeded, give them back when it failed.
957async function settle($: Api, ids: string[], failed: boolean): Promise<void> {
958  const io = makeIo($)
959  if (failed) {
960    for (const id of ids) reservedDecisions.delete(id)
961  } else {
962    // An id stays in `reservedDecisions` once spent: a call that read its snapshot before the spend
963    // must still not use the event again.
964    for (const id of ids) await consumeDecision(io, id)
965  }
966  await refresh($)
967}
968
969// The files `git add` staged in this session (see guard). Unknown at the start: not an artifact commit.
970const staged: { all: boolean; paths: string[] } = { all: false, paths: [] }
971
972// Decision events a tool call has taken (see guard). An id stays here once spent: event ids are
973// unique, and a call that read its snapshot before the spend must still not use the event again.
974const reservedDecisions = new Set<string>()
975// `/temper:temper <reserved word>`: null means "not mine", and the prompt based command runs.
976async function handleTemper($: Api, parsed: Parsed, originKind: string): Promise<CommandRunResult | null> {
977  const snap = await ensure($)
978  if (snap.inert) return null
979  // Every word that changes state needs the person's own composer: the decisions, pause and
980  // resume, and a mode or enforcement change. Checked before anything else, so a refusal never
981  // depends on the arguments. The read only words (status, timeline, help, report) and
982  // showing the current mode stay open to any origin.
983  // With enforcement off the person has switched the guard off, so a decision word is accepted from any origin (a
984  // headless run) and its event keeps that origin. Pause, resume, mode and enforcement stay the person's own.
985  const changes: readonly string[] = [...DECISION_WORDS, 'pause', 'resume']
986  const changesState = changes.includes(parsed.word) || ((parsed.word === 'mode' || parsed.word === 'enforcement') && parsed.rest.trim() !== '')
987  const fromPerson = originKind === 'composer'
988  const unattended = !fromPerson && snap.enforcement === 'off' && DECISION_WORDS.includes(parsed.word)
989  if (changesState && !fromPerson && !unattended) return { text: NEEDS_INTERACTIVE }
990
991  const plan = planCommand(parsed, pendingDrift)
992  if (plan.kind === 'error') return { text: plan.text }
993
994  if (plan.kind === 'local') {
995    switch (plan.word) {
996      case 'help':
997        return { text: HELP }
998      case 'status':
999        return { text: statusText(await refresh($)) }
1000      case 'timeline':
1001        return { text: timelineText(await refresh($)) }
1002      case 'report': {
1003        // Decided on a fresh read of the run, not the cached one: the Commit steps' state clear, or a run
1004        // started since, changes nothing the cache sees until a turn ends with an answer.
1005        const fresh = await refresh($)
1006        if (fresh.slug === null) {
1007          // A finished run's Commit steps clear its run state; the report it kept is still shown.
1008          const kept = await keptReport(makeIo($))
1009          return { text: kept === null ? 'No run is active. There is no report to show.' : `No run is active. The last report kept for this project:\n\n${kept}` }
1010        }
1011        // A run held from memory (its state file is gone) is shown but not kept, so its report never
1012        // replaces the one a finished run kept.
1013        if (fresh.sync.line === LOST_STATE) return { text: `${LOST_STATE}\n\n${reportText(fresh)}` }
1014        return { text: await writeReport(makeIo($), fresh) }
1015      }
1016      case 'pr':
1017      case 'discuss':
1018      case 'continue':
1019        // Claude writes the description, answers the message, or does the On Continue steps of a stage the
1020        // person already approved (the decision event exists): the prompt based command handles it. The
1021        // mod records nothing here.
1022        return null
1023      case 'play':
1024        // Only the person opens the game. It works in every mode, because the person asked.
1025        if (originKind !== 'composer') return { text: 'Only the user can open the game. Next: ask the user to run /temper:temper play.' }
1026        return { text: await toggleGame($) }
1027      case 'mode': {
1028        const wanted = plan.rest.trim().toLowerCase()
1029        if (wanted === '') return { text: interactive ? await askMode($) : `Temper mode: ${snap.mode}` }
1030        if (wanted !== 'full' && wanted !== 'minimal' && wanted !== 'off') return { text: 'Usage: /temper:temper mode <full|minimal|off>' }
1031        return { text: await switchMode($, parseUiMode(wanted)) }
1032      }
1033      case 'enforcement': {
1034        const wanted = plan.rest.trim().toLowerCase()
1035        if (wanted === '') return { text: `Temper enforcement: ${snap.enforcement}` }
1036        if (wanted !== 'on' && wanted !== 'off') return { text: 'Usage: /temper:temper enforcement <on|off>' }
1037        return { text: await switchEnforcement($, wanted) }
1038      }
1039      default:
1040        if (snap.mode !== 'full') return { text: 'The pane shows in full mode only. Use /temper:temper mode full.' }
1041        return { text: await togglePane($) }
1042    }
1043  }
1044
1045  // A decision: only the person's own composer creates one. The event is the record the
1046  // commit gate and the rules trust; Claude's part (mirroring it in the CLI, continuing
1047  // the phase) is the prompt based /temper:temper, which runs next. prompt.submit is not used
1048  // here: the engine refuses it from inside command.run.
1049  const done = await decideAs($, plan.command, fromPerson ? 'person' : 'model', undefined, unattended ? originKind || 'unknown' : undefined)
1050  if (done.error) return { text: done.error }
1051  if (plan.command.type === 'pause' || plan.command.type === 'resume') {
1052    return { text: `The run is ${plan.command.type === 'pause' ? 'paused. You have control' : 'resumed'}.` }
1053  }
1054  return null
1055}
1056
1057async function afterGateCheck($: Api): Promise<void> {
1058  adopt($, await syncCheck(makeIo($), options, await refresh($), true))
1059}
1060
1061const REVIEWERS = ['temper-review', 'temper:temper-review']
1062
1063// Whether each subagent seen at turn.step is the Temper review agent, by its id. The agent list is read once per id.
1064const reviewerIds = new Map<string, boolean>()
1065
1066async function isReviewer($: Api, agentId: string): Promise<boolean> {
1067  const known = reviewerIds.get(agentId)
1068  if (known !== undefined) return known
1069  const agents = await $.agent.list()
1070  const found = agents.find(a => a.id === agentId)
1071  if (found === undefined) return false
1072  const yes = REVIEWERS.includes(found.type)
1073  reviewerIds.set(agentId, yes)
1074  return yes
1075}
1076
1077// The model and effort a step runs with: on the main loop from the phaseModels option, in the Temper review agent
1078// from the reviewerModel option. Null leaves the step exactly as the engine built it. The spawn of an agent is never
1079// changed: the reviewer model applies to the steps of that agent only.
1080async function phasePick($: Api, agentId: string | undefined): Promise<{ model?: string; effort?: 'low' | 'medium' | 'high' | 'xhigh' | 'max' } | null> {
1081  if (agentId !== undefined) {
1082    const model = typeof options.reviewerModel === 'string' ? options.reviewerModel.trim() : ''
1083    return model && (await isReviewer($, agentId)) ? { model } : null
1084  }
1085  const raw = options.phaseModels
1086  if (typeof raw !== 'string' || raw.trim() === '') return null
1087  const snap = await ensure($)
1088  if (snap.inert || snap.state.phase === null || snap.state.phase === 'done') return null
1089  const pick = parsePhaseModel(parsePhaseModels(raw)[snap.state.phase])
1090  return pick.model || pick.effort ? pick : null
1091}
1092
1093// Wiring only. Every hook fails open: an exception passes the call through, except the
1094// detected violation, which is the one place this module denies.
1095export const register: Register = (on, opts) => {
1096  options = opts
1097  current = null
1098  root = ''
1099  lastRun = null
1100  rootVerified = false
1101  sessionCwd = ''
1102  bashCwd = ''
1103  loopedDecisions.clear()
1104  pendingDrift = null
1105  lastMoveAt = 0
1106  pressing.clear()
1107  staged.all = false
1108  staged.paths = []
1109  interactive = false
1110  paneOpen = false
1111  lastPhase = undefined
1112  gameOpen = false
1113  gameBest = 0
1114  gameOver = false
1115  gameBanner = null
1116  working = false
1117  drawSurface = null
1118  live.mode = undefined
1119  live.enforcement = undefined
1120  live.paneExpanded = undefined
1121  reviewerIds.clear()
1122
1123  on('session.start', async ($, e, next) => {
1124    await settleRoot($, e.cwd)
1125    current = null
1126    drawSurface = e.surface
1127    interactive = e.isInteractive
1128    reservedDecisions.clear()
1129    await refresh($)
1130      .then(snap => autoOpenPane($, snap))
1131      .catch(() => undefined)
1132    return next(e)
1133  })
1134
1135  on('classic.SessionStart', async ($, e, next) => {
1136    await settleRoot($, e.cwd)
1137    await refresh($).catch(() => undefined)
1138    return next(e)
1139  })
1140
1141  // Only the reserved first words of /temper:temper are handled here; anything else (a feature
1142  // description) goes on to the prompt based command unchanged. Bare /temper:temper toggles the
1143  // pane while a run is active.
1144  on('command.run', async ($, e, next) => {
1145    if (e.command !== 'temper' && e.command !== 'temper:temper') return next(e)
1146    const parsed = parseArgs(e.args)
1147    // Only the person's own composer counts here: a command from any other origin never marks
1148    // the question answered and never opens the dialog. /temper:temper mode handles the
1149    // question itself (an explicit mode applies at once, a bare one asks once) and either way
1150    // counts as the first run answer.
1151    const isPerson = e.origin?.kind === 'composer'
1152    if (isPerson) {
1153      if (parsed?.word === 'mode') await markModeAsked($, parsed.rest.trim() !== '').catch(() => undefined)
1154      else await firstRunAsk($).catch(() => undefined)
1155    }
1156    if (parsed === null) {
1157      const snap = await ensure($).catch(() => null)
1158      const isBare = e.args.trim() === ''
1159      // Only the person's own bare /temper:temper toggles the pane. The same words from a plugin (the
1160      // button that launches a stage) go on to the orchestrator.
1161      if (isBare && isPerson && snap && !snap.inert && snap.mode === 'full' && snap.state.phase !== null && snap.state.phase !== 'done') {
1162        return { text: await togglePane($) }
1163      }
1164      return next(e)
1165    }
1166    const out = await handleTemper($, parsed, e.origin?.kind ?? '').catch(() => null)
1167    return out ?? next(e)
1168  })
1169
1170  // One Temper line on a generated pull request description, while a run is on and the
1171  // prAttribution option is on.
1172  on('attribution.text', { kind: 'pr' }, async ($, e, next) => {
1173    const result = await next(e)
1174    try {
1175      const snap = await ensure($)
1176      if (snap.inert || snap.prAttribution !== 'on' || snap.slug === null) return result
1177      return { text: `${result.text}\n\nMade with Temper. The phases have gates. /temper:temper report shows the report.` }
1178    } catch {
1179      return result
1180    }
1181  })
1182
1183  // Subagent calls arrive here too (e.agentId names the loop); they are held to the same
1184  // phase rules as the main loop.
1185  on('tool.call', async ($, e, next) => {
1186    const g = await guard($, e.tool, { ...e }).catch((): Guarded => failedGuard(e.tool))
1187    if (g.deny !== null) return { deny: g.deny }
1188    let result
1189    try {
1190      result = await next(e)
1191    } catch (err) {
1192      const failed = !(await tookEffect($, g).catch(() => false))
1193      await settle($, g.ids, failed).catch(() => undefined)
1194      releaseLoops(g, failed)
1195      restoreStaged(g, true)
1196      await afterTrivialExit($, g).catch(() => undefined)
1197      throw err
1198    }
1199    const errored = (result as { isError?: boolean }).isError === true
1200    restoreStaged(g, errored)
hooks/temper-mod/adapter.ts 604 lines
1// Adapter helpers: everything that touches Claude Code's `$` (files, store, state,
2// version) so register.tsx can stay wiring only. The pure rules live in core/.
3//
4// Where state lives (mods-plan 3.3): phase history is append only event files under
5// `.temper/specs/{slug}/events/`, one file per event, written once with a unique name.
6// Verdicts and criteria status are read, never written, from `.temper/gates.json` and
7// `.temper/status.json`. The folded RunState is cached in this module and mirrored to
8// `$.state` so drawing and compaction see it.
9
10import type { PluginOptions } from 'claude-code'
11
12import { parseEnforcement, parseMaxLoops, parseOnOff, parseUiMode, readConfigValue, versionAtLeast } from './core/config'
13import type { UiMode } from './core/config'
14import { mergeCriteria, parseCriteria, parseStatus, parseTitle, progress } from './core/criteria'
15import type { MergedCriterion } from './core/criteria'
16import { encodeEvent, eventFileName, readEvents, stamp } from './core/events'
17import type { Draft, Phase, TemperEvent } from './core/events'
18import { stageOf } from './core/cli'
19import { cliPhase, parseBuildState, parseFindings, parseGates, phaseFromStage, samePhase } from './core/gates'
20import type { Finding } from './core/gates'
21import { buildView } from './core/view'
22import type { View } from './core/view'
23import { decide, initialState, phaseLabel, reduce } from './core/machine'
24import type { Command, RunState, Verdicts } from './core/machine'
25import { planFileList, taskProgress, tasksLeft } from './core/planfiles'
26import { renderReport } from './core/report'
27import { sectionText } from './core/section'
28import type { DecisionKind } from './core/bash'
29import type { CommitFacts, HumanDecision } from './core/rules'
30import { normalizePath } from './core/paths'
31import type { TemperRun } from '../../types'
32
33// Everything the adapter needs from Claude Code, as plain functions. register.tsx builds
34// one from `$` (the module scan only follows `$` inside the file that spells it), which
35// also lets this file be tested without an engine.
36export type Io = {
37  // File text, or null when the file is missing or unreadable.
38  read: (path: string) => Promise<string | null>
39  list: (path: string) => Promise<Array<{ name: string; kind: string }>>
40  write: (path: string, text: string) => Promise<void>
41  storeGet: (key: string) => Promise<unknown>
42  storeSet: (key: string, value: unknown) => Promise<void>
43  version: () => Promise<string | undefined>
44  setRun: (run: TemperRun) => Promise<void>
45  // The live mode, mirrored so a redraw sees a change at once.
46  setMode: (mode: UiMode) => Promise<void>
47  // Waits ms milliseconds (the engine's clock: the mod reads no global timer).
48  pause: (ms: number) => Promise<void>
49  // Where the Temper script is (pluginCliFrom): the button prompts of the published view name the CLI by it.
50  // Left out: the plain `scripts/temper`.
51  cli?: string
52}
53
54// A choice of the person (a move) that no mirror call has recorded in the CLI yet.
55export type PendingMove = { id: string; draft: Draft }
56
57// How the mod's picture of the run compares with the CLI state (build-state.json). The CLI is the truth
58// for WHERE the run is: the phase of the folded state is always derived from it. Events only say WHO decided.
59export type Sync = {
60  // The phase the CLI is at; null when the mod cannot tell.
61  cli: Phase | 'done' | null
62  // The one line shown under the bar when something does not agree; null when all is well.
63  line: string | null
64  // The person's move that is not mirrored yet (key 1 records it again). Reused, never recreated.
65  pending: PendingMove | null
66  // The mod cannot tell where the run is: it blocks nothing and writes nothing.
67  failOpen: boolean
68  // The CLI looks reset (it is earlier than checks that passed): phase rules do not block writes.
69  looksReset: boolean
70}
71
72export const NO_SYNC: Sync = { cli: null, line: null, pending: null, failOpen: false, looksReset: false }
73
74export type Snapshot = {
75  // True below the minimum Claude Code version or when it cannot be read: nothing acts.
76  inert: boolean
77  sync: Sync
78  enforcement: 'on' | 'off'
79  mode: UiMode
80  prAttribution: 'on' | 'off'
81  slug: string | null
82  // autonomy.enabled in .claude/temper.config.
83  autonomyEnabled: boolean
84  // phases.design: true in .claude/temper.config (design is switched on for medium and complex runs).
85  designRequired: boolean
86  // The run's complexity from build-state.json (decides whether Plan is followed by design).
87  complexity: string | null
88  // The CLI's next stage, as written in build-state.json (the mod's phases are coarser: design is Plan).
89  cliNext: string | null
90  // The run's own branch and the command that started it (build-state.json): the CLI commit gate's
91  // Build checkpoint carve-out reads both.
92  branch: string | null
93  runCommand: string | null
94  // intent.md exists in the spec folder (the commit gate then wants an intent verdict).
95  hasIntent: boolean
96  specDir: string
97  state: RunState
98  verdicts: Verdicts
99  planFiles: string[]
100  title: string | null
101  criteria: MergedCriterion[]
102  findings: Finding[]
103  // "task N of M": see taskProgress in core/planfiles.ts (ticked rows in tasks.md, or
104  // the numeric `task` in build-state.json when the orchestrator sets one).
105  task: { n: number; of: number } | null
106  // Tasks of tasks.md that are not done; null when tasks.md has no task headings.
107  tasksLeft: number | null
108  unreadable: string[]
109  // Check wrote config-suggestions.json in the spec folder (offers "Review config suggestions").
110  configSuggestions: boolean
111  // Unconsumed human decision events per kind, for the CLI decision guard.
112  humanDecisions: HumanDecision[]
113}
114
115const OWN_PREFIX = 'ev:'
116const USED_PREFIX = 'used:'
117const STATE_ROOT = '.temper'
118const REPORT_PATH = `${STATE_ROOT}/report.md`
119
120const str = (options: PluginOptions, key: string): string | undefined => {
121  const v = options[key]
122  return typeof v === 'string' ? v : undefined
123}
124
125// Values changed live (`/temper:temper mode`, `/temper:temper enforcement`) win over the options of this
126// load until the config change reloads the module with the new options.
127export const live: { mode?: UiMode; enforcement?: 'on' | 'off'; paneExpanded?: boolean; paneOpen?: boolean } = {}
128
129export function settingsFrom(options: PluginOptions) {
130  return {
131    mode: live.mode ?? parseUiMode(str(options, 'uiMode')),
132    enforcement: live.enforcement ?? parseEnforcement(str(options, 'enforcement')),
133    prAttribution: parseOnOff(str(options, 'prAttribution'), 'on'),
134  }
135}
136
137export function idleSnapshot(options: PluginOptions, inert: boolean): Snapshot {
138  return {
139    inert,
140    sync: NO_SYNC,
141    ...settingsFrom(options),
142    slug: null,
143    autonomyEnabled: false,
144    designRequired: false,
145    complexity: null,
146    cliNext: null,
147    branch: null,
148    runCommand: null,
149    hasIntent: false,
150    specDir: '',
151    state: initialState(parseMaxLoops('', str(options, 'fixMaxLoops'))),
152    verdicts: {},
153    planFiles: [],
154    title: null,
155    criteria: [],
156    findings: [],
157    task: null,
158    tasksLeft: null,
159    unreadable: [],
160    configSuggestions: false,
161    humanDecisions: [],
162  }
163}
164
165const readText = (io: Io, path: string): Promise<string | null> => io.read(path).catch(() => null)
166
167async function readEventFiles(io: Io, dir: string): Promise<Array<{ name: string; text: string }>> {
168  const entries = await io.list(dir).catch(() => [])
169  const out: Array<{ name: string; text: string }> = []
170  for (const entry of entries) {
171    if (entry.kind !== 'file' || !entry.name.endsWith('.json')) continue
172    const text = await readText(io, `${dir}/${entry.name}`)
173    out.push({ name: entry.name, text: text ?? '' })
174  }
175  return out
176}
177
178// A decision event the person made that no CLI call has matched yet.
179function humanKind(ev: TemperEvent): DecisionKind | null {
180  if (ev.origin !== 'person') return null
181  if (ev.type === 'override') return 'override'
182  if (ev.type === 'accept') return 'accept'
183  // Every advance the person makes pays for its own mirror call (`state advance <stage>_complete`):
184  // the guard checks every advance since the third review, so every one needs its event in the pool.
185  // Fix has no CLI stage of its own, so a move out of Fix has nothing to mirror.
186  if (ev.type === 'advance' && ev.from !== 'fix') return 'advance'
187  if (ev.type === 'back') return 'back'
188  return null
189}
190
191// A digest of the exact text of an event file. The store keeps it under `ev:{id}`, so trust
192// follows the content, not just the name: rewriting a trusted file in place (same name,
193// different text) makes it untrusted. SHA-256 where the environment has it, else FNV-1a.
194export async function digestText(text: string): Promise<string> {
195  try {
196    const buf = await crypto.subtle.digest('SHA-256', new TextEncoder().encode(text))
197    return 'sha256:' + [...new Uint8Array(buf)].map(b => b.toString(16).padStart(2, '0')).join('')
198  } catch {
199    let h = 0x811c9dc5
200    for (let i = 0; i < text.length; i++) h = Math.imul(h ^ text.charCodeAt(i), 0x01000193) >>> 0
201    return 'fnv:' + h.toString(16)
202  }
203}
204
205async function isOwn(io: Io, id: string, text: string): Promise<boolean> {
206  try {
207    return (await io.storeGet(OWN_PREFIX + id)) === (await digestText(text))
208  } catch {
209    return false
210  }
211}
212
213// Event names are `{ts}-{session}-{seq}.json`. This module has no session id call, so a
214// random id per load stands in: two loads never share a name.
215const SESSION = (Math.random().toString(16).slice(2) + '00000000').slice(0, 8)
216// The phase an event decided, so a CLI call is matched only to a decision made for it.
217function decisionOf(ev: TemperEvent, kind: DecisionKind): HumanDecision {
218  // The CLI names Fix by its Check stage, as the follow up command does.
219  if (ev.type === 'override') return { id: ev.id, kind, phase: stageOf(ev.phase) }
220  if (ev.type === 'back') return { id: ev.id, kind, phase: stageOf(ev.to) }
221  if (ev.type === 'accept') return { id: ev.id, kind, phase: 'review', findingId: ev.findingId }
222  if (ev.type === 'advance') return { id: ev.id, kind, phase: ev.from }
223  return { id: ev.id, kind }
224}
225
226const FLOW_ORDER: readonly Phase[] = ['intent', 'plan', 'build', 'review', 'check']
227
228// Makes the phase the CLI's. The events say who decided what; they never say where the run is. When
229// the mod's phase and the CLI's differ, the CLI wins, for the display and for every deny:
230//  - the person's last move is not mirrored yet: the CLI phase stays, one line says so, key 1 records it;
231//  - the CLI moved ahead of the events (the orchestrator advanced it): follow it, write nothing;
232//  - the CLI name is not one the mod knows: fail open, say so, block nothing;
233//  - the CLI looks reset (earlier than a check that passed): do not block work those checks passed.
234export function reconcile(state: RunState, nextStage: string | null, pending: PendingMove | null): { state: RunState; sync: Sync } {
235  if (!state.started) return { state, sync: NO_SYNC }
236  const cli = cliPhase(nextStage)
237  if (cli === null) {
238    return {
239      state,
240      sync: { cli: null, line: 'Temper state: the mod cannot tell where the run is. It does not block anything until the state can be read.', pending: null, failOpen: true, looksReset: false },
241    }
242  }
243  let next = state
244  let line: string | null = null
245  let held: PendingMove | null = null
246  let behind = false
247  if (!samePhase(state.phase, cli)) {
248    const ahead = state.phase === null ? -1 : state.phase === 'done' ? 5 : state.phase === 'fix' ? 4 : FLOW_ORDER.indexOf(state.phase)
249    const at = cli === 'done' ? 5 : cli === 'fix' ? 4 : FLOW_ORDER.indexOf(cli)
250    // The CLI is earlier than the person's own choices, with nothing waiting to be mirrored and no skip
251    // that the orchestrator has yet to advance past: nobody asked for that. It looks reset.
252    // Design is part of Plan in the mod: the person approved the plan, and the CLI is at its design stage.
253    // That is the run going on (found live), not a reset.
254    const designing = nextStage === 'design' && state.phase === 'build'
255    behind = at >= 0 && ahead > at && pending === null && !designing && state.history[state.history.length - 1]?.kind !== 'override'
256    next = { ...state, phase: cli, since: { ...state.since, ...(cli !== 'done' && state.since[cli] === undefined ? { [cli]: 0 } : {}) } }
257    next.loopLimitReached = false
258    if (pending !== null) {
259      held = pending
260      line = `Temper state: the run is at ${phaseLabel(cli)}. Your last choice is not recorded yet. Press 1 to record it.`
261    }
262  }
263  // A check that passed later than the phase the CLI is at, or choices that went further than the CLI
264  // says with nothing pending: the CLI state looks reset.
265  let looksReset = behind
266  if (cli !== 'done' && cli !== 'fix') {
267    const at = FLOW_ORDER.indexOf(cli)
268    looksReset = looksReset || (at >= 0 && FLOW_ORDER.some((p, i) => i > at && next.gate[p] === 'fresh'))
269  }
270  if (looksReset && line === null) line = `Temper state looks reset: the run is at ${phaseLabel(cli)}, but it went further before. Temper does not block work that already passed.`
271  return { state: next, sync: { cli, line, pending: held, failOpen: false, looksReset } }
272}
273
274const BOOTSTRAP = { ts: 0, session: 'bootstrap', seq: 1 }
275let seq = 0
276
277export async function loadSnapshot(io: Io, options: PluginOptions): Promise<Snapshot> {
278  const version = await io.version().catch(() => undefined)
279  if (!versionAtLeast(version)) return idleSnapshot(options, true)
280
281  const cfg = settingsFrom(options)
282  const configText = (await readText(io, '.claude/temper.config')) ?? ''
283  const maxLoops = parseMaxLoops(configText, str(options, 'fixMaxLoops'))
284  const autonomyEnabled = readConfigValue(configText, 'autonomy.enabled') === 'true'
285  const designRequired = readConfigValue(configText, 'phases.design') === 'true'
286  const statePath = `${STATE_ROOT}/build-state.json`
287  let raw = await readText(io, statePath)
288  let bs = parseBuildState(raw ?? '')
289  // The CLI rewrites build-state.json in place, so a read in that instant sees an empty or cut file (found live: the
290  // bar said "No Temper run is active" at Done). A file that exists but does not read as a run is read again
291  // before the mod decides there is no run. A file that is missing is a run that ended.
292  for (let i = 0; i < 4 && bs === null && raw !== null; i++) {
293    await io.pause(60).catch(() => undefined)
294    raw = await readText(io, statePath)
295    bs = parseBuildState(raw ?? '')
296  }
297  if (!bs) return { ...idleSnapshot(options, false), state: initialState(maxLoops) }
298
299  const specDir = bs.specPath
300  const verdicts = parseGates((await readText(io, `${STATE_ROOT}/gates.json`)) ?? '')
301  const intentText = (await readText(io, `${specDir}/intent.md`)) ?? ''
302  const title = parseTitle(intentText) ?? bs.spec
303
304  const own = new Set<string>()
305  const fold = (events: readonly TemperEvent[]) => reduce(events, verdicts, { maxLoops, isTrusted: ev => own.has(ev.id) })
306
307  const files = await readEventFiles(io, `${specDir}/events`)
308  const textOf = new Map(files.map(f => [f.name, f.text]))
309  const read = readEvents(files)
310  const unreadable = read.unreadable
311  let events = read.events
312  // One store key per event id: an event file whose id is not here was not written by
313  // this mod, so it is unverified and never counts as an approval.
314  for (const ev of events) if (await isOwn(io, ev.id, textOf.get(eventFileName(ev)) ?? '')) own.add(ev.id)
315  let state = fold(events)
316
317  // A run the CLI began before the mod saw it: enter it once, at the phase the CLI
318  // recorded. The start event is written by the mod itself, so it is a trusted record.
319  // The bootstrap start has one fixed name (ts 0, session "bootstrap", seq 1), so entering a run
320  // again overwrites that file instead of adding another: a lost or unwritable store cannot pile
321  // up start files, and a planted or edited file under that name is replaced with a genuine one.
322  // A next_stage the CLI does not use gives no phase to start at: nothing is guessed (never Intent).
323  if (!state.started) {
324    const phase = cliPhase(bs.nextStage)
325    if (phase !== null && phase !== 'done') {
326      const start = await writeEvent(io, specDir, { type: 'start', slug: bs.spec, title, phase, origin: 'system' }, BOOTSTRAP)
327      events = [...events.filter(e => e.id !== start.id), start]
328      own.add(start.id)
329      state = fold(events)
330    }
331  }
332
333  const planText = (await readText(io, `${specDir}/plan.md`)) ?? ''
334  const tasksText = (await readText(io, `${specDir}/tasks.md`)) ?? ''
335  const status = parseStatus((await readText(io, `${STATE_ROOT}/status.json`)) ?? '')
336
337  // A decision belongs to the plan it approved: when the run later goes back to that stage or an earlier one (a step
338  // back the mod trusts), a decision for that stage that was never spent is stale. It is not an approval for the
339  // plan that is made again.
340  let open: Array<{ hd: HumanDecision; ev: TemperEvent; stage: number; stalable: boolean }> = []
341  for (const ev of events) {
342    if (ev.type === 'back' && own.has(ev.id)) {
343      const t = FLOW_ORDER.indexOf(ev.to === 'fix' ? 'check' : ev.to)
344      open = open.filter(o => !o.stalable || o.stage < t)
345    }
346    const kind = humanKind(ev)
347    if (kind && own.has(ev.id) && !(await io.storeGet(USED_PREFIX + ev.id))) {
348      const stage = ev.type === 'advance' ? ev.from : ev.type === 'override' ? ev.phase : ev.type === 'accept' ? 'review' : null
349      open.push({ hd: decisionOf(ev, kind), ev, stage: stage === null ? -1 : FLOW_ORDER.indexOf(stage === 'fix' ? 'check' : stage), stalable: kind !== 'back' })
350    }
351  }
352  const humanDecisions: HumanDecision[] = open.map(o => o.hd)
353  // A move of the person that no mirror call has recorded yet: the latest one is the one to record.
354  let pendingMove: PendingMove | null = null
355  for (const o of open) if (o.ev.type === 'advance' || o.ev.type === 'back' || o.ev.type === 'override') pendingMove = { id: o.ev.id, draft: o.ev }
356
357  const sync = reconcile(state, bs.nextStage, pendingMove)
358  state = sync.state
359
360  return {
361    inert: false,
362    sync: sync.sync,
363    ...cfg,
364    slug: bs.spec,
365    autonomyEnabled,
366    designRequired,
367    complexity: bs.complexity,
368    cliNext: bs.nextStage,
369    branch: bs.branch,
370    runCommand: bs.command,
371    hasIntent: intentText !== '',
372    specDir,
373    state,
374    verdicts,
375    planFiles: planFileList(planText, tasksText),
376    title,
377    criteria: mergeCriteria(parseCriteria(intentText), status),
378    findings: parseFindings((await readText(io, `${STATE_ROOT}/evidence/review.json`)) ?? ''),
379    task: taskProgress(tasksText, bs.task),
380    tasksLeft: tasksLeft(tasksText),
381    unreadable: unreadable.map(u => u.name),
382    configSuggestions: (await io.list(specDir).catch(() => [])).some(e => e.kind === 'file' && e.name === 'config-suggestions.json'),
383    humanDecisions,
384  }
385}
386
387// The CLI gate wants a design verdict (PASS, or a person's override of design) exactly when the spec has a design.md
388// (scripts/temper gate_commit, the checkpoint carve-out). Without a design.md nothing is asked.
389async function designSatisfied(io: Io, snap: Snapshot): Promise<boolean> {
390  if ((await readText(io, `${snap.specDir}/design.md`)) === null) return true
391  try {
392    const gates = JSON.parse((await readText(io, `${STATE_ROOT}/gates.json`)) ?? '{}') as { design?: { verdict?: unknown } }
393    if (gates.design?.verdict === 'PASS') return true
394  } catch {
395    // an unreadable gates.json holds no verdict
396  }
397  try {
398    const rows = JSON.parse((await readText(io, `${STATE_ROOT}/overrides.json`)) ?? '[]') as unknown
399    return Array.isArray(rows) && rows.some(r => typeof r === 'object' && r !== null && (r as { stage?: unknown }).stage === 'design')
400  } catch {
401    return false
402  }
403}
404
405// The spec files a run writes once it has an intent; a run that holds one of them has something to lose.
406const RUN_ARTIFACTS = ['intent.md', 'plan.md', 'tasks.md', 'design.md']
407
408// The TRIVIAL exit of the orchestrator: a `state clear` of a run that never left Intent. Read fresh for the call, so a
409// file written a moment ago counts. True only when the CLI's files and the spec folder all say so:
410//  - build-state.json: Intent is next, and its stage shows no completed step (`state init` writes `started`);
411//  - gates.json holds no verdict, and overrides.json is an empty list (each may also be missing);
412//  - the spec folder holds none of intent.md, plan.md, tasks.md and design.md, as a text or as a name in its listing.
413// A gates.json or overrides.json that does not read counts as missing (the guard refuses a command that removes or locks
414// either while a run is active); one that reads but is not what the CLI writes keeps the run. The mod's own record is
415// checked by the rule (core/rules.ts).
416export async function nothingToLose(io: Io, specDir: string): Promise<boolean> {
417  // The parsed text of a file, or undefined when it does not read.
418  const json = async (path: string): Promise<unknown> => {
419    const text = await readText(io, path)
420    return text === null ? undefined : JSON.parse(text)
421  }
422  try {
423    const bs = await json(`${STATE_ROOT}/build-state.json`)
424    if (typeof bs !== 'object' || bs === null) return false
425    const { stage, next_stage: next } = bs as { stage?: unknown; next_stage?: unknown }
426    if (next !== 'intent' || typeof stage !== 'string' || stage.includes('_complete')) return false
427    const gates = await json(`${STATE_ROOT}/gates.json`)
428    if (gates !== undefined) {
429      if (typeof gates !== 'object' || gates === null || Array.isArray(gates)) return false
430      if (Object.values(gates).some(row => typeof row === 'object' && row !== null && 'verdict' in row)) return false
431    }
432    const rows = await json(`${STATE_ROOT}/overrides.json`)
433    if (rows !== undefined && (!Array.isArray(rows) || rows.length > 0)) return false
434  } catch {
435    return false
436  }
437  for (const name of RUN_ARTIFACTS) if ((await readText(io, `${specDir}/${name}`)) !== null) return false
438  // A spec folder that is not there yet (the TRIVIAL run wrote nothing) lists as empty.
439  const entries = await io.list(specDir).catch(() => [])
440  return Array.isArray(entries) && !entries.some(e => RUN_ARTIFACTS.includes(e.name))
441}
442
443// The facts of the CLI commit gate that the mod can read (see CommitFacts in core/rules.ts). `root` is the
444// project folder; the current branch comes from .git/HEAD. Never throws: a fact that cannot be read is false.
445export async function commitFacts(io: Io, snap: Snapshot, root: string, staged: { all: boolean; paths: string[] }): Promise<CommitFacts> {
446  // Compared as the CLI gate compares (a case sensitive grep for ^.temper/specs/): `.TEMPER/SPECS` is not it.
447  const specs = !staged.all && staged.paths.length > 0 && staged.paths.every(p => normalizePath(p, root).startsWith('.temper/specs/'))
448  const dir = root.replace(/\/$/, '')
449  const head = (await readText(io, `${dir}/.git/HEAD`)) ?? ''
450  const cur = /^ref:\s*refs\/heads\/(.+?)\s*$/.exec(head)?.[1] ?? null
451  const s = snap.state
452  const satisfied = (p: 'intent' | 'plan') => snap.verdicts[p]?.verdict === 'PASS' || s.overrides.some(o => o.phase === p)
453  let checkpoint = false
454  let hint: string | undefined
455  if (snap.sync.cli === 'build' && snap.runCommand === 'temper' && snap.branch !== null) {
456    if (cur !== snap.branch) {
457      hint = `A Build checkpoint commit needs the run's branch ${snap.branch}, and you are on ${cur ?? 'no branch'}.`
458    } else if (satisfied('plan') && (!snap.hasIntent || satisfied('intent')) && (await designSatisfied(io, snap))) {
459      let rows: unknown = []
460      try {
461        rows = JSON.parse((await readText(io, `${STATE_ROOT}/evidence/build.json`)) ?? '[]')
462      } catch {
463        rows = []
464      }
465      const hits = (Array.isArray(rows) ? rows : []).filter(r => typeof r === 'object' && r !== null && (r as { phase?: unknown }).phase === 'green' && String((r as { claim?: unknown }).claim ?? '').toLowerCase().includes('test'))
466      const last = hits[hits.length - 1] as { exit_code?: unknown } | undefined
467      checkpoint = last !== undefined && Number(last.exit_code) === 0
468      if (!checkpoint) hint = 'A Build checkpoint commit needs a green test run on record.'
469    }
470  }
471  return { stagedSpecsOnly: specs, checkpoint, ...(hint ? { hint } : {}) }
472}
473
474// Writes one event file, once, under a unique name, and records a digest of its text in
475// `$.store` (one key per id) so a later load can tell its own events from files it did not
476// write, or from one that was rewritten afterwards.
477export async function writeEvent(io: Io, specDir: string, draft: Draft, fixed?: { ts: number; session: string; seq: number }): Promise<TemperEvent> {
478  seq += 1
479  const ev = stamp(draft, fixed ?? { ts: Date.now(), session: SESSION, seq })
480  const text = encodeEvent(ev)
481  await io.write(`${specDir}/events/${eventFileName(ev)}`, text)
482  await io.storeSet(OWN_PREFIX + ev.id, await digestText(text))
483  return ev
484}
485
486// The CLI's own files that a decision call changes (the state, the overrides, the ledger, the verdicts), as one text.
487// A call that ends in an error may still have done its work (`state advance ...; exit 1`): when this text differs
488// from the one taken before the call, the decision was used.
489export async function runFingerprint(io: Io): Promise<string> {
490  const parts: string[] = []
491  for (const name of ['build-state.json', 'overrides.json', 'gates.json', 'feedback-loops.json']) parts.push(`${name}:${(await readText(io, `${STATE_ROOT}/${name}`)) ?? '-'}`)
492  const entries = await io.list(`${STATE_ROOT}/evidence`).catch(() => [])
493  for (const e of entries.filter(x => x.kind === 'file').sort((a, b) => a.name.localeCompare(b.name))) parts.push(`evidence/${e.name}:${(await readText(io, `${STATE_ROOT}/evidence/${e.name}`)) ?? '-'}`)
494  return parts.join('\n')
495}
496
497// Marks one human event as matched by a CLI call, so it authorizes that call only once.
498export async function consumeDecision(io: Io, eventId: string): Promise<void> {
499  await io.storeSet(USED_PREFIX + eventId, 1)
500}
501
502export type Applied = { snap: Snapshot; error?: string; events: Draft[] }
503
504// Runs one command through the machine; on success writes its events, reloads the
505// snapshot, and keeps the report when the run reached Done.
506export async function apply(io: Io, options: PluginOptions, snap: Snapshot, command: Command): Promise<Applied> {
507  const decision = decide(snap.state, command)
508  if ('error' in decision) return { snap, error: decision.error, events: [] }
509  for (const draft of decision.events) await writeEvent(io, snap.specDir, draft)
510  const next = await loadSnapshot(io, options)
511  await publish(io, next)
512  if (next.state.phase === 'done' && snap.state.phase !== 'done') await writeReport(io, next)
513  return { snap: next, events: decision.events }
514}
515
516// The report of the run as it stands, not kept.
517export function reportText(snap: Snapshot): string {
518  return renderReport({
519    state: snap.state,
520    criteria: snap.criteria,
521    generatedAt: Date.now(),
522    unreadable: snap.unreadable,
523  })
524}
525
526// Keeps the report (the mod writes no file: see makeIo) and returns its text.
527export async function writeReport(io: Io, snap: Snapshot): Promise<string> {
528  const md = reportText(snap)
529  await io.write(REPORT_PATH, md)
530  return md
531}
532
533// The report kept last in this project, for `/temper:temper report` with no run active: a finished
534// run's Commit steps clear its run state, and its report stays readable. A report file that a 9.6.0
535// or 9.6.1 run wrote is read too. Null when there is none.
536export async function keptReport(io: Io): Promise<string | null> {
537  const text = await readText(io, REPORT_PATH)
538  return text !== null && text.trim() !== '' ? text : null
539}
540
541export function composeText(snap: Snapshot): string {
542  const s = snap.state
543  return sectionText({
544    enforcement: snap.enforcement,
545    phase: s.phase,
546    title: snap.title,
547    task: snap.task,
548    progress: snap.criteria.length > 0 ? progress(snap.criteria) : null,
549    paused: s.paused,
550    loopLimitReached: s.loopLimitReached,
551    stale: s.stale,
552    sync: snap.sync.line,
553    actionContext: {
554      ready: s.phase !== null && s.phase !== 'done' ? s.gate[s.phase] === 'fresh' : false,
555      allChecksPass: s.gate.check === 'fresh',
556      hasFindings: snap.findings.length > 0,
557    },
558  })
559}
560
561export const viewOf = (snap: Snapshot, cli?: string): View =>
562  buildView({ state: snap.state, title: snap.title, criteria: snap.criteria, findings: snap.findings, task: snap.task, tasksLeft: snap.tasksLeft, sync: snap.sync.line, pending: snap.sync.pending !== null, enforcement: snap.enforcement, configSuggestions: snap.configSuggestions, expanded: live.paneExpanded ?? false, paneOpen: live.paneOpen ?? false, ...(cli !== undefined ? { cli } : {}) })
563
564// Mirrors the folded state into `$.state` for drawing and compaction.
565export async function publish(io: Io, snap: Snapshot): Promise<void> {
566  await io.setRun({ slug: snap.slug, phase: snap.state.phase, title: snap.title, summary: composeText(snap), view: viewOf(snap, io.cli) })
567  await io.setMode(snap.mode)
568}
569
570// Check results are read from the CLI's verdict, never decided here. A fresh PASS ends
571// the run; a fresh FAIL (only when `includeFail`, after `temper gate check` ran) enters Fix.
572export async function syncCheck(io: Io, options: PluginOptions, snap: Snapshot, includeFail: boolean): Promise<Snapshot> {
573  if (snap.state.phase !== 'check') return snap
574  const gate = snap.state.gate.check
575  const result = gate === 'fresh' ? 'pass' : gate === 'fail' && includeFail ? 'fail' : null
576  if (result === null) return snap
577  return (await apply(io, options, snap, { type: 'checkResult', result, origin: 'system' })).snap
578}
579
580const iso = (ts: number): string => new Date(ts).toISOString()
581
582// `/temper:temper status`: the section text plus the counts a person wants at a glance.
583export function statusText(snap: Snapshot): string {
584  if (snap.state.phase === null) return 'No Temper run is active. Start one with /temper:temper <feature description>.'
585  const s = snap.state
586  const lines = [composeText(snap)]
587  lines.push(`Mode: ${snap.mode}, enforcement: ${snap.enforcement}`)
588  if (s.loops > 0) lines.push(`Check to Fix loops: ${s.loops} (limit ${s.maxLoops})`)
589  if (s.overrides.length > 0) lines.push(`Overrides: ${s.overrides.length}`)
590  if (s.accepted.length > 0) lines.push(`Accepted findings: ${s.accepted.length}`)
591  if (s.drift.length > 0) lines.push(`Scope drift decisions: ${s.drift.length}`)
592  if (s.unverified.length > 0) lines.push(`Event files that Temper does not trust: ${s.unverified.length}`)
593  if (snap.unreadable.length > 0) lines.push(`Unreadable event files: ${snap.unreadable.join(', ')}`)
594  return lines.join('\n')
595}
596
597export function timelineText(snap: Snapshot): string {
598  const h = snap.state.history
599  if (h.length === 0) return 'No phases yet.'
600  return h
601    .map(r => `${iso(r.ts)}  ${r.kind}: ${r.from ? phaseLabel(r.from) : 'start'} to ${phaseLabel(r.to)}`)
602    .join('\n')
603}
604
hooks/temper-mod/core/actions.ts 394 lines
1// Per-phase actions and hotkeys (mods-plan 3.7). One flow, two views: the orchestrator
2// (commands/temper.md) still runs the stages and the CLI still judges the gates. These actions are
3// the options the original orchestrator asks as questions, with the same words, as buttons. Nothing
4// else is a button: the person can still ask for anything else by typing (key 4, Discuss).
5//
6//   1  the main step: "Continue to <next>" once the check passed, "Loop back to <upstream>" when it
7//      failed, else "Start <phase>" / "Run <phase>" (the orchestrator runs the stage)
8//   2  3  two more original options of the phase (Grill me, Teach me, Walk through step by step ...)
9//   4  Discuss: the original "Other": the person types any message
10//   9  Skip with a reason (the original "Override and continue")
11//   0  More: the other original options as a numbered menu (1 to 9, 0 goes back)
12//
13// A label is a short phrase that says the result. Each action has one short line (`desc`, 10 words
14// at most) that the pane shows under the label. No internal word ("gate", "override") appears in a
15// label or a line. scripts/check-original-options.sh and tests/mod/actions.test.ts refuse a label that
16// is not an original option or one of the explicit extras.
17
18import type { Phase } from './events'
19import { CLI, IN_PLUGIN, pluginRootOf } from './cli'
20
21// 1 to 9 and 0 are digits because only a digit works from an empty prompt (a letter would type into
22// the composer). The More menu reuses the digits 1 to 9; 0 leaves the menu.
23export type ActionKey = '1' | '2' | '3' | '4' | '5' | '6' | '7' | '8' | '9' | '0'
24
25export type Action = {
26  key: ActionKey
27  id: string
28  label: string
29  // A shorter label for a narrow band (under 100 columns).
30  short?: string
31  // What the action does, in 10 words at most. The pane shows it under the label.
32  desc: string
33  // Text submitted to Claude as a prompt.
34  prompt?: string
35  // A reserved subcommand typed as `/temper:temper <command>`.
36  command?: string
37  asksReason?: boolean
38  // After the command is recorded and mirrored, run `/temper:temper` (no arguments) so the
39  // orchestrator launches the stage. Also set alone: the action only launches the stage.
40  resume?: boolean
41  // Put this draft in the prompt box and let the person type the rest (Discuss, Change).
42  fill?: string
43  // Record the person's last choice in the CLI again: the mirror prompt for the pending decision is
44  // submitted once more. No new decision is made.
45  record?: boolean
46}
47
48export type ActionContext = {
49  // The current phase's check passed (fresh PASS verdict).
50  ready?: boolean
51  // The state of the current phase's check: fresh, fail, stale (a step back made it old) or none.
52  gate?: 'fresh' | 'stale' | 'fail' | 'none'
53  tasksDone?: boolean
54  hasFindings?: boolean
55  allChecksPass?: boolean
56  loopLimitReached?: boolean
57  // The run is paused (saved for later).
58  paused?: boolean
59  // The task the Build phase works on, from tasks.md, and how many tasks are not done (null: unknown).
60  task?: { n: number; of: number } | null
61  tasksLeft?: number | null
62  // Check wrote config-suggestions.json in the spec folder.
63  configSuggestions?: boolean
64  // The person's last move is not recorded in the CLI yet (Snapshot.sync.pending).
65  pending?: boolean
66  // Where the Temper script is: its full path in the plugin folder, or the plain `scripts/temper` when that is not
67  // known (see pluginCliFrom). A prompt names the CLI and the plugin's files by it.
68  cli?: string
69}
70
71export type ActionSet = { primary: Action[]; discuss: Action; override: Action | null; more: Action[] }
72
73// Every prompt ends with one of these two, so Claude acts and does not narrate.
74export const ACT = 'Do this now. Reply with one short line.'
75export const SHOW = 'Do this now. Keep the answer short.'
76
77const prompt = (key: ActionKey, id: string, label: string, desc: string, text: string, tail: string = ACT, short?: string): Action => ({
78  key,
79  id,
80  label,
81  desc,
82  prompt: `${text} ${tail}`,
83  ...(short ? { short } : {}),
84})
85const command = (key: ActionKey, id: string, label: string, desc: string, cmd: string, asksReason = false, resume = false, short?: string): Action => ({
86  key,
87  id,
88  label,
89  desc,
90  command: cmd,
91  ...(asksReason ? { asksReason } : {}),
92  ...(resume ? { resume } : {}),
93  ...(short ? { short } : {}),
94})
95const launch = (key: ActionKey, id: string, label: string, desc: string): Action => ({ key, id, label, desc, resume: true })
96const draft = (key: ActionKey, id: string, label: string, desc: string, fill: string, short?: string): Action => ({ key, id, label, desc, fill, ...(short ? { short } : {}) })
97
98const OVERRIDE: Action = command('9', 'override', 'Skip with a reason', 'Skip this step. Temper writes your reason in the report.', 'override', true, true)
99
100// The original "Other": a free message about this step. It changes nothing by itself.
101export const DISCUSS: Action = draft('4', 'discuss', 'Discuss', 'Type your own message about this step.', 'Discuss this step: ')
102
103const cap = (p: string): string => `${p.charAt(0).toUpperCase()}${p.slice(1)}`
104
105// What follows each phase, and what the orchestrator loops back to when a check fails.
106const NEXT: Record<Phase, string> = { intent: 'Plan', plan: 'Build', build: 'Review', review: 'Check', check: 'Commit', fix: 'Check' }
107const UPSTREAM: Partial<Record<Phase, Phase>> = { plan: 'intent', build: 'plan', review: 'build' }
108
109const READY_DESC: Record<Phase, string> = {
110  intent: 'The intent is checked. Plan opens and Claude plans.',
111  plan: 'The plan is checked. Build opens and Claude starts building.',
112  build: 'All tasks are done. Review opens and Claude reviews.',
113  review: 'The review has no open problems. Check opens.',
114  check: 'All checks pass. You can then commit.',
115  fix: 'Run the checks again.',
116}
117
118// Build is gated at every checkpoint (one task per launch): a check that fails while tasks are still
119// open is no failure, only a checkpoint. The next step is the next task.
120const isCheckpoint = (phase: Phase, ctx: ActionContext): boolean => {
121  const left = ctx.tasksLeft
122  return phase === 'build' && !ctx.ready && Boolean(ctx.task) && (left === undefined ? ctx.gate !== 'fail' : (left ?? 0) > 0)
123}
124
125// Key 1, by what the check said. The orchestrator runs every stage and its check itself, so a person
126// arriving at a phase finds the result already there.
127function main(phase: Phase, ctx: ActionContext): Action {
128  // The person chose a move that the CLI has not recorded: key 1 records it again. It is not a new choice.
129  if (ctx.pending) return { key: '1', id: 'record', label: 'Record my choice', desc: 'Temper records your last choice again.', record: true }
130  const next = NEXT[phase]
131  if (ctx.ready) {
132    return command('1', 'continue', `Continue to ${next}`, READY_DESC[phase], phase === 'intent' || phase === 'plan' ? 'approve' : 'next', false, true, `To ${next}`)
133  }
134  if (isCheckpoint(phase, ctx) && ctx.task) {
135    return launch('1', 'run-stage', `Continue with task ${ctx.task.n}`, 'Claude builds the task and stops for you.')
136  }
137  const up = UPSTREAM[phase]
138  if (ctx.gate === 'fail' && up) return loopBack(up, '1')
139  if (phase === 'intent') return launch('1', 'run-stage', 'Start Intent', 'Claude writes the intent and checks it. Then you decide.')
140  return launch('1', 'run-stage', ctx.gate === 'fail' ? `Run ${cap(phase)} again` : `Run ${cap(phase)}`, `Claude runs ${cap(phase)} and checks it. Then you decide.`)
141}
142
143// The original "Loop back to {upstream}": it asks for a reason, records the step back, and the orchestrator
144// launches the earlier stage again. The loop budget (loops.max-per-type) is kept by the CLI.
145const loopBack = (to: Phase, key: ActionKey): Action =>
146  command(key, 'loop-back', `Loop back to ${cap(to)}`, `Redo ${cap(to)}. You give a reason.`, `back ${to}`, true, true, `Back to ${cap(to)}`)
147
148const noun = (phase: Phase): string => ({ intent: 'intent', plan: 'plan', build: 'work', review: 'review', check: 'checks', fix: 'fix' })[phase]
149
150const grill = (phase: Phase, key: ActionKey = '1'): Action =>
151  prompt(key, 'grill-me', 'Grill me', 'Claude asks hard questions about it.', `Use the grill-me skill on the current ${noun(phase)}.`, SHOW)
152const teach = (phase: Phase, key: ActionKey = '1'): Action =>
153  prompt(key, 'teach-me', 'Teach me', 'Claude explains it so you learn it.', `Use the teach-me skill on the current ${noun(phase)}.`, SHOW)
154// "Save for later" is the original pause. A paused run offers Resume in its place.
155const save = (paused: boolean, key: ActionKey = '1'): Action =>
156  paused ? command(key, 'resume', 'Resume', 'Give the run back to Claude.', 'resume', false, true) : command(key, 'pause', 'Save for later', 'Pause the run. You come back to it.', 'pause', false, false)
157
158const walk = (key: ActionKey): Action =>
159  prompt(
160    key,
161    'walk-through',
162    'Walk through step by step',
163    'Claude explains the plan one part at a time.',
164    'Walk me through the plan step by step: the scenarios, the files, the tasks. Stop after each part and wait for me.',
165    SHOW,
166    'Walk through',
167  )
168// The plan review files are in the plugin folder, not in the project: a prompt names them by full path. When the
169// plugin folder is not known, it names the plain paths and says where they are.
170const htmlReview = (key: ActionKey, cli: string): Action => {
171  const root = pluginRootOf(cli)
172  return prompt(
173    key,
174    'html-review',
175    'Open HTML review',
176    'See the plan in a web page. Add comments.',
177    root === null
178      ? 'Open the HTML review of the plan as written in reference/plan-review.md: render it with scripts/plan_review.py, open it, wait for me, then apply review-comments.json. Both files are in the Temper plugin folder, not in the project.'
179      : `Open the HTML review of the plan as written in ${root}/reference/plan-review.md: render it with ${root}/scripts/plan_review.py, open it, wait for me, then apply review-comments.json.`,
180    SHOW,
181  )
182}
183const shareReview = (key: ActionKey, cli: string): Action => {
184  const root = pluginRootOf(cli)
185  return prompt(
186    key,
187    'share-review',
188    'Share HTML review',
189    'Publish the plan page so others can comment.',
190    root === null
191      ? 'Share the HTML review of the plan as written in reference/plan-review.md in the Temper plugin folder. Ask me before anything leaves this machine.'
192      : `Share the HTML review of the plan as written in ${root}/reference/plan-review.md. Ask me before anything leaves this machine.`,
193    SHOW,
194    'Share review',
195  )
196}
197const archDepth = (key: ActionKey): Action =>
198  prompt(
199    key,
200    'arch-depth',
201    'Architecture depth review',
202    'Check the changes for seams, adapters and locality.',
203    'Run the Architecture Depth Review on the changed files. Add its ARCH-DEPTH findings to the review summary.',
204    SHOW,
205    'Depth review',
206  )
207const configSuggestions = (key: ActionKey): Action =>
208  prompt(
209    key,
210    'config-suggestions',
211    'Review config suggestions',
212    'Claude shows each suggested setting. You choose.',
213    'Show each item in config-suggestions.json. For each one, ask me to accept, reject or defer it.',
214    SHOW,
215    'Review config',
216  )
217// The Temper script by its full path; the plain path and where it is when the plugin folder is not known.
218const where = (cli: string): string => (pluginRootOf(cli) === null ? ` ${IN_PLUGIN}` : '')
219
220const stop = (key: ActionKey, cli: string): Action => ({
221  ...command(key, 'stop', 'Stop', 'Run the build check. Then save for later.', 'pause'),
222  prompt: `Run the build check (${cli} gate build).${where(cli)} Then wait. ${SHOW}`,
223})
224
225// The two buttons after Continue, by phase, and the rest of the original options for the menu.
226function pair(phase: Phase, ctx: ActionContext): [Action, Action] {
227  switch (phase) {
228    case 'intent':
229      return [grill(phase, '2'), teach(phase, '3')]
230    case 'plan':
231      return [walk('2'), htmlReview('3', ctx.cli ?? CLI)]
232    case 'build':
233      return isCheckpoint(phase, ctx)
234        ? [draft('2', 'change', 'Change', 'Tell Claude what to change in this task.', 'Change this task: '), stop('3', ctx.cli ?? CLI)]
235        : [teach(phase, '2'), grill(phase, '3')]
236    case 'review':
237      return [archDepth('2'), grill(phase, '3')]
238    case 'check':
239      return ctx.configSuggestions ? [configSuggestions('2'), teach(phase, '3')] : [grill(phase, '2'), teach(phase, '3')]
240    case 'fix':
241      return [grill(phase, '2'), teach(phase, '3')]
242  }
243}
244
245// The other original options, as a numbered menu (1 to 9). An option already on a main button is not repeated.
246function menu(phase: Phase, ctx: ActionContext): Action[] {
247  const paused = ctx.paused ?? false
248  const loopMain = main(phase, ctx).id === 'loop-back'
249  const shown = new Set([...pair(phase, ctx).map(a => a.id)])
250  const extra = (a: Action[]): Action[] => a.filter(x => !shown.has(x.id))
251  switch (phase) {
252    case 'intent':
253      return [save(paused)]
254    case 'plan':
255      return [...extra([grill(phase), teach(phase), shareReview('1', ctx.cli ?? CLI)]), save(paused)]
256    case 'build':
257      return isCheckpoint(phase, ctx) ? [grill(phase), teach(phase), save(paused)] : [...(loopMain ? [] : [loopBack('plan', '1')]), save(paused)]
258    case 'review':
259      return [...extra([teach(phase)]), ...(loopMain ? [] : [loopBack('build', '1')]), save(paused)]
260    case 'check':
261      return [...(ctx.configSuggestions ? extra([grill(phase)]) : []), save(paused)]
262    case 'fix':
263      // At the loop limit Save for later is already key 3.
264      return ctx.loopLimitReached ? [grill(phase), teach(phase)] : [grill(phase), teach(phase), save(paused)]
265  }
266}
267
268// The menu keys: 1 to 9, in order. A tenth entry would have no key, so a phase never has one.
269export function numbered(list: readonly Action[]): Action[] {
270  const keys: ActionKey[] = ['1', '2', '3', '4', '5', '6', '7', '8', '9']
271  return list.slice(0, keys.length).map((a, i) => ({ ...a, key: keys[i] ?? '9' }))
272}
273
274// The phase's menu, with its keys.
275export function moreActions(phase: Phase, ctx: ActionContext): Action[] {
276  return numbered(menu(phase, ctx))
277}
278
279export function actionsFor(phase: Phase, ctx: ActionContext): ActionSet {
280  const set = baseActions(phase, ctx)
281  // A move of the person that the CLI has not recorded: key 1 records it again, in every phase.
282  return ctx.pending ? { ...set, primary: [main(phase, ctx), ...set.primary.slice(1)] } : set
283}
284
285function baseActions(phase: Phase, ctx: ActionContext): ActionSet {
286  const more = moreActions(phase, ctx)
287  if (phase === 'fix') {
288    if (ctx.loopLimitReached) {
289      return {
290        primary: [loopBack('plan', '1'), command('2', 'override-limit', 'Skip with a reason', 'Go on without a pass. Temper records it.', 'override', true, true), { ...save(ctx.paused ?? false, '3') }],
291        discuss: DISCUSS,
292        override: OVERRIDE,
293        more,
294      }
295    }
296    // The check passes again (the orchestrator re-ran it): going back to Check is the next step, and key 1.
297    if (ctx.allChecksPass) {
298      return {
299        primary: [
300          command('1', 'continue', 'Continue to Check', 'The checks pass again. Run the check step.', 'next', false, true, 'To Check'),
301          prompt('2', 'fix-failures', 'Fix the failures', 'Claude makes the smallest change that fixes them.', 'Fix the failed checks with the smallest change. Then stop for the check run.'),
302          prompt('3', 'fix-findings', 'Fix the findings', 'Claude fixes the open review problems.', 'Fix the open review findings, one at a time.'),
303        ],
304        discuss: DISCUSS,
305        override: OVERRIDE,
306        more,
307      }
308    }
309    return {
310      primary: [
311        prompt('1', 'fix-failures', 'Fix the failures', 'Claude makes the smallest change that fixes them.', 'Fix the failed checks with the smallest change. Then stop for the check run.'),
312        prompt('2', 'fix-findings', 'Fix the findings', 'Claude fixes the open review problems.', 'Fix the open review findings, one at a time.'),
313        command('3', 'to-check', 'Continue to Check', 'Run the checks again.', 'next', false, true, 'To Check'),
314      ],
315      discuss: DISCUSS,
316      override: OVERRIDE,
317      more,
318    }
319  }
320  const [two, three] = pair(phase, ctx)
321  return { primary: [main(phase, ctx), two, three], discuss: DISCUSS, override: OVERRIDE, more }
322}
323
324// The actions of a finished run: the original Commit question ("Commit" / "Save for later" / "Other").
325// The Commit steps are named by the slash command, and the CLI by `cli` (see ActionContext.cli).
326export function doneActions(cli: string = CLI): ActionSet {
327  return {
328    primary: [
329      prompt(
330        '1',
331        'commit',
332        'Commit',
333        'Claude commits the work. It does not push.',
334        'The user pressed Commit. Do the Commit steps of /temper:temper now. ' +
335          `In short: run ${cli} gate commit. If it passes, set intent.md to Status completed, run ${cli} state archive, ` +
336          `stage the diff and the .temper/specs artifacts, make one conventional commit, then run ${cli} state clear.${where(cli)} ` +
337          'Do not ask the Commit question again. Do not push.',
338      ),
339      { key: '2', id: 'save-done', label: 'Save for later', desc: 'Leave the work as it is. Commit later.', command: 'saved' },
340    ],
341    discuss: DISCUSS,
342    override: null,
343    more: [],
344  }
345}
346
347// Per finding actions, shown in the pane.
348export function findingActions(id: string): Action[] {
349  return [
350    prompt('1', `fix-${id}`, 'Fix', 'Claude fixes this problem, with a test.', `Fix review finding ${id}. Write a regression test.`),
351    { ...command('2', `accept-${id}`, 'Accept', 'Keep it as it is. You give a reason.', `accept ${id}`, true) },
352    prompt('3', `explain-${id}`, 'Explain', 'Claude says what is wrong and why it matters.', `Explain review finding ${id}. Say what is wrong, where it is, and why it matters.`, SHOW),
353  ]
354}
355
356// The sentence that goes after "Next:" in the system prompt section and in denials.
357export function nextStep(phase: Phase | 'done', ctx: ActionContext): string {
358  switch (phase) {
359    case 'done':
360      return 'the run is done. Commit is allowed'
361    case 'intent':
362      return ctx.ready
363        ? 'ask the user to continue to Plan (key 1 or /temper:temper approve)'
364        : 'finish intent.md until the intent check passes. Then ask the user to continue to Plan (key 1 or /temper:temper approve)'
365    case 'plan':
366      return ctx.ready
367        ? 'ask the user to continue to Build (key 1 or /temper:temper approve)'
368        : 'finish plan.md and tasks.md. Then ask the user to continue to Build (key 1 or /temper:temper approve)'
369    case 'build':
370      return ctx.tasksDone
371        ? 'send the work to Review (key 1 or /temper:temper next)'
372        : 'do the next task in tasks.md. Write a failing test first. Stay inside the plan files'
373    case 'review':
374      return ctx.hasFindings
375        ? 'fix the open findings, or ask the user to accept them with a reason (/temper:temper accept <id> <reason>)'
376        : 'run the review (key 1)'
377    case 'check':
378      return ctx.allChecksPass ? 'mark the run done (key 1 or /temper:temper next)' : 'run the checks (key 1 in Check or /temper:check)'
379    case 'fix':
380      return ctx.loopLimitReached
381        ? 'the fix loop limit is reached. Ask the user to loop back to Plan, skip with a reason, or save the run for later'
382        : 'fix the failed checks. Then continue to Check (key 3)'
383  }
384}
385
386// The sentence under the bar starts with the key and the label of the main action ("1 Continue to
387// Build."), then its line. It reads the same in the band, the pane, the hint and the line under an answer.
388export function nowText(phase: Phase | 'done', ctx: ActionContext): string {
389  if (phase === 'done') return 'The run is done. 1 Commit when you are ready.'
390  if (phase === 'fix' && ctx.loopLimitReached) return 'Fixing did not work after several tries. Choose: 1 loop back to Plan, 2 skip with a reason, or 3 save for later.'
391  const first = actionsFor(phase, ctx).primary[0]
392  return first ? `1 ${first.label}. ${first.desc}` : ''
393}
394
hooks/temper-mod/core/bash.ts 1583 lines
1// Bash classifier. Structural and conservative: the command is split into statements (quote,
2// heredoc and substitution aware), shell variables are resolved statement by statement in order,
3// brace expansion is expanded, and every write capable construct is checked against the guarded
4// Temper paths. A target that cannot be resolved fails closed when the command names Temper
5// state, or when the file name itself cannot be known. Reading a guarded file is never flagged.
6//
7// Limits, stated honestly: a variable set in an earlier call, a profile or the environment cannot
8// be seen, and Bash can write files in ways no classifier catches (an interpreter that builds its
9// path at run time, for one). The hard guarantee is the tool layer and `git commit`; the native
10// pre-commit hook is the backstop.
11
12import { normalizePath } from './paths'
13
14// `config` and `hooks` are protected while a run is active only (the person writes the config with
15// /temper:init, and installs the hooks, when no run is on). `hooks` is the native commit gate: the git hooks, the git
16// config (core.hooksPath) and the Temper commit hook (the temper-gate folder in the git folder, the older
17// temper-pre-commit file there, and the older folders hooks-temper and temper-git-hooks, which git still runs when
18// core.hooksPath stays on one because it holds hooks of other tools).
19export type ProtectedKind = 'events' | 'gates' | 'status' | 'overrides' | 'state' | 'folder' | 'evidence' | 'loops' | 'config' | 'hooks'
20
21export type DecisionKind = 'override' | 'accept' | 'advance' | 'back'
22
23// A decision CLI call as the command spells it: which kind, the phase (stage) and finding id it
24// names, and `invalid` when a flag the CLI reads (`--id`, `--stage`, `--reason`) is repeated, so
25// the CLI and the mod could read different values. An invalid call is never authorised.
26// `next` is the stage a `state advance <stage>_complete <next>` call names as next.
27export type DecisionCall = { kind: DecisionKind; stage?: string; id?: string; next?: string; invalid?: boolean }
28
29// A `scripts/temper state ...` call that moves or removes run state.
30export type StateOp = { op: 'set'; key: string; value?: string } | { op: 'clear' } | { op: 'archive' } | { op: 'init' } | { op: 'loop'; from?: string; to?: string }
31
32export type BashClass = {
33  commits: boolean
34  decisions: DecisionKind[]
35  calls: DecisionCall[]
36  stateOps: StateOp[]
37  // Every write onto a guarded path, and every write whose target cannot be checked.
38  protectedWrites: string[]
39  // The subset of protectedWrites that could not be resolved (fail closed).
40  uncheckable: string[]
41  // What `git add` stages in this command: `all` when it stages the whole tree (-A, ., -u), else the
42  // paths. A commit that stages only files under .temper/specs/ is the artifact chain (the CLI commit gate
43  // lets it through in every phase).
44  staged: { all: boolean; paths: string[] }
45  // The command may run the Temper CLI in a way this classifier cannot read as a plain call (a
46  // glob, an unresolved variable or substitution, a launcher, an interpreter, a shell fed by a
47  // pipe) AND it holds a decision verb. Fail closed: only the person decides.
48  opaque: boolean
49  // Why `opaque` is set, so the refusal tells the true reason. `dynamic`: the subcommand (or the verb) of a Temper
50  // call is written as a variable, a substitution or an escaped string, so the call cannot be read. `decision`:
51  // otherwise, the text holds a decision word. `hidden`: the command may run the script and hides part of what it
52  // runs (a substitution, an expansion, a here-string), with no decision word in the text. Null when `opaque` is false.
53  opaqueWhy: 'dynamic' | 'decision' | 'hidden' | null
54  // The command makes another name or copy of the Temper script, or sources it.
55  alias: boolean
56  // A shell, or a builtin that runs text as commands (source and the like), is given a program the text does not show (a pipe from an unknown command, a file on
57  // stdin, a substitution) or one that is written to hide a word (quote splits, `$`, backslashes, braces, globs).
58  hidden: boolean
59  // A word names a guarded file (or a glob that can stand for one) in a command that is not a plain read.
60  guardedUse: string[]
61  // TEMPER_DIR or TEMPER_CONFIG is set for a command: the CLI would read other files than the run's.
62  envTamper: boolean
63  // `git commit --no-verify` / -n, or a change of core.hooksPath: the native pre-commit hook would not run.
64  noVerify: boolean
65  hookTamper: boolean
66  // A commit is made in a way that is not a plain `git commit` of the index: a pathspec, -i, -o, merge,
67  // cherry-pick, am, pull, revert, commit-tree. The artifact-only carve-out never applies to it.
68  unplainCommit: boolean
69  // The working directory after the command (relative to where it started, or absolute); null when unknown.
70  cwdAfter: string | null
71}
72
73// Every name is compared without regard to case: macOS (APFS) and Windows folders are case
74// insensitive, so `.TEMPER/Gates.json` is the same file.
75const PROTECTED: ReadonlyArray<readonly [ProtectedKind, RegExp]> = [
76  ['events', /(^|\/)\.temper\/specs\/[^/\s]+\/events(\/|$)/i],
77  ['gates', /(^|\/)\.temper\/gates\.json$/i],
78  ['status', /(^|\/)\.temper\/status\.json$/i],
79  ['overrides', /(^|\/)\.temper\/overrides\.json$/i],
80  ['state', /(^|\/)\.temper\/build-state\.json$/i],
81  // The evidence ledger and the loop counter are written by the CLI only.
82  // (The ledger files only: a report a command keeps next to them, `coverage-report.txt`, is no ledger.)
83  ['evidence', /(^|\/)\.temper\/evidence\/[^/]+\.json$/i],
84  ['loops', /(^|\/)\.temper\/feedback-loops\.json$/i],
85  // What decides how the run is checked (autonomy, thresholds, blocking), and the native commit gate: the git hooks, the
86  // git config, and the Temper commit hook (`.git/temper-gate`, the folder core.hooksPath names, with its pre-commit, and
87  // `.git/temper-pre-commit`, the file an older Temper line runs). A path under `temper-pre-commit` counts too: the line
88  // runs the hook only when it is a file, so a folder made in its place would switch the hook off. Temper's older
89  // folders (`.git/hooks-temper` from --global up to 9.6.4, `.git/temper-git-hooks` from 9.6.5) count as well: the
90  // installer leaves core.hooksPath on one that holds hooks of other tools, and git then runs its pre-commit.
91  ['config', /(^|\/)\.claude\/temper\.config$/i],
92  ['hooks', /(^|\/)\.git\/(?:hooks(\/|$)|hooks-temper(\/|$)|temper-git-hooks(\/|$)|config$|temper-gate(\/|$)|temper-pre-commit(\/|$))/i],
93  // Folders that hold guarded files: removing or replacing one removes them too.
94  ['folder', /(^|\/)\.temper(\/specs(\/[^/\s]+)?)?\/?$/i],
95]
96
97// Which guarded Temper path a path names, or null.
98export function protectedKind(path: string): ProtectedKind | null {
99  // `..` and `.` segments are collapsed first, so specs/a/../a/events is the events folder.
100  const p = normalizePath(path)
101  for (const [kind, re] of PROTECTED) if (re.test(p)) return kind
102  return null
103}
104
105const MENTION = /\.temper\/(?:specs\/[^\s'"`]+\/events[^\s'"`]*|gates\.json|status\.json|overrides\.json|build-state\.json|feedback-loops\.json|evidence\/[^\s'"`]*\.json)|\.claude\/temper\.config(?![\w.])|\.git\/(?:hooks|hooks-temper|temper-git-hooks|config|temper-gate|temper-pre-commit)(?![\w.-])/gi
106// An interpreter program that holds a guarded file name on its own (`os.path.join('.temper', 'gates.json')`), or the
107// folder itself next to a call that removes or moves things (`shutil.rmtree('.temper')`).
108const BARE_NAMES = /(?<![\w.-])(?:gates|status|overrides)\.json(?![\w])|(?<![\w.-])build-state\.json|(?<![\w.-])feedback-loops\.json/gi
109const FOLDER_QUOTED = /(?<=['"`])\.temper\/?(?=['"`])/
110const REMOVER = /\b(?:rmtree|rmdir|removedirs|unlink|rimraf|rmSync|unlinkSync|renameSync|os\.rename|os\.replace|os\.remove|shutil\.move|truncate|chmod|chown)\w*/i
111
112// The command names Temper state: strict mode, where an unresolvable write target is refused.
113const NAMED = /\.temper|gates\.json|status\.json|overrides\.json|build-state\.json|\bevents\b|temper\.config|\.git\/(?:hooks|temper-git-hooks|config|temper-gate|temper-pre-commit)/i
114
115// A path whose text names a guarded thing even when the rest cannot be resolved.
116const NAMES_GUARDED = /gates\.json|status\.json|overrides\.json|build-state\.json|temper\.config|feedback-loops\.json|\.git\/(?:hooks|temper-git-hooks|config|temper-gate|temper-pre-commit)|(^|\/)events(\/|$)|(^|\/)\.temper(\/|$)/i
117
118const PROTECTED_NAMES = ['.temper', 'events', 'gates.json', 'status.json', 'overrides.json', 'build-state.json', 'feedback-loops.json', 'temper.config']
119
120// A command that names one of these files is a plain read, or it is refused while a run is active.
121const GUARDED_FILE = /gates\.json|status\.json|overrides\.json|build-state\.json|feedback-loops\.json|(?:^|\/)temper\.config$|\.temper\/specs\/[^/\s'"`]+\/events|\.temper\/evidence\/[^\s'"`]*\.json|\.git\/(?:hooks|temper-git-hooks|config$|temper-gate|temper-pre-commit)/i
122
123// Whether one word (quotes already removed, variables filled in) names a guarded file, or a glob that
124// can stand for one: `.tem*/gates.js*`. A glob counts when its segment has three literal characters or
125// starts with a dot (`.t*`), so `*` and `build/*` do not.
126function mentionsGuarded(word: string, globs = true): boolean {
127  if (GUARDED_FILE.test(word)) return true
128  if (!globs) return false
129  for (const seg of word.split(/[\s/=:,'"`;()&|<>]+/)) {
130    if (!GLOB.test(seg)) continue
131    const literals = seg.replace(/[*?[\]]/g, '').length
132    if (literals < 3 && !seg.startsWith('.')) continue
133    if (PROTECTED_NAMES.some(n => globRegExp(seg.replace(/[[\]]/g, '?')).test(n))) return true
134  }
135  return false
136}
137
138// ---- Statements and words ----------------------------------------------------------------
139
140type Word = { text: string; dynamic: boolean }
141
142// Splits a command into top level statements on `;`, `&`, `&&`, `||`, `|` and newlines, outside
143// quotes and substitutions, and leaves heredoc bodies out (they are data or code for another
144// interpreter, which the interpreter rule reads from the whole text).
145function topStatements(cmd: string): string[] {
146  return statementsOf(cmd).map(x => x.stmt)
147}
148
149// The same, with the statement a single `|` pipes into this one (null when there is none): a shell that
150// reads its program from a pipe needs to know what produced it.
151function statementsOf(cmd: string): Array<{ stmt: string; from: string | null }> {
152  const out: Array<{ stmt: string; from: string | null }> = []
153  let cur = ''
154  let quote: '' | "'" | '"' = ''
155  let depth = 0
156  let heredoc: { tag: string; dash: boolean } | null = null
157  const pending: Array<{ tag: string; dash: boolean }> = []
158  let piped = false
159  let prev: string | null = null
160  const flush = () => {
161    if (cur.trim()) {
162      out.push({ stmt: cur.trim(), from: piped ? prev : null })
163      prev = cur.trim()
164      piped = false
165    }
166    cur = ''
167  }
168  for (let i = 0; i < cmd.length; i++) {
169    const c = cmd[i] ?? ''
170    if (heredoc) {
171      // Skip lines until the terminator.
172      const nl = cmd.indexOf('\n', i)
173      const line = (nl < 0 ? cmd.slice(i) : cmd.slice(i, nl)).replace(/\r$/, '')
174      const body = heredoc.dash ? line.replace(/^\t+/, '') : line
175      if (body === heredoc.tag) heredoc = pending.shift() ?? null
176      i = nl < 0 ? cmd.length : nl
177      continue
178    }
179    if (quote) {
180      cur += c
181      if (c === '\\' && quote === '"' && i + 1 < cmd.length) cur += cmd[++i] ?? ''
182      else if (c === quote) quote = ''
183      continue
184    }
185    if (c === '\\' && i + 1 < cmd.length) {
186      cur += c + (cmd[++i] ?? '')
187      continue
188    }
189    if (c === "'" || c === '"') {
190      quote = c
191      cur += c
192      continue
193    }
194    // A comment (`# ...` at the start of a word) is not part of any command: `npm test  # see .git/hooks`.
195    if (c === '#' && (i === 0 || /[\s;&|(]/.test(cmd[i - 1] ?? ''))) {
196      const nl = cmd.indexOf('\n', i)
197      i = nl < 0 ? cmd.length : nl - 1
198      continue
199    }
200    if (c === '$' && cmd[i + 1] === '(') {
201      depth++
202      cur += '$('
203      i++
204      continue
205    }
206    if ((c === '<' || c === '>') && cmd[i + 1] === '(') {
207      depth++
208      cur += c + '('
209      i++
210      continue
211    }
212    if (c === '(') depth++
213    if (c === ')' && depth > 0) depth--
214    if (c === '`') {
215      // A backtick span is one substitution.
216      const end = cmd.indexOf('`', i + 1)
217      const stop = end < 0 ? cmd.length : end + 1
218      cur += cmd.slice(i, stop)
219      i = stop - 1
220      continue
221    }
222    if (c === '<' && cmd[i + 1] === '<' && cmd[i + 2] !== '<') {
223      const m = /^<<(-?)\s*(['"]?)([\w.-]+)\2/.exec(cmd.slice(i))
224      if (m) {
225        const tag = { tag: m[3] ?? '', dash: m[1] === '-' }
226        if (cmd.indexOf('\n', i) >= 0) pending.push(tag)
227        cur += m[0]
228        i += m[0].length - 1
229        continue
230      }
231    }
232    if (depth === 0 && (c === ';' || c === '|' || c === '\n' || c === '&')) {
233      if (c === '&' && (cmd[i - 1] === '>' || cmd[i + 1] === '>' || /\d/.test(cmd[i + 1] ?? ''))) {
234        cur += c
235        continue
236      }
237      const single = c === '|' && cmd[i + 1] !== '|'
238      if ((c === '&' || c === '|') && cmd[i + 1] === c) i++
239      flush()
240      if (single) piped = true
241      else prev = null
242      if (c === '\n' && pending.length > 0) heredoc = pending.shift() ?? null
243      continue
244    }
245    cur += c
246  }
247  flush()
248  return out
249}
250
251// Splits a statement into words (quotes removed, backslashes processed outside single quotes).
252// `dynamic` marks a word holding a command or process substitution.
253function wordsOf(stmt: string): Word[] {
254  const words: Word[] = []
255  let cur = ''
256  let dynamic = false
257  let started = false
258  let quote: '' | "'" | '"' = ''
259  let depth = 0
260  const push = () => {
261    if (started) words.push({ text: cur, dynamic })
262    cur = ''
263    dynamic = false
264    started = false
265  }
266  for (let i = 0; i < stmt.length; i++) {
267    const c = stmt[i] ?? ''
268    if (quote === "'") {
269      if (c === "'") quote = ''
270      else cur += c
271      continue
272    }
273    if (quote === '"') {
274      if (c === '"') quote = ''
275      else if (c === '\\' && i + 1 < stmt.length) cur += stmt[++i] ?? ''
276      else {
277        if (c === '$' && stmt[i + 1] === '(') dynamic = true
278        if (c === '`') dynamic = true
279        cur += c
280      }
281      continue
282    }
283    if (depth > 0) {
284      cur += c
285      if (c === '(') depth++
286      if (c === ')') depth--
287      continue
288    }
289    if (c === "'" || c === '"') {
290      quote = c
291      started = true
292      continue
293    }
294    if (c === '\\' && i + 1 < stmt.length) {
295      cur += stmt[++i] ?? ''
296      started = true
297      continue
298    }
299    if ((c === '$' || c === '<' || c === '>') && stmt[i + 1] === '(') {
300      dynamic = true
301      started = true
302      depth = 1
303      cur += c + '('
304      i++
305      continue
306    }
307    if (c === '`') {
308      const end = stmt.indexOf('`', i + 1)
309      const stop = end < 0 ? stmt.length : end + 1
310      cur += stmt.slice(i, stop)
311      dynamic = true
312      started = true
313      i = stop - 1
314      continue
315    }
316    if (/\s/.test(c)) {
317      push()
318      continue
319    }
320    cur += c
321    started = true
322  }
323  push()
324  return words
325}
326
327// Command and process substitutions inside a statement, as nested commands to analyse.
328function substitutions(stmt: string): string[] {
329  const out: string[] = []
330  let single = false
331  for (let i = 0; i < stmt.length; i++) {
332    // Inside single quotes nothing is substituted.
333    if (stmt[i] === "'" && stmt[i - 1] !== '\\') single = !single
334    if (single) continue
335    if ((stmt[i] === '$' || stmt[i] === '<' || stmt[i] === '>') && stmt[i + 1] === '(') {
336      let d = 1
337      let j = i + 2
338      for (; j < stmt.length && d > 0; j++) {
339        if (stmt[j] === '(') d++
340        if (stmt[j] === ')') d--
341      }
342      out.push(stmt.slice(i + 2, j - 1))
343      i = j - 1
344    } else if (stmt[i] === '`') {
345      const end = stmt.indexOf('`', i + 1)
346      const stop = end < 0 ? stmt.length : end
347      out.push(stmt.slice(i + 1, stop))
348      i = stop
349    }
350  }
351  return out
352}
353
354// ---- Variables and braces -----------------------------------------------------------------
355
356type Vars = Map<string, string>
357
358// Replaces $NAME and ${NAME} with known values, repeatedly (a fixed cap) so nested references
359// resolve; an unknown name is left in place.
360function expandVars(text: string, vars: Vars): string {
361  let t = text
362  for (let n = 0; n < 8; n++) {
363    const next = t.replace(/\$\{(\w+)\}|\$(\w+)/g, (m, a, b) => vars.get(a ?? b ?? '') ?? m)
364    if (next === t) break
365    t = next
366  }
367  return t
368}
369
370// Expands `{a,b}` and `{1..3}` (a fixed cap); null when the expansion is too large.
371function braceExpand(word: string): string[] | null {
372  const out: string[] = []
373  const walk = (w: string, depth: number): boolean => {
374    if (depth > 6 || out.length > 64) return false
375    const open = w.search(/\{[^{}]*(,|\.\.)[^{}]*\}/)
376    if (open < 0) {
377      out.push(w)
378      return true
379    }
380    const close = w.indexOf('}', open)
381    const body = w.slice(open + 1, close)
382    const head = w.slice(0, open)
383    const tail = w.slice(close + 1)
384    const alts = body.includes(',') ? body.split(',') : seq(body)
385    if (alts === null) return false
386    for (const a of alts) if (!walk(head + a + tail, depth + 1)) return false
387    return true
388  }
389  return walk(word, 0) && out.length <= 64 ? out : null
390}
391
392function seq(body: string): string[] | null {
393  const m = /^(-?\d+)\.\.(-?\d+)$/.exec(body)
394  if (!m) return null
395  const a = Number(m[1])
396  const b = Number(m[2])
397  if (Math.abs(b - a) > 32) return null
398  const r: string[] = []
399  for (let i = a; a <= b ? i <= b : i >= b; i += a <= b ? 1 : -1) r.push(String(i))
400  return r
401}
402
403const GLOB = /[*?[\]]/
404
405// A file name that is `temper`, or a glob (`tempe[r]`, `temp*`, `*`) that can match it.
406function namesTemper(text: string): boolean {
407  const base = BASE(text).toLowerCase()
408  if (base === 'temper') return true
409  if (!GLOB.test(base)) return false
410  let re = ''
411  for (let i = 0; i < base.length; i++) {
412    const c = base[i] ?? ''
413    if (c === '*') re += '.*'
414    else if (c === '?') re += '.'
415    else if (c === '[' && base.indexOf(']', i + 2) > 0) {
416      const end = base.indexOf(']', i + 2)
417      re += `[${base.slice(i + 1, end).replace(/^!/, '^').replace(/\\/g, '\\\\')}]`
418      i = end
419    } else re += c.replace(/[.+^${}()|[\]\\]/g, '\\$&')
420  }
421  try {
422    return new RegExp(`^${re}$`).test('temper')
423  } catch {
424    return true
425  }
426}
427
428// ---- Standard input by any of its names -----------------------------------------------------
429
430// Standard input is /dev/stdin, /dev/fd/N, /proc/self/fd/N, /proc/thread-self/fd/N, /proc/<pid>/fd/N or
431// /proc/<pid>/task/<tid>/fd/N. Case is ignored (a folder on macOS may not tell case apart).
432const STDIN_PATH = /^\/(?:dev\/(?:stdin|fd\/\d+)|proc\/(?:self|thread-self|\d+)\/(?:task\/\d+\/)?fd\/\d+)$/i
433// What a glob, or a path with a part that cannot be read, is tried against.
434const STDIN_SAMPLES: string[] = ['/dev/stdin']
435for (let n = 0; n < 10; n++) {
436  STDIN_SAMPLES.push(`/dev/fd/${n}`)
437  for (const p of ['self', 'thread-self', '1', '42']) STDIN_SAMPLES.push(`/proc/${p}/fd/${n}`, `/proc/${p}/task/1/fd/${n}`)
438}
439// A part of a word the shell fills in that this command did not set: ${..}, $(..), `..`, $NAME (the whole name), $$
440// and the like.
441const UNREAD = /\$\{[^}]*\}|\$\([^)]*\)|`[^`]*`|\$(?:\w+|[$!#?*@-])/g
442const MARK = '\u0000'
443
444// A glob (with MARK for a part that cannot be read) as the text of a regular expression.
445function stdinGlob(p: string): string {
446  let re = ''
447  for (let i = 0; i < p.length; i++) {
448    const c = p[i] ?? ''
449    if (c === MARK) re += '.*'
450    else if (c === '*') re += '[^/]*'
451    else if (c === '?') re += '[^/]'
452    else if (c === '[' && p.indexOf(']', i + 2) > 0) {
453      const end = p.indexOf(']', i + 2)
454      re += `[${p.slice(i + 1, end).replace(/^!/, '^').replace(/\\/g, '\\\\')}]`
455      i = end
456    } else re += c.replace(/[.+^${}()|[\]\\]/g, '\\$&')
457  }
458  return re
459}
460
461// Whether one of the names of standard input can be matched by a regular expression.
462function anyStdin(re: string, anchored: boolean): boolean {
463  try {
464    const glob = new RegExp(`${anchored ? '^' : ''}${re}$`, 'i')
465    return STDIN_SAMPLES.some(s => glob.test(s))
466  } catch {
467    return true
468  }
469}
470
471// The end of a path that is all that is known of it: what follows the last part that cannot be read (MARK), or the
472// whole of a relative path, and of that only what follows its last `..` step (such a step can undo anything before
473// it). `.` steps and doubled slashes are taken out. The result is matched against the end of the names of standard
474// input; a path that ends in a part that cannot be read can be any name.
475function stdinTail(text: string): boolean {
476  const parts = text.split('/')
477  const glued = parts[0] ?? ''
478  let rest = parts.slice(1).filter(x => x !== '' && x !== '.')
479  const up = rest.lastIndexOf('..')
480  const head = up >= 0 ? '' : glued
481  if (up >= 0) rest = rest.slice(up + 1)
482  const tail = rest.length > 0 || head === '' ? `${head}/${rest.join('/')}` : head
483  if (tail === '/' || tail === '') return true
484  return anyStdin(stdinGlob(tail), false)
485}
486
487// The last part of a path is `stdin`, or the last two are `fd/<n>`.
488const STDIN_NAME = /(?:^|\/)(?:stdin|fd\/\d+)$/i
489
490// Whether a path names standard input, in any spelling. The variables this command set are filled in first, and
491// `.`, `..`, doubled slashes and a /proc/<pid>/root prefix (the root folder again) are taken out, so /dev/./stdin,
492// //dev/stdin, /dev//stdin and /dev/../dev/stdin all count. Of a brace expansion the first word counts (it is the
493// script). A glob counts when it can match one of the names.
494// A relative path is read against `cwd`, the folder the shell is in (see classifyBash): a full folder gives the
495// full path. A folder in the project (relative to its root, whose place is not known here) gives a file of the
496// project, unless `..` steps leave the project: they may reach the root folder, so ../../../dev/stdin counts.
497// When `fed` says the command is given input (a pipe, a heredoc, a here-string or a file), that input is what may
498// run, so more counts:
499//  - a part that cannot be read (an unknown variable, a substitution) stands for any text, `..` steps too: the path
500//    counts when what follows that part can end a name of standard input (`$D/stdin`, `$S`);
501//  - a relative path counts when it can end a name of standard input (stdin, fd/0, dev/stdin), whatever the folder:
502//    the folder may not be known (`cd "$X"`), and a link in the project can lead anywhere;
503//  - a full path counts when it ends in stdin or fd/<n>, since a link can make any folder /dev.
504function readsStdin(text: string, vars: Vars, fed: boolean, cwd: string | null): boolean {
505  const t = expandVars(text, vars)
506  let marked = (braceExpand(t)?.[0] ?? t).replace(UNREAD, MARK)
507  // `$'..'` and `$".."` leave a `$` before the text. An escape inside `$'..'` cannot be read here, so it stands for any text.
508  if (marked.includes('$')) marked = /\\/.test(marked) ? MARK : marked.replace(/\$/g, '')
509  if (marked.includes(MARK)) return fed && stdinTail(marked.slice(marked.lastIndexOf(MARK) + 1))
510  if (!marked.startsWith('/')) {
511    if (cwd !== null && cwd.startsWith('/')) marked = `${cwd}/${marked}`
512    else if (fed) return STDIN_NAME.test(normalizePath(marked)) || stdinTail(`/${marked}`)
513    else if (cwd === null) return false
514    else {
515      const rel = normalizePath(cwd ? `${cwd}/${marked}` : marked)
516      if (!rel.startsWith('..')) return false
517      marked = `/${rel}`
518    }
519  }
520  let p = normalizePath(marked).replace(/^\/(?:\.\.\/)+/, '/')
521  for (let n = 0; n < 8 && /^\/proc\/[^/]*\/root\//i.test(p); n++) p = p.replace(/^\/proc\/[^/]*\/root/i, '')
522  if (fed && STDIN_NAME.test(p)) return true
523  if (!GLOB.test(p)) return STDIN_PATH.test(p)
524  return anyStdin(stdinGlob(p), true)
525}
526
527// The text of a statement with its quoted parts ('..', "..", $'..') and its escaped characters taken out.
528function unquoted(stmt: string): string {
529  let out = ''
530  let quote: '' | "'" | '"' | '$' = ''
531  for (let i = 0; i < stmt.length; i++) {
532    const c = stmt[i] ?? ''
533    if (quote === "'") {
534      if (c === "'") quote = ''
535    } else if (quote === '"' || quote === '$') {
536      if (c === '\\') i++
537      else if (c === (quote === '"' ? '"' : "'")) quote = ''
538    } else if (c === '\\') i++
539    else if (c === '$' && stmt[i + 1] === "'") {
540      quote = '$'
541      i++
542    } else if (c === "'" || c === '"') quote = c
543    else out += c
544  }
545  return out
546}
547
548// Whether a statement is given input: a pipe into it, a heredoc, a here-string or a file on standard input. A `<` inside
549// a quoted argument (`--grep '<title>'`) is no redirect, and input from /dev/null (`<`, `0<`) is no input.
550const fedInput = (stmt: string, pipedFrom: string | null): boolean =>
551  pipedFrom !== null || /(?:^|[^<>&\d])\d*<(?![(&])/.test(unquoted(stmt).replace(/(^|[^<>&\d])0?<\s*\/dev\/null(?![^\s;&|)])/g, '$1'))
552
553// Words in a command that name a decision on the run. Read on the whole text of an opaque launch.
554const VERB = /\b(?:override|accept|advance|next_stage|run_mode|clear|archive|init|loop)\b/i
555const verbIn = (text: string): boolean => VERB.test(text) || (/\bstate\b/i.test(text) && /\bset\b/i.test(text))
556
557// `temper` as a word or a path component inside any text (a quoted string holds several words).
558const TEMPER_WORD = /(?:^|[\s/'"=;(&|<>`])temper(?=$|[\s/'"`;)&|<>])/i
559// A construct that hides what the shell will really run: ANSI-C string, substitution, parameter
560// expansion, here-string, process substitution, brace expansion.
561const HIDES = /\$'|\$\(|`|\$\{|<<<|<\(|>\(|\{[^{}\s]*,[^{}\s]*\}/
562const pieceNamesTemper = (text: string): boolean => TEMPER_WORD.test(text) || text.split(/[\s'"`;()&|<>=]+/).some(p => p !== '' && namesTemper(p))
563
564// Commands that only read or print: they never run the Temper script. awk and sed read too, unless
565// the program runs something (see readerDanger).
566const READERS = new Set([
567  'cat', 'grep', 'egrep', 'fgrep', 'rg', 'head', 'tail', 'less', 'more', 'wc', 'ls', 'stat', 'file', 'diff', 'cmp',
568  'git', 'echo', 'printf', 'cd', 'pushd', 'test', '[', 'shellcheck', 'basename', 'dirname', 'realpath', 'readlink', 'which', 'type',
569  'sed', 'awk', 'gawk', 'nl', 'tr', 'cut', 'sort', 'uniq', 'bat', 'strings', 'xxd', 'od', 'pytest', 'find',
570])
571
572// Whether a text holds a git command that creates a commit: commit, cherry-pick, merge, revert, am,
573// commit-tree, rebase --continue, and a pull that merges. The read only and stopping forms
574// (--abort, --quit, --skip, --show-current-patch) and `git push` are not. The pre-commit hook of
575// git does not run for most of these, so the guard has to refuse them while the commit gate is open.
576export function gitCreatesCommit(text: string, anchored = false): boolean {
577  for (const m of text.matchAll(/\bgit\s+((?:(?:-c|-C)\s+\S+\s+|--[\w-]+(?:=\S+)?\s+|-\w\s+)*)(commit-tree|commit|cherry-pick|merge|revert|am|rebase|pull)\b([^;&|\n]*)/gi)) {
578    if (anchored && m.index !== 0) break
579    const sub = (m[2] ?? '').toLowerCase()
580    const rest = m[3] ?? ''
581    if (/--(?:abort|quit|skip|show-current-patch|edit-todo)\b/.test(rest) && sub !== 'commit') continue
582    if (sub === 'rebase') {
583      if (/--continue\b/.test(rest)) return true
584      continue
585    }
586    if (sub === 'pull') {
587      if (/--ff-only\b|--rebase\b|-r\b/.test(rest)) continue
588      return true
589    }
590    if (sub === 'merge' && /--(?:abort|quit)\b/.test(rest)) continue
591    return true
592  }
593  return false
594}
595
596// The arguments after a command word that cannot be read are the arguments of a plain Temper read or
597// record call: the words that say what it does are literal (no $, no substitution) and name a call that
598// decides nothing (gate, report, status, model, config, evidence add|run|list|resolve, state get, state set of
599// a bookkeeping key).
600function plainTemperArgs(args: Word[], argText: string[]): boolean {
601  const lit = (i: number): string | null => (i < args.length && !args[i]?.dynamic && !/[$`]/.test(argText[i] ?? '') ? (argText[i] ?? null) : null)
602  const a0 = lit(0)
603  if (a0 === null) return false
604  if (['gate', 'report', 'status', 'model', 'config', 'bands', 'metrics'].includes(a0)) return true
605  const a1 = lit(1)
606  if (a0 === 'evidence') return a1 !== null && ['add', 'run', 'list', 'resolve'].includes(a1)
607  if (a0 === 'state') {
608    if (a1 === 'get') return true
609    if (a1 === 'set') {
610      const key = lit(2)
611      return key !== null && ['complexity', 'base_sha', 'regression_test', 'task'].includes(key)
612    }
613  }
614  return false
615}
616
617// An awk program that runs a command (system, getline, a pipe) or a sed program that does (the e command or flag).
618function readerDanger(cmd: string, args: Word[]): boolean {
619  const text = args.map(a => a.text).join(' ')
620  if (cmd === 'find') return args.some(a => /^-(?:exec|execdir|ok|okdir|delete|fprint\w*|fls)$/.test(a.text))
621  if (cmd === 'awk' || cmd === 'gawk') return /system\s*\(|getline|\|\s*["']?\w|\|&/.test(text)
622  if (cmd === 'sed') return /(?:^|[;{}\s])[0-9$,]*e(?:\s|;|$)|s(.)(?:(?!\1).)*\1(?:(?!\1).)*\1[a-z]*e/.test(text)
623  return false
624}
625
626// Shells that take a program as a string (-c) or from a file.
627const SHELLS = new Set(['bash', 'sh', 'zsh', 'dash', 'ksh', 'fish', 'ash', 'tcsh', 'csh'])
628// Commands that run another command or program.
629const EXEC = new Set([...SHELLS, 'make', 'gmake', 'awk', 'gawk', 'watch', 'script', 'parallel', 'busybox', 'xargs', 'env', 'eval', 'exec', 'find', 'nohup', 'timeout', 'sudo'])
630// A file name that is a script by its extension.
631const SCRIPT_FILE = /\.(?:sh|bash|zsh|ksh|fish|py|pl|rb|js|mjs|cjs|ts|php|mk|awk)$|(?:^|\/)makefile$/i
632// A program given on the command line to an interpreter (python -c, node -e, perl -e, php -r).
633const PROGRAM_FLAG = /^-[a-zA-Z]*[cer]$|^--eval$/
634
635function globRegExp(glob: string): RegExp {
636  let re = ''
637  for (const c of glob) {
638    if (c === '*') re += '.*'
639    else if (c === '?') re += '.'
640    else re += c.replace(/[.+^${}()|[\]\\]/g, '\\$&')
641  }
642  return new RegExp(`^${re}$`, 'i')
643}
644
645// A glob such as `ev*`, `.temper/g*` or `status.js?n` could match a guarded name.
646function globCouldMatch(path: string): boolean {
647  return path.split('/').some(seg => GLOB.test(seg) && PROTECTED_NAMES.some(n => globRegExp(seg.replace(/[[\]]/g, '?')).test(n)))
648}
649
650// ---- Commands -------------------------------------------------------------------------------
651
652const BASE = (w: string): string => w.replace(/^.*\//, '')
653
654const KEYWORDS = new Set(['then', 'do', 'else', 'elif', 'if', 'while', 'until', '!', '{', '}', 'time', 'command', 'builtin', 'exec', 'nohup', 'stdbuf', 'setsid'])
655const DECLARE = new Set(['export', 'declare', 'typeset', 'local', 'readonly'])
656
657// Removes wrapper words from the front: env, timeout, nice, ionice, xargs, sudo and keywords.
658function unwrap(ws: Word[]): Word[] {
659  let w = ws
660  for (let guard = 0; guard < 20 && w.length > 0; guard++) {
661    const first = BASE(w[0]?.text ?? '')
662    const rest = w.slice(1)
663    const skipOpts = (list: Word[], withArg: string[] = []): Word[] => {
664      let i = 0
665      while (i < list.length) {
666        const t = list[i]?.text ?? ''
667        if (withArg.includes(t)) i += 2
668        else if (t.startsWith('-') && t !== '--') i += 1
669        else break
670      }
671      return list.slice(i)
672    }
673    if (KEYWORDS.has(first)) {
674      w = skipOpts(rest)
675    } else if (first === 'sudo') {
676      w = skipOpts(rest, ['-u', '-g', '-h', '-p', '-C', '-D', '-R', '-T'])
677    } else if (first === 'env') {
678      let i = 0
679      while (i < rest.length) {
680        const t = rest[i]?.text ?? ''
681        if (['-u', '-C', '-S'].includes(t)) i += 2
682        else if (t.startsWith('-') || /^[A-Za-z_]\w*=/.test(t)) i += 1
683        else break
684      }
685      w = rest.slice(i)
686    } else if (first === 'timeout') {
687      const r = skipOpts(rest, ['-s', '-k', '--signal', '--kill-after'])
688      w = /^\d/.test(r[0]?.text ?? '') ? r.slice(1) : r
689    } else if (first === 'caffeinate') {
690      w = skipOpts(rest, ['-t', '-w'])
691    } else if (first === 'nice') {
692      w = skipOpts(rest, ['-n'])
693    } else if (first === 'ionice') {
694      w = skipOpts(rest, ['-c', '-n', '-p'])
695    } else if (first === 'xargs') {
696      w = skipOpts(rest, ['-n', '-P', '-I', '-L', '-s', '-d', '-E'])
697    } else if (first === '--') {
698      w = rest
699    } else break
700  }
701  return w
702}
703
704// The files an in place editor (`sed -i`, `perl -i`, `awk -i`) rewrites: not its options, not an
705// empty backup suffix (`-i ''`), and not the script, which is the first plain word unless it was
706// given with -e, -f or a combined flag such as -pe.
707function inPlaceFiles(args: Word[]): Word[] {
708  const files: Word[] = []
709  let scriptGiven = false
710  for (let i = 0; i < args.length; i++) {
711    const t = args[i]?.text ?? ''
712    if (/^-[a-zA-Z]*[ef]$/.test(t) || t === '--expression' || t === '--file') {
713      scriptGiven = true
714      i++
715    } else if (t.startsWith('-')) continue
716    else if (t === '') continue
717    else if (!scriptGiven) scriptGiven = true
718    else files.push(args[i] as Word)
719  }
720  return files
721}
722
723const WRITES_ANY = new Set(['tee', 'rm', 'touch', 'truncate', 'shred', 'unlink'])
724const WRITES_LAST = new Set(['cp', 'mv', 'install', 'ln', 'rsync'])
725const INTERPRETERS = /^(?:python[\d.]*|node|deno|bun|perl|ruby|php|osascript)$/
726
727const GUARD_KEYS = new Set(['stage', 'next_stage', 'branch', 'spec_path', 'run_mode'])
728
729type Flags = Map<string, string[]>
730
731// Reads `--flag value` / `--flag=value` pairs the way the CLI does: a flag takes the next word
732// whatever it is. Returns every value per flag so a repeat can be seen.
733function readFlags(ws: Word[], names: readonly string[]): Flags {
734  const flags: Flags = new Map()
735  for (let i = 0; i < ws.length; i++) {
736    const t = ws[i]?.text ?? ''
737    const eq = t.indexOf('=')
738    const name = eq > 0 ? t.slice(0, eq) : t
739    if (!names.includes(name)) continue
740    const value = eq > 0 ? t.slice(eq + 1) : (ws[++i]?.text ?? '')
741    flags.set(name, [...(flags.get(name) ?? []), value])
742  }
743  return flags
744}
745
746// Shell setup tools: what one of them prints may be run as commands (sourced, or run from a substitution). Anything else
747// built by a substitution is a program the text does not show.
748const ENV_NAMES = new Set(['ssh-agent', 'pyenv', 'rbenv', 'nodenv', 'jenv', 'goenv', 'direnv', 'fnm', 'mise', 'asdf', 'brew', 'conda', 'minikube', 'docker-machine', 'starship', 'zoxide', 'dircolors', 'opam', 'keychain', 'gpg-agent', 'thefuck', 'register-python-argcomplete'])
749
750// Substitutions of a plain lookup that cannot build a command: `$(pwd)`, `$(git rev-parse HEAD)`.
751const TRIVIAL_SUBST = /^\s*(?:pwd|date|nproc|uname|whoami|hostname|git\s+rev-parse|which|command\s+-v|basename|dirname|realpath|mktemp)\b/
752// Programs that print any text they are given: a substitution of one of them can build a command from pieces.
753const GENERATORS = new Set([
754  // Text tools.
755  'echo', 'printf', 'cat', 'awk', 'gawk', 'sed', 'tr', 'xxd', 'rev', 'head', 'tail', 'cut', 'jq',
756  // Transfer and encoding tools.
757  'base64', 'curl', 'wget', 'openssl',
758  // Shells, language runtimes and other producers.
759  'python', 'python3', 'node', 'perl', 'ruby', 'sh', 'bash', 'zsh', 'env', 'printenv', 'yes', 'seq', 'tee', 'dd', 'php', 'deno', 'bun',
760])
761// A substitution whose command is a shell setup tool, a plain lookup, or a program run by its path
762// (`$(scripts/ensure-jdk.sh --export)`): the text shows what produces the program. A text generator is not that.
763function safeSubst(inner: string): boolean {
764  const first = inner.trim().split(/\s+/)[0] ?? ''
765  const base = first.replace(/^.*\//, '').toLowerCase()
766  if (ENV_NAMES.has(base) || TRIVIAL_SUBST.test(inner)) return true
767  return first.includes('/') && !GENERATORS.has(base.replace(/[\d.]+$/, ''))
768}
769const stripSafe = (t: string): string => t.replace(/\$\(([^()`$]*)\)/g, (m: string, inner: string) => (safeSubst(inner) ? '' : m))
770// A program text that still holds a substitution or a variable after the shell setup idioms and the plain lookups.
771const hiddenText = (t: string): boolean => /[$`]/.test(stripSafe(t))
772
773// Commands that only read when they are given a guarded file.
774const PLAIN_READERS = new Set([
775  'cat', 'grep', 'egrep', 'fgrep', 'rg', 'head', 'tail', 'less', 'more', 'wc', 'ls', 'stat', 'file', 'diff', 'cmp', 'jq',
776  'echo', 'printf', 'test', '[', 'basename', 'dirname', 'realpath', 'readlink', 'which', 'type', 'nl', 'cut', 'tr', 'strings',
777  'xxd', 'od', 'bat', 'md5', 'md5sum', 'shasum', 'sha1sum', 'sha256sum', 'sha512sum', 'cd', 'pushd', 'temper', 'true', 'false',
778])
779// git subcommands that only read.
780const GIT_READS = new Set([
781  'diff', 'log', 'show', 'status', 'ls-files', 'blame', 'grep', 'cat-file', 'rev-parse', 'show-ref', 'check-ignore', 'ls-tree',
782  'shortlog', 'describe', 'rev-list', 'diff-tree', 'diff-files', 'diff-index', 'whatchanged', 'reflog', 'name-rev', 'version', 'help',
783])
784// git subcommands that never change the index.
785const GIT_NO_INDEX = new Set([
786  ...GIT_READS, 'branch', 'push', 'fetch', 'remote', 'tag', 'config', 'init', 'clone', 'switch', 'worktree', 'gc', 'fsck', 'bisect',
787  'submodule', 'prune', 'remote-show', 'notes', 'archive', 'bundle', 'apply',
788])
789const PERMS = new Set(['chmod', 'chown', 'chgrp', 'chflags', 'setfacl', 'chattr', 'xattr'])
790// Commands that change files, as they would be run through xargs.
791const XARGS_WRITERS = new Set(['tee', 'rm', 'touch', 'truncate', 'shred', 'unlink', 'cp', 'mv', 'install', 'ln', 'rsync', 'dd', 'sed', 'perl', 'tar', 'unzip', 'patch', ...PERMS])
792// `-exec` commands of find that only read.
793const FIND_SAFE_EXEC = new Set(['cat', 'grep', 'egrep', 'fgrep', 'rg', 'head', 'tail', 'wc', 'ls', 'stat', 'file', 'echo', 'printf', 'basename', 'dirname', 'shasum', 'md5', 'md5sum', 'sha256sum', 'jq', 'diff', 'cmp', 'test', '['])
794const FIND_PATH_FILTERS = new Set(['-name', '-iname', '-path', '-ipath', '-wholename', '-iwholename', '-regex', '-iregex', '-lname', '-ilname'])
795// Paths find looks at that no command reads from: a temporary folder cannot hold the run.
796const FIND_ELSEWHERE = /^\/(?:private\/)?(?:tmp|var\/folders|var\/tmp)(?:\/|$)|^\/dev(?:\/|$)/
797const FIND_SAMPLES = ['.temper', '.temper/gates.json', '.temper/status.json', '.temper/overrides.json', '.temper/build-state.json', '.temper/feedback-loops.json', '.temper/evidence/build.json', '.temper/specs/x/events/1.json']
798
799// Whether a find filter (-name '*.pyc', -path '*/node_modules/*', -regex ...) can match a guarded file.
800function findFilterCouldMatch(flag: string, value: string): boolean {
801  if (/^-i?regex$/.test(flag)) {
802    try {
803      const re = new RegExp(`^(?:${value})$`, 'i')
804      return FIND_SAMPLES.some(p => re.test(`./${p}`) || re.test(p))
805    } catch {
806      return true
807    }
808  }
809  if (/name$/.test(flag)) {
810    const re = globRegExp(value.replace(/\[[^\]]*\]/g, '?'))
811    return PROTECTED_NAMES.some(n => re.test(n))
812  }
813  const re = globRegExp(value.replace(/\[[^\]]*\]/g, '?'))
814  return FIND_SAMPLES.some(p => re.test(`./${p}`) || re.test(p))
815}
816
817// A find that deletes, writes a file, or runs something that is not a plain reader.
818function findWrites(args: Word[]): boolean {
819  for (let i = 0; i < args.length; i++) {
820    const t = args[i]?.text ?? ''
821    if (t === '-delete' || /^-f(?:print0?|printf|ls)$/.test(t)) return true
822    if (/^-(?:exec|execdir|ok|okdir)$/.test(t)) {
823      if (!FIND_SAFE_EXEC.has(BASE(args[i + 1]?.text ?? '').toLowerCase())) return true
824    }
825  }
826  return false
827}
828
829// Index of the git subcommand in the words after `git`, skipping the options that take a value.
830function gitSubAt(list: string[]): number {
831  let i = 0
832  while (i < list.length) {
833    const t = list[i] ?? ''
834    if (['-C', '-c', '--git-dir', '--work-tree', '--namespace', '--super-prefix', '--exec-path'].includes(t)) i += 2
835    else if (t.startsWith('-')) i += 1
836    else break
837  }
838  return i
839}
840
841export function classifyBash(command: string, startCwd: string | null = ''): BashClass {
842  const text = command.replace(/>\|/g, '>')
843  // The same text without the quotes and backslashes that split a word (`te""mper`, `ov\erride`).
844  const squash = text.replace(/["'\\]/g, '')
845  const strict = NAMED.test(text)
846  const decisions: DecisionKind[] = []
847  const calls: DecisionCall[] = []
848  const stateOps: StateOp[] = []
849  const writes: string[] = []
850  const uncheckable: string[] = []
851  let commits = false
852  let interpreter = false
853  let sawSed = false
854  // Launch facts for the fail closed rule: see BashClass.opaque.
855  let alias = false
856  let mentionAny = false
857  let mentionLoud = false
858  let dynamicCommand = false
859  let shellStdin = false
860  // A Temper call whose subcommand or verb cannot be read from the text ($'..', ${..}, $(..)).
861  let opaqueCall = false
862  // The same for a script found by a substitution: `$T $c`, where the command names the script.
863  let dynamicSub = false
864  // Files this command writes, so a script that is written and then run can be seen.
865  const created: string[] = []
866  // A file this command wrote is run by it (as a command, or as the script of a shell or interpreter).
867  let ranCreated = false
868  // A script file (by its extension) is written.
869  let wroteScript = false
870  // An interpreter reads its program from standard input (a heredoc or a pipe).
871  let stdinProgram = false
872  const staged = { all: false, paths: [] as string[] }
873  // `cat scripts/temper` was read: a following `tee` or redirect makes a copy.
874  let readsScript = false
875  const vars: Vars = new Map()
876  // The working directory the command has moved to (relative to where it started); null when unknown.
877  let cwd: string | null = startCwd
878  // Hardening facts (see BashClass).
879  let hidden = false
880  let envTamper = false
881  let noVerify = false
882  let hookTamper = false
883  let unplainCommit = false
884  const guardedUse: string[] = []
885  // Any statement names a guarded file; a write through xargs is read against it after the walk.
886  let anyMention = false
887  let xargsWriter = false
888  // A shell reads its program from standard input.
889  let stdinShell = false
890
891  const flag = (path: string, unknown = false) => {
892    writes.push(path)
893    if (unknown) uncheckable.push(path)
894  }
895
896  // The paths a word can stand for: every brace expansion with variables filled in, resolved
897  // against the working directory; `null` for one that cannot be resolved.
898  const candidates = (w: Word): Array<string | null> => {
899    const expanded = expandVars(w.text, vars)
900    const alts = braceExpand(expanded)
901    if (alts === null || w.dynamic) return [null]
902    return alts.map(a => {
903      if (/[$`]|^~|[<>]\(/.test(a)) return null
904      if (a.startsWith('/')) return normalizePath(a)
905      if (cwd === null) return null
906      return normalizePath(cwd ? `${cwd}/${a}` : a)
907    })
908  }
909
910  // One write target. `dest` is a directory the sources land in (cp, mv, install, ln, rsync):
911  // such a directory is only refused when a source landing in it would be a guarded file.
912  const check = (w: Word, sources: Word[] = [], isDest = false) => {
913    const raw = expandVars(w.text, vars)
914    const list = candidates(w)
915    for (const r of list) {
916      if (r === null) {
917        const unknownName = /[$`]|\{|[<>]\(/.test(raw.split('/').pop() ?? '') || w.dynamic || (raw.split('/').pop() ?? '') === ''
918        if (strict || unknownName || NAMES_GUARDED.test(raw)) flag(raw, true)
919        continue
920      }
921      if (GLOB.test(r)) {
922        if (strict || globCouldMatch(r)) flag(r, true)
923        continue
924      }
925      const k = protectedKind(r)
926      if (k === null) continue
927      if (isDest && (k === 'folder' || k === 'events')) {
928        const lands = sources.length === 0 ? [''] : sources.map(s => BASE(expandVars(s.text, vars)))
929        const bad = lands.some(b => b === '' || b === '.' || b === '..' || /[$`*?[{]/.test(b) || protectedKind(`${r}/${b}`) !== null && protectedKind(`${r}/${b}`) !== 'folder')
930        if (bad || k === 'events') flag(r)
931        continue
932      }
933      flag(r)
934    }
935  }
936
937  const analyse = (stmt: string, depth: number, pipedFrom: string | null = null): void => {
938    if (depth > 6) return
939    // A substitution or a subshell runs in a copy of the shell: a `cd` inside does not move this one.
940    for (const sub of substitutions(stmt)) {
941      const keep = cwd
942      for (const s of statementsOf(sub)) analyse(s.stmt, depth + 1, s.from)
943      cwd = keep
944    }
945    let ws = wordsOf(stmt)
946    // A subshell or group: analyse what is inside.
947    if (ws.length > 0 && /^\(.*\)$/.test(stmt.trim())) {
948      const keep = cwd
949      for (const s of statementsOf(stmt.trim().slice(1, -1))) analyse(s.stmt, depth + 1, s.from)
950      cwd = keep
951      return
952    }
953
954    // Assignments, in order, left to right, including after export/declare and in an env prefix.
955    for (let guard = 0; guard < 40 && ws.length > 0; guard++) {
956      const first = ws[0]?.text ?? ''
957      if (DECLARE.has(first)) {
958        ws = ws.slice(1).filter((w, i, a) => !(w.text.startsWith('-') && !a.slice(0, i).some(x => !x.text.startsWith('-'))))
959        continue
960      }
961      const m = /^([A-Za-z_]\w*)\+?=(.*)$/.exec(first)
962      if (!m) break
963      if (/^TEMPER_(?:DIR|CONFIG)$/i.test(m[1] ?? '')) envTamper = true
964      vars.set(m[1] ?? '', expandVars(m[2] ?? '', vars))
965      ws = ws.slice(1)
966    }
967    if (ws.length === 0) return
968
969    // Redirects first: they apply whatever the command is.
970    const argv: Word[] = []
971    const targets: string[] = []
972    for (let i = 0; i < ws.length; i++) {
973      const t = ws[i]?.text ?? ''
974      const m = /^(?:\d*|&)>{1,2}(?!&)(.*)$/.exec(t)
975      if (m && !/^\d*>&/.test(t)) {
976        const target = m[1] ? { text: m[1], dynamic: ws[i]?.dynamic ?? false } : ws[++i]
977        if (target) {
978          check(target)
979          targets.push(expandVars(target.text, vars))
980        }
981      } else if (/^\d*<</.test(t) || /^\d*<(?!\()/.test(t)) {
982        if (t === '<' || /^\d*<<<?-?$/.test(t)) i++
983      } else argv.push(ws[i] as Word)
984    }
985    for (const t of targets) {
986      created.push(t)
987      if (SCRIPT_FILE.test(t)) wroteScript = true
988    }
989    // A copy of the Temper script by redirect: `cat scripts/temper > /tmp/t`.
990    if (targets.length > 0 && argv.some(x => namesTemper(expandVars(x.text, vars))) && ['cat', 'head', 'tail', 'dd', 'tee', 'sed', 'awk'].includes(BASE(argv[0]?.text ?? '').toLowerCase())) alias = true
991    // Words that name the script, before any wrapper is taken off (env -S 'scripts/temper ...').
992    const named = argv.some(x => pieceNamesTemper(expandVars(x.text, vars)))
993
994    // The command word, with variables filled in (`T=scripts/temper; $T gate x` runs the script).
995    const fill = (list: Word[]): Word[] => (list[0] ? [{ ...list[0], text: expandVars(list[0].text, vars) }, ...list.slice(1)] : list)
996    const viaXargs = argv.some(x => BASE(x.text).toLowerCase() === 'xargs')
997    let w = fill(unwrap(argv))
998    // TEMPER_DIR / TEMPER_CONFIG set for the command (`env TEMPER_DIR=x ...`, `env -S "TEMPER_CONFIG=x ..."`).
999    if (argv.some(x => /(?:^|[\s;&|(])TEMPER_(?:DIR|CONFIG)\+?=/i.test(x.text))) envTamper = true
1000    // A guarded file named by this command is read, or the command is refused while a run is active.
1001    {
1002      const first = BASE(expandVars(w[0]?.text ?? '', vars)).toLowerCase()
1003      const shellC = SHELLS.has(first) && w.slice(1).some(x => /^-\w*c$/.test(x.text))
1004      if (!shellC && (w.length > 0 || argv.length > 0)) {
1005        let words = (w.length > 0 ? w.slice(1) : argv).map(x => expandVars(x.text, vars))
1006        // A commit message is text, not a file: `git commit -m "fix gates.json"`.
1007        if (first === 'git') {
1008          const skip = new Set<number>()
1009          words.forEach((t, k) => {
1010            if (t === '-m' || t === '--message' || /^-[a-zA-Z]*m$/.test(t)) skip.add(k + 1)
1011            if (/^--message=/.test(t)) skip.add(k)
1012            if (/^-m./.test(t)) skip.add(k)
1013          })
1014          words = words.filter((_, k) => !skip.has(k))
1015        }
1016        // The program of sed and awk is a program, not a path: only a name written out counts there, not a glob.
1017        const progLike = first === 'sed' || first === 'awk' || first === 'gawk'
1018        const hits = words.filter(t => mentionsGuarded(t, !progLike))
1019        if (hits.length > 0) {
1020          anyMention = true
1021          let reads = PLAIN_READERS.has(first)
1022          if (first === 'git') {
1023            const sub = words[gitSubAt(words)]
1024            reads = sub !== undefined && (GIT_READS.has(sub) || sub === 'commit')
1025          }
1026          if (first === 'find') reads = !findWrites(w.slice(1))
1027          // sed reads unless it edits in place or runs a command; its `w file` is read from the whole text below.
1028          if (first === 'sed') reads = !words.some(t => /^-[a-zA-Z]*i|^--in-place/.test(t)) && !readerDanger('sed', w.slice(1))
1029          // awk reads unless a guarded name sits in its program (a redirect inside it writes) or it edits in place.
1030          if (first === 'awk' || first === 'gawk') {
1031            const progAt = words.findIndex((t, k) => !t.startsWith('-') && !['-v', '-F', '-f', '-i'].includes(words[k - 1] ?? ''))
1032            reads = !readerDanger(first, w.slice(1)) && !words.some(t => /^-[a-zA-Z]*i$|^--in-place$/.test(t)) && (progAt < 0 || !mentionsGuarded(words[progAt] ?? '', false))
1033          }
1034          // The plugin's own acceptance checker reads the evidence ledger and prints.
1035          if (INTERPRETERS.test(first) && /(^|\/)scripts\/acceptance\.py$/.test(words.find(t => !t.startsWith('-')) ?? '')) reads = true
1036          if (!reads) guardedUse.push(hits[0] ?? '')
1037        }
1038      }
1039    }
1040    if (XARGS_WRITERS.has(BASE(w[0]?.text ?? '').toLowerCase()) && viaXargs) xargsWriter = true
1041    if (viaXargs && BASE(w[0]?.text ?? '').toLowerCase() === 'git') staged.all = true
1042    // A file this command wrote is run: as the command itself, or as the script of a shell,
1043    // interpreter, make or source.
1044    {
1045      const first = expandVars(w[0]?.text ?? '', vars)
1046      const base = BASE(first).toLowerCase()
1047      const runner = SHELLS.has(base) || INTERPRETERS.test(base) || EXEC.has(base) || base === 'source' || base === '.'
1048      if (created.includes(first) || (runner && w.slice(1).some(x => created.includes(expandVars(x.text, vars))))) ranCreated = true
1049    }
1050    // `source scripts/temper ...` and `. scripts/temper ...` run the script in this shell.
1051    if (['source', '.'].includes(w[0]?.text ?? '') && w.slice(1).some(a => namesTemper(expandVars(a.text, vars)))) alias = true
1052    // A shell given the script as a file (after options such as -o), -s, or a program string after -c.
1053    while (w.length > 0 && SHELLS.has(BASE(w[0]?.text ?? '').toLowerCase())) {
1054      const rest = w.slice(1)
1055      // The shell reads its commands from standard input: a heredoc or a here-string is in the text; a pipe is shown
1056      // only when an echo or a printf (or a cat of a heredoc) feeds it.
1057      const fromStdin = (): void => {
1058        shellStdin = true
1059        stdinShell = true
1060        const shown = /<</.test(stmt) || (pipedFrom !== null && (/^(?:echo|printf)\s/.test(pipedFrom) || (/^cat\b/.test(pipedFrom) && /<</.test(pipedFrom))))
1061        if (!shown) hidden = true
1062      }
1063      // The startup files the shell reads before its program or script: the value after --rcfile or --init-file among
1064      // its options, BASH_ENV for bash, and ENV for an interactive shell (-i). BASH_ENV and ENV count when this command set
1065      // them, before the shell or through env. One that names standard input runs the program from the pipe.
1066      {
1067        const startup: string[] = []
1068        let interactive = false
1069        for (let i = 0; i < rest.length; i++) {
1070          const t = rest[i]?.text ?? ''
1071          if (/^-[a-zA-Z]*i[a-zA-Z]*$/.test(t)) interactive = true
1072          if (t === '--rcfile' || t === '--init-file') startup.push(rest[i + 1]?.text ?? '')
1073          if (['-o', '+o', '-O', '+O', '--rcfile', '--init-file'].includes(t)) i++
1074          else if (!/^[-+]/.test(t)) break
1075        }
1076        const setHere = (name: string): string[] => [
1077          ...(vars.has(name) ? [vars.get(name) ?? ''] : []),
1078          ...argv.flatMap(x => (x.text.startsWith(`${name}=`) ? [x.text.slice(name.length + 1)] : [])),
1079        ]
1080        if (BASE(w[0]?.text ?? '').toLowerCase() === 'bash') startup.push(...setHere('BASH_ENV'))
1081        if (interactive) startup.push(...setHere('ENV'))
1082        const fed = fedInput(stmt, pipedFrom)
1083        if (startup.some(f => f !== '' && readsStdin(f, vars, fed, cwd))) {
1084          fromStdin()
1085          return
1086        }
1087      }
1088      const ci = rest.findIndex(x => /^-\w*c$/.test(x.text))
1089      if (ci >= 0) {
1090        const script = rest[ci + 1]
1091        const scriptText = expandVars(script?.text ?? '', vars)
1092        if (script?.dynamic || /[$`]/.test(script?.text ?? '')) shellStdin = true
1093        // The program is built by a substitution (or is a variable nothing set): it is not shown by the text.
1094        if (script?.dynamic ? hiddenText(scriptText) : /^\s*\$/.test(scriptText)) hidden = true
1095        // xargs hands the words it reads to the shell as its program: a -c program that is only the {} placeholder, or
1096        // no program word at all (the first word xargs reads is the program).
1097        if (viaXargs && (script === undefined || /^\s*(?:\{\}|"?\$(?:@|\*|\d)"?)\s*$/.test(scriptText))) hidden = true
1098        // A string given to a shell: git commit inside it counts, and so does a name of the script.
1099        if (/\bgit\b/i.test(scriptText) && gitCreatesCommit(scriptText)) commits = true
1100        for (const s of statementsOf(scriptText)) analyse(s.stmt, depth + 1, s.from)
1101        if (named || pieceNamesTemper(scriptText)) mentionAny = true
1102        return
1103      }
1104      // The first word that is not an option (an option such as -o, --rcfile or --init-file takes a value) is the script.
1105      let fileAt = -1
1106      let sSeen = false
1107      for (let i = 0; i < rest.length; i++) {
1108        const t = rest[i]?.text ?? ''
1109        if (t === '-s') {
1110          shellStdin = true
1111          sSeen = true
1112        }
1113        if (['-o', '+o', '-O', '+O', '--rcfile', '--init-file'].includes(t)) i++
1114        else if (!/^[-+]/.test(t)) {
1115          fileAt = i
1116          break
1117        }
1118      }
1119      // A script argument that names standard input in any spelling (see readsStdin) is a program read from the
1120      // pipe, the same as no script file at all.
1121      const stdinFile = fileAt >= 0 && readsStdin(rest[fileAt]?.text ?? '', vars, fedInput(stmt, pipedFrom), cwd)
1122      if (fileAt < 0 || sSeen || stdinFile) {
1123        // No script file (or -s with arguments, the words after the options are arguments): the shell reads its
1124        // commands from standard input.
1125        fromStdin()
1126        return
1127      }
1128      w = fill(unwrap(rest.slice(fileAt)))
1129      if (BASE(w[0]?.text ?? '').toLowerCase() === 'temper') break
1130      // A script that is not named temper: a glob or a variable could still stand for it, and a
1131      // script given by a process substitution cannot be read.
1132      if (namesTemper(w[0]?.text ?? '') || named) {
1133        mentionAny = true
1134        mentionLoud = true
1135      }
1136      if (w[0]?.dynamic || /[$`]|[<>]\(/.test(w[0]?.text ?? '')) dynamicCommand = true
1137      return
1138    }
1139    if (w.length === 0) {
1140      // Everything was a wrapper: `env -S 'scripts/temper ...'` holds its command in a string.
1141      if (named) {
1142        mentionAny = true
1143        mentionLoud = true
1144      }
1145      return
1146    }
1147    const cmd = BASE(w[0]?.text ?? '').toLowerCase()
1148    const args = w.slice(1)
1149    const argText = args.map(a => expandVars(a.text, vars))
1150
1151    // Facts for the fail closed rule: a statement that is not a plain Temper call but names the
1152    // script (a word, a glob that can match it, a name inside a string, a directory called temper in
1153    // command position) or runs a command that cannot be read. Reading commands are exempt.
1154    // A command word that cannot be read, followed by a plain Temper read or record call, is a script found by a
1155    // command substitution, then used for reads (`T=$(command -v temper); $T state get next_stage`): it is not read
1156    // as a mention of the script at all.
1157    const plainDynamic = (w[0]?.dynamic || /[$`]|[<>]\(/.test(w[0]?.text ?? '')) && plainTemperArgs(args, argText)
1158    if ((cmd !== 'temper' || viaXargs) && !plainDynamic) {
1159      const cmdText = w[0]?.text ?? ''
1160      // `python3 -m pytest -k "temper and accept"` runs tests: it only reads the word.
1161      const testRun = INTERPRETERS.test(cmd) && args.some((a, k) => a.text === '-m' && /^(?:pytest|unittest|coverage)$/.test(argText[k + 1] ?? ''))
1162      const safeReader = (READERS.has(cmd) && !readerDanger(cmd, args)) || testRun
1163      if ((cmd === 'make' || cmd === 'gmake') && args.some(a => a.text === '-f') && !args.some(a => /^[^-]/.test(a.text) && a.text !== '')) stdinProgram = true
1164      const names = named || w.some(x => namesTemper(expandVars(x.text, vars))) || cmdText.toLowerCase().split('/').includes('temper')
1165      if (names) {
1166        mentionAny = true
1167        if (!safeReader) mentionLoud = true
1168        if (cmd === 'cat') readsScript = true
1169      }
1170      if (viaXargs && cmd === 'temper') mentionLoud = true
1171      // A command word that cannot be read (`T=$(command -v temper); $T ...`) may be a script found by a command
1172      // substitution. It is fine for plain reads; anything else about it fails closed.
1173      if ((w[0]?.dynamic || /[$`]|[<>]\(/.test(cmdText)) && !plainTemperArgs(args, argText)) {
1174        dynamicCommand = true
1175        // The script found that way, given a subcommand that cannot be read either (`T=$(ls scripts/temper); $T $c`).
1176        if (names && args[0] !== undefined && (args[0].dynamic || /[$`]/.test(argText[0] ?? ''))) dynamicSub = true
1177      }
1178      // An interpreter program (python -c, node -e, perl -e) or a program on standard input that names
1179      // the script. A file run by an interpreter, and a test run (-m pytest -k "temper"), are not read here.
1180      if (INTERPRETERS.test(cmd)) {
1181        const flagAt = args.findIndex(a => PROGRAM_FLAG.test(a.text))
1182        const program = flagAt >= 0 ? argText.slice(flagAt + 1) : []
1183        if (program.some(a => /temper|subprocess/i.test(a))) mentionLoud = true
1184        if (!args.some(a => !a.text.startsWith('-')) && flagAt < 0) stdinProgram = true
1185        // A program file that is standard input in any spelling (python3 /dev/stdin) is a program on standard input too.
1186        const fileArg = args.find(a => !a.text.startsWith('-'))
1187        if (flagAt < 0 && fileArg !== undefined && readsStdin(fileArg.text, vars, fedInput(stmt, pipedFrom), cwd)) stdinProgram = true
1188      }
1189    }
1190    // A git command that creates a commit, wherever it sits (find -exec, env, watch, ...).
1191    if (cmd !== 'git' && cmd !== 'temper' && (!READERS.has(cmd) || cmd === 'find') && gitCreatesCommit(argText.join(' '))) {
1192      commits = true
1193      unplainCommit = true
1194    }
1195
1196    if (cmd === 'eval') {
1197      const joined = args.map(a => expandVars(a.text, vars)).join(' ')
1198      // A program built by a substitution (not a shell setup idiom from ENV_NAMES) is not shown.
1199      if ((args.some(a => a.dynamic) || /[$`]/.test(joined)) && hiddenText(joined)) hidden = true
1200      for (const s of statementsOf(stripSafe(joined))) analyse(s.stmt, depth + 1, s.from)
hooks/temper-mod/core/cli.ts 90 lines
1// The exact `scripts/temper` invocations the mod asks Claude to run to mirror a person's
2// decision in the CLI state. Pure. The stage names are those of STAGE_SEQ_TEMPER in
3// scripts/temper (a test there keeps the two in step).
4
5import type { Phase } from './events'
6
7export const CLI = 'scripts/temper'
8
9// Where the mod's module file sits in the plugin folder: the plugin folder is its module URL with this taken off.
10export const MOD_FILE = '/hooks/temper-mod/register.tsx'
11
12// The Temper script in the plugin folder as a full path, from the module URL of the mod (<plugin> followed by
13// MOD_FILE). The path is decoded FIRST and then checked, so an encoded space or quote (%20, %27) cannot reach a
14// command. Any odd location gives `scripts/temper`.
15export function pluginCliFrom(url: string | undefined): string {
16  try {
17    const here = decodeURIComponent(new URL(url ?? '').pathname)
18    if (here.endsWith(MOD_FILE) && !/[\s'"`$;&|<>()\\]/.test(here)) return `${here.slice(0, -MOD_FILE.length)}/scripts/temper`
19  } catch {
20    // no module URL here
21  }
22  return CLI
23}
24
25// The plugin folder, from where the Temper script is (what pluginCliFrom gave). Null when only the plain
26// `scripts/temper` is known: a prompt then names the plain path and says where it is (IN_PLUGIN).
27export function pluginRootOf(cli: string): string | null {
28  const tail = '/scripts/temper'
29  return cli !== CLI && cli.startsWith('/') && cli.endsWith(tail) ? cli.slice(0, -tail.length) : null
30}
31
32// What a prompt or a deny text adds when the Temper script is only known by its plain name.
33export const IN_PLUGIN = 'The script is in the Temper plugin folder, not in the project.'
34
35// STAGE_SEQ_TEMPER, in order. `state advance` takes `<stage>_complete <next stage>`.
36export const CLI_STAGES = 'intent plan design build review check'
37
38const STAGES = CLI_STAGES.split(' ')
39
40export const isCliStage = (s: string): boolean => STAGES.includes(s)
41
42// `state advance` for a move the person made, one command per CLI stage the move covers. Plan
43// is followed by design when the run's complexity is medium or complex (design belongs to the
44// Plan phase here), and Check is followed by commit. A move out of Fix has no CLI stage.
45export function advanceCommands(from: Phase, to: Phase | 'done', complexity: string | null): string[] {
46  const state = (stage: string, next: string) => `${CLI} state advance ${stage}_complete ${next}`
47  switch (from) {
48    case 'intent':
49      return [state('intent', 'plan')]
50    case 'plan':
51      return complexity === 'medium' || complexity === 'complex'
52        ? [state('plan', 'design'), state('design', 'build')]
53        : [state('plan', 'build')]
54    case 'build':
55      return [state('build', 'review')]
56    case 'review':
57      return [state('review', 'check')]
58    case 'check':
59      return to === 'done' ? [state('check', 'commit')] : []
60    default:
61      return []
62  }
63}
64
65// The gate stage an override or an advance names, for a phase of the mod.
66export const stageOf = (phase: Phase): string => (phase === 'fix' ? 'check' : phase)
67
68// A reason is person typed text inside a shell command: single quoted, quotes escaped, newlines
69// flattened, so nothing in it (quotes, $(...), backticks) is read as shell.
70export const shellQuote = (text: string): string => `'${text.replace(/\r?\n/g, ' ').replace(/'/g, "'\\''")}'`
71
72export const overrideCommand = (phase: Phase, reason: string): string => `${CLI} override ${stageOf(phase)} --reason ${shellQuote(reason)}`
73
74export const acceptCommand = (findingId: string, reason: string): string => `${CLI} evidence accept --stage review --id ${findingId} --reason ${shellQuote(reason)}`
75
76// Sending a run back: the CLI resumes from the stage it is pointed at.
77export const backCommand = (to: Phase): string => `${CLI} state set next_stage ${stageOf(to)}`
78
79// The loop the CLI keeps for a step back: it counts against loops.max-per-type and clears the evidence of
80// the stage that is redone and every later one.
81export const loopCommand = (from: Phase, to: Phase, reason: string): string => `${CLI} state loop ${stageOf(from)} ${stageOf(to)} --reason ${shellQuote(reason)}`
82
83// The stage of a `state advance <stage>_complete <next>` call, as a phase of the mod, only for
84// the two approvals that need a person: leaving Intent and leaving Plan. Design belongs to
85// Plan but its own advance needs no second approval, so it is not mapped.
86export function guardedPhase(stageArg: string | undefined): Phase | null {
87  const stage = (stageArg ?? '').replace(/_complete$/, '')
88  return stage === 'intent' || stage === 'plan' ? stage : null
89}
90
hooks/temper-mod/core/commands.ts 182 lines
1// The reserved first words of `/temper:temper` and what each one means. Pure: this file turns
2// typed arguments into a machine command, a local answer, or an error; the adapter runs
3// the result. Anything that is not a reserved word is left to the prompt based /temper:temper.
4
5import type { Draft, DriftChoice, Phase } from './events'
6import { PHASES } from './events'
7import { ACT } from './actions'
8import { CLI, IN_PLUGIN, acceptCommand, advanceCommands, backCommand, loopCommand, overrideCommand } from './cli'
9import type { Command } from './machine'
10
11export const RESERVED = [
12  'status',
13  'timeline',
14  'back',
15  'override',
16  'approve',
17  'accept',
18  'drift',
19  'pause',
20  'resume',
21  'help',
22  'report',
23  'pr',
24  'mode',
25  'enforcement',
26  'next',
27  'pane',
28  'play',
29  'discuss',
30  'continue',
31] as const
32
33export type Reserved = (typeof RESERVED)[number]
34
35export const isReserved = (w: string): w is Reserved => (RESERVED as readonly string[]).includes(w)
36
37export type Parsed = { word: Reserved; rest: string }
38
39// null when the first word is not reserved (a feature description, or nothing).
40export function parseArgs(args: string): Parsed | null {
41  const text = args.trim()
42  const m = /^(\S+)\s*([\s\S]*)$/.exec(text)
43  const word = (m?.[1] ?? '').toLowerCase()
44  return m && isReserved(word) ? { word, rest: (m[2] ?? '').trim() } : null
45}
46
47// What the machine is asked to do, before the origin is stamped on.
48export type Bare = Command extends infer C ? (C extends unknown ? Omit<C, 'origin' | 'author'> : never) : never
49
50export type Plan =
51  | { kind: 'command'; command: Bare }
52  | { kind: 'local'; word: Reserved; rest: string }
53  | { kind: 'error'; text: string }
54
55const FLOW_PHASES: readonly string[] = PHASES.filter(p => p !== 'fix')
56
57const split = (rest: string): [string, string] => {
58  const m = /^(\S+)\s*([\s\S]*)$/.exec(rest.trim())
59  return [m?.[1] ?? '', (m?.[2] ?? '').trim()]
60}
61
62export function planCommand(parsed: Parsed, pendingDrift: string | null): Plan {
63  const { word, rest } = parsed
64  switch (word) {
65    case 'approve':
66      return { kind: 'command', command: { type: 'approve' } }
67    case 'next':
68      return { kind: 'command', command: { type: 'advance' } }
69    case 'pause':
70      return { kind: 'command', command: { type: 'pause' } }
71    case 'resume':
72      return { kind: 'command', command: { type: 'resume' } }
73    case 'override':
74      return { kind: 'command', command: { type: 'override', reason: rest } }
75    case 'back': {
76      const [to, reason] = split(rest)
77      if (!FLOW_PHASES.includes(to.toLowerCase())) return { kind: 'error', text: 'Usage: /temper:temper back <intent|plan|build|review|check> <reason>' }
78      return { kind: 'command', command: { type: 'back', to: to.toLowerCase() as Phase, reason } }
79    }
80    case 'accept': {
81      const [id, reason] = split(rest)
82      if (!id) return { kind: 'error', text: 'Usage: /temper:temper accept <finding id> <reason>' }
83      return { kind: 'command', command: { type: 'acceptFinding', id, reason } }
84    }
85    case 'drift': {
86      const [choiceWord, reason] = split(rest)
87      const choices: Record<string, DriftChoice> = { add: 'add', revert: 'revert', allow: 'allow-once', 'allow-once': 'allow-once' }
88      const choice = choices[choiceWord.toLowerCase()]
89      if (!choice) return { kind: 'error', text: 'Usage: /temper:temper drift <add|revert|allow> <reason>' }
90      if (!pendingDrift) return { kind: 'error', text: 'No scope drift waits for a decision.' }
91      return { kind: 'command', command: { type: 'drift', path: pendingDrift, choice, reason } }
92    }
93    default:
94      return { kind: 'local', word, rest }
95  }
96}
97
98// The command a button runs after the mirror prompt (`$.command.run`, as `/temper:temper` with no
99// arguments), so the orchestrator (commands/temper.md, its Resume path) launches the next stage with
100// its own brief. No arguments: the decision is already recorded and mirrored.
101export const RESUME = 'temper:temper'
102
103// What Claude is asked to do once the person decided with a button (the command path runs the
104// prompt based command instead). The decision is already recorded by the mod: the prompt says so
105// and names the exact CLI command that mirrors it in the CLI state (the commit gate reads that
106// state). It never says what to run next: RESUME does that, through the orchestrator.
107// Every command named here is a valid `scripts/temper` invocation.
108// `cli` is where the Temper script really is (the plugin folder, not the project). Without it the
109// prompt names `scripts/temper` and says the script lives in the plugin folder, so Claude does not
110// waste the one allowed run on a path that does not exist in the project.
111// `from` is the phase a back step leaves (the loop the CLI keeps is from -> to). Without it, only the step is mirrored.
112export function followUp(draft: Draft, complexity: string | null = null, cli: string = CLI, from: Phase | null = null): string | null {
113  const text = followUpText(draft, complexity, from)
114  if (text === null) return null
115  if (cli !== CLI) return text.split(`\`${CLI} `).join(`\`${cli} `)
116  return text.includes(`\`${CLI} `) ? text.replace(` ${ACT}`, ` ${IN_PLUGIN} ${ACT}`) : text
117}
118
119function followUpText(draft: Draft, complexity: string | null, from: Phase | null): string | null {
120  // Short on purpose: one or two lines, and no narration.
121  const recorded = 'The decision is already recorded.'
122  switch (draft.type) {
123    case 'advance': {
124      const cmds = advanceCommands(draft.from, draft.to, complexity)
125      if (cmds.length === 0) return null
126      const run = ` Run ${cmds.map(c => `\`${c}\``).join(' then ')}.`
127      if (draft.to === 'done') return `Temper: the user finished the run. ${recorded}${run} Do not commit. ${ACT}`
128      return `Temper: the user moved the run from ${draft.from} to ${draft.to}. ${recorded}${run} Do not start the next stage yourself. ${ACT}`
129    }
130    case 'override':
131      return `Temper: the user skipped ${draft.phase} (reason: ${draft.reason}). ${recorded} Run \`${overrideCommand(draft.phase, draft.reason)}\`. Do not start the next stage yourself. ${ACT}`
132    case 'accept':
133      return `Temper: the user accepted finding ${draft.findingId} (reason: ${draft.reason}). ${recorded} Run \`${acceptCommand(draft.findingId, draft.reason)}\`. ${ACT}`
134    case 'back':
135      // A loop back is a loop of the CLI: `state loop` keeps the budget (loops.max-per-type) and clears the evidence of
136      // the stages that are redone. The guard lets it through once, for this decision. The step itself follows it.
137      if (from !== null && from !== draft.to) {
138        return (
139          `Temper: the user went back to ${draft.to} (reason: ${draft.reason}). ${recorded} ` +
140          `Run \`${loopCommand(from, draft.to, draft.reason)}\`. If it prints BLOCKED, the loop budget is spent: say so and stop. ` +
141          `Otherwise run \`${backCommand(draft.to)}\`. Do not start the stage yourself. ${ACT}`
142        )
143      }
144      return `Temper: the user went back to ${draft.to} (reason: ${draft.reason}). ${recorded} Run \`${backCommand(draft.to)}\`. Do not start the stage yourself. ${ACT}`
145    case 'drift':
146      return draft.choice === 'revert'
147        ? `Temper: the user chose to revert ${draft.path}. Restore it to its committed state. Then stay inside the plan. ${ACT}`
148        : `Temper: the user decided the scope drift for ${draft.path} (${draft.choice}). Go on. ${ACT}`
149    default:
150      return null
151  }
152}
153
154// The answer to a word that changes the run when it did not come from the person's prompt box (a `claude -p` run, the
155// Agent SDK, Claude itself). It says where a decision is made, and never sends the person back to the same command.
156export const NEEDS_INTERACTIVE =
157  'Temper: decisions need an interactive session. Only the user decides there, with the Temper bar or with /temper:temper approve (or back, override, accept) typed in the prompt box. ' +
158  'This command came from somewhere else, such as a claude -p run, the Agent SDK or Claude. Next: continue the run in an interactive claude session.'
159
160// The words that record a decision of the person.
161export const DECISION_WORDS: readonly string[] = ['approve', 'next', 'back', 'override', 'accept', 'drift']
162
163export const HELP = [
164  'Temper subcommands (type them after /temper:temper):',
165  '  status               show where the run is',
166  '  timeline             show the phases of the run',
167  '  approve              approve the current phase (Intent or Plan)',
168  '  next                 move on when the step has passed its check',
169  '  back <phase> <why>   go back (later steps need a new check)',
170  '  override <reason>    skip the current step once (give a reason)',
171  '  accept <id> <why>    accept a review finding (give a reason)',
172  '  drift <add|revert|allow> <reason>   decide a scope drift',
173  '  pause / resume       take the run over, give it back',
174  '  report               show the run report (with no run active, the last one kept)',
175  '  pr                   ask Claude for a pull request description',
176  '  play                 play Temper Run while you wait (key 8 too; r runs, w jumps, s ducks, q leaves)',
177  '  discuss <text>       send a message about the step you are at (the same as key 4)',
178  '  continue <stage>     the Temper bar sends this after you chose Continue: Claude does the On Continue steps of that stage',
179  '  mode, enforcement, pane   show or change what Temper shows and enforces',
180  'Any other text after /temper:temper is a feature description. It starts or resumes a run.',
181].join('\n')
182
hooks/temper-mod/core/config.ts 112 lines
1// The few keys the mod reads: `.claude/temper.config` (a small YAML subset) and the
2// plain-string userConfig fields. userConfig declares no `options` (that would stop the
3// whole plugin loading before Claude Code 2.1.271), so every value is validated here.
4
5export type UiMode = 'full' | 'minimal' | 'off'
6
7// Value of a dotted key such as `fix.max-loops`, read by indentation. Comments, inline
8// comments and surrounding quotes are dropped. Null when the key is absent.
9export function readConfigValue(text: string, dotted: string): string | null {
10  const want = dotted.split('.')
11  const stack: Array<{ indent: number; key: string }> = []
12  for (const raw of text.split('\n')) {
13    if (raw.trim() === '' || raw.trimStart().startsWith('#')) continue
14    const m = /^(\s*)([A-Za-z0-9_-]+):\s*(.*)$/.exec(raw)
15    if (!m) continue
16    const indent = (m[1] ?? '').length
17    const key = m[2] ?? ''
18    for (let top = stack[stack.length - 1]; top !== undefined && top.indent >= indent; top = stack[stack.length - 1]) stack.pop()
19    const path = [...stack.map(s => s.key), key]
20    const value = (m[3] ?? '').replace(/\s+#.*$/, '').trim()
21    if (value === '') {
22      stack.push({ indent, key })
23      continue
24    }
25    if (path.length === want.length && path.every((k, i) => k === want[i])) {
26      return value.replace(/^(['"])(.*)\1$/, '$2')
27    }
28  }
29  return null
30}
31
32const norm = (v: string | undefined): string => (v ?? '').trim().toLowerCase()
33
34export function parseUiMode(v: string | undefined): UiMode {
35  const n = norm(v)
36  return n === 'minimal' || n === 'off' || n === 'full' ? n : 'full'
37}
38
39export function parseOnOff(v: string | undefined, fallback: 'on' | 'off'): 'on' | 'off' {
40  const n = norm(v)
41  return n === 'on' || n === 'off' ? n : fallback
42}
43
44// The game setting: `on` offers the game while Claude works and answers the command, `command`
45// answers the command only, `off` hides it all. Checked here, never as a picker.
46export type GameMode = 'on' | 'command' | 'off'
47
48export function parseGameMode(v: string | undefined): GameMode {
49  const n = norm(v)
50  // Off has synonyms, so a person who writes false or no gets what they mean. Anything that is
51  // not a known word (a typo, an empty value) is on, the default, and shows no message.
52  if (['off', 'false', 'no', '0', 'disabled', 'none'].includes(n)) return 'off'
53  return n === 'command' ? 'command' : 'on'
54}
55
56export const parseEnforcement =(v: string | undefined): 'on' | 'off' => parseOnOff(v, 'on')
57
58const positiveInt = (v: string | null | undefined): number | null => {
59  if (v === null || v === undefined || !/^\s*\d+\s*$/.test(v)) return null
60  const n = Number(v)
61  return n >= 1 ? n : null
62}
63
64// fix.max-loops in temper.config wins, then the userConfig field, then 3.
65export function parseMaxLoops(configText: string, userValue: string | undefined): number {
66  return positiveInt(readConfigValue(configText, 'fix.max-loops')) ?? positiveInt(userValue) ?? 3
67}
68
69// "build=sonnet, plan=opus" -> { build: 'sonnet', plan: 'opus' }; malformed pairs skipped.
70export function parsePhaseModels(v: string | undefined): Record<string, string> {
71  const out: Record<string, string> = {}
72  for (const pair of (v ?? '').split(',')) {
73    const m = /^\s*([A-Za-z-]+)\s*=\s*(\S+)\s*$/.exec(pair)
74    if (m?.[1] && m[2]) out[m[1].toLowerCase()] = m[2]
75  }
76  return out
77}
78
79export const MIN_VERSION = '2.1.287'
80
81// "2.1.288" or "2.1.280-dev.2026..." -> [2, 1, 288 or 280]; null when not a release.
82export function parseVersion(v: string): [number, number, number] | null {
83  const m = /^(\d+)\.(\d+)\.(\d+)/.exec(v.trim())
84  return m ? [Number(m[1]), Number(m[2]), Number(m[3])] : null
85}
86
87// True when `v` is at least `min`. An unparseable version is not supported: the mod
88// stays inert rather than guess.
89export function versionAtLeast(v: string | undefined, min: string = MIN_VERSION): boolean {
90  const have = parseVersion(v ?? '')
91  const want = parseVersion(min)
92  if (!have || !want) return false
93  for (let i = 0; i < 3; i++) {
94    const a = have[i] ?? 0
95    const b = want[i] ?? 0
96    if (a !== b) return a > b
97  }
98  return true
99}
100
101export type Effort = 'low' | 'medium' | 'high' | 'xhigh' | 'max'
102const EFFORTS: readonly string[] = ['low', 'medium', 'high', 'xhigh', 'max']
103
104// "sonnet", "sonnet:high" or ":high" -> a model, an effort, or both; empty parts stay out.
105export function parsePhaseModel(value: string | undefined): { model?: string; effort?: Effort } {
106  const [model = '', effort = ''] = (value ?? '').split(':').map(s => s.trim())
107  const out: { model?: string; effort?: Effort } = {}
108  if (model) out.model = model
109  if (EFFORTS.includes(effort)) out.effort = effort as Effort
110  return out
111}
112
hooks/temper-mod/core/events.ts 132 lines
1// Event codec: the shape of one Temper event and how it is written and read back.
2// Pure. Nothing here touches a file; the adapter hands over file names and text and
3// writes the text this module returns.
4
5export type Phase = 'intent' | 'plan' | 'build' | 'review' | 'check' | 'fix'
6
7export const PHASES: readonly Phase[] = ['intent', 'plan', 'build', 'review', 'check', 'fix']
8
9export type Origin = 'person' | 'model' | 'system'
10
11export type DriftChoice = 'add' | 'revert' | 'allow-once'
12
13export type EventBody =
14  | { type: 'start'; slug: string; title: string; phase?: Phase }
15  | { type: 'advance'; from: Phase; to: Phase | 'done' }
16  | { type: 'back'; to: Phase; reason: string }
17  | { type: 'override'; phase: Phase; reason: string }
18  | { type: 'accept'; findingId: string; reason: string }
19  | { type: 'drift'; path: string; choice: DriftChoice; reason: string }
20  | { type: 'driftUsed'; path: string }
21  | { type: 'checkResult'; result: 'pass' | 'fail' }
22  | { type: 'pause' }
23  | { type: 'resume' }
24
25export type Draft = EventBody & { origin: Origin; author?: string }
26
27export type StampMeta = { ts: number; session: string; seq: number }
28
29export type TemperEvent = Draft & StampMeta & { id: string }
30
31export type DecodeResult = { ok: true; event: TemperEvent } | { ok: false; reason: string }
32
33export type EventFile = { name: string; text: string }
34
35export type Unreadable = { name: string; reason: string }
36
37export function stamp(draft: Draft, meta: StampMeta): TemperEvent {
38  return { ...draft, ...meta, id: `${meta.ts}-${meta.session}-${meta.seq}` }
39}
40
41export function eventFileName(ev: { id: string }): string {
42  return `${ev.id}.json`
43}
44
45export function encodeEvent(ev: TemperEvent): string {
46  return JSON.stringify(ev)
47}
48
49const ORIGINS: readonly string[] = ['person', 'model', 'system']
50const CHOICES: readonly string[] = ['add', 'revert', 'allow-once']
51
52const isString = (v: unknown): v is string => typeof v === 'string'
53const isPhase = (v: unknown): v is Phase => isString(v) && (PHASES as readonly string[]).includes(v)
54
55// Required string/phase fields per event type, checked on decode.
56function bodyProblem(o: Record<string, unknown>): string | null {
57  switch (o.type) {
58    case 'start':
59      return isString(o.slug) && isString(o.title) && (o.phase === undefined || isPhase(o.phase))
60        ? null
61        : 'start needs slug and title'
62    case 'advance':
63      return isPhase(o.from) && (o.to === 'done' || isPhase(o.to)) ? null : 'advance needs from and to'
64    case 'back':
65      return isPhase(o.to) && isString(o.reason) ? null : 'back needs to and reason'
66    case 'override':
67      return isPhase(o.phase) && isString(o.reason) ? null : 'override needs phase and reason'
68    case 'accept':
69      return isString(o.findingId) && isString(o.reason) ? null : 'accept needs findingId and reason'
70    case 'drift':
71      return isString(o.path) && isString(o.reason) && CHOICES.includes(o.choice as string)
72        ? null
73        : 'drift needs path, choice and reason'
74    case 'driftUsed':
75      return isString(o.path) ? null : 'driftUsed needs path'
76    case 'checkResult':
77      return o.result === 'pass' || o.result === 'fail' ? null : 'checkResult needs result'
78    case 'pause':
79    case 'resume':
80      return null
81    default:
82      return 'unknown event type'
83  }
84}
85
86export function decodeEvent(text: string): DecodeResult {
87  let raw: unknown
88  try {
89    raw = JSON.parse(text)
90  } catch {
91    return { ok: false, reason: 'not valid JSON (torn write?)' }
92  }
93  if (typeof raw !== 'object' || raw === null || Array.isArray(raw)) {
94    return { ok: false, reason: 'not an event object' }
95  }
96  const o = raw as Record<string, unknown>
97  if (!isString(o.id) || typeof o.ts !== 'number' || !isString(o.session) || typeof o.seq !== 'number') {
98    return { ok: false, reason: 'missing id, ts, session or seq' }
99  }
100  if (!isString(o.origin) || !ORIGINS.includes(o.origin)) {
101    return { ok: false, reason: 'missing or unknown origin' }
102  }
103  const problem = bodyProblem(o)
104  if (problem !== null) return { ok: false, reason: problem }
105  return { ok: true, event: o as unknown as TemperEvent }
106}
107
108export function compareEvents(a: TemperEvent, b: TemperEvent): number {
109  if (a.ts !== b.ts) return a.ts - b.ts
110  if (a.session !== b.session) return a.session < b.session ? -1 : 1
111  return a.seq - b.seq
112}
113
114// Only `*.json` files are considered events; other names are ignored, not reported.
115export function readEvents(files: readonly EventFile[]): { events: TemperEvent[]; unreadable: Unreadable[] } {
116  const events: TemperEvent[] = []
117  const unreadable: Unreadable[] = []
118  for (const f of files) {
119    if (!f.name.endsWith('.json')) continue
120    const out = decodeEvent(f.text)
121    if (!out.ok) {
122      unreadable.push({ name: f.name, reason: out.reason })
123    } else if (eventFileName(out.event) !== f.name) {
124      unreadable.push({ name: f.name, reason: 'file name does not match the event id' })
125    } else {
126      events.push(out.event)
127    }
128  }
129  events.sort(compareEvents)
130  return { events, unreadable }
131}
132
hooks/temper-mod/core/machine.ts 315 lines
1// The phase machine. Event sourced and pure: `reduce` folds the event log (plus the
2// verdicts the CLI wrote) into a RunState, `decide` turns one command into drafts of
3// new events or an error. Nothing here reads a file or knows how events are stored.
4
5import { compareEvents } from './events'
6import type { Draft, DriftChoice, Origin, Phase, TemperEvent } from './events'
7
8export type Verdict = { verdict: 'PASS' | 'FAIL'; ts: number }
9
10// The CLI's verdicts per stage (from .temper/gates.json); the mod only reads them.
11export type Verdicts = Partial<Record<Phase, Verdict>>
12
13export type GateState = 'fresh' | 'stale' | 'fail' | 'none'
14
15export type ReduceOptions = {
16  maxLoops?: number
17  // Decision events failing this test are listed as unverified and not folded.
18  isTrusted?: (ev: TemperEvent) => boolean
19}
20
21export type OverrideRecord = { id: string; phase: Phase; reason: string; author?: string; ts: number }
22export type AcceptRecord = { id: string; findingId: string; reason: string; author?: string; ts: number }
23export type DriftRecord = { id: string; path: string; choice: DriftChoice; reason: string; author?: string; ts: number }
24export type HistoryRecord = {
25  ts: number
26  kind: 'start' | 'advance' | 'back' | 'override' | 'check'
27  from: Phase | null
28  to: Phase | 'done'
29}
30
31export type RunState = {
32  started: boolean
33  slug: string | null
34  title: string | null
35  phase: Phase | 'done' | null
36  paused: boolean
37  // Entry time of each phase, and the time a back step invalidated it.
38  since: Partial<Record<Phase, number>>
39  invalidated: Partial<Record<Phase, number>>
40  gate: Record<Phase, GateState>
41  stale: Phase[]
42  loops: number
43  maxLoops: number
44  loopLimitReached: boolean
45  checkOverridden: boolean
46  overrides: OverrideRecord[]
47  accepted: AcceptRecord[]
48  drift: DriftRecord[]
49  addedPaths: string[]
50  allowOnce: string[]
51  history: HistoryRecord[]
52  unverified: string[]
53}
54
55export const ONLY_USER = 'Only the user can approve this. Next: ask the user to press 1 or run /temper:temper approve.'
56
57const FLOW: readonly Phase[] = ['intent', 'plan', 'build', 'review', 'check']
58
59export const phaseLabel = (p: Phase | 'done'): string => (p === 'done' ? 'Done' : p.charAt(0).toUpperCase() + p.slice(1))
60
61// Position in the forward flow; Fix sits between Check and Done.
62function order(p: Phase): number {
63  return p === 'fix' ? FLOW.indexOf('check') + 0.5 : FLOW.indexOf(p)
64}
65
66function nextOf(p: Phase): Phase | 'done' {
67  if (p === 'fix') return 'done'
68  const i = FLOW.indexOf(p)
69  return FLOW[i + 1] ?? 'done'
70}
71
72function emptyGate(): Record<Phase, GateState> {
73  return { intent: 'none', plan: 'none', build: 'none', review: 'none', check: 'none', fix: 'none' }
74}
75
76export function initialState(maxLoops = 3): RunState {
77  return {
78    started: false,
79    slug: null,
80    title: null,
81    phase: null,
82    paused: false,
83    since: {},
84    invalidated: {},
85    gate: emptyGate(),
86    stale: [],
87    loops: 0,
88    maxLoops,
89    loopLimitReached: false,
90    checkOverridden: false,
91    overrides: [],
92    accepted: [],
93    drift: [],
94    addedPaths: [],
95    allowOnce: [],
96    history: [],
97    unverified: [],
98  }
99}
100
101export function reduce(events: readonly TemperEvent[], verdicts: Verdicts = {}, opts: ReduceOptions = {}): RunState {
102  const s = initialState(opts.maxLoops ?? 3)
103  const sorted = [...events].sort(compareEvents)
104
105  const enter = (to: Phase | 'done', ts: number) => {
106    s.phase = to
107    if (to !== 'done') s.since[to] = ts
108  }
109
110  for (const ev of sorted) {
111    // Every kind of event changes what is enforced (pause lifts the rules, checkResult
112    // ends the run, start resets it), so an event file the mod did not write counts for
113    // none of them. The adapter trusts only ids it recorded itself.
114    if (opts.isTrusted && !opts.isTrusted(ev)) {
115      s.unverified.push(ev.id)
116      continue
117    }
118    switch (ev.type) {
119      case 'start': {
120        const keepUnverified = s.unverified
121        Object.assign(s, initialState(s.maxLoops), { unverified: keepUnverified })
122        s.started = true
123        s.slug = ev.slug
124        s.title = ev.title
125        enter(ev.phase ?? 'intent', ev.ts)
126        s.history.push({ ts: ev.ts, kind: 'start', from: null, to: ev.phase ?? 'intent' })
127        break
128      }
129      case 'advance':
130        if (s.phase === ev.from) {
131          const failedAt = s.since.fix
132          enter(ev.to, ev.ts)
133          // Back from Fix to Check: the check that failed is the one this Check answers. A verdict written after
134          // that failure (the orchestrator ran the check again while the run was in Fix) counts for this Check, so
135          // the person is not sent through the same check run again.
136          if (ev.from === 'fix' && ev.to === 'check' && failedAt !== undefined) s.since.check = failedAt
137          s.history.push({ ts: ev.ts, kind: 'advance', from: ev.from, to: ev.to })
138        }
139        break
140      case 'back':
141        if (s.phase !== null && s.phase !== 'done') {
142          const from = s.phase
143          for (const p of FLOW) if (order(p) >= order(ev.to)) s.invalidated[p] = ev.ts
144          s.loops = 0
145          enter(ev.to, ev.ts)
146          s.history.push({ ts: ev.ts, kind: 'back', from, to: ev.to })
147        }
148        break
149      case 'override':
150        if (s.phase === ev.phase) {
151          s.overrides.push({ id: ev.id, phase: ev.phase, reason: ev.reason, author: ev.author, ts: ev.ts })
152          if (ev.phase === 'check' || ev.phase === 'fix') s.checkOverridden = true
153          const to = nextOf(ev.phase)
154          enter(to, ev.ts)
155          s.history.push({ ts: ev.ts, kind: 'override', from: ev.phase, to })
156        }
157        break
158      case 'accept':
159        s.accepted.push({ id: ev.id, findingId: ev.findingId, reason: ev.reason, author: ev.author, ts: ev.ts })
160        break
161      case 'drift':
162        s.drift.push({ id: ev.id, path: ev.path, choice: ev.choice, reason: ev.reason, author: ev.author, ts: ev.ts })
163        if (ev.choice === 'add') s.addedPaths.push(ev.path)
164        if (ev.choice === 'allow-once') s.allowOnce.push(ev.path)
165        break
166      case 'driftUsed': {
167        const i = s.allowOnce.indexOf(ev.path)
168        if (i >= 0) s.allowOnce.splice(i, 1)
169        break
170      }
171      case 'checkResult':
172        if (s.phase === 'check') {
173          if (ev.result === 'pass') {
174            enter('done', ev.ts)
175            s.history.push({ ts: ev.ts, kind: 'check', from: 'check', to: 'done' })
176          } else {
177            s.loops += 1
178            enter('fix', ev.ts)
179            s.history.push({ ts: ev.ts, kind: 'check', from: 'check', to: 'fix' })
180          }
181        }
182        break
183      case 'pause':
184        s.paused = true
185        break
186      case 'resume':
187        s.paused = false
188        break
189    }
190  }
191
192  s.loopLimitReached = s.phase === 'fix' && s.loops > s.maxLoops
193
194  for (const p of FLOW) {
195    const v = verdicts[p]
196    const boundary = Math.max(s.since[p] ?? -Infinity, s.invalidated[p] ?? -Infinity)
197    // gates.json times have whole second resolution and phase entry times are in
198    // milliseconds, so compare whole seconds and treat equal as fresh.
199    if (!v) s.gate[p] = 'none'
200    else if (Math.floor(v.ts / 1000) < Math.floor(boundary / 1000)) s.gate[p] = 'stale'
201    else s.gate[p] = v.verdict === 'PASS' ? 'fresh' : 'fail'
202  }
203  s.stale = FLOW.filter(p => s.invalidated[p] !== undefined && p !== s.phase && s.gate[p] !== 'fresh')
204  return s
205}
206
207// `anyOrigin`: the person turned enforcement off, so a decision is accepted from any origin (a `claude -p` run, the
208// Agent SDK). Its events keep the origin they came from.
209export type Command = { origin: Origin; author?: string; anyOrigin?: boolean } & (
210  | { type: 'start'; slug: string; title: string; phase?: Phase }
211  | { type: 'approve' }
212  | { type: 'advance' }
213  | { type: 'back'; to: Phase; reason: string }
214  | { type: 'override'; reason: string; phase?: Phase }
215  | { type: 'acceptFinding'; id: string; reason: string }
216  | { type: 'drift'; path: string; choice: DriftChoice; reason: string }
217  | { type: 'useDrift'; path: string }
218  | { type: 'checkResult'; result: 'pass' | 'fail' }
219  | { type: 'pause' }
220  | { type: 'resume' }
221)
222
223export type Decision = { events: Draft[] } | { error: string }
224
225const fail = (error: string): Decision => ({ error })
226
227function loopLimitMessage(s: RunState): string {
228  return (
229    `The fix limit is reached (${s.loops} failed check runs, limit ${s.maxLoops}). ` +
230    'Next: make a new plan (/temper:temper back plan <reason>), skip with a reason (/temper:temper override <reason>), ' +
231    'or take over (/temper:temper pause).'
232  )
233}
234
235export function decide(state: RunState, cmd: Command): Decision {
236  const who = { origin: cmd.origin, ...(cmd.author !== undefined ? { author: cmd.author } : {}) }
237  const phase = state.phase
238  const isPerson = cmd.origin === 'person' || cmd.anyOrigin === true
239
240  if (cmd.type === 'start') {
241    if (phase !== null && phase !== 'done') return fail('A Temper run is active. Finish it, or run /temper:temper pause.')
242    return { events: [{ type: 'start', slug: cmd.slug, title: cmd.title, ...(cmd.phase ? { phase: cmd.phase } : {}), ...who }] }
243  }
244  if (phase === null) return fail('No Temper run is active. Next: start one with /temper:temper <feature description>.')
245
246  if (cmd.type === 'resume') {
247    if (!isPerson) return fail(ONLY_USER)
248    return { events: [{ type: 'resume', ...who }] }
249  }
250  if (state.paused) return fail('The run is paused. Next: run /temper:temper resume.')
251  if (phase === 'done') return fail('The run is done. Next: commit, or start a new run with /temper:temper <feature description>.')
252
253  // At the fix loop limit only re-plan, override and pause remain.
254  if (state.loopLimitReached) {
255    const legal =
256      cmd.type === 'override' || cmd.type === 'pause' || (cmd.type === 'back' && cmd.to === 'plan')
257    if (!legal) return fail(loopLimitMessage(state))
258  }
259
260  switch (cmd.type) {
261    case 'approve':
262    case 'advance': {
263      const needsPerson = cmd.type === 'approve' || phase === 'intent' || phase === 'plan'
264      if (needsPerson && !isPerson) return fail(ONLY_USER)
265      if (phase !== 'fix') {
266        const g = state.gate[phase]
267        const name = phaseLabel(phase)
268        if (g === 'stale' && state.invalidated[phase] !== undefined) {
269          return fail(`${name} needs a new check. A step back made the old check invalid.`)
270        }
271        if (g === 'stale' || g === 'none') {
272          return fail(`${name} has not passed its check yet. Next: run the ${phase} check, the gate command of the Temper CLI.`)
273        }
274        if (g === 'fail') return fail(`${name} did not pass its check. Next: fix the problems. Then run the ${phase} check again.`)
275      }
276      return { events: [{ type: 'advance', from: phase, to: phase === 'check' ? 'done' : phase === 'fix' ? 'check' : nextOf(phase), ...who }] }
277    }
278    case 'back': {
279      if (!isPerson) return fail(ONLY_USER)
280      if (!cmd.reason.trim()) return fail('A back step needs a reason. Use /temper:temper back <phase> <reason>.')
281      if (cmd.to === 'fix' || !(order(cmd.to) < order(phase))) {
282        return fail(`Go back to a phase before ${phaseLabel(phase)}.`)
283      }
284      return { events: [{ type: 'back', to: cmd.to, reason: cmd.reason.trim(), ...who }] }
285    }
286    case 'override': {
287      if (!isPerson) return fail(ONLY_USER)
288      if (!cmd.reason.trim()) return fail('A skip needs a reason. Use /temper:temper override <reason>.')
289      if (cmd.phase !== undefined && cmd.phase !== phase) return fail(`You can skip only the current step, ${phaseLabel(phase)}.`)
290      return { events: [{ type: 'override', phase, reason: cmd.reason.trim(), ...who }] }
291    }
292    case 'acceptFinding': {
293      if (!isPerson) return fail(ONLY_USER)
294      if (phase !== 'review' && phase !== 'fix') return fail('You can accept findings in Review or Fix.')
295      if (!cmd.reason.trim()) return fail('Accept needs a reason. Use /temper:temper accept <id> <reason>.')
296      return { events: [{ type: 'accept', findingId: cmd.id, reason: cmd.reason.trim(), ...who }] }
297    }
298    case 'drift': {
299      if (!isPerson) return fail(ONLY_USER)
300      if (phase !== 'build' && phase !== 'fix') return fail('Scope drift decisions work in Build or Fix.')
301      if (cmd.choice === 'allow-once' && !cmd.reason.trim()) return fail('Allow once needs a reason.')
302      return { events: [{ type: 'drift', path: cmd.path, choice: cmd.choice, reason: cmd.reason.trim(), ...who }] }
303    }
304    case 'useDrift':
305      if (!state.allowOnce.includes(cmd.path)) return fail(`No allowance for ${cmd.path}.`)
306      return { events: [{ type: 'driftUsed', path: cmd.path, ...who }] }
307    case 'checkResult':
308      if (phase !== 'check') return fail('Check results count only in the Check phase.')
309      return { events: [{ type: 'checkResult', result: cmd.result, ...who }] }
310    case 'pause':
311      if (!isPerson) return fail(ONLY_USER)
312      return { events: [{ type: 'pause', ...who }] }
313  }
314}
315
hooks/temper-mod/core/paths.ts 53 lines
1// Path helpers shared by the rules and the plan file reader. Pure string work: the
2// module environment has no Node `path`.
3
4// Collapse `.`, `..`, doubled slashes and backslashes; strip `root` when the path is
5// absolute and inside it. The result never starts with `./` or `/` for in-root paths.
6export function normalizePath(input: string, root = ''): string {
7  let p = input.replace(/\\/g, '/')
8  const r = root.replace(/\\/g, '/').replace(/\/+$/, '')
9  // The engine reports the real folder (/private/tmp/x on macOS). A tool call may name it through
10  // the link (/tmp/x): that is the same folder, so both spellings are inside the root.
11  const alias = r.replace(/^\/private(?=\/)/, '')
12  if (r && (p === r || p.startsWith(r + '/'))) p = p.slice(r.length).replace(/^\/+/, '')
13  else if (alias !== r && (p === alias || p.startsWith(alias + '/'))) p = p.slice(alias.length).replace(/^\/+/, '')
14  const out: string[] = []
15  for (const part of p.split('/')) {
16    if (part === '' || part === '.') continue
17    if (part === '..') {
18      if (out.length > 0 && out[out.length - 1] !== '..') out.pop()
19      else out.push('..')
20    } else out.push(part)
21  }
22  return (p.startsWith('/') ? '/' : '') + out.join('/')
23}
24
25function globToRegExp(glob: string): RegExp {
26  let re = ''
27  for (let i = 0; i < glob.length; i++) {
28    const c = glob[i] ?? ''
29    if (c === '*') {
30      if (glob[i + 1] === '*') {
31        re += '.*'
32        i++
33        if (glob[i + 1] === '/') i++
34      } else re += '[^/]*'
35    } else if (c === '?') re += '[^/]'
36    else re += c.replace(/[.+^${}()|[\]\\]/g, '\\$&')
37  }
38  return new RegExp('^' + re + '$')
39}
40
41// True when `path` is named by one plan entry: an exact file, a `dir/` prefix, or a glob.
42export function matchesPlanEntry(path: string, entry: string): boolean {
43  const e = normalizePath(entry)
44  const p = normalizePath(path)
45  if (entry.endsWith('/')) return p === e || p.startsWith(e + '/')
46  if (e.includes('*') || e.includes('?')) return globToRegExp(e).test(p)
47  return p === e
48}
49
50export function matchesPlan(path: string, entries: readonly string[]): boolean {
51  return entries.some(e => matchesPlanEntry(path, e))
52}
53
hooks/temper-mod/core/rules.ts 513 lines
1// Deny rules: `evaluate(state, ctx, toolCall)` answers allow or deny for one tool call.
2// Pure. Every deny reason ends with what to do next (mods-plan 3.4).
3
4import { classifyBash, protectedKind } from './bash'
5import type { DecisionKind, ProtectedKind } from './bash'
6import type { Phase } from './events'
7import { ONLY_USER, phaseLabel } from './machine'
8import type { RunState } from './machine'
9import { CLI, guardedPhase } from './cli'
10import { matchesPlan, normalizePath } from './paths'
11
12// A decision the person made that no CLI call has matched yet: the event id, its kind, and
13// the phase (or finding id) it was made for.
14export type HumanDecision = { id: string; kind: DecisionKind; phase?: string; findingId?: string }
15
16export type RuleContext = {
17  // Absolute project root; absolute tool paths inside it are made relative.
18  root?: string
19  // The spec directory, relative to the root: `.temper/specs/{slug}`.
20  specDir: string
21  // File entries from plan.md and tasks.md (exact paths, `dir/` prefixes or globs).
22  planFiles: readonly string[]
23  // Human decision events not yet matched by a CLI call.
24  humanDecisions?: readonly HumanDecision[]
25  // The run's complexity (build-state.json): medium and complex runs have a design stage after plan.
26  complexity?: string | null
27  // `phases.design: true` is written in .claude/temper.config. When it is not, the orchestrator may go from a
28  // medium or complex plan straight to Build (the project never switched design on), so that step is accepted too.
29  designRequired?: boolean
30  // The CLI state looks reset (it is earlier than checks that passed): the phase rules do not block a
31  // write. Protected paths stay protected.
32  failOpenWrites?: boolean
33  // What the CLI commit gate (`temper gate commit`) would let through, read from the same facts. The mod's
34  // commit rule is never stricter than it.
35  commit?: CommitFacts
36  // `autonomy.enabled: true` in .claude/temper.config: the person has opted in to autonomous runs.
37  autonomyEnabled?: boolean
38  // Files a Fix finding action is currently active for (Review phase writes).
39  fixFiles?: readonly string[]
40  // The folder the shell is in (relative to the project root, or absolute), carried from the earlier Bash calls;
41  // null when it is not known. Staged paths and write targets are read against it.
42  cwd?: string | null
43  // Back decisions a `state loop` call has already used: the loop is spent once, the `state set next_stage` that follows
44  // spends the decision itself.
45  loopedDecisions?: readonly string[]
46  // Where the Temper script is: its full path in the plugin folder, or the plain `scripts/temper` when that is not
47  // known (see pluginCliFrom). Every deny text names the CLI by it.
48  cli?: string
49  // The CLI's files and the spec folder show a run that never left Intent, read fresh for this call (see nothingToLose
50  // in the adapter): no completed stage, no verdict, no skip, and none of intent.md, plan.md, tasks.md and design.md.
51  // With the mod's own record holding only the start, a `state clear` then loses nothing (the TRIVIAL exit of the
52  // orchestrator). Left out: not known, and the clear is refused.
53  nothingToLose?: boolean
54}
55
56// The carve-outs of the CLI commit gate (scripts/temper gate_commit, docs/decisions/0009):
57//  - artifact only: every staged file is under .temper/specs/ (the intent accept commit, the plan commit);
58//  - Build checkpoint: next_stage is build, command temper, the current branch is the run's branch, the
59//    plan (and intent, design) gates are satisfied, and the last build test row is green.
60export type CommitFacts = {
61  // Every file this commit stages is under .temper/specs/ (as far as the mod saw it staged).
62  stagedSpecsOnly: boolean
63  // The Build checkpoint carve-out holds.
64  checkpoint: boolean
65  // Why a checkpoint commit would be refused, in the CLI's words, for the deny text (a branch problem).
66  hint?: string
67}
68
69export type ToolCall = { tool: string; input: Record<string, unknown> }
70
71export type RuleResult =
72  | { allow: true; consume?: DecisionKind | 'drift'; driftPath?: string; eventId?: string; eventIds?: string[]; loopIds?: string[] }
73  | { deny: string; drift?: string }
74
75const ALLOW: RuleResult = { allow: true }
76const WRITE_TOOLS = new Set(['Write', 'Edit', 'MultiEdit', 'NotebookEdit'])
77
78const TEST_FILE = [/(^|\/)(?:tests?|__tests__|specs?)\//, /\.(?:test|spec)\.[^/]+$/, /(^|\/)test_[^/]*$/, /_test\.[^/]+$/]
79
80export const isTestFile = (path: string): boolean => TEST_FILE.some(re => re.test(path))
81
82function targetPath(call: ToolCall): string | null {
83  const raw = call.tool === 'NotebookEdit' ? call.input.notebook_path : call.input.file_path
84  return typeof raw === 'string' && raw.length > 0 ? raw : null
85}
86
87const inDir = (path: string, dir: string): boolean => path === dir || path.startsWith(dir + '/')
88
89function protectedDeny(kind: ProtectedKind, cli: string): RuleResult {
90  if (kind === 'events' || kind === 'overrides') return { deny: ONLY_USER }
91  if (kind === 'folder') {
92    return {
93      deny:
94        'Temper: the .temper folders hold the run, its verdicts and its decisions. Do not remove or replace them by hand. ' +
95        `Next: use ${cli} state archive after the run, or name one file.`,
96    }
97  }
98  if (kind === 'config') {
99    return {
100      deny:
101        'Temper: .claude/temper.config sets how this run is checked (autonomy, thresholds, what blocks a review). Only the user changes it while a run is active. ' +
102        'Next: ask the user to edit it themselves, or to end the run first.',
103    }
104  }
105  if (kind === 'evidence' || kind === 'loops') {
106    return {
107      deny:
108        'Temper: the evidence ledger and the loop counter belong to the temper CLI. Do not write them by hand. ' +
109        `Next: use ${cli} evidence add, run or resolve, and ${cli} state loop.`,
110    }
111  }
112  if (kind === 'hooks') {
113    return {
114      deny:
115        'Temper: the git hooks, the Temper commit hook and core.hooksPath are the native commit gate. Do not change them while a run is active. ' +
116        'Next: ask the user, or finish the run first.',
117    }
118  }
119  if (kind === 'state') {
120    return {
121      deny:
122        'Temper: use the temper CLI to change run state. Do not write it by hand. ' +
123        `Next: use ${cli} state set or ${cli} state advance.`,
124    }
125  }
126  return {
127    deny:
128      'Temper: the temper CLI makes the gate verdicts. Do not write them by hand. ' +
129      `Next: run ${cli} gate <stage> and read the verdict.`,
130  }
131}
132
133const COMMIT_NEXT: Record<Phase, string> = {
134  intent: 'finish the phases up to Check (key 1 or /temper:temper next)',
135  plan: 'finish the phases up to Check (key 1 or /temper:temper next)',
136  build: 'finish Build, then do Review and Check (key 1 or /temper:temper next)',
137  review: 'finish Review, then do Check (key 1 or /temper:temper next)',
138  check: 'run the checks (key 1 in Check or /temper:check)',
139  fix: 'fix the failed checks. Then run the checks again (key 1 in Fix)',
140}
141
142const isActive = (s: RunState): s is RunState & { phase: Phase } => s.phase !== null && s.phase !== 'done'
143
144function phaseWriteRule(s: RunState & { phase: Phase }, ctx: RuleContext, path: string): RuleResult {
145  const spec = normalizePath(ctx.specDir)
146  const inSpec = inDir(path, spec)
147  const specFile = (name: string) => path === `${spec}/${name}`
148  const label = phaseLabel(s.phase)
149
150  switch (s.phase) {
151    case 'intent':
152      if (specFile('intent.md') || specFile('intent-context.json')) return ALLOW
153      return {
154        deny:
155          `Temper: ${label} phase. Writing ${path} is not allowed until the user approves the intent. ` +
156          'Do not look for another way. Do not offer to turn Temper off. ' +
157          'Next: finish intent.md. Then ask the user to approve it (key 1 or /temper:temper approve).',
158      }
159    case 'plan': {
160      const ok =
161        ['intent.md', 'plan.md', 'tasks.md', 'design.md', 'config-suggestions.json'].some(specFile) ||
162        (inSpec && /-context\.json$/.test(path)) ||
163        (path.startsWith('docs/decisions/') && path.endsWith('.md'))
164      if (ok) return ALLOW
165      return {
166        deny:
167          `Temper: ${label} phase. Writing ${path} is not allowed until the user approves the plan. ` +
168          'Do not look for another way. Do not offer to turn Temper off. ' +
169          'Next: finish plan.md and tasks.md. Then ask the user to approve them (key 1 or /temper:temper approve).',
170      }
171    }
172    case 'review':
173      if (inSpec || (ctx.fixFiles ?? []).some(f => normalizePath(f, ctx.root) === path)) return ALLOW
174      return {
175        deny:
176          `Temper: ${label} phase. Writing ${path} is not allowed. Review changes the spec folder only. ` +
177          'Next: write the finding in the spec folder. Fix it when the user starts Fix (key 1 in Review, Fix all).',
178      }
179    case 'check':
180      if (inSpec) return ALLOW
181      return {
182        deny:
183          `Temper: ${label} phase. Writing ${path} is not allowed. Check only runs checks. ` +
184          'Next: run the checks (key 1 in Check or /temper:check).',
185      }
186    case 'build':
187    case 'fix': {
188      if (inSpec || isTestFile(path) || matchesPlan(path, ctx.planFiles) || s.addedPaths.some(p => normalizePath(p) === path)) {
189        return ALLOW
190      }
191      if (s.allowOnce.some(p => normalizePath(p) === path)) return { allow: true, consume: 'drift', driftPath: path }
192      return {
193        deny:
194          `Temper: scope drift. ${path} is not in the plan. ` +
195          'Next: ask the user to choose: add to plan, revert, or allow once with a reason ' +
196          '(/temper:temper drift add|revert|allow <reason>).',
197        drift: path,
198      }
199    }
200  }
201}
202
203// The CLI has no Fix stage: a decision made in Fix is a decision about the Check stage.
204const stageName = (p: string): string => (p === 'fix' ? 'check' : p)
205
206// The person skipped this stage with a reason (the override event) and has not stepped back since. A skip is the
207// person's own decision to go on, so the `state advance` that follows it needs no second approval, also for Intent
208// and Plan. A later step back ends it: a re-planned Plan must be approved again.
209function skippedStage(s: RunState, stage: string): boolean {
210  const lastBack = s.history.reduce((t, h) => (h.kind === 'back' ? Math.max(t, h.ts) : t), -Infinity)
211  return s.overrides.some(o => stageName(o.phase) === stage && o.ts >= lastBack)
212}
213
214// The Plan was approved by the person in this run: a trusted advance out of Plan, not undone by a
215// later back step to Intent or Plan. Untrusted event files never reach the state, so they never count.
216const planApproved = (s: RunState): boolean => {
217  let approved = false
218  for (const h of s.history) {
219    if (h.kind === 'advance' && h.from === 'plan') approved = true
220    if (h.kind === 'back' && (h.to === 'intent' || h.to === 'plan')) approved = false
221  }
222  return approved
223}
224
225const AUTONOMY_DENY =
226  'Temper: only the user can turn on autonomous mode, at the plan gate. ' +
227  'Next: ask the user to approve the plan (key 1 or /temper:temper approve) and to set autonomy.enabled: true in .claude/temper.config. Then try again.'
228
229const GUARD_KEYS = new Set(['stage', 'next_stage', 'branch', 'spec_path', 'run_mode', 'command'])
230
231const UNCHECKABLE = (cli: string): string =>
232  'Temper: this command writes to a path that Temper cannot check, and it names Temper state. ' +
233  'Next: write the exact file path with no variables, globs, braces or substitutions. Or use ' +
234  `${cli} gate <stage>, ${cli} evidence or ${cli} state.`
235
236const REPEATED_FLAG =
237  'Temper: this decision call repeats a flag (--id, --stage or --reason). ' +
238  'Temper cannot match it to the decision of the user. Next: run the call again. Give each flag one time.'
239
240const STATE_RESTART =
241  'Temper: state init and state loop restart or move the run. Only the user decides that. ' +
242  'Next: ask the user to use the Temper bar buttons (Loop back, Go back) or /temper:temper back <phase> <reason>.'
243
244// The stage that follows a stage in the CLI sequence (STAGE_SEQ_TEMPER in scripts/temper), with design
245// between plan and build for a medium or complex run. Null for a name that is not a stage.
246export function nextStage(stage: string, complexity: string | null | undefined): string | null {
247  switch (stage) {
248    case 'intent':
249      return 'plan'
250    case 'plan':
251      return complexity === 'medium' || complexity === 'complex' ? 'design' : 'build'
252    case 'design':
253      return 'build'
254    case 'build':
255      return 'review'
256    case 'review':
257      return 'check'
258    case 'check':
259      return 'commit'
260    default:
261      return null
262  }
263}
264
265// The mod's commit rule defers to the CLI commit gate: a commit passes when the gate would pass it. Every
266// stage that the gate checks (plan, build, review, check) passed or was overridden; or one of its two
267// carve-outs holds (an artifact only commit, or a Build checkpoint on the run's branch).
268function commitAllowed(s: RunState, facts: CommitFacts | undefined): boolean {
269  if (facts?.stagedSpecsOnly || facts?.checkpoint) return true
270  const passed = (p: Phase) => s.gate[p] === 'fresh' || s.overrides.some(o => o.phase === p)
271  return (['plan', 'build', 'review', 'check'] as const).every(passed)
272}
273
274// The check of a stage passed (fresh PASS) or the person overrode it, and the run is at that stage.
275// Design has no verdict of its own in the mod: it follows an approved plan whose check passed.
276function followsVerdict(s: RunState, stage: string, complexity?: string | null): boolean {
277  if (stage === 'design') return (complexity === 'medium' || complexity === 'complex') && (planApproved(s) || skippedStage(s, 'plan')) && (s.gate.plan === 'fresh' || s.overrides.some(o => o.phase === 'plan'))
278  if (!(stage in s.gate)) return false
279  const here = s.phase === stage || (s.phase === 'fix' && stage === 'check')
280  return here && (s.gate[stage as Phase] === 'fresh' || s.overrides.some(o => o.phase === stage))
281}
282
283const STATE_END = (cli: string): string =>
284  'Temper: do not clear or archive the run state during a run. Next: finish the run, ' +
285  `commit, then run ${cli} state archive.`
286
287const stateSetDeny = (key: string): string =>
288  `Temper: state set ${key} moves the run. Only the user can do this. ` +
289  'Next: ask the user to run /temper:temper back <phase> <reason>.'
290
291const OPAQUE_DENY =
292  'Temper: this command runs the Temper script in a way Temper cannot read, and it holds a decision word. ' +
293  'Only the user decides. Next: ask the user to use the buttons or the /temper:temper subcommands (approve, override, accept, back).'
294
295const DYNAMIC_DENY = (cli: string): string =>
296  'Temper: this command writes what a Temper call does (its subcommand) in a form Temper cannot read: a variable, a substitution or an escaped string. ' +
297  'It could be a decision, and only the user decides, with the buttons or the /temper:temper subcommands. ' +
298  `Next: run each ${cli} call with its words written out, one per Bash call.`
299
300const OPAQUE_HIDDEN_DENY = (cli: string): string =>
301  'Temper: this command may run the Temper script (it names the script, or starts a program that can run it) and hides part of what it runs behind a substitution, an expansion or a here-string, so Temper cannot read it. ' +
302  `Next: run ${cli} with its words written out, in a Bash call of its own.`
303
304const ALIAS_DENY = (cli: string): string =>
305  'Temper: do not link, copy or source the Temper script. A second name for it hides the decision calls. ' +
306  `Next: run ${cli} by its own path. The user decides with the buttons or /temper:temper.`
307
308const HIDDEN_DENY =
309  'Temper: a shell, or a builtin that runs text as commands (source and the like), is given a program that this command does not show (a pipe from another command, a file on stdin, a substitution, or a word split by quotes, backslashes, braces or globs). ' +
310  'While a run is active only a program that is written out plainly passes. Next: run each command in its own Bash call, with the words written out.'
311
312const GUARDED_USE_DENY = (word: string, cli: string): string =>
313  `Temper: this command names ${word}, a file of the run, and it is not a plain read. The CLI writes those files; nothing else does. ` +
314  `Next: read it with cat, grep, jq, head or git diff, or use ${cli} gate, evidence or state.`
315
316const ENV_DENY = (cli: string): string =>
317  'Temper: TEMPER_DIR and TEMPER_CONFIG point the CLI at other files than the run\'s, so its verdicts would be written for a run that is not this one. ' +
318  `Next: run ${cli} with no TEMPER_DIR or TEMPER_CONFIG.`
319
320const HOOKS_DENY =
321  'Temper: --no-verify, -n and core.hooksPath switch the native pre-commit hook off. The hook is the commit gate for every commit. ' +
322  'Next: commit without them. If the hook blocks the commit, finish the stages it names.'
323
324const BACK_DENY =
325  'Temper: state advance cannot move the run to an earlier stage. Only the user steps back. ' +
326  'Next: ask the user to use Go back or /temper:temper back <phase> <reason>.'
327
328const KEY_DENY = (key: string): string =>
329  `Temper: state set ${key} is allowed only in the form and the phase the orchestrator uses ` +
330  '(complexity while the plan is open, base_sha as the current commit in Plan or Build, command never). ' +
331  'Next: ask the user if the run needs another value.'
332
333// Where each stage sits in the CLI sequence; the mod's phase of the same name for a run that is at it.
334const STAGE_INDEX: Record<string, number> = { intent: 0, plan: 1, design: 2, build: 3, review: 4, check: 5, commit: 6 }
335const PHASE_INDEX: Record<string, number> = { intent: 0, plan: 1, build: 3, review: 4, check: 5, fix: 5, done: 6 }
336
337function evaluateBash(s: RunState, ctx: RuleContext, command: string): RuleResult {
338  const c = classifyBash(command, ctx.cwd === undefined ? '' : ctx.cwd)
339  const cli = ctx.cli ?? CLI
340
341  if (c.alias) return { deny: ALIAS_DENY(cli) }
342  if (c.opaque && isActive(s)) return { deny: c.opaqueWhy === 'dynamic' ? DYNAMIC_DENY(cli) : c.opaqueWhy === 'hidden' ? OPAQUE_HIDDEN_DENY(cli) : OPAQUE_DENY }
343  if (isActive(s)) {
344    if (c.hidden) return { deny: HIDDEN_DENY }
345    if (c.envTamper) return { deny: ENV_DENY(cli) }
346    if (c.hookTamper || c.noVerify) return { deny: HOOKS_DENY }
347  }
348
349  if (c.protectedWrites.length > 0) {
350    const known = c.protectedWrites.filter(p => !c.uncheckable.includes(p))
351    if (known.length === 0) return { deny: UNCHECKABLE(cli) }
352    // The config, the git hooks and the Temper commit hook are the run's only while a run is active (/temper:init writes
353    // the config and installs the hook before).
354    // A path of no known kind (a link made inside .temper) gets the text of the .temper folders: it is no decision.
355    const kinds = known.map(p => protectedKind(p) ?? 'folder').filter(k => isActive(s) || (k !== 'config' && k !== 'hooks'))
356    if (kinds.length > 0) {
357      const forged = kinds.find(k => k === 'events' || k === 'overrides')
358      return protectedDeny(forged ?? kinds[0] ?? 'folder', cli)
359    }
360    if (c.uncheckable.length > 0) return { deny: UNCHECKABLE(cli) }
361  }
362  if (isActive(s) && c.guardedUse.length > 0) return { deny: GUARDED_USE_DENY(c.guardedUse[0] ?? 'a guarded file', cli) }
363
364  // Removing or archiving the run's state is for after the run: while a run is active it would
365  // take the gate ledger and the overrides with it. One exception, the TRIVIAL exit of the orchestrator: a clear of a run
366  // that never left Intent loses nothing. The mod's own record holds only the start (no advance, no step back, no skip),
367  // and the CLI's files and the spec folder agree (ctx.nothingToLose).
368  const trivialExit = s.phase === 'intent' && s.history.every(h => h.kind === 'start') && ctx.nothingToLose === true
369  if (isActive(s) && c.stateOps.some(o => o.op === 'archive' || (o.op === 'clear' && !trivialExit))) return { deny: STATE_END(cli) }
370  if (isActive(s) && c.stateOps.some(o => o.op === 'init')) return { deny: STATE_RESTART }
371  // `state loop <from> <to>` keeps the loop budget and clears the evidence of the stages that are redone.
372  // While a run is active it is for the person's Loop back only: it passes when the person's back decision
373  // for that stage waits unspent. The decision is spent by the `state set next_stage` call that follows.
374  // The loop leaves the stage the run is at (a loop from another stage only burns the budget), and it uses the back
375  // decision once: a second loop call needs another decision. The step that follows still spends it.
376  const loopIds: string[] = []
377  if (isActive(s)) {
378    const here = stageName(s.phase)
379    for (const op of c.stateOps) {
380      if (op.op !== 'loop') continue
381      const to = op.to
382      const fromOk = op.from !== undefined && (op.from === here || (here === 'plan' && op.from === 'design'))
383      const used = ctx.loopedDecisions ?? []
384      const hit =
385        to === undefined || !fromOk
386          ? undefined
387          : (ctx.humanDecisions ?? []).find(h => h.kind === 'back' && h.phase !== undefined && stageName(h.phase) === stageName(to) && !used.includes(h.id) && !loopIds.includes(h.id))
388      if (hit === undefined) return { deny: STATE_RESTART }
389      loopIds.push(hit.id)
390    }
391  }
392
393  if (isActive(s)) {
394    if (c.calls.some(call => call.invalid)) return { deny: REPEATED_FLAG }
395    for (const op of c.stateOps) {
396      // Keys that move the run. next_stage is matched to the person's back decision below.
397      if (op.op === 'set' && op.key === 'run_mode' && op.value === 'autonomous') {
398        // Armed only after the person approved the Plan, and only where autonomy is switched on.
399        if (!(planApproved(s) && ctx.autonomyEnabled === true)) return { deny: AUTONOMY_DENY }
400        continue
401      }
402      if (op.op === 'set' && GUARD_KEYS.has(op.key) && op.key !== 'next_stage' && !(op.key === 'run_mode' && op.value === 'interactive')) {
403        return { deny: op.key === 'command' ? KEY_DENY(op.key) : stateSetDeny(op.key) }
404      }
405      // The keys the orchestrator sets itself, only in the form and the phase it sets them: the complexity while the
406      // plan is open (a later change would drop Design), base_sha as the current commit in Plan or Build (a later one
407      // would shrink what Review and Check look at).
408      if (op.op === 'set' && op.key === 'complexity') {
409        const open = (s.phase === 'intent' || s.phase === 'plan') && !planApproved(s) && !skippedStage(s, 'plan')
410        if (!(open && /^(?:trivial|simple|medium|complex)$/.test(op.value ?? ''))) return { deny: KEY_DENY(op.key) }
411      }
412      if (op.op === 'set' && op.key === 'base_sha') {
413        const v = op.value ?? ''
414        const form = /^[0-9a-f]{7,40}$/i.test(v) || v === '$(git rev-parse HEAD)' || v === '`git rev-parse HEAD`'
415        if (!((s.phase === 'plan' || s.phase === 'build') && form)) return { deny: KEY_DENY(op.key) }
416      }
417    }
418    // Each guarded CLI call needs its own unconsumed human decision made for that phase (and
419    // that finding, when the call names one).
420    const pool = [...(ctx.humanDecisions ?? [])]
421    let first: { kind: DecisionKind; eventId: string } | null = null
422    const matched: string[] = []
423    for (const call of c.calls) {
424      // An advance needs a person only when it leaves Intent or Plan (`intent_complete`,
425      // `plan_complete`); the phase it names is the phase the person approved, whatever phase the
426      // run is in by now. Later advances follow a verdict and are not guarded.
427      // Every advance is checked, not only the two approvals: the CLI stores any next stage it is given.
428      // A call passes with (a) a matching human decision, or (b) when it is the exact next stage of the
429      // run and the check of the stage it completes passed (the autonomous run, and a stage that follows
430      // its own verdict). Intent and Plan always need the person.
431      let approved: string | null = null
432      if (call.kind === 'advance') {
433        const stage = (call.stage ?? '').replace(/_complete$/, '')
434        const expected = nextStage(stage, ctx.complexity)
435        const skipsDesign = stage === 'plan' && expected === 'design' && ctx.designRequired !== true && call.next === 'build'
436        const wellFormed = /_complete$/.test(call.stage ?? '') && expected !== null && (call.next === expected || skipsDesign)
437        // Design belongs to Plan: the person's Continue at the design check spends an advance decision of Plan.
438        approved = wellFormed ? (stage === 'design' ? 'plan' : stage) : null
439        // No advance lowers the stage the run is at, whoever approved it: only a step back does (the user's).
440        const target = STAGE_INDEX[call.next ?? '']
441        if (target !== undefined && target < (PHASE_INDEX[s.phase] ?? 0)) return { deny: BACK_DENY }
442        const hasHuman = approved !== null && pool.some(h => h.kind === 'advance' && (h.phase === undefined || h.phase === approved))
443        if (!hasHuman) {
444          if (wellFormed && stage !== 'intent' && stage !== 'plan' && followsVerdict(s, stage, ctx.complexity)) continue
445          // A skip is for the stage the run is at: it is not an approval for a stage the run went past.
446          if (wellFormed && stage !== 'design' && skippedStage(s, stage) && stage === stageName(s.phase)) continue
447          // The person's decision for this stage waits, but the call names another next stage (found live: a
448          // medium run with design on, advanced straight to Build). Say which stage is next; do not send the
449          // model back to the person for a decision that was already made.
450          const waits = pool.some(h => h.kind === 'advance' && (h.phase === undefined || h.phase === stage))
451          if (!wellFormed && waits && expected !== null && /_complete$/.test(call.stage ?? '')) {
452            return {
453              deny:
454                `Temper: after ${stage} the next stage of this run is ${expected}, not ${call.next ?? 'none'}. ` +
455                `Next: run ${cli} state advance ${stage}_complete ${expected}. The user's approval is already recorded.`,
456            }
457          }
458          return { deny: ONLY_USER }
459        }
460      }
461      // A step to another stage by `state set next_stage`: the person's back decision, or the exact next
462      // stage after a check that passed.
463      if (call.kind === 'back' && !pool.some(h => h.kind === 'back' && (call.stage === undefined || h.phase === undefined || stageName(h.phase) === stageName(call.stage)))) {
464        const cur = s.phase === 'fix' ? 'check' : s.phase
465        if (call.stage !== undefined && call.stage === nextStage(cur, ctx.complexity) && followsVerdict(s, cur)) continue
466      }
467      const i = pool.findIndex(
468        h =>
469          h.kind === call.kind &&
470          // An accept is for one stage's ledger: --stage must be the stage the person accepted in.
471          (call.kind === 'accept' ? call.stage !== undefined && h.phase === call.stage : (call.kind === 'advance' ? h.phase === undefined || h.phase === approved : call.stage === undefined || h.phase === undefined || stageName(h.phase) === stageName(call.stage))) &&
472          (call.id === undefined || h.findingId === undefined || h.findingId === call.id),
473      )
474      const hit = i >= 0 ? pool[i] : undefined
475      if (!hit) return { deny: ONLY_USER }
476      pool.splice(i, 1)
477      matched.push(hit.id)
478      first ??= { kind: call.kind, eventId: hit.id }
479    }
480    if (first && !(c.commits && !s.paused)) return { allow: true, consume: first.kind, eventId: first.eventId, eventIds: matched, ...(loopIds.length > 0 ? { loopIds } : {}) }
481  }
482
483  // The artifact only carve-out is for a plain `git commit` of the index: a pathspec commit, a merge, a cherry-pick,
484  // an am, a pull or a revert brings in files that no `git add` named.
485  const facts = ctx.commit && c.unplainCommit ? { ...ctx.commit, stagedSpecsOnly: false } : ctx.commit
486  if (c.commits && isActive(s) && !s.paused && !commitAllowed(s, facts)) {
487    return {
488      deny: `Temper: commit blocked. Check has not passed. ${ctx.commit?.hint ? `${ctx.commit.hint} ` : ''}Next: ${COMMIT_NEXT[s.phase]}. The native pre-commit hook is the backstop.`,
489    }
490  }
491  return loopIds.length > 0 ? { allow: true, loopIds } : ALLOW
492}
493
494export function evaluate(state: RunState, ctx: RuleContext, call: ToolCall): RuleResult {
495  if (call.tool === 'Bash') {
496    const command = typeof call.input.command === 'string' ? call.input.command : ''
497    return evaluateBash(state, ctx, command)
498  }
499  if (!WRITE_TOOLS.has(call.tool)) return ALLOW
500
501  const raw = targetPath(call)
502  if (raw === null) return ALLOW
503  const path = normalizePath(raw, ctx.root)
504
505  const kind = protectedKind(path)
506  // The config, the git hooks and the Temper commit hook are guarded while a run is active (the person writes them with
507  // no run on).
508  if (kind !== null && (isActive(state) || (kind !== 'config' && kind !== 'hooks'))) return protectedDeny(kind, ctx.cli ?? CLI)
509
510  if (!isActive(state) || state.paused || ctx.failOpenWrites) return ALLOW
511  return phaseWriteRule(state, ctx, path)
512}
513
hooks/temper-mod/core/section.ts 57 lines
1// The text of the `temper:phase` system prompt section (mods-plan 3.5). Pure: the
2// adapter appends the returned text as a session scope section on every request.
3
4import type { Phase } from './events'
5import { nextStep, type ActionContext } from './actions'
6import { phaseLabel } from './machine'
7
8export const SECTION_ID = 'temper:phase'
9
10// What a message from the user at a gate means. It is the original "Other" choice of the orchestrator:
11// the Temper bar (key 4, Discuss) and the prompt box both send it. Also in commands/temper.md.
12export const GATE_MESSAGE = 'If the user writes a message at a gate, answer it. If it asks for a change, make the change, run the gate again, then wait for the user again.'
13
14export type SectionInput = {
15  enforcement: 'on' | 'off'
16  phase: Phase | 'done' | null
17  title: string | null
18  task?: { n: number; of: number } | null
19  progress: { passed: number; total: number; passedIds: readonly string[] } | null
20  paused?: boolean
21  loopLimitReached?: boolean
22  stale?: readonly Phase[]
23  actionContext?: ActionContext
24  // The line that says the bar and the CLI do not agree (Snapshot.sync.line).
25  sync?: string | null
26}
27
28export function sectionText(input: SectionInput): string {
29  const lines: string[] = [input.enforcement === 'on' ? 'Temper enforcement: active' : 'Temper enforcement: off (UI only)']
30
31  if (input.phase === null) {
32    lines.push('Phase: none. No Temper run is active.')
33    return lines.join('\n')
34  }
35
36  let phase = `Phase: ${phaseLabel(input.phase)}`
37  if (input.task) phase += ` (task ${input.task.n} of ${input.task.of})`
38  if (input.paused) phase += ' (paused)'
39  if (input.title) phase += ` · Intent: "${input.title}"`
40  lines.push(phase)
41
42  const p = input.progress
43  if (p && p.total > 0) {
44    const ids = p.passedIds.length > 0 ? ` (${p.passedIds.join(', ')})` : ''
45    lines.push(`Criteria: ${p.passed} of ${p.total} passed${ids}`)
46  }
47
48  if (input.stale && input.stale.length > 0) {
49    lines.push(`Stale: ${input.stale.map(phaseLabel).join(', ')}. A back step made them invalid. Each needs a new verdict.`)
50  }
51
52  lines.push(`Next: ${nextStep(input.phase, { ...input.actionContext, loopLimitReached: input.loopLimitReached })}`)
53  if (input.sync) lines.push(input.sync)
54  lines.push(GATE_MESSAGE)
55  return lines.join('\n')
56}
57