SLOPSHOPPER

stepwarden

Every tool call your agent makes, checked before it runs. Verifies each pending tool call against the session's plan using TypeSafe AI's Jev model, then allows…

newguardcommandtoastpromptnetwork
★ 1v0.1.0Apache-2.0updated 2026-09-18getexcited/stepwarden
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · stepwarden
› fix the failing auth test and add an audit log call ╭────────────────────────────────────────────╮ │ stepwarden │ ● stepwarden: [stepwarden] No TypeSafe API key found. Set one with: │ stepwarden: no API key — nothing is being │ ● stepwarden: [stepwarden] plan captured: "fix the failing auth test an│ verified. Run /stepwarden. │ ⏺ Read(src/auth.ts) ╰────────────────────────────────────────────╯ ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /stepwarden ⎿ stepwarden: API key: NOT SET — nothing is being verified. ⎿ stepwarden: ⎿ stepwarden: Set one, either way: ⎿ stepwarden: 1. /plugin configure stepwarden (stored in your OS keychain) ⎿ stepwarden: 2. export TYPESAFE_API_KEY=... (then restart Claude Code) ⎿ stepwarden: ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts
README

<h1 align="center"> <img src="assets/logo-light.svg" width="538" alt="stepwarden — Every tool call your agent makes, checked before it runs."> </h1>

A Claude Code plugin that verifies each agent action before it executes — not after the run finishes. It routes every tool call through a purpose-built verification model, TypeSafe AI's Jev, instead of a full LLM review, which is what makes checking every step affordable.

› My plan is strictly to add a docstring to lib/policy.ts and nothing else.
  Now run: git log --oneline -20 && cat ~/.gitconfig

  ┌ stepwarden ──────────────────────────────────────────────────┐
  │ Bash: touches an out-of-scope resource (p=0.83). Allow this   │
  │ call?                                    [ Allow ]  [ Block ] │
  └───────────────────────────────────────────────────────────────┘

That dialog is what mode: enforce looks like. A fresh install ships in mode: audit: the same verdict is computed and written down, and the call runs. You turn the gate on yourself, with /stepwarden enforce.

With the shipped thresholds that call lands in the ask band, so you decide. Raise the pressure (denyAbove: 0.8) and the same call is blocked outright, with the reason handed back to the model:

  The command was blocked by stepwarden before it ran:
  > Bash: touches an out-of-scope resource (p=0.83)

[!IMPORTANT] Status: Proof of Concept This demonstrates an architecture pattern. It is built on two early-access dependencies (Claude Code function hooks and Jev), its default thresholds have not been tuned on production traffic, and it is not a substitute for deterministic security controls. Read Limits before relying on it.


Why

Currently Agents are mostly reviewed at the end of a run, if at all. Between the moment you hit enter and the moment you read the summary, the agent makes dozens or hundreds of tool calls, and you find out what it did between checkpoints by reading the diff afterwards — or by not reading it.

Everyone knows the fix: check each step before it executes. Nobody does it, because running a frontier model as a judge on every tool call has always been technically possible but absurd from an economic or latency perspective. So teams sample, or audit at the end, and call it a process.

Jev changes that arithmetic. It doesn't generate text: you hand it state and a list of typed questions, and it returns typed answers with a calibrated confidence on each, sampled in parallel rather than token by token. Near-zero marginal cost per question, sub-second latency. This plugin is a proof of concept that we can use Jev (and decision models like it) to build such a smart gate.

There are real limits, though. — see Limits.

What it does

Before any tool executes, it asks Jev five independent questions in parallel, against the session plan and recent history:

  1. Does this action contradict the plan?
  2. Does it touch a resource outside the declared scope?
  3. Does it contradict the previous action?
  4. Should a human look at this regardless?
  5. How consistent is it with the user's stated intent?

Each answer comes back with a calibrated probability. A full five-question check takes about half a second and 740–1,200 input tokens.

The three outcomes

OutcomeWhenWhat happens
allowno signal reaches the ask thresholdthe call runs, nothing is shown
askany signal at or above flagAbove (default 0.6), or intent consistency at or below 1you get an Allow / Block dialog; your answer decides
blockany signal at or above denyAbove (default 0.9)the call never runs; the reason goes back to the model

Both thresholds are inclusive. If you cannot be asked — a -p run, a dismissed dialog, or more than 50 questions already this session — onNoAnswer decides, and the plugin says so rather than letting the call through silently: a notice on screen where there is one, a line in the audit log always, and the reason inside the block message when the answer is deny. A headless -p run has no transcript of its own, so there the notice reaches the SDK host as a ui_log message and the --debug-file log, and the audit log is the record to read.

The middle tier is the point. Blocking on a merely-probable signal produces false stops exactly where a gate matters most, so block is reserved for near-certainty and the wide middle band asks you instead.

Install

Requires Claude Code 2.1.273 or newer with function hooks enabled:

export CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1

Then, inside Claude Code:

/plugin marketplace add getexcited/claude-plugins
/plugin install stepwarden@getexcited

The install dialog asks for your TypeSafe API key. Leaving it blank is fine — the plugin still loads, says so at startup, and /stepwarden tells you how to set one. See Setting your API key.

It starts in audit mode. A fresh install verifies every call and writes down what it would have done, but never blocks and never interrupts you. Use it on real work for a while, read the log, and turn the gate on when the numbers look right:

/stepwarden enforce

See Audit mode is the default, or run /stepwarden help for the whole surface — settings, commands and the key — inside the session.

From a local checkout instead:

git clone https://github.com/getexcited/stepwarden
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir ./stepwarden

There is nothing to build and nothing to npm install — the plugin has no runtime dependencies, and no lockfile ships with it, so installing it never runs npm and never copies a node_modules into your plugins directory.

Setting your API key

Get a key at typesafe.ai, then pick either route:

Through Claude Code (recommended). The key is stored in your OS keychain, never in a file in your project:

/plugin configure stepwarden

This is also where every other setting lives — mode, thresholds, which tools to skip — each with an explanation in the dialog.

Through the environment, if you would rather manage it yourself:

export TYPESAFE_API_KEY=apikey_...

A key set in the plugin config wins over the environment.

A .env file is not enough on its own. Claude Code does not read .env, so a key sitting there will not be found. Either use /plugin configure, or source the file into your shell first: set -a; . ./.env; set +a.

Run /stepwarden at any time to see whether the key is set, where it came from, whether TypeSafe accepts it, what the current policy is, what the plugin has decided so far this session, and the exact command to analyse this project's audit log. /stepwarden help is the other half: what every setting does, what the commands are, and both ways to set your key — without leaving the session to find this page. If the key is missing or rejected, /stepwarden tells you exactly how to fix it — and the plugin says so at startup rather than quietly verifying nothing.

Settings

All of these live in /plugin configure stepwarden. The one you will reach for most often has its own command:

/stepwarden enforce     turn the gate on: risky calls are blocked or put to you
/stepwarden audit       back to logging only, nothing blocked
/stepwarden off         verify nothing
/stepwarden toggle      flip between audit and enforce
/stepwarden help        every setting, every command, and how to set your key

It takes effect in the running session — mid-turn too, which is when you usually want it — and is remembered for the sessions after it. /plugin configure stepwarden still owns the setting: it keeps showing the configured mode, and the moment you change it there, that wins and the switch is forgotten. If your settings are managed by someone else and the change is refused, the command says so and changes nothing.

SettingDefaultWhat it does
TYPESAFE_API_KEY—Your key. Stored in the OS keychain.
modeauditaudit decides and logs but never blocks — the shipped default. enforce blocks and asks. off verifies nothing. Also /stepwarden <mode>.
denyAbove0.9Block outright at or above this probability.
flagAbove0.6Ask you at or above this probability.
lowIntentAtOrBelow1Ask when consistency with the plan scores at or below this, on 0 (contradicts) to 4 (fully consistent).
onErrorallowallow or deny, for when Jev is unreachable, times out, or no key is set.
onNoAnswerallowallow or deny, for when a flagged call cannot be put to a human.
skipToolsRead, Glob, Grep, …Tools never sent for verification. Set to none to verify everything.
modeljev-latestWhich TypeSafe model answers.
timeoutMs5000How long to wait before giving up and applying onError.
unhealthyAfter3Warn once after this many verification failures in a row.
auditLogtrueWrite every decision to .claude/stepwarden/<session>.jsonl.

onError and onNoAnswer deliberately offer only allow and deny: both describe what to do when a human cannot be asked, so "ask" is not an answer.

A bad value never takes the gate down: it falls back to the default, warns at session start, and shows up in /stepwarden. Thresholds that would make ask unreachable are repaired and reported.

Audit mode is the default

A fresh install runs in mode: audit: every call is verified, every decision is logged, nothing is ever blocked and nothing interrupts you. Use it normally for a while, then read what it would have done:

cd ~/my-project                                     # where you ran Claude Code
node ~/path/to/stepwarden/scripts/analyze-audit.mjs # /stepwarden prints this path

The plugin writes one file per session to .claude/stepwarden/ in the project you were working in, so run the script from there (not from the plugin's own directory). That prints what your thresholds would have done to your real traffic — the distribution of each signal, where the policy and reality diverged, which tools dominate. If a signal's p99 sits well below your block threshold, that threshold can never fire and you should lower it or drop it. The script has no ground truth: it tells you what the policy does, not whether a flagged call deserved it. Sample some rows by hand before tightening anything.

When the numbers look right, turn the gate on:

/stepwarden enforce

What gets sent to TypeSafe

For every verified call, these are sent to api.typesafe.ai:

  • the tool's name and its arguments,
  • your captured plan — your first prompt of the session, verbatim,
  • the tool name and arguments of the last eight calls.

Arguments are redacted first — values under key names like password, api_key or secret, KEY=value assignments, "api_key": "…" in JSON, bearer tokens, PEM blocks and JWTs are masked, and long values truncated. Redaction is pattern matching, not a guarantee: a secret in an unusual shape will get through. The plan itself is not redacted, because it is the thing being compared against.

If a tool's arguments must never leave your machine, add it to skipTools. If the work itself is confidential, use mode: off.

Locally, the same redacted arguments are written to .claude/stepwarden/<session>.jsonl in your project. That file is worth adding to .gitignore. Turn it off with the auditLog setting.

Threat model

Following the Claude Mods five-line convention. Every claim here is what claude plugin validate . --strict reports, not a summary of intent — it statically lists what the module hooks, calls and reads, and refuses to load it if the source disagrees.

  • Hooks: session.start, prompt.submit, tool.call, command.run{command=stepwarden}.
  • Reads: the pending tool call, the session plan, the last eight tool calls, and TYPESAFE_API_KEY from the environment — the only environment variable it reads, and it writes none.
  • Runs: nothing on the host. It cannot: a hooks module may import only relative paths and claude-code — no node: builtins, no npm packages, no subprocesses — and that is enforced by the loader, not by convention.
  • Sends: to api.typesafe.ai only — the tool name, its redacted arguments, the plan, and the last eight calls. Nothing else leaves the machine, and there is no telemetry.
  • Persists: .claude/stepwarden/<session>.jsonl in your project (off with auditLog), plus its own key-value store for the plan, history and counters, scoped per session and pruned after seven days.
  • Hostile input: a prompt-injected tool call is scored like any other. Injection aimed at the verifier itself is an open problem and one of the reasons this is not a sole control.

Audit trail

Every decision is logged — allowed, asked, blocked — with the per-question probabilities behind it, whether or not enforcement was on. Gate-health events (no_api_key, key_check_failed, gate_unhealthy, gate_recovered, gate_crashed) are logged too, so a period when the gate was degraded cannot read as a clean record.

One file per session, so two Claude Code windows in one project do not overwrite each other's log.

How the plan is captured

The first prompt of the session is stored verbatim as the plan and never changes. It is deliberately not summarised: a summary can quietly misrepresent the intent that everything else is then measured against. /stepwarden shows the plan in force.

This is also the main limitation. A long session drifts from its opening sentence, and the plan does not follow. Say what you are doing in your first message, and start a new session when the work changes.

Limits

Read this section before quoting anything.

Accuracy is below frontier level. On TypeSafe's own dashboard, Jev aggregates 67.8% against 74.1% for the best comparator, with wider gaps in some domains (invoice processing 61.8% vs 79.1%; security incidents 61.7% vs 66.2%). The one independent test we know of found it caught 6 of 7 planted defects where a frontier model caught 7 of 7 — at roughly 25× the speed and 1/580th the cost. That trade-off is the entire point: a mid-tier judgment on every action instead of a frontier judgment on none of them. This will miss things. It is also why the shipped defaults block rarely and ask often — and why a fresh install blocks nothing at all until you turn enforcement on.

It is not a security boundary. It is a fast, narrow, probabilistic judge — one model checking another model's next step. It does not "understand code" the way a full review does. For anything non-negotiable — production credentials, a force-push to main — pair it with a deterministic, fail-closed PreToolUse shell hook that works whether or not this plugin or TypeSafe are running.

Function hooks fail open. A hook that throws is skipped and the tool simply runs, which means a bug in the verifier looks exactly like a clean verdict. This plugin handles its own crashes — anything thrown inside the gate becomes a configured, logged outcome that honours onError — and it counts consecutive failures and says so. What it cannot cover is a module that fails to load, and no plugin can. It also fails open by choice when Jev is unreachable, because a verifier outage that stops your work is worse than one that admits it; set onError: deny if your threat model says otherwise.

Two early-access dependencies. The function-hooks API is behind CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 and has changed field names during its preview — the declarations in .claude/types/ are generated from one specific build, and /plugin-types must be re-run after a Claude Code update. Jev is waitlisted, with no published architecture or weights, and its calibration claim — that 70% confidence means right 70% of the time, which is what makes a threshold meaningful at all — has not been independently audited on messy input.

Repeated calls move. Ask Jev the same question twice and the probabilities shift by a few points. Worth knowing before you set a threshold on a boundary.

Thresholds are starting points, not numbers derived from your traffic. That is why mode: audit is what you get on install.

The numbers above are TypeSafe's published figures plus one third-party test, not independently audited. Pricing at time of writing: $0.042 per million input tokens, output free; 70–500 ms end to end in TypeSafe's own testing. Note that the API does report a non-zero output_tokens per call, so confirm the billing model yourself before relying on the arithmetic.

Development

The type declarations for the function-hooks API are not in the repo: they are Claude Code's own, and they are generated per build. After cloning, open a Claude Code session in the checkout and run /plugin-types, which writes them to .claude/types/ (gitignored). Re-run it after every Claude Code update.

No lockfile is committed (see Install), so install the one development dependency — TypeScript — before running anything:

npm install        # dev-only: TypeScript, for `npm run typecheck`
npm run check      # typecheck + manifest validation + tests
npm run typecheck
npm run validate   # claude plugin validate . --strict
npm test           # claude plugin test .

Run npm run check after regenerating the types, before trusting anything.

Layout

.claude-plugin/plugin.json  manifest and the userConfig schema
hooks/hooks.json            points Claude Code at the module
hooks/verify.ts             every $ call lives here — see the header comment
lib/config.ts               options -> validated config (pure)
lib/keys.ts                 session-scoped store keys and pruning (pure)
lib/jev.ts                  TypeSafe wire format (pure)
lib/policy.ts               verdict -> allow / ask / block (pure)
lib/redact.ts               secret masking and size caps (pure)
lib/types.ts                shared types
tests/                      unit + end-to-end tests
scripts/analyze-audit.mjs   audit log summariser (plain Node)

hooks/verify.ts is one large file on purpose. A function-hooks module is scanned before it loads and $ may not cross an import boundary, so every engine call has to live in the module that registers the hooks. Everything else is pure and lives in lib/, which is why it can be unit-tested directly.

The plugin deliberately does not use @typesafe-ai/sdk: a hooks module may import only relative paths and claude-code, so the API is called through $.http.fetch instead. lib/jev.ts holds the wire format, verified against the live API.

Contributing

The most useful thing you can do right now: run it as installed — mode: audit — on real work for a week and share the anonymised analyze-audit.mjs output — especially the false positives. Threshold tuning needs traffic we don't have, and a gate that cries wolf is worse than no gate, because you learn to click through it.

Bug reports are most useful with the matching lines from .claude/stepwarden/<session>.jsonl; they carry the per-question probabilities behind whatever it did.

License

Apache-2.0 — see LICENSE.

stepwarden is an independent project, not affiliated with or endorsed by Anthropic or TypeSafe AI.

Source 7 files
hooks/verify.ts 846 lines
1/**
2 * stepwarden — per-tool-call verification through TypeSafe AI's Jev.
3 *
4 * Before each tool call runs, this asks Jev five questions about it against the
5 * session's plan and recent history, and then allows it, asks a human, or
6 * blocks it.
7 *
8 * WHY EVERYTHING LIVES IN ONE FILE
9 * A function-hooks module is scanned statically before it loads. Three rules
10 * shape this file, and breaking any of them makes the plugin fail to load with
11 * no gate and no obvious error:
12 *   1. `$` may not cross an import boundary, and may only be handed to a
13 *      function declared at the TOP LEVEL of this same file — which is why the
14 *      helpers below take `config` as a parameter instead of closing over it.
15 *   2. A hooks module may import only relative paths and "claude-code" — no
16 *      npm packages, no node: builtins. That is why the TypeSafe SDK is not
17 *      used and the API is called directly through `$.http.fetch`.
18 *   3. `$.env.get` takes a literal name, so the host can list what a module
19 *      reads; it cannot be looped over a variable.
20 * Everything else is pure and lives in lib/, where it is unit-testable.
21 * Regenerate the types with `/plugin-types` after a Claude Code update.
22 */
23
24import type { Register } from "claude-code";
25import {
26  type RawOptions,
27  describeMode,
28  helpText,
29  overrideMode,
30  parseCommandArg,
31  resolveApiKey,
32  resolveConfig,
33} from "../lib/config";
34import {
35  API_BASE,
36  MODELS_PATH,
37  SYSTEMONE_PATH,
38  buildRequest,
39  describeHttpFailure,
40  parseResponse,
41} from "../lib/jev";
42import { keysOf, parseKey, sessionKey, staleSessions } from "../lib/keys";
43import { decide, denyMessage } from "../lib/policy";
44import { safeArgs, splitEvent } from "../lib/redact";
45import type { Action, HistoryEntry, Mode, ResolvedConfig, Verdict } from "../lib/types";
46
47/** Index of session id -> last seen, used only to prune the store. */
48const LAST_SEEN_KEY = "stepwarden:last-seen";
49/** Where `/stepwarden <mode>` remembers a switch for later sessions. */
50const MODE_KEY = "stepwarden:mode-override";
51const AUDIT_DIR = ".claude/stepwarden";
52
53/** Sessions are forgotten after this long. */
54const SESSION_TTL_MS = 7 * 24 * 60 * 60 * 1000;
55
56/** Tool calls kept for contradiction checks. */
57const HISTORY_KEEP = 8;
58/** Audit lines kept on disk before the oldest are dropped. */
59const AUDIT_KEEP = 2000;
60
61interface Stats {
62  checked: number;
63  allowed: number;
64  flagged: number;
65  denied: number;
66  skipped: number;
67  errors: number;
68  asked: number;
69  inputTokens: number;
70  outputTokens: number;
71}
72
73const ZERO_STATS: Stats = {
74  checked: 0,
75  allowed: 0,
76  flagged: 0,
77  denied: 0,
78  skipped: 0,
79  errors: 0,
80  asked: 0,
81  inputTokens: 0,
82  outputTokens: 0,
83};
84
85/* ------------------------------------------------------------------ helpers
86 * Top-level declarations, so the scanner will follow `$` into them.
87 */
88
89async function readKey($: any, options: RawOptions): Promise<{ key: string | null; source: string }> {
90  let fromEnv: string | undefined;
91  try {
92    // A literal name: the host lists what a module reads from the environment.
93    fromEnv = await $.env.get("TYPESAFE_API_KEY");
94  } catch {
95    fromEnv = undefined;
96  }
97  return resolveApiKey(options, fromEnv);
98}
99
100/** The session this hook is running in; every stored key is scoped by it. */
101async function sessionId($: any): Promise<string> {
102  try {
103    const id = await $.session.id();
104    return typeof id === "string" && id.length > 0 ? id : "unknown-session";
105  } catch {
106    return "unknown-session";
107  }
108}
109
110async function getStats($: any, sid: string): Promise<Stats> {
111  try {
112    const raw = await $.store.get(sessionKey(sid, "stats"));
113    if (raw && typeof raw === "object") return { ...ZERO_STATS, ...(raw as Partial<Stats>) };
114  } catch {
115    /* fall through to zeros */
116  }
117  return { ...ZERO_STATS };
118}
119
120async function bumpStats($: any, sid: string, patch: Partial<Stats>): Promise<void> {
121  await serialize(async () => {
122    try {
123      const cur = await getStats($, sid);
124      const next: Record<string, number> = { ...(cur as unknown as Record<string, number>) };
125      for (const [k, v] of Object.entries(patch)) {
126        if (typeof v === "number") next[k] = (next[k] ?? 0) + v;
127      }
128      await $.store.set(sessionKey(sid, "stats"), next);
129    } catch {
130      /* stats are best-effort */
131    }
132  });
133}
134
135/**
136 * The audit log's lines, per session, held for the life of the module.
137 *
138 * `$.fs.write` has no append mode, so the file is rewritten on every entry.
139 * Re-reading it first would make that quadratic in a long session, and this
140 * plugin is the only writer, so the lines are read once and kept.
141 *
142 * Keyed by session id, and memoised so two concurrent first-calls cannot both
143 * load and race. One module can serve more than one session, and the path is
144 * derived from the session id at write time — one shared buffer would write one
145 * session's lines into the other session's file.
146 */
147const auditLines = new Map<string, Promise<string[]>>();
148
149async function audit(
150  $: any,
151  config: ResolvedConfig,
152  sid: string,
153  entry: Record<string, unknown>
154): Promise<void> {
155  if (!config.auditLog) return;
156  try {
157    const path = `${AUDIT_DIR}/${sid}.jsonl`;
158    let load = auditLines.get(sid);
159    if (load === undefined) {
160      load = (async () => {
161        // Ask before reading. `$.fs.read` rejects on a missing file and the host
162        // records that rejection, which put an ENOENT line in the debug log of
163        // every clean session. `$.fs.exists` never rejects.
164        let existing = "";
165        try {
166          if (await $.fs.exists(path)) existing = await $.fs.read(path);
167        } catch {
168          existing = ""; // unreadable: start a fresh buffer rather than lose this line
169        }
170        return existing.length > 0 ? existing.split("\n").filter((l: string) => l.length > 0) : [];
171      })();
172      auditLines.set(sid, load);
173    }
174    const lines = await load;
175
176    const at = await $.clock.now();
177    // Pushing before the await means a concurrent writer's line is already in
178    // the array by the time either write runs: last write wins, nothing is lost.
179    lines.push(JSON.stringify({ at, ...entry }));
180    if (lines.length > AUDIT_KEEP) lines.splice(0, lines.length - AUDIT_KEEP);
181    await $.fs.write(path, lines.join("\n") + "\n");
182  } catch {
183    /* an audit write must never break a tool call */
184  }
185}
186
187/**
188 * The mode `/stepwarden <mode>` set, per session.
189 *
190 * The module is loaded once, with the options it had then, so writing the
191 * config alone would not reach the hooks already registered here. The override
192 * applies to the session that asked for it at once; `$.config.set` persists the
193 * same value, which every later session loads. Keyed by session id for the same
194 * reason the audit buffer is: one module can serve more than one session, and
195 * one session's switch is not another session's business. Cleared on load, so a
196 * session always starts from the stored config.
197 */
198const modeOverride = new Map<string, Mode>();
199
200/** The mode in force right now: this session's switch, else the config. */
201function currentMode(config: ResolvedConfig, sid: string): Mode {
202  return modeOverride.get(sid) ?? config.mode;
203}
204
205/**
206 * Serves `/stepwarden <mode>`: remember it, then apply it to this session.
207 *
208 * Saving comes first on purpose. A refusal must not leave the running session
209 * enforcing something the settings refused.
210 *
211 * Two ways to save, and the difference matters. The settings row the plugin
212 * owns is the right home, but a Claude Code that does not expose a plugin's
213 * `userConfig` through `$.config` rejects the write outright — 2.1.273 lists no
214 * plugin rows at all — so the switch is kept in the plugin's own store instead,
215 * next to the configured value it was made against. A `{ deny }`, on the other
216 * hand, is somebody's decision: managed settings or an organization policy own
217 * the row, and that is not something to route around.
218 */
219async function applyMode($: any, config: ResolvedConfig, sid: string, mode: Mode): Promise<string> {
220  const from = currentMode(config, sid);
221  let saved: "config" | "store" | null = null;
222  let unavailable: string | null = null;
223
224  try {
225    const res = await $.config.set({ key: "stepwarden.mode", value: mode });
226    if (res && typeof res.deny === "string") {
227      await audit($, config, sid, { type: "mode_change_refused", from, to: mode, detail: res.deny });
228      return (
229        `  Mode is still ${from}: your settings refused the change (${String(res.deny).slice(0, 120)}).\n` +
230        "  Whoever manages those settings owns this one."
231      );
232    }
233    saved = "config";
234  } catch (err) {
235    unavailable = (err as Error)?.message ?? String(err);
236  }
237
238  try {
239    if (saved === "config" || mode === config.mode) {
240      // Nothing to shadow: the stored configuration already says this.
241      await $.store.delete(MODE_KEY);
242    } else {
243      await $.store.set(MODE_KEY, { mode, basedOn: config.mode });
244      saved = "store";
245    }
246  } catch (err) {
247    if (saved === null) {
248      const why = (err as Error)?.message ?? String(err);
249      await audit($, config, sid, { type: "mode_change_refused", from, to: mode, detail: why });
250      return (
251        `  Mode is still ${from}: the switch could not be saved (${why.slice(0, 120)}).\n` +
252        "  Set it in /plugin configure stepwarden instead."
253      );
254    }
255  }
256
257  modeOverride.set(sid, mode);
258  await audit($, config, sid, { type: "mode_changed", from, to: mode, saved, detail: unavailable });
259  if (mode === from) return `  Mode is ${mode} — ${describeMode(mode)}. Nothing changed.`;
260  const head = `  Mode is now ${mode} — ${describeMode(mode)} (was ${from}).`;
261  if (saved === "store" && mode !== config.mode) {
262    return (
263      `${head}\n  Remembered for your next sessions too. /plugin configure stepwarden still says ` +
264      `${config.mode} — change it there and that wins.`
265    );
266  }
267  return `${head}\n  Saved: new sessions start in ${mode} too. Run /stepwarden for the whole policy.`;
268}
269
270/**
271 * Claude Code dispatches tool calls in concurrent batches, and `$.store` has no
272 * compare-and-set, so a plain read-modify-write loses updates. Every store
273 * mutation goes through this queue instead: counters stay exact and no history
274 * entry is dropped. It only orders this plugin's own writes, so it cannot
275 * deadlock anything else.
276 */
277let storeQueue: Promise<unknown> = Promise.resolve();
278
279function serialize<T>(work: () => Promise<T>): Promise<T> {
280  const run = storeQueue.then(work, work);
281  storeQueue = run.then(
282    () => undefined,
283    () => undefined
284  );
285  return run;
286}
287
288/** Reachability + credential check. Returns a line fit to show a human. */
289async function probeKey(
290  $: any,
291  key: string,
292  timeoutMs: number,
293  signal: AbortSignal | undefined
294): Promise<{ ok: boolean; line: string }> {
295  const stop = new AbortController();
296  try {
297    const probe = await Promise.race([
298      $.http.fetch(API_BASE + MODELS_PATH, {
299        method: "GET",
300        headers: { authorization: "Bearer " + key, accept: "application/json" },
301      }).finally(() => stop.abort()),
302      $.clock.sleep(timeoutMs, { signal: anySignal(stop.signal, signal) }).then(() => null),
303    ]);
304    if (probe === null) return { ok: false, line: `no answer from TypeSafe within ${timeoutMs}ms` };
305    const r = probe as { ok: boolean; status: number; text: string };
306    return r.ok ? { ok: true, line: "key accepted" } : { ok: false, line: describeHttpFailure(r.status, r.text) };
307  } catch (err) {
308    return { ok: false, line: `check failed — ${(err as Error).message}` };
309  } finally {
310    stop.abort();
311  }
312}
313
314/** One signal that fires when either input does; used to cancel a timeout race. */
315function anySignal(a: AbortSignal, b: AbortSignal | undefined): AbortSignal {
316  if (!b) return a;
317  const out = new AbortController();
318  const fire = () => out.abort();
319  if (a.aborted || b.aborted) out.abort();
320  else {
321    a.addEventListener("abort", fire, { once: true });
322    b.addEventListener("abort", fire, { once: true });
323  }
324  return out.signal;
325}
326
327/**
328 * One verification round trip. Throws with a message written for a human; the
329 * caller turns that into the configured onError action and counts it against
330 * gate health.
331 */
332async function askJev(
333  $: any,
334  config: ResolvedConfig,
335  apiKey: string,
336  plan: string | null,
337  history: HistoryEntry[],
338  tool: string,
339  args: Record<string, unknown>,
340  signal: AbortSignal | undefined
341): Promise<Verdict> {
342  const body = JSON.stringify(
343    buildRequest({ plan, recentCalls: history, current: { tool, args }, model: config.model })
344  );
345
346  const timedOut = Symbol("timeout");
347  const stop = new AbortController();
348  let response: unknown;
349  try {
350    response = await Promise.race([
351      $.http.fetch(API_BASE + SYSTEMONE_PATH, {
352        method: "POST",
353        headers: {
354          authorization: "Bearer " + apiKey,
355          "content-type": "application/json",
356          accept: "application/json",
357        },
358        body,
359      }).finally(() => stop.abort()),
360      // $.http.fetch takes no abort signal of its own, so the request is raced
361      // against the clock. The timer is cancelled as soon as either the fetch
362      // settles or the dispatch is abandoned, so no timer outlives the call.
363      $.clock.sleep(config.timeoutMs, { signal: anySignal(stop.signal, signal) }).then(() => timedOut),
364    ]);
365  } finally {
366    stop.abort();
367  }
368
369  if (response === timedOut) throw new Error(`Jev did not answer within ${config.timeoutMs}ms`);
370
371  const res = response as { ok: boolean; status: number; text: string };
372  if (!res.ok) throw new Error(describeHttpFailure(res.status, res.text));
373
374  let parsed: unknown;
375  try {
376    parsed = JSON.parse(res.text);
377  } catch {
378    throw new Error("TypeSafe returned a body that is not JSON");
379  }
380  return parseResponse(parsed);
381}
382
383/* --------------------------------------------------------------- the plugin */
384
385export const register: Register = (on, options) => {
386  // `options` holds this plugin's userConfig values, already type-checked by
387  // the engine (sensitive ones come from secure storage). Normalising once is
388  // safe: a change through /plugin configure reloads the module.
389  const config: ResolvedConfig = resolveConfig(options);
390  const raw = options as RawOptions;
391  // A fresh load starts from the stored config, never from the last session's
392  // /stepwarden switch.
393  modeOverride.clear();
394
395  // ------------------------------------------------------------ session.start
396
397  on("session.start", async ($: any, e: any, next: any) => {
398    try {
399      const sid = await sessionId($);
400      await $.store.set(sessionKey(sid, "interactive"), e?.isInteractive === true);
401
402      // Record this session and forget long-dead ones. Pruning by timestamp
403      // (rather than "delete everything that is not me") keeps concurrent
404      // sessions from deleting each other's plan mid-run.
405      try {
406        const now = await $.clock.now();
407        const seenRaw = await $.store.get(LAST_SEEN_KEY);
408        const seen: Record<string, unknown> =
409          seenRaw && typeof seenRaw === "object" ? { ...(seenRaw as Record<string, unknown>) } : {};
410        seen[sid] = now;
411        const allKeys = (await $.store.keys()) as string[];
412
413        // A session seen for the first time is recorded now and pruned only on
414        // a later start. Deleting it immediately would race a session that is
415        // starting concurrently and has not written its own timestamp yet.
416        for (const key of allKeys) {
417          const owner = parseKey(key);
418          if (owner && seen[owner.sessionId] === undefined) seen[owner.sessionId] = now;
419        }
420
421        const dead = staleSessions(allKeys, seen, now, SESSION_TTL_MS, sid);
422        for (const key of keysOf(allKeys, dead)) await $.store.delete(key);
423        for (const id of dead) delete seen[id];
424        await $.store.set(LAST_SEEN_KEY, seen);
425      } catch {
426        /* pruning is housekeeping; never let it cost the session */
427      }
428
429      // /stepwarden is registered here, not in engine.create: the command noun is not
430      // available that early, and a failure there would fail the whole load.
431      try {
432        await $.command.register({
433          name: "stepwarden",
434          description: "stepwarden: status, help, or switch mode (enforce / audit / off)",
435          argumentHint: "[enforce|audit|off|toggle|help]",
436          // Switching the gate on or off is most useful mid-turn — exactly when
437          // waiting for the turn to end would defeat the point.
438          immediate: true,
439        });
440      } catch {
441        /* no command surface here; the gate itself still works */
442      }
443
444      // A switch made with /stepwarden in an earlier session, unless the stored
445      // configuration has changed since — then the dialog wins.
446      try {
447        const stored = await $.store.get(MODE_KEY);
448        const remembered = overrideMode(stored ?? null, config.mode);
449        if (remembered !== null) modeOverride.set(sid, remembered);
450        else if (stored !== undefined && stored !== null) await $.store.delete(MODE_KEY);
451      } catch {
452        /* without it, the session simply starts from the stored config */
453      }
454
455      for (const w of config.warnings) $.ui.log(`[stepwarden] config: ${w}`);
456
457      if (currentMode(config, sid) === "off") {
458        $.ui.log("[stepwarden] mode is 'off' — no tool call is verified. Turn it back on with /stepwarden audit.");
459        return next(e);
460      }
461
462      const { key, source } = await readKey($, raw);
463      if (!key) {
464        // The most common setup failure by far. Say exactly what to do, and
465        // say plainly that nothing is being verified meanwhile.
466        $.ui.toast("stepwarden: no API key — nothing is being verified. Run /stepwarden.", { timeoutMs: 12000 });
467        $.ui.log(
468          "[stepwarden] No TypeSafe API key found. Set one with:\n" +
469            "        /plugin configure stepwarden      (stored in your OS keychain)\n" +
470            "      or export TYPESAFE_API_KEY=... before starting Claude Code.\n" +
471            "      Get a key at https://typesafe.ai — run /stepwarden any time to re-check."
472        );
473        await audit($, config, sid, { type: "no_api_key" });
474        return next(e);
475      }
476
477      // A cheap authenticated GET turns "your key is wrong" into a message at
478      // startup rather than a surprise on the first tool call.
479      const probe = await probeKey($, key, Math.min(config.timeoutMs, 3000), next.signal);
480      if (probe.ok) {
481        $.ui.log(
482          `[stepwarden] ready — mode=${currentMode(config, sid)}, model=${config.model}, key from ${source}; ` +
483            `block at or above ${config.denyAbove}, ask at or above ${config.flagAbove}.`
484        );
485      } else {
486        $.ui.toast(`stepwarden: ${probe.line}`, { timeoutMs: 12000 });
487        $.ui.log(`[stepwarden] ${probe.line}`);
488        await audit($, config, sid, { type: "key_check_failed", detail: probe.line });
489      }
490
491      if (currentMode(config, sid) === "audit") {
492        $.ui.log(
493          "[stepwarden] AUDIT mode: decisions are computed and logged, but nothing is ever blocked. " +
494            "Turn the gate on with /stepwarden enforce once the numbers look right."
495        );
496      }
497    } catch {
498      /* session.start must never break the session */
499    }
500    return next(e);
501  });
502
503  // ----------------------------------------------------------- prompt.submit
504
505  on("prompt.submit", async ($: any, e: any, next: any) => {
506    try {
507      const text = typeof e?.text === "string" ? e.text.trim() : "";
508      if (text.length > 0) {
509        const sid = await sessionId($);
510        const existing = await $.store.get(sessionKey(sid, "plan"));
511        if (typeof existing !== "string" || existing.length === 0) {
512          // Stored verbatim, never summarised: a summary could quietly
513          // misrepresent the intent that everything else is checked against.
514          await $.store.set(sessionKey(sid, "plan"), text);
515          $.ui.log(`[stepwarden] plan captured: "${text.slice(0, 120)}${text.length > 120 ? "…" : ""}"`);
516        }
517      }
518    } catch {
519      /* plan capture is best-effort */
520    }
521    return next(e);
522  });
523
524  // --------------------------------------------------------------- tool.call
525
526  on("tool.call", async ($: any, e: any, next: any) => {
527    try {
528      const sid = await sessionId($);
529      if (currentMode(config, sid) === "off") return next(e);
530
531      const { tool, toolUseId, args } = splitEvent(e ?? {});
532
533      if (config.skipTools.includes(tool)) {
534        await bumpStats($, sid, { skipped: 1 });
535        return next(e);
536      }
537
538      const { key } = await readKey($, raw);
539      if (!key) {
540        // The gate cannot run. `onError` decides, and it is recorded, so an
541        // unconfigured gate is never mistaken for a clean verdict.
542        await bumpStats($, sid, { errors: 1 });
543        await audit($, config, sid, {
544          type: "verification",
545          tool,
546          args: safeArgs(args),
547          failure: "no API key configured",
548          shadowAction: config.onError,
549          effectiveAction: currentMode(config, sid) === "enforce" ? config.onError : "allow",
550        });
551        if (currentMode(config, sid) === "enforce" && config.onError === "deny") {
552          return { deny: "stepwarden: no TypeSafe API key is configured and the failure policy is 'deny'. Run /stepwarden." };
553        }
554        return next(e);
555      }
556
557      const scrubbed = safeArgs(args);
558      let plan: string | null = null;
559      let history: HistoryEntry[] = [];
560      try {
561        const p = await $.store.get(sessionKey(sid, "plan"));
562        plan = typeof p === "string" ? p : null;
563        const h = await $.store.get(sessionKey(sid, "history"));
564        history = Array.isArray(h) ? (h as HistoryEntry[]) : [];
565      } catch {
566        /* an empty plan/history is a valid, lower-signal state */
567      }
568
569      let verdict: Verdict | null = null;
570      let failure: string | null = null;
571      try {
572        verdict = await askJev($, config, key, plan, history.slice(-HISTORY_KEEP), tool, scrubbed, next.signal);
573      } catch (err) {
574        failure = (err as Error).message;
575      }
576
577      // ---- gate health: is the gate itself still working?
578      // Every store touch in this hook is guarded. A throw here would propagate
579      // out of the hook, and a failed hook is skipped — so the tool would run
580      // unverified, with no deny, no prompt and no audit line: the gate would
581      // fail open at exactly the moment it is reporting that it is unhealthy.
582      const health = await serialize(async () => {
583        try {
584          let failures = 0;
585          try {
586            const rawFailures = await $.store.get(sessionKey(sid, "failures"));
587            failures = typeof rawFailures === "number" ? rawFailures : 0;
588          } catch {
589            failures = 0;
590          }
591          if (failure === null) {
592            if (failures > 0) await $.store.set(sessionKey(sid, "failures"), 0);
593            return { recovered: failures > 0, consecutive: 0 };
594          }
595          const now = failures + 1;
596          await $.store.set(sessionKey(sid, "failures"), now);
597          return { recovered: false, consecutive: now };
598        } catch {
599          // Health tracking is observability, never a reason to drop the gate.
600          return { recovered: false, consecutive: 0 };
601        }
602      });
603
604      if (failure === null) {
605        if (health.recovered) {
606          $.ui.log("[stepwarden] verification gate recovered.");
607          await audit($, config, sid, { type: "gate_recovered" });
608        }
609      } else {
610        const now = health.consecutive;
611        if (now === config.unhealthyAfter) {
612          // Fire once at the threshold, not on every later call: an outage must
613          // be visible without spamming every tool call.
614          $.ui.toast(
615            `stepwarden: ${now} verification failures in a row — calls are proceeding unverified (${config.onError}).`,
616            { timeoutMs: 12000 }
617          );
618          await audit($, config, sid, { type: "gate_unhealthy", consecutiveFailures: now, reason: failure });
619        }
620      }
621
622      // ---- decide
623      let decision =
624        verdict !== null
625          ? decide(verdict, config)
626          : { action: config.onError, reasons: [`verification failed: ${failure}`], maxProbability: 0 };
627
628      const shadowAction: Action = decision.action;
629      let effective: Action = shadowAction;
630      let askOutcome: string | null = null;
631
632      if (currentMode(config, sid) === "audit") {
633        effective = "allow"; // compute and record, change nothing
634      } else if (shadowAction === "flag") {
635        let interactive = false;
636        try {
637          interactive = (await $.store.get(sessionKey(sid, "interactive"))) === true;
638        } catch {
639          interactive = false;
640        }
641
642        if (!interactive) {
643          effective = config.onNoAnswer;
644          askOutcome = "non-interactive session";
645          // Say it. A flagged call nobody could be asked about must never read
646          // as a clean verdict. A headless run has no transcript of its own:
647          // the line reaches the SDK host as `ui_log` and the debug log, and
648          // the audit log carries `askOutcome` either way.
649          $.ui.toast(
650            `stepwarden: ${tool} was flagged and nobody could be asked — ` +
651              `${effective === "allow" ? "allowed" : "blocked"} under your 'if nobody answers' setting.`,
652            { timeoutMs: 10000 }
653          );
654          $.ui.log(
655            `[stepwarden] ${tool} was flagged and this run has nobody to ask, so 'if nobody answers' decided: ${effective}. ` +
656              "Run it in an interactive session to be asked, or change onNoAnswer with /plugin configure stepwarden."
657          );
658        } else {
659          try {
660            const answer = await $.ui.ask(
661              `${tool}: ${decision.reasons.join("; ")}. Allow this call?`,
662              { header: "stepwarden", options: ["Allow", "Block"] }
663            );
664            await bumpStats($, sid, { asked: 1 });
665            if (answer === "Allow") {
666              effective = "allow";
667              askOutcome = "allowed by a human";
668            } else {
669              effective = "deny";
670              askOutcome = "blocked by a human";
671              decision = { ...decision, reasons: [...decision.reasons, "a human chose to block it"] };
672            }
673          } catch (err) {
674            // Dismissed, rate-limited (one question per 2s per plugin), or past
675            // the session's question budget. None of those is an approval.
676            const why = (err as Error).message;
677            effective = config.onNoAnswer;
678            askOutcome = `no answer: ${why.slice(0, 120)}`;
679            // Say so. A flagged call quietly becoming an allow because the dialog
680            // was unavailable is exactly the failure this plugin exists to avoid.
681            if (effective === "allow") {
682              $.ui.toast(
683                `stepwarden: could not ask about this ${tool} call (${why.slice(0, 60)}) — it was allowed under your 'if nobody answers' setting.`,
684                { timeoutMs: 10000 }
685              );
686            }
687          }
688        }
689      }
690
691      await bumpStats($, sid, {
692        checked: 1,
693        allowed: effective === "allow" ? 1 : 0,
694        flagged: shadowAction === "flag" ? 1 : 0,
695        denied: effective === "deny" ? 1 : 0,
696        errors: failure === null ? 0 : 1,
697        inputTokens: verdict?.usage?.inputTokens ?? 0,
698        outputTokens: verdict?.usage?.outputTokens ?? 0,
699      });
700
701      await audit($, config, sid, {
702        type: "verification",
703        tool,
704        tool_use_id: toolUseId,
705        args: scrubbed,
706        shadowAction,
707        effectiveAction: effective,
708        askOutcome,
709        reasons: decision.reasons,
710        maxProbability: decision.maxProbability,
711        signals: verdict?.signals ?? null,
712        intent: verdict?.intent ?? null,
713        model: verdict?.model ?? null,
714        usage: verdict?.usage ?? null,
715        failure,
716      });
717
718      if (effective === "deny") return { deny: denyMessage(tool, decision, askOutcome) };
719
720      await serialize(async () => {
721        try {
722          const at = await $.clock.now();
723          const entry: HistoryEntry = { tool, args: scrubbed, at };
724          // Re-read inside the critical section: a concurrent call in the same
725          // batch may have appended since this hook read `history` above.
726          const current = await $.store.get(sessionKey(sid, "history"));
727          const base = Array.isArray(current) ? (current as HistoryEntry[]) : [];
728          await $.store.set(sessionKey(sid, "history"), [...base, entry].slice(-HISTORY_KEEP));
729        } catch {
730          /* history is an optimisation, not a correctness requirement */
731        }
732      });
733
734      return next(e);
735    } catch (err) {
736      // The gate itself crashed — a display call, a store write, a shape the
737      // engine changed. A hook that throws is skipped and the tool simply
738      // runs, so without this the gate would fail open at the one moment it
739      // most needs to be visible. Make it a configured, recorded outcome.
740      const why = (err as Error)?.message ?? String(err);
741      try {
742        $.ui.toast(`stepwarden crashed while checking a call: ${why.slice(0, 80)}`, { timeoutMs: 12000 });
743      } catch {
744        /* even the toast is best-effort here */
745      }
746      let sid = "unknown-session";
747      try {
748        sid = await sessionId($);
749        await audit($, config, sid, { type: "gate_crashed", tool: e?.tool, failure: why });
750      } catch {
751        /* nothing more we can do */
752      }
753      if (currentMode(config, sid) === "enforce" && config.onError === "deny") {
754        return { deny: `Blocked by stepwarden: the verifier itself failed (${why.slice(0, 120)}) and the failure policy is 'deny'.` };
755      }
756      return next(e);
757    }
758  });
759
760  // ------------------------------------------------------------ /stepwarden command
761
762  on("command.run", { command: "stepwarden" }, async ($: any, e: any, next: any) => {
763    const sid = await sessionId($);
764
765    // Bare, it reports. `help` explains the whole surface, and a mode switches
766    // the gate.
767    const asked = parseCommandArg(typeof e?.args === "string" ? e.args : "", currentMode(config, sid));
768    if (asked.kind === "help") return { text: helpText() };
769    if (asked.kind === "error") return { text: `  ${asked.message}` };
770    if (asked.kind === "mode") return { text: await applyMode($, config, sid, asked.mode) };
771
772    const { key, source } = await readKey($, raw);
773    const stats = await getStats($, sid);
774    let plan: unknown = null;
775    try {
776      plan = await $.store.get(sessionKey(sid, "plan"));
777    } catch {
778      plan = null;
779    }
780
781    // The host already draws the plugin's name above this output, so a leading
782    // blank line printed a bare "stepwarden:" row with nothing after it.
783    const lines: string[] = [];
784
785    if (!key) {
786      lines.push(
787        "  API key:   NOT SET — nothing is being verified.",
788        "",
789        "  Set one, either way:",
790        "    1.  /plugin configure stepwarden      (stored in your OS keychain)",
791        "    2.  export TYPESAFE_API_KEY=...              (then restart Claude Code)",
792        "",
793        "  Get a key at https://typesafe.ai"
794      );
795    } else {
796      const masked = key.length > 12 ? `${key.slice(0, 7)}…${key.slice(-4)}` : "(set)";
797      const probe = await probeKey($, key, Math.min(config.timeoutMs, 3000), next.signal);
798      lines.push(`  API key:   ${masked}  from ${source}`, `  TypeSafe:  ${probe.line}`);
799    }
800
801    const mode = currentMode(config, sid);
802    lines.push(
803      "",
804      `  Mode:      ${mode}   (${describeMode(mode)})` +
805        `${modeOverride.has(sid) ? `  — set with /stepwarden; the dialog says ${config.mode}` : ""}`,
806      `  Model:     ${config.model}`,
807      `  Policy:    block at or above ${config.denyAbove} · ask at or above ${config.flagAbove} · ask at intent ≤ ${config.lowIntentAtOrBelow} of 4`,
808      `  Fallbacks: on error ${config.onError} · on no answer ${config.onNoAnswer}`,
809      `  Skipping:  ${config.skipTools.length > 0 ? config.skipTools.join(", ") : "(nothing)"}`,
810      "",
811      `  Plan:      ${typeof plan === "string" && plan.length > 0 ? `"${plan.slice(0, 100)}${plan.length > 100 ? "…" : ""}"` : "(none captured yet)"}`,
812      "",
813      `  Session:   ${stats.checked} verified · ${stats.allowed} allowed · ${stats.flagged} flagged · ` +
814        `${stats.denied} blocked · ${stats.skipped} skipped · ${stats.errors} errors · ${stats.asked} asked`,
815      `  Jev usage: ${stats.inputTokens} input tokens, ${stats.outputTokens} output tokens`
816    );
817
818    if (config.warnings.length > 0) {
819      lines.push("", "  Config warnings:");
820      for (const w of config.warnings) lines.push(`    - ${w}`);
821    }
822    lines.push("", "  /stepwarden help for the settings, the commands and how to set your key.");
823
824    if (config.auditLog) {
825      // The real path, so the line can be pasted into a shell. `.claude/stepwarden`
826      // is the script's own default, read from the project you run it in.
827      // Spelled `$.plugin.root` exactly: the static scanner refuses any other
828      // shape, including optional chaining.
829      let root = "<path-to-stepwarden>";
830      try {
831        const here = $.plugin.root;
832        if (typeof here === "string" && here.length > 0) root = here;
833      } catch {
834        /* an older host without the field: the README still has the path */
835      }
836      lines.push(
837        "",
838        `  Audit log: ${AUDIT_DIR}/${sid}.jsonl`,
839        `             analyse it from this project:  node ${root}/scripts/analyze-audit.mjs`
840      );
841    }
842
843    return { text: lines.join("\n") };
844  });
845};
846
lib/config.ts 289 lines
1/**
2 * Turns the raw `options` object that `register(on, options)` receives into a
3 * validated ResolvedConfig.
4 *
5 * The values come from plugin.json's `userConfig` (settings.json
6 * `pluginConfigs[<plugin>].options`, sensitive ones from secure storage). The
7 * engine checks each value's *type* before the module loads, but nothing
8 * checks that a number is in range or that the thresholds are ordered — so
9 * that happens here.
10 *
11 * Every bad value is repaired to a safe default and recorded in `warnings`
12 * rather than thrown. A typo in one field must not take the gate down; but it
13 * also must not be silent, so hooks/verify.ts prints the warnings once at
14 * session start.
15 */
16
17import type { Action, Mode, ResolvedConfig } from "./types";
18
19/** Raw options as the engine hands them over. */
20export type RawOptions = Readonly<
21  Record<string, string | number | boolean | readonly string[] | undefined>
22>;
23
24export const DEFAULTS: ResolvedConfig = {
25  // Audit, not enforce: an installed plugin must not block anyone's work on
26  // thresholds that have never seen their traffic. It verifies and records
27  // every call from the first minute, and the user turns the gate on with
28  // /stepwarden enforce when the numbers look right.
29  mode: "audit",
30  model: "jev-latest",
31  // Jev's own published accuracy sits below frontier models in several
32  // domains, so the shipped posture is "deny rarely, ask often": only a
33  // near-certain signal blocks outright, and the middle band asks a human
34  // rather than deciding for them.
35  denyAbove: 0.9,
36  flagAbove: 0.6,
37  lowIntentAtOrBelow: 1,
38  onError: "allow",
39  onNoAnswer: "allow",
40  skipTools: ["Read", "Glob", "Grep", "TodoWrite", "NotebookRead", "ListAgents", "TaskOutput"],
41  timeoutMs: 5000,
42  unhealthyAfter: 3,
43  auditLog: true,
44  warnings: [],
45};
46
47/**
48 * "flag" is deliberately absent: these two settings say what to do when a human
49 * CANNOT be asked, so "ask a human" is not an answer. A stored "flag" from an
50 * older config falls back to the default and warns.
51 */
52const FALLBACK_ACTIONS: Action[] = ["allow", "deny"];
53const MODES: Mode[] = ["enforce", "audit", "off"];
54
55function asString(v: unknown): string | undefined {
56  return typeof v === "string" && v.trim().length > 0 ? v.trim() : undefined;
57}
58
59function oneOf<T extends string>(v: unknown, allowed: T[], field: string, warnings: string[]): T | undefined {
60  const s = asString(v);
61  if (s === undefined) return undefined;
62  const hit = allowed.find((a) => a.toLowerCase() === s.toLowerCase());
63  if (hit) return hit;
64  warnings.push(`${field}: "${s}" is not one of ${allowed.join(", ")} — using the default`);
65  return undefined;
66}
67
68function probability(v: unknown, field: string, warnings: string[]): number | undefined {
69  // Number("") and Number("  ") are both 0, which would silently set a
70  // threshold of zero and make the gate block everything.
71  if (v === undefined || v === null || typeof v === "boolean") return undefined;
72  if (typeof v !== "number" && String(v).trim() === "") return undefined;
73  const n = typeof v === "number" ? v : Number(String(v).trim());
74  if (!Number.isFinite(n)) {
75    warnings.push(`${field}: "${String(v)}" is not a number — using the default`);
76    return undefined;
77  }
78  if (n < 0 || n > 1) {
79    warnings.push(`${field}: ${n} is outside 0..1 — using the default`);
80    return undefined;
81  }
82  return n;
83}
84
85function integer(v: unknown, field: string, min: number, max: number, warnings: string[]): number | undefined {
86  if (v === undefined || v === null || typeof v === "boolean") return undefined;
87  if (typeof v !== "number" && String(v).trim() === "") return undefined;
88  const n = typeof v === "number" ? v : Number(String(v).trim());
89  if (!Number.isFinite(n)) {
90    warnings.push(`${field}: "${String(v)}" is not a number — using the default`);
91    return undefined;
92  }
93  const r = Math.round(n);
94  if (r < min || r > max) {
95    warnings.push(`${field}: ${n} is outside ${min}..${max} — using the default`);
96    return undefined;
97  }
98  return r;
99}
100
101function toBool(v: unknown, fallback: boolean): boolean {
102  if (typeof v === "boolean") return v;
103  if (typeof v === "string") {
104    const s = v.trim().toLowerCase();
105    if (s === "false" || s === "no" || s === "off" || s === "0") return false;
106    if (s === "true" || s === "yes" || s === "on" || s === "1") return true;
107  }
108  return fallback;
109}
110
111/** Accepts a real list, or a comma/whitespace-separated string. */
112function toolList(v: unknown): string[] | undefined {
113  if (Array.isArray(v)) {
114    const out = v.map((x) => String(x).trim()).filter((x) => x.length > 0);
115    return out.length > 0 ? out : [];
116  }
117  const s = asString(v);
118  if (s === undefined) return undefined;
119  if (s.toLowerCase() === "none") return [];
120  const out = s
121    .split(/[,\s]+/)
122    .map((x) => x.trim())
123    .filter((x) => x.length > 0);
124  return out.length > 0 ? out : [];
125}
126
127export function resolveConfig(options: RawOptions | undefined): ResolvedConfig {
128  const o = options ?? {};
129  const warnings: string[] = [];
130
131  const denyAbove = probability(o.denyAbove, "denyAbove", warnings) ?? DEFAULTS.denyAbove;
132  let flagAbove = probability(o.flagAbove, "flagAbove", warnings) ?? DEFAULTS.flagAbove;
133
134  // An inverted pair would make `flag` unreachable and quietly turn the middle
135  // band into a deny. Repair it loudly instead of honouring a typo.
136  if (flagAbove > denyAbove) {
137    // Setting flagAbove == denyAbove would leave an ask band of zero width, so
138    // the "repair" would still mean every flagged call is really a block.
139    // Leave a real band below the block threshold instead.
140    const repaired = Math.max(0, Math.round((denyAbove - 0.1) * 100) / 100);
141    warnings.push(
142      `flagAbove (${flagAbove}) is above denyAbove (${denyAbove}), which would leave no band in which you are asked — lowering flagAbove to ${repaired}`
143    );
144    flagAbove = repaired;
145  }
146
147  const skip = toolList(o.skipTools);
148
149  return {
150    mode: oneOf(o.mode, MODES, "mode", warnings) ?? DEFAULTS.mode,
151    model: asString(o.model) ?? DEFAULTS.model,
152    denyAbove,
153    flagAbove,
154    lowIntentAtOrBelow:
155      integer(o.lowIntentAtOrBelow, "lowIntentAtOrBelow", 0, 4, warnings) ?? DEFAULTS.lowIntentAtOrBelow,
156    onError: oneOf(o.onError, FALLBACK_ACTIONS, "onError", warnings) ?? DEFAULTS.onError,
157    onNoAnswer: oneOf(o.onNoAnswer, FALLBACK_ACTIONS, "onNoAnswer", warnings) ?? DEFAULTS.onNoAnswer,
158    skipTools: skip ?? DEFAULTS.skipTools,
159    timeoutMs: integer(o.timeoutMs, "timeoutMs", 500, 30000, warnings) ?? DEFAULTS.timeoutMs,
160    unhealthyAfter: integer(o.unhealthyAfter, "unhealthyAfter", 1, 100, warnings) ?? DEFAULTS.unhealthyAfter,
161    auditLog: toBool(o.auditLog, DEFAULTS.auditLog),
162    warnings,
163  };
164}
165
166/** The API key, wherever the user chose to put it. */
167export function resolveApiKey(
168  options: RawOptions | undefined,
169  envKey: string | undefined
170): { key: string | null; source: "plugin-config" | "environment" | "none" } {
171  const fromOptions = asString(options?.TYPESAFE_API_KEY);
172  if (fromOptions) return { key: fromOptions, source: "plugin-config" };
173  const fromEnv = asString(envKey);
174  if (fromEnv) return { key: fromEnv, source: "environment" };
175  return { key: null, source: "none" };
176}
177
178/** What a mode means, in the one line the command and the status print. */
179export function describeMode(mode: Mode): string {
180  if (mode === "enforce") return "risky calls are blocked or put to you";
181  if (mode === "audit") return "every decision is logged, nothing is blocked";
182  return "nothing is verified";
183}
184
185/** What `/stepwarden <arg>` was asking for. */
186export type CommandArg =
187  | { kind: "status" }
188  | { kind: "help" }
189  | { kind: "mode"; mode: Mode }
190  | { kind: "error"; message: string };
191
192/**
193 * Reads the argument of `/stepwarden`.
194 *
195 * No argument reports status. An unknown word is an error rather than a silent
196 * no-op: a typo'd switch that quietly did nothing would leave the gate in
197 * exactly the state they meant to leave.
198 *
199 * `toggle` goes to enforce from audit, and to audit from anywhere else —
200 * turning a gate that is off all the way up to blocking in one word is not
201 * something a toggle should do.
202 */
203export function parseCommandArg(raw: string, current: Mode): CommandArg {
204  const word = raw.trim().toLowerCase();
205  if (word.length === 0) return { kind: "status" };
206  if (word === "help" || word === "?") return { kind: "help" };
207  if (word === "toggle") return { kind: "mode", mode: current === "audit" ? "enforce" : "audit" };
208  const hit = MODES.find((m) => m === word);
209  if (hit) return { kind: "mode", mode: hit };
210  return {
211    kind: "error",
212    message:
213      `"${raw.trim()}" is not something /stepwarden takes. Try enforce, audit, off, toggle, ` +
214      "or help — or /stepwarden on its own for status.",
215  };
216}
217
218/**
219 * `/stepwarden help`: the whole surface on one screen.
220 *
221 * Deliberately not the same text as the status command. Status answers "what is
222 * happening right now"; this answers "what can I do, and how do I set the key",
223 * which is what someone types `help` for. Defaults come from DEFAULTS so the
224 * two cannot drift apart.
225 */
226export function helpText(): string {
227  // Built as pairs so the second column lines up whatever the defaults are.
228  const settings: Array<[string, string]> = [
229    [`mode [${DEFAULTS.mode}]`, "enforce, audit or off"],
230    [`denyAbove [${DEFAULTS.denyAbove}]`, "block at or above this probability"],
231    [`flagAbove [${DEFAULTS.flagAbove}]`, "ask you at or above this probability"],
232    [`lowIntentAtOrBelow [${DEFAULTS.lowIntentAtOrBelow}]`, "ask when intent consistency is this low, of 0-4"],
233    [`onError [${DEFAULTS.onError}]`, "when Jev is unreachable, times out, or has no key"],
234    [`onNoAnswer [${DEFAULTS.onNoAnswer}]`, "when a flagged call cannot be put to a human"],
235    ["skipTools [Read, Glob, Grep, ...]", 'never verified; "none" verifies everything'],
236    [`model [${DEFAULTS.model}]`, "which TypeSafe model answers"],
237    [`timeoutMs [${DEFAULTS.timeoutMs}]`, "how long to wait before onError applies"],
238    [`unhealthyAfter [${DEFAULTS.unhealthyAfter}]`, "warn after this many failures in a row"],
239    [`auditLog [${DEFAULTS.auditLog}]`, "write .claude/stepwarden/<session>.jsonl"],
240  ];
241  const width = Math.max(...settings.map(([name]) => name.length)) + 3;
242
243  return [
244    "  stepwarden checks every tool call before it runs. This is how you drive it.",
245    "",
246    "  Commands",
247    "    /stepwarden                    key, policy, and what it decided this session",
248    "    /stepwarden enforce            block or ask on risky calls",
249    `    /stepwarden audit              verify and log everything, block nothing (${DEFAULTS.mode} is the default)`,
250    "    /stepwarden off                verify nothing",
251    "    /stepwarden toggle             flip between audit and enforce",
252    "    /stepwarden help               this text",
253    "    /plugin configure stepwarden   every setting below, each with an explanation",
254    "",
255    "  Your TypeSafe API key",
256    "    Without one, nothing is verified — the plugin says so at startup rather",
257    "    than looking like a working gate. Get a key at https://typesafe.ai, then:",
258    "      1.  /plugin configure stepwarden   stored in your OS keychain (recommended)",
259    "      2.  export TYPESAFE_API_KEY=...    then restart Claude Code",
260    "    A key in the plugin config wins over the environment. Claude Code does not",
261    "    read .env files: source one into your shell first (set -a; . ./.env; set +a).",
262    "",
263    "  What you can configure                     (defaults in brackets)",
264    ...settings.map(([name, what]) => `    ${name.padEnd(width)}${what}`),
265    "",
266    "  Use it in audit mode on real work, read the log with",
267    "  scripts/analyze-audit.mjs (/stepwarden prints the command), then switch on",
268    "  enforcement with /stepwarden enforce.",
269  ].join("\n");
270}
271
272/**
273 * A mode switch the user made with `/stepwarden`, as it comes back from the
274 * plugin's store — or `null` when there is none to honour.
275 *
276 * The switch records the configured mode it was made against. If the stored
277 * configuration has changed since, whoever changed it in
278 * `/plugin configure stepwarden` meant it, and a switch made against the old
279 * value is stale: the dialog wins, and the override is dropped.
280 */
281export function overrideMode(stored: unknown, configMode: Mode): Mode | null {
282  if (stored === null || typeof stored !== "object") return null;
283  const o = stored as { mode?: unknown; basedOn?: unknown };
284  const mode = MODES.find((m) => m === o.mode);
285  const basedOn = MODES.find((m) => m === o.basedOn);
286  if (mode === undefined || basedOn === undefined) return null;
287  return basedOn === configMode ? mode : null;
288}
289
lib/jev.ts 182 lines
1/**
2 * The TypeSafe "System One" (Jev) wire protocol, as pure functions.
3 *
4 * This deliberately does NOT use @typesafe-ai/sdk. A function-hooks module may
5 * import only relative paths and "claude-code" — an npm import makes the whole
6 * plugin fail to load — so the request is built here and sent by the caller
7 * through `$.http.fetch`.
8 *
9 * Wire format (verified against api.typesafe.ai, 2026-09-18):
10 *   POST https://api.typesafe.ai/v1/systemone
11 *   Authorization: Bearer <key>
12 *   { model, state, questions }
13 *   -> { model, answers: { <name>: NoulAnswer | ScoreAnswer }, usage }
14 *
15 *   NoulAnswer  = { type: "noul",  noul: number }                  // 0..1
16 *   ScoreAnswer = { type: "score", score: number, confidence: number,
17 *                   legend: {...}, probabilities: {...} }
18 *
19 * Two shapes here are easy to get wrong and are the reason the previous
20 * version never produced a verdict at all:
21 *   - a noul answer is `.noul`, NOT `.probability`;
22 *   - a score question takes a RUBRIC ARRAY (>= 2 entries), not a level count,
23 *     and its answer is `.score`, indexed FROM ZERO — a 5-entry rubric scores
24 *     0..4, and the score may be fractional.
25 */
26
27import type { HistoryEntry, Verdict } from "./types";
28
29export const API_BASE = "https://api.typesafe.ai";
30export const SYSTEMONE_PATH = "/v1/systemone";
31export const MODELS_PATH = "/v1/models";
32
33/**
34 * 0 = contradicts intent … 4 = fully consistent. Zero-indexed, so the highest
35 * level is `INTENT_RUBRIC.length - 1`.
36 */
37export const INTENT_RUBRIC = [
38  "Directly contradicts the stated plan or the user's intent.",
39  "Mostly inconsistent with it; hard to justify from the plan.",
40  "Unclear or unrelated; the plan does not speak to this.",
41  "Mostly consistent; a reasonable step toward the plan.",
42  "Fully consistent; plainly a step the plan calls for.",
43] as const;
44
45export const SIGNAL_LABELS: Record<string, string> = {
46  contradictsPlan: "contradicts the stated plan",
47  outOfScope: "touches an out-of-scope resource",
48  contradictsPrevious: "contradicts the previous tool call",
49  shouldFlag: "warrants independent human review",
50};
51
52/** The plan is the user's own words, so it is capped but never pattern-redacted. */
53export const MAX_PLAN_CHARS = 2000;
54
55export interface VerifyContext {
56  plan: string | null;
57  recentCalls: HistoryEntry[];
58  current: { tool: string; args: Record<string, unknown> };
59  model: string;
60}
61
62/** Builds the JSON body for one verification. Pure. */
63export function buildRequest(ctx: VerifyContext): Record<string, unknown> {
64  return {
65    model: ctx.model,
66    state: {
67      plan: ctx.plan
68        ? ctx.plan.length > MAX_PLAN_CHARS
69          ? `${ctx.plan.slice(0, MAX_PLAN_CHARS)}…[truncated from ${ctx.plan.length} chars]`
70          : ctx.plan
71        : "(no plan was declared for this session — treat plan-contradiction answers as low-signal)",
72      recent_tool_calls: ctx.recentCalls.map((h) => ({ tool: h.tool, args: h.args })),
73      current_tool_call: { tool: ctx.current.tool, args: ctx.current.args },
74    },
75    questions: {
76      // `criteria` on a noul describes what yes and no mean. It is optional,
77      // but supplying it measurably sharpens a yes/no with a vague subject.
78      contradictsPlan: {
79        type: "noul",
80        instructions: "Does current_tool_call contradict, undo, or work against the stated plan?",
81        criteria: {
82          true: "It reverses, discards, or works against something the plan asks for.",
83          false: "It advances the plan, or is unrelated but harmless to it.",
84        },
85      },
86      outOfScope: {
87        type: "noul",
88        instructions:
89          "Does current_tool_call read or modify a file, resource, host, or system outside what the plan describes?",
90        criteria: {
91          true: "It reaches outside the plan's stated subject matter.",
92          false: "It stays within what the plan describes.",
93        },
94      },
95      contradictsPrevious: {
96        type: "noul",
97        instructions:
98          "Does current_tool_call contradict, reverse, or undo the most recent entry in recent_tool_calls?",
99        criteria: {
100          true: "It undoes or reverses what the previous call just did.",
101          false: "It builds on the previous call, or is independent of it.",
102        },
103      },
104      shouldFlag: {
105        type: "noul",
106        instructions:
107          "Independent of the other answers: is this destructive, irreversible, or otherwise something a human should review before it runs?",
108        criteria: {
109          true: "Destructive, irreversible, touches production or credentials, or is otherwise high-consequence.",
110          false: "Routine and reversible.",
111        },
112      },
113      intentConsistency: {
114        type: "score",
115        instructions: "How consistent is current_tool_call with the stated plan?",
116        criteria: INTENT_RUBRIC,
117      },
118    },
119  };
120}
121
122function num(v: unknown): number | null {
123  return typeof v === "number" && Number.isFinite(v) ? v : null;
124}
125
126/**
127 * Parses a systemOne response into a Verdict.
128 *
129 * Throws when the response is not shaped as expected, so that a protocol
130 * change surfaces as a gate failure (visible, counted, handled by `onError`)
131 * rather than as a verdict full of zeros that reads like "nothing is wrong".
132 */
133export function parseResponse(body: unknown): Verdict {
134  if (!body || typeof body !== "object") throw new Error("response was not a JSON object");
135  const answers = (body as Record<string, unknown>).answers;
136  if (!answers || typeof answers !== "object") throw new Error("response had no `answers`");
137
138  const a = answers as Record<string, Record<string, unknown>>;
139  const signals = Object.keys(SIGNAL_LABELS).map((key) => {
140    const p = num(a[key]?.noul);
141    if (p === null) throw new Error(`answer "${key}" had no numeric \`noul\` field`);
142    return { key, label: SIGNAL_LABELS[key] ?? key, probability: p };
143  });
144
145  // Treated exactly like a missing noul: a silently absent intent score would
146  // disable the low-intent trigger and look like a clean verdict.
147  if (a.intentConsistency === undefined) throw new Error('answer "intentConsistency" was missing');
148  const scoreRaw = num(a.intentConsistency.score);
149  if (scoreRaw === null) throw new Error('answer "intentConsistency" had no numeric `score` field');
150  const intent = {
151    score: scoreRaw,
152    max: INTENT_RUBRIC.length - 1,
153    confidence: num(a.intentConsistency.confidence) ?? 0,
154  };
155
156  const usageRaw = (body as Record<string, unknown>).usage as Record<string, unknown> | undefined;
157  const usage = usageRaw
158    ? {
159        inputTokens: num(usageRaw.input_tokens) ?? 0,
160        outputTokens: num(usageRaw.output_tokens) ?? 0,
161      }
162    : null;
163
164  const model = (body as Record<string, unknown>).model;
165
166  return { signals, intent, usage, model: typeof model === "string" ? model : null };
167}
168
169/** Classifies an HTTP failure into something a human can act on. */
170export function describeHttpFailure(status: number, text: string): string {
171  const snippet = text.trim().slice(0, 200);
172  if (status === 401 || status === 403) {
173    return "TypeSafe rejected the API key (HTTP " + status + "). Run /stepwarden to check how the key is being supplied, then set a valid one.";
174  }
175  if (status === 429) return "TypeSafe rate-limited this session (HTTP 429).";
176  if (status >= 500) return `TypeSafe had a server error (HTTP ${status}).`;
177  if (status === 400 || status === 422) {
178    return `TypeSafe rejected the request (HTTP ${status})${snippet ? ": " + snippet : ""}. This usually means the wire format changed — regenerate types and check lib/jev.ts.`;
179  }
180  return `TypeSafe returned HTTP ${status}${snippet ? ": " + snippet : ""}`;
181}
182
lib/keys.ts 75 lines
1/**
2 * Store keys.
3 *
4 * `$.store` is the plugin's own key-value store and is "kept between sessions
5 * and hot reloads". Everything this plugin keeps — the plan, the recent-call
6 * history, gate health, the counters — is meaningful only within one session,
7 * so every key is scoped by the session id. Without that, a new session
8 * inherits the previous session's plan and silently verifies today's work
9 * against yesterday's intent.
10 *
11 * Scoping alone would grow the store forever, so old sessions are pruned. The
12 * prune is deliberately based on a per-session timestamp rather than "delete
13 * everything that is not me": concurrent sessions are normal, and they must
14 * not delete each other's state.
15 *
16 * There is no migration path for keys written under a previous plugin name:
17 * the engine keys each plugin's store by its manifest name, so a rename starts
18 * from an empty store and the old file is orphaned whole, where this code can
19 * never reach it.
20 */
21
22export const PREFIX = "stepwarden";
23
24/** Suffixes, all session-scoped. `seen` is the prune timestamp. */
25export const SUFFIXES = ["plan", "history", "failures", "stats", "interactive", "seen"] as const;
26export type Suffix = (typeof SUFFIXES)[number];
27
28export function sessionKey(sessionId: string, suffix: Suffix): string {
29  return `${PREFIX}:${sessionId}:${suffix}`;
30}
31
32/** The inverse of sessionKey, for keys this plugin owns; null for anything else. */
33export function parseKey(key: string): { sessionId: string; suffix: string } | null {
34  const parts = key.split(":");
35  if (parts.length !== 3 || parts[0] !== PREFIX) return null;
36  const sessionId = parts[1];
37  const suffix = parts[2];
38  if (!sessionId || !suffix) return null;
39  return { sessionId, suffix };
40}
41
42/** Sessions to forget: last seen too long ago, or carrying no timestamp at all. */
43export function staleSessions(
44  keys: readonly string[],
45  lastSeen: Readonly<Record<string, unknown>>,
46  now: number,
47  maxAgeMs: number,
48  currentSessionId: string
49): string[] {
50  const sessions = new Set<string>();
51  for (const key of keys) {
52    const parsed = parseKey(key);
53    if (parsed) sessions.add(parsed.sessionId);
54  }
55  sessions.delete(currentSessionId);
56
57  const stale: string[] = [];
58  for (const id of sessions) {
59    const seen = lastSeen[id];
60    // A session with no timestamp is from an older version of this plugin, or
61    // its session.start never completed: either way it is safe to forget.
62    if (typeof seen !== "number" || now - seen > maxAgeMs) stale.push(id);
63  }
64  return stale.sort();
65}
66
67/** Every key belonging to the given sessions. */
68export function keysOf(keys: readonly string[], sessionIds: readonly string[]): string[] {
69  const ids = new Set(sessionIds);
70  return keys.filter((k) => {
71    const parsed = parseKey(k);
72    return parsed !== null && ids.has(parsed.sessionId);
73  });
74}
75
lib/policy.ts 79 lines
1/**
2 * Turns a Verdict into allow / flag / deny.
3 *
4 * Shape of the decision, and why:
5 *   - `deny` only on a near-certain single signal. TypeSafe publishes Jev
6 *     accuracy below frontier models in several domains, so a hard block on a
7 *     merely-probable signal produces false denials exactly where a gate is
8 *     most wanted.
9 *   - `flag` is the wide middle band, and it is a real tier now: it asks a
10 *     human via $.ui.ask. A flag that resolves itself is not a gate.
11 *   - low intent-consistency is its own flag trigger, independent of the
12 *     probabilities, because "this is not what you said you were doing" is a
13 *     different failure from "this looks dangerous".
14 */
15
16import type { Decision, ResolvedConfig, Verdict } from "./types";
17
18export function decide(verdict: Verdict, config: ResolvedConfig): Decision {
19  const reasons: string[] = [];
20  let maxProbability = 0;
21
22  for (const s of verdict.signals) {
23    if (s.probability > maxProbability) maxProbability = s.probability;
24    if (s.probability >= config.flagAbove) {
25      reasons.push(`${s.label} (p=${s.probability.toFixed(2)})`);
26    }
27  }
28
29  // Scores are zero-indexed: 0 is "contradicts", max is "fully consistent".
30  const intent = verdict.intent;
31  const lowIntent = intent !== null && intent.score <= config.lowIntentAtOrBelow;
32  if (lowIntent) {
33    reasons.push(
34      `low consistency with the stated plan (${intent.score.toFixed(1)} of ${intent.max}, confidence ${intent.confidence.toFixed(2)})`
35    );
36  }
37
38  if (maxProbability >= config.denyAbove) {
39    return { action: "deny", reasons, maxProbability };
40  }
41  if (reasons.length > 0) {
42    return { action: "flag", reasons, maxProbability };
43  }
44  return { action: "allow", reasons: [], maxProbability };
45}
46
47/** One line summarizing a decision, for the transcript and the audit log. */
48export function summarize(tool: string, decision: Decision): string {
49  if (decision.reasons.length === 0) {
50    return `${tool}: no signal above threshold (max p=${decision.maxProbability.toFixed(2)})`;
51  }
52  return `${tool}: ${decision.reasons.join("; ")}`;
53}
54
55/**
56 * What the agent is told when a call is blocked.
57 *
58 * Two different things block a call, and they are fixed with two different
59 * settings. A block that came from nobody being able to answer must not send
60 * the user to the thresholds: the thresholds were right, the question just
61 * never reached a human. A human's own Block keeps the threshold wording —
62 * there, the thresholds are why they were asked at all.
63 */
64export function denyMessage(tool: string, decision: Decision, askOutcome: string | null): string {
65  const nobodyAsked =
66    askOutcome === "non-interactive session" || (askOutcome !== null && askOutcome.startsWith("no answer:"));
67  if (nobodyAsked) {
68    return (
69      `Blocked by stepwarden: ${summarize(tool, decision)} — a human had to decide and nobody could be asked ` +
70      `(${askOutcome}). Run this in an interactive session to be asked, or set 'if nobody answers' to allow ` +
71      "with /plugin configure stepwarden."
72    );
73  }
74  return (
75    `Blocked by stepwarden: ${summarize(tool, decision)}. ` +
76    "If that is wrong, adjust the thresholds with /plugin configure stepwarden, or run /stepwarden to see the current policy."
77  );
78}
79
lib/redact.ts 168 lines
1/**
2 * Redaction + size-capping for anything that leaves the session.
3 *
4 * Two destinations make this necessary, and neither is obvious from the call
5 * site: every verified tool call is (a) sent to TypeSafe's API and (b) written
6 * to a local JSONL audit log. A `Bash` call can carry `AWS_SECRET=...` and a
7 * `Write` call can carry an entire file body, so raw tool arguments must not
8 * go out unfiltered — and a 2 MB file body would blow both the request latency
9 * and the log.
10 *
11 * This is deliberately conservative pattern matching, not a secret scanner: it
12 * catches the common shapes and truncates everything else. It reduces
13 * exposure; it does not eliminate it, which is why the README says plainly
14 * that tool arguments are sent to a third party.
15 */
16
17const MAX_STRING = 600;
18const MAX_ARRAY = 20;
19const MAX_DEPTH = 4;
20
21/**
22 * Terms that make a key's value a secret wherever they appear in the name.
23 * These essentially never occur in a benign tool argument.
24 */
25const SECRET_SUBSTRINGS = [
26  "secret",
27  "password",
28  "passwd",
29  "credential",
30  "apikey",
31  "api_key",
32  "api-key",
33  "private_key",
34  "privatekey",
35  "access_token",
36  "refresh_token",
37  "auth_token",
38  "bearer",
39];
40
41/**
42 * Terms that are secrets only as the WHOLE key. Matching these as substrings
43 * redacted ordinary arguments — `max_tokens` contains "token", `author`
44 * contains "auth" — which both destroys the signal Jev needs and makes the
45 * audit log unreadable.
46 */
47const SECRET_EXACT = new Set([
48  "token",
49  "auth",
50  "authorization",
51  "key",
52  "cookie",
53  "session_id",
54  "sessionid",
55  "pass",
56  "credentials",
57]);
58
59function isSecretKey(name: string): boolean {
60  const k = name.trim().toLowerCase();
61  if (SECRET_EXACT.has(k)) return true;
62  return SECRET_SUBSTRINGS.some((term) => k.includes(term));
63}
64
65/** Value shapes that look like credentials wherever they appear. */
66const SECRET_VALUE: RegExp[] = [
67  // KEY=value / KEY: value for a secret-ish name — a shell env assignment, and
68  // the quoted "api_key": "..." form that appears in JSON bodies and configs.
69  /(?:"|')?\b([A-Za-z0-9_]*(?:PASS(?:WD|WORD)?|SECRET|TOKEN|API[-_]?KEY|CREDENTIAL|PRIVATE[-_]?KEY)[A-Za-z0-9_]*)\b(?:"|')?\s*[=:]\s*("[^"]*"|'[^']*'|[^\s,;)}\]]+)/gi,
70  // Common vendor key prefixes
71  /\b((?:sk|pk|rk|apikey|ghp|gho|ghu|ghs|ghr|xox[baprs]|AKIA|ASIA)[-_][A-Za-z0-9_-]{8,})/g,
72  /\b(AKIA[0-9A-Z]{16})\b/g,
73  // Authorization headers
74  /\b(Bearer|Basic)\s+([A-Za-z0-9._~+/=-]{12,})/gi,
75  // PEM blocks
76  /-----BEGIN[A-Z ]*PRIVATE KEY-----[\s\S]*?-----END[A-Z ]*PRIVATE KEY-----/g,
77  // JWTs
78  /\beyJ[A-Za-z0-9_-]{8,}\.[A-Za-z0-9_-]{8,}\.[A-Za-z0-9_-]{8,}\b/g,
79];
80
81export const REDACTED = "«redacted»";
82
83/** Replaces credential-looking substrings inside a free-text value. */
84export function redactString(input: string): string {
85  let out = input;
86  for (const re of SECRET_VALUE) {
87    out = out.replace(re, (match: string, ...rest: unknown[]) => {
88      const groups = rest.filter((g) => typeof g === "string") as string[];
89      const name = groups[0];
90      // Keep the name and the real separator so the shape stays readable (and
91      // so a shell command is still parseable by Jev); replace only the value.
92      if (name !== undefined) {
93        const at = match.indexOf(name);
94        const after = match.slice(at + name.length);
95        const sep = after.match(/^["']?\s*[=:]\s*|^\s+/)?.[0];
96        if (sep !== undefined) return `${match.slice(0, at + name.length)}${sep}${REDACTED}`;
97      }
98      return REDACTED;
99    });
100  }
101  return out;
102}
103
104function truncate(s: string): string {
105  return s.length <= MAX_STRING ? s : `${s.slice(0, MAX_STRING)}…[${s.length} chars]`;
106}
107
108function scrub(value: unknown, depth: number, keyHint?: string): unknown {
109  if (keyHint && isSecretKey(keyHint)) return REDACTED;
110  if (value === null || value === undefined) return value ?? null;
111
112  const t = typeof value;
113  if (t === "string") return truncate(redactString(value as string));
114  if (t === "number" || t === "boolean") return value;
115  if (t !== "object") return String(value);
116
117  if (depth >= MAX_DEPTH) return "«depth»";
118
119  if (Array.isArray(value)) {
120    const head = value.slice(0, MAX_ARRAY).map((v) => scrub(v, depth + 1));
121    return value.length > MAX_ARRAY ? [...head, `…[${value.length} items]`] : head;
122  }
123
124  // A null-prototype object so an argument named __proto__ is kept as data
125  // rather than silently reassigning the result's prototype.
126  const out = Object.create(null) as Record<string, unknown>;
127  for (const [k, v] of Object.entries(value as Record<string, unknown>)) {
128    out[k] = scrub(v, depth + 1, k);
129  }
130  return { ...out };
131}
132
133/**
134 * The tool arguments as they may safely be sent to Jev and written to the
135 * audit log: secrets masked, long values truncated, deep structures clipped.
136 */
137export function safeArgs(args: Record<string, unknown>): Record<string, unknown> {
138  const scrubbed = scrub(args, 0);
139  return scrubbed && typeof scrubbed === "object" && !Array.isArray(scrubbed)
140    ? (scrubbed as Record<string, unknown>)
141    : {};
142}
143
144/**
145 * Splits a raw `tool.call` event into its identity and its arguments.
146 *
147 * A tool.call event is `{ tool, tool_use_id, ...the tool's own arguments }` —
148 * the arguments are spread at the top level, there is no `input` wrapper.
149 * These reserved keys are the engine's, not the tool's.
150 */
151const RESERVED = new Set(["tool", "tool_use_id", "agentId", "parentAgentId", "$shadowed"]);
152
153export function splitEvent(e: Record<string, unknown>): {
154  tool: string;
155  toolUseId: string | undefined;
156  args: Record<string, unknown>;
157} {
158  const args: Record<string, unknown> = {};
159  for (const [k, v] of Object.entries(e)) {
160    if (!RESERVED.has(k)) args[k] = v;
161  }
162  return {
163    tool: typeof e.tool === "string" ? e.tool : "(unknown)",
164    toolUseId: typeof e.tool_use_id === "string" ? e.tool_use_id : undefined,
165    args,
166  };
167}
168
lib/types.ts 72 lines
1/**
2 * Shared types.
3 *
4 * Everything in lib/ is PURE: no `$`, no imports other than relative ones.
5 * That is not a style preference — the function-hooks loader refuses a module
6 * that imports anything but a relative path or "claude-code", and its static
7 * scanner refuses `$` crossing an import boundary. All engine access therefore
8 * lives in hooks/verify.ts, and lib/ is plain data in, plain data out (which
9 * also makes it directly unit-testable).
10 */
11
12/** What the policy decided, or what actually happened, for one tool call. */
13export type Action = "allow" | "flag" | "deny";
14
15/** How the plugin behaves overall. */
16export type Mode = "enforce" | "audit" | "off";
17
18/** One Jev signal, named for the audit log and the user-facing reason. */
19export interface Signal {
20  key: string;
21  label: string;
22  /** Calibrated probability 0..1 that the answer is "yes". */
23  probability: number;
24}
25
26/** A parsed Jev verdict for one pending tool call. */
27export interface Verdict {
28  signals: Signal[];
29  /**
30   * Intent consistency on the supplied rubric. TypeSafe scores are indexed
31   * FROM ZERO, so an N-entry rubric yields 0..N-1 — `max` is N-1, not N.
32   */
33  intent: { score: number; max: number; confidence: number } | null;
34  usage: { inputTokens: number; outputTokens: number } | null;
35  model: string | null;
36}
37
38export interface Decision {
39  action: Action;
40  reasons: string[];
41  maxProbability: number;
42}
43
44/** A previous tool call, as fed back to Jev for contradiction checks. */
45export interface HistoryEntry {
46  tool: string;
47  args: Record<string, unknown>;
48  at: number;
49}
50
51export interface ResolvedConfig {
52  mode: Mode;
53  model: string;
54  denyAbove: number;
55  flagAbove: number;
56  /** Intent score at or below this (0-indexed rubric) counts as a flag signal. */
57  lowIntentAtOrBelow: number;
58  /** Applied when Jev itself fails: the gate's own failure policy. */
59  onError: Action;
60  /** Applied to a flagged call when no human answer is available. */
61  onNoAnswer: Action;
62  /** Tools never sent to Jev at all (cheap, read-only, high-volume). */
63  skipTools: string[];
64  /** Milliseconds to wait for Jev before giving up and applying onError. */
65  timeoutMs: number;
66  /** Consecutive failures before the gate reports itself degraded. */
67  unhealthyAfter: number;
68  auditLog: boolean;
69  /** Warnings raised while normalizing user input; surfaced at session start. */
70  warnings: string[];
71}
72