Every tool call your agent makes, checked before it runs. Verifies each pending tool call against the session's plan using TypeSafe AI's Jev model, then allows…

<h1 align="center"> <img src="assets/logo-light.svg" width="538" alt="stepwarden — Every tool call your agent makes, checked before it runs."> </h1>
A Claude Code plugin that verifies each agent action before it executes — not after the run finishes. It routes every tool call through a purpose-built verification model, TypeSafe AI's Jev, instead of a full LLM review, which is what makes checking every step affordable.
› My plan is strictly to add a docstring to lib/policy.ts and nothing else.
Now run: git log --oneline -20 && cat ~/.gitconfig
┌ stepwarden ──────────────────────────────────────────────────┐
│ Bash: touches an out-of-scope resource (p=0.83). Allow this │
│ call? [ Allow ] [ Block ] │
└───────────────────────────────────────────────────────────────┘
That dialog is what mode: enforce looks like. A fresh install ships in mode: audit: the same verdict is computed and written down, and the call runs. You turn the gate on yourself, with /stepwarden enforce.
With the shipped thresholds that call lands in the ask band, so you decide. Raise the pressure (denyAbove: 0.8) and the same call is blocked outright, with the reason handed back to the model:
The command was blocked by stepwarden before it ran:
> Bash: touches an out-of-scope resource (p=0.83)
[!IMPORTANT] Status: Proof of Concept This demonstrates an architecture pattern. It is built on two early-access dependencies (Claude Code function hooks and Jev), its default thresholds have not been tuned on production traffic, and it is not a substitute for deterministic security controls. Read Limits before relying on it.
Currently Agents are mostly reviewed at the end of a run, if at all. Between the moment you hit enter and the moment you read the summary, the agent makes dozens or hundreds of tool calls, and you find out what it did between checkpoints by reading the diff afterwards — or by not reading it.
Everyone knows the fix: check each step before it executes. Nobody does it, because running a frontier model as a judge on every tool call has always been technically possible but absurd from an economic or latency perspective. So teams sample, or audit at the end, and call it a process.
Jev changes that arithmetic. It doesn't generate text: you hand it state and a list of typed questions, and it returns typed answers with a calibrated confidence on each, sampled in parallel rather than token by token. Near-zero marginal cost per question, sub-second latency. This plugin is a proof of concept that we can use Jev (and decision models like it) to build such a smart gate.
There are real limits, though. — see Limits.
Before any tool executes, it asks Jev five independent questions in parallel, against the session plan and recent history:
Each answer comes back with a calibrated probability. A full five-question check takes about half a second and 740–1,200 input tokens.
| Outcome | When | What happens |
|---|---|---|
| allow | no signal reaches the ask threshold | the call runs, nothing is shown |
| ask | any signal at or above flagAbove (default 0.6), or intent consistency at or below 1 | you get an Allow / Block dialog; your answer decides |
| block | any signal at or above denyAbove (default 0.9) | the call never runs; the reason goes back to the model |
Both thresholds are inclusive. If you cannot be asked — a -p run, a dismissed dialog, or more than 50 questions already this session — onNoAnswer decides, and the plugin says so rather than letting the call through silently: a notice on screen where there is one, a line in the audit log always, and the reason inside the block message when the answer is deny. A headless -p run has no transcript of its own, so there the notice reaches the SDK host as a ui_log message and the --debug-file log, and the audit log is the record to read.
The middle tier is the point. Blocking on a merely-probable signal produces false stops exactly where a gate matters most, so block is reserved for near-certainty and the wide middle band asks you instead.
Requires Claude Code 2.1.273 or newer with function hooks enabled:
export CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1
Then, inside Claude Code:
/plugin marketplace add getexcited/claude-plugins
/plugin install stepwarden@getexcited
The install dialog asks for your TypeSafe API key. Leaving it blank is fine — the plugin still loads, says so at startup, and /stepwarden tells you how to set one. See Setting your API key.
It starts in audit mode. A fresh install verifies every call and writes down what it would have done, but never blocks and never interrupts you. Use it on real work for a while, read the log, and turn the gate on when the numbers look right:
/stepwarden enforce
See Audit mode is the default, or run /stepwarden help for the whole surface — settings, commands and the key — inside the session.
From a local checkout instead:
git clone https://github.com/getexcited/stepwarden
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude --plugin-dir ./stepwarden
There is nothing to build and nothing to npm install — the plugin has no runtime dependencies, and no lockfile ships with it, so installing it never runs npm and never copies a node_modules into your plugins directory.
Get a key at typesafe.ai, then pick either route:
Through Claude Code (recommended). The key is stored in your OS keychain, never in a file in your project:
/plugin configure stepwarden
This is also where every other setting lives — mode, thresholds, which tools to skip — each with an explanation in the dialog.
Through the environment, if you would rather manage it yourself:
export TYPESAFE_API_KEY=apikey_...
A key set in the plugin config wins over the environment.
A
.envfile is not enough on its own. Claude Code does not read.env, so a key sitting there will not be found. Either use/plugin configure, or source the file into your shell first:set -a; . ./.env; set +a.
Run /stepwarden at any time to see whether the key is set, where it came from, whether TypeSafe accepts it, what the current policy is, what the plugin has decided so far this session, and the exact command to analyse this project's audit log. /stepwarden help is the other half: what every setting does, what the commands are, and both ways to set your key — without leaving the session to find this page. If the key is missing or rejected, /stepwarden tells you exactly how to fix it — and the plugin says so at startup rather than quietly verifying nothing.
All of these live in /plugin configure stepwarden. The one you will reach for most often has its own command:
/stepwarden enforce turn the gate on: risky calls are blocked or put to you
/stepwarden audit back to logging only, nothing blocked
/stepwarden off verify nothing
/stepwarden toggle flip between audit and enforce
/stepwarden help every setting, every command, and how to set your key
It takes effect in the running session — mid-turn too, which is when you usually want it — and is remembered for the sessions after it. /plugin configure stepwarden still owns the setting: it keeps showing the configured mode, and the moment you change it there, that wins and the switch is forgotten. If your settings are managed by someone else and the change is refused, the command says so and changes nothing.
| Setting | Default | What it does |
|---|---|---|
TYPESAFE_API_KEY | — | Your key. Stored in the OS keychain. |
mode | audit | audit decides and logs but never blocks — the shipped default. enforce blocks and asks. off verifies nothing. Also /stepwarden <mode>. |
denyAbove | 0.9 | Block outright at or above this probability. |
flagAbove | 0.6 | Ask you at or above this probability. |
lowIntentAtOrBelow | 1 | Ask when consistency with the plan scores at or below this, on 0 (contradicts) to 4 (fully consistent). |
onError | allow | allow or deny, for when Jev is unreachable, times out, or no key is set. |
onNoAnswer | allow | allow or deny, for when a flagged call cannot be put to a human. |
skipTools | Read, Glob, Grep, … | Tools never sent for verification. Set to none to verify everything. |
model | jev-latest | Which TypeSafe model answers. |
timeoutMs | 5000 | How long to wait before giving up and applying onError. |
unhealthyAfter | 3 | Warn once after this many verification failures in a row. |
auditLog | true | Write every decision to .claude/stepwarden/<session>.jsonl. |
onError and onNoAnswer deliberately offer only allow and deny: both describe what to do when a human cannot be asked, so "ask" is not an answer.
A bad value never takes the gate down: it falls back to the default, warns at session start, and shows up in /stepwarden. Thresholds that would make ask unreachable are repaired and reported.
A fresh install runs in mode: audit: every call is verified, every decision is logged, nothing is ever blocked and nothing interrupts you. Use it normally for a while, then read what it would have done:
cd ~/my-project # where you ran Claude Code
node ~/path/to/stepwarden/scripts/analyze-audit.mjs # /stepwarden prints this path
The plugin writes one file per session to .claude/stepwarden/ in the project you were working in, so run the script from there (not from the plugin's own directory). That prints what your thresholds would have done to your real traffic — the distribution of each signal, where the policy and reality diverged, which tools dominate. If a signal's p99 sits well below your block threshold, that threshold can never fire and you should lower it or drop it. The script has no ground truth: it tells you what the policy does, not whether a flagged call deserved it. Sample some rows by hand before tightening anything.
When the numbers look right, turn the gate on:
/stepwarden enforce
For every verified call, these are sent to api.typesafe.ai:
Arguments are redacted first — values under key names like password, api_key or secret, KEY=value assignments, "api_key": "…" in JSON, bearer tokens, PEM blocks and JWTs are masked, and long values truncated. Redaction is pattern matching, not a guarantee: a secret in an unusual shape will get through. The plan itself is not redacted, because it is the thing being compared against.
If a tool's arguments must never leave your machine, add it to skipTools. If the work itself is confidential, use mode: off.
Locally, the same redacted arguments are written to .claude/stepwarden/<session>.jsonl in your project. That file is worth adding to .gitignore. Turn it off with the auditLog setting.
Following the Claude Mods five-line convention. Every claim here is what claude plugin validate . --strict reports, not a summary of intent — it statically lists what the module hooks, calls and reads, and refuses to load it if the source disagrees.
session.start, prompt.submit, tool.call, command.run{command=stepwarden}.TYPESAFE_API_KEY from the environment — the only environment variable it reads, and it writes none.claude-code — no node: builtins, no npm packages, no subprocesses — and that is enforced by the loader, not by convention.api.typesafe.ai only — the tool name, its redacted arguments, the plan, and the last eight calls. Nothing else leaves the machine, and there is no telemetry..claude/stepwarden/<session>.jsonl in your project (off with auditLog), plus its own key-value store for the plan, history and counters, scoped per session and pruned after seven days.Every decision is logged — allowed, asked, blocked — with the per-question probabilities behind it, whether or not enforcement was on. Gate-health events (no_api_key, key_check_failed, gate_unhealthy, gate_recovered, gate_crashed) are logged too, so a period when the gate was degraded cannot read as a clean record.
One file per session, so two Claude Code windows in one project do not overwrite each other's log.
The first prompt of the session is stored verbatim as the plan and never changes. It is deliberately not summarised: a summary can quietly misrepresent the intent that everything else is then measured against. /stepwarden shows the plan in force.
This is also the main limitation. A long session drifts from its opening sentence, and the plan does not follow. Say what you are doing in your first message, and start a new session when the work changes.
Read this section before quoting anything.
Accuracy is below frontier level. On TypeSafe's own dashboard, Jev aggregates 67.8% against 74.1% for the best comparator, with wider gaps in some domains (invoice processing 61.8% vs 79.1%; security incidents 61.7% vs 66.2%). The one independent test we know of found it caught 6 of 7 planted defects where a frontier model caught 7 of 7 — at roughly 25× the speed and 1/580th the cost. That trade-off is the entire point: a mid-tier judgment on every action instead of a frontier judgment on none of them. This will miss things. It is also why the shipped defaults block rarely and ask often — and why a fresh install blocks nothing at all until you turn enforcement on.
It is not a security boundary. It is a fast, narrow, probabilistic judge — one model checking another model's next step. It does not "understand code" the way a full review does. For anything non-negotiable — production credentials, a force-push to main — pair it with a deterministic, fail-closed PreToolUse shell hook that works whether or not this plugin or TypeSafe are running.
Function hooks fail open. A hook that throws is skipped and the tool simply runs, which means a bug in the verifier looks exactly like a clean verdict. This plugin handles its own crashes — anything thrown inside the gate becomes a configured, logged outcome that honours onError — and it counts consecutive failures and says so. What it cannot cover is a module that fails to load, and no plugin can. It also fails open by choice when Jev is unreachable, because a verifier outage that stops your work is worse than one that admits it; set onError: deny if your threat model says otherwise.
Two early-access dependencies. The function-hooks API is behind CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 and has changed field names during its preview — the declarations in .claude/types/ are generated from one specific build, and /plugin-types must be re-run after a Claude Code update. Jev is waitlisted, with no published architecture or weights, and its calibration claim — that 70% confidence means right 70% of the time, which is what makes a threshold meaningful at all — has not been independently audited on messy input.
Repeated calls move. Ask Jev the same question twice and the probabilities shift by a few points. Worth knowing before you set a threshold on a boundary.
Thresholds are starting points, not numbers derived from your traffic. That is why mode: audit is what you get on install.
The numbers above are TypeSafe's published figures plus one third-party test, not independently audited. Pricing at time of writing: $0.042 per million input tokens, output free; 70–500 ms end to end in TypeSafe's own testing. Note that the API does report a non-zero output_tokens per call, so confirm the billing model yourself before relying on the arithmetic.
The type declarations for the function-hooks API are not in the repo: they are Claude Code's own, and they are generated per build. After cloning, open a Claude Code session in the checkout and run /plugin-types, which writes them to .claude/types/ (gitignored). Re-run it after every Claude Code update.
No lockfile is committed (see Install), so install the one development dependency — TypeScript — before running anything:
npm install # dev-only: TypeScript, for `npm run typecheck`
npm run check # typecheck + manifest validation + tests
npm run typecheck
npm run validate # claude plugin validate . --strict
npm test # claude plugin test .
Run npm run check after regenerating the types, before trusting anything.
.claude-plugin/plugin.json manifest and the userConfig schema
hooks/hooks.json points Claude Code at the module
hooks/verify.ts every $ call lives here — see the header comment
lib/config.ts options -> validated config (pure)
lib/keys.ts session-scoped store keys and pruning (pure)
lib/jev.ts TypeSafe wire format (pure)
lib/policy.ts verdict -> allow / ask / block (pure)
lib/redact.ts secret masking and size caps (pure)
lib/types.ts shared types
tests/ unit + end-to-end tests
scripts/analyze-audit.mjs audit log summariser (plain Node)
hooks/verify.ts is one large file on purpose. A function-hooks module is scanned before it loads and $ may not cross an import boundary, so every engine call has to live in the module that registers the hooks. Everything else is pure and lives in lib/, which is why it can be unit-tested directly.
The plugin deliberately does not use @typesafe-ai/sdk: a hooks module may import only relative paths and claude-code, so the API is called through $.http.fetch instead. lib/jev.ts holds the wire format, verified against the live API.
The most useful thing you can do right now: run it as installed — mode: audit — on real work for a week and share the anonymised analyze-audit.mjs output — especially the false positives. Threshold tuning needs traffic we don't have, and a gate that cries wolf is worse than no gate, because you learn to click through it.
Bug reports are most useful with the matching lines from .claude/stepwarden/<session>.jsonl; they carry the per-question probabilities behind whatever it did.
Apache-2.0 — see LICENSE.
stepwarden is an independent project, not affiliated with or endorsed by Anthropic or TypeSafe AI.
hooks/verify.ts 846 lines1/**
2 * stepwarden — per-tool-call verification through TypeSafe AI's Jev.
3 *
4 * Before each tool call runs, this asks Jev five questions about it against the
5 * session's plan and recent history, and then allows it, asks a human, or
6 * blocks it.
7 *
8 * WHY EVERYTHING LIVES IN ONE FILE
9 * A function-hooks module is scanned statically before it loads. Three rules
10 * shape this file, and breaking any of them makes the plugin fail to load with
11 * no gate and no obvious error:
12 * 1. `$` may not cross an import boundary, and may only be handed to a
13 * function declared at the TOP LEVEL of this same file — which is why the
14 * helpers below take `config` as a parameter instead of closing over it.
15 * 2. A hooks module may import only relative paths and "claude-code" — no
16 * npm packages, no node: builtins. That is why the TypeSafe SDK is not
17 * used and the API is called directly through `$.http.fetch`.
18 * 3. `$.env.get` takes a literal name, so the host can list what a module
19 * reads; it cannot be looped over a variable.
20 * Everything else is pure and lives in lib/, where it is unit-testable.
21 * Regenerate the types with `/plugin-types` after a Claude Code update.
22 */
23
24import type { Register } from "claude-code";
25import {
26 type RawOptions,
27 describeMode,
28 helpText,
29 overrideMode,
30 parseCommandArg,
31 resolveApiKey,
32 resolveConfig,
33} from "../lib/config";
34import {
35 API_BASE,
36 MODELS_PATH,
37 SYSTEMONE_PATH,
38 buildRequest,
39 describeHttpFailure,
40 parseResponse,
41} from "../lib/jev";
42import { keysOf, parseKey, sessionKey, staleSessions } from "../lib/keys";
43import { decide, denyMessage } from "../lib/policy";
44import { safeArgs, splitEvent } from "../lib/redact";
45import type { Action, HistoryEntry, Mode, ResolvedConfig, Verdict } from "../lib/types";
46
47/** Index of session id -> last seen, used only to prune the store. */
48const LAST_SEEN_KEY = "stepwarden:last-seen";
49/** Where `/stepwarden <mode>` remembers a switch for later sessions. */
50const MODE_KEY = "stepwarden:mode-override";
51const AUDIT_DIR = ".claude/stepwarden";
52
53/** Sessions are forgotten after this long. */
54const SESSION_TTL_MS = 7 * 24 * 60 * 60 * 1000;
55
56/** Tool calls kept for contradiction checks. */
57const HISTORY_KEEP = 8;
58/** Audit lines kept on disk before the oldest are dropped. */
59const AUDIT_KEEP = 2000;
60
61interface Stats {
62 checked: number;
63 allowed: number;
64 flagged: number;
65 denied: number;
66 skipped: number;
67 errors: number;
68 asked: number;
69 inputTokens: number;
70 outputTokens: number;
71}
72
73const ZERO_STATS: Stats = {
74 checked: 0,
75 allowed: 0,
76 flagged: 0,
77 denied: 0,
78 skipped: 0,
79 errors: 0,
80 asked: 0,
81 inputTokens: 0,
82 outputTokens: 0,
83};
84
85/* ------------------------------------------------------------------ helpers
86 * Top-level declarations, so the scanner will follow `$` into them.
87 */
88
89async function readKey($: any, options: RawOptions): Promise<{ key: string | null; source: string }> {
90 let fromEnv: string | undefined;
91 try {
92 // A literal name: the host lists what a module reads from the environment.
93 fromEnv = await $.env.get("TYPESAFE_API_KEY");
94 } catch {
95 fromEnv = undefined;
96 }
97 return resolveApiKey(options, fromEnv);
98}
99
100/** The session this hook is running in; every stored key is scoped by it. */
101async function sessionId($: any): Promise<string> {
102 try {
103 const id = await $.session.id();
104 return typeof id === "string" && id.length > 0 ? id : "unknown-session";
105 } catch {
106 return "unknown-session";
107 }
108}
109
110async function getStats($: any, sid: string): Promise<Stats> {
111 try {
112 const raw = await $.store.get(sessionKey(sid, "stats"));
113 if (raw && typeof raw === "object") return { ...ZERO_STATS, ...(raw as Partial<Stats>) };
114 } catch {
115 /* fall through to zeros */
116 }
117 return { ...ZERO_STATS };
118}
119
120async function bumpStats($: any, sid: string, patch: Partial<Stats>): Promise<void> {
121 await serialize(async () => {
122 try {
123 const cur = await getStats($, sid);
124 const next: Record<string, number> = { ...(cur as unknown as Record<string, number>) };
125 for (const [k, v] of Object.entries(patch)) {
126 if (typeof v === "number") next[k] = (next[k] ?? 0) + v;
127 }
128 await $.store.set(sessionKey(sid, "stats"), next);
129 } catch {
130 /* stats are best-effort */
131 }
132 });
133}
134
135/**
136 * The audit log's lines, per session, held for the life of the module.
137 *
138 * `$.fs.write` has no append mode, so the file is rewritten on every entry.
139 * Re-reading it first would make that quadratic in a long session, and this
140 * plugin is the only writer, so the lines are read once and kept.
141 *
142 * Keyed by session id, and memoised so two concurrent first-calls cannot both
143 * load and race. One module can serve more than one session, and the path is
144 * derived from the session id at write time — one shared buffer would write one
145 * session's lines into the other session's file.
146 */
147const auditLines = new Map<string, Promise<string[]>>();
148
149async function audit(
150 $: any,
151 config: ResolvedConfig,
152 sid: string,
153 entry: Record<string, unknown>
154): Promise<void> {
155 if (!config.auditLog) return;
156 try {
157 const path = `${AUDIT_DIR}/${sid}.jsonl`;
158 let load = auditLines.get(sid);
159 if (load === undefined) {
160 load = (async () => {
161 // Ask before reading. `$.fs.read` rejects on a missing file and the host
162 // records that rejection, which put an ENOENT line in the debug log of
163 // every clean session. `$.fs.exists` never rejects.
164 let existing = "";
165 try {
166 if (await $.fs.exists(path)) existing = await $.fs.read(path);
167 } catch {
168 existing = ""; // unreadable: start a fresh buffer rather than lose this line
169 }
170 return existing.length > 0 ? existing.split("\n").filter((l: string) => l.length > 0) : [];
171 })();
172 auditLines.set(sid, load);
173 }
174 const lines = await load;
175
176 const at = await $.clock.now();
177 // Pushing before the await means a concurrent writer's line is already in
178 // the array by the time either write runs: last write wins, nothing is lost.
179 lines.push(JSON.stringify({ at, ...entry }));
180 if (lines.length > AUDIT_KEEP) lines.splice(0, lines.length - AUDIT_KEEP);
181 await $.fs.write(path, lines.join("\n") + "\n");
182 } catch {
183 /* an audit write must never break a tool call */
184 }
185}
186
187/**
188 * The mode `/stepwarden <mode>` set, per session.
189 *
190 * The module is loaded once, with the options it had then, so writing the
191 * config alone would not reach the hooks already registered here. The override
192 * applies to the session that asked for it at once; `$.config.set` persists the
193 * same value, which every later session loads. Keyed by session id for the same
194 * reason the audit buffer is: one module can serve more than one session, and
195 * one session's switch is not another session's business. Cleared on load, so a
196 * session always starts from the stored config.
197 */
198const modeOverride = new Map<string, Mode>();
199
200/** The mode in force right now: this session's switch, else the config. */
201function currentMode(config: ResolvedConfig, sid: string): Mode {
202 return modeOverride.get(sid) ?? config.mode;
203}
204
205/**
206 * Serves `/stepwarden <mode>`: remember it, then apply it to this session.
207 *
208 * Saving comes first on purpose. A refusal must not leave the running session
209 * enforcing something the settings refused.
210 *
211 * Two ways to save, and the difference matters. The settings row the plugin
212 * owns is the right home, but a Claude Code that does not expose a plugin's
213 * `userConfig` through `$.config` rejects the write outright — 2.1.273 lists no
214 * plugin rows at all — so the switch is kept in the plugin's own store instead,
215 * next to the configured value it was made against. A `{ deny }`, on the other
216 * hand, is somebody's decision: managed settings or an organization policy own
217 * the row, and that is not something to route around.
218 */
219async function applyMode($: any, config: ResolvedConfig, sid: string, mode: Mode): Promise<string> {
220 const from = currentMode(config, sid);
221 let saved: "config" | "store" | null = null;
222 let unavailable: string | null = null;
223
224 try {
225 const res = await $.config.set({ key: "stepwarden.mode", value: mode });
226 if (res && typeof res.deny === "string") {
227 await audit($, config, sid, { type: "mode_change_refused", from, to: mode, detail: res.deny });
228 return (
229 ` Mode is still ${from}: your settings refused the change (${String(res.deny).slice(0, 120)}).\n` +
230 " Whoever manages those settings owns this one."
231 );
232 }
233 saved = "config";
234 } catch (err) {
235 unavailable = (err as Error)?.message ?? String(err);
236 }
237
238 try {
239 if (saved === "config" || mode === config.mode) {
240 // Nothing to shadow: the stored configuration already says this.
241 await $.store.delete(MODE_KEY);
242 } else {
243 await $.store.set(MODE_KEY, { mode, basedOn: config.mode });
244 saved = "store";
245 }
246 } catch (err) {
247 if (saved === null) {
248 const why = (err as Error)?.message ?? String(err);
249 await audit($, config, sid, { type: "mode_change_refused", from, to: mode, detail: why });
250 return (
251 ` Mode is still ${from}: the switch could not be saved (${why.slice(0, 120)}).\n` +
252 " Set it in /plugin configure stepwarden instead."
253 );
254 }
255 }
256
257 modeOverride.set(sid, mode);
258 await audit($, config, sid, { type: "mode_changed", from, to: mode, saved, detail: unavailable });
259 if (mode === from) return ` Mode is ${mode} — ${describeMode(mode)}. Nothing changed.`;
260 const head = ` Mode is now ${mode} — ${describeMode(mode)} (was ${from}).`;
261 if (saved === "store" && mode !== config.mode) {
262 return (
263 `${head}\n Remembered for your next sessions too. /plugin configure stepwarden still says ` +
264 `${config.mode} — change it there and that wins.`
265 );
266 }
267 return `${head}\n Saved: new sessions start in ${mode} too. Run /stepwarden for the whole policy.`;
268}
269
270/**
271 * Claude Code dispatches tool calls in concurrent batches, and `$.store` has no
272 * compare-and-set, so a plain read-modify-write loses updates. Every store
273 * mutation goes through this queue instead: counters stay exact and no history
274 * entry is dropped. It only orders this plugin's own writes, so it cannot
275 * deadlock anything else.
276 */
277let storeQueue: Promise<unknown> = Promise.resolve();
278
279function serialize<T>(work: () => Promise<T>): Promise<T> {
280 const run = storeQueue.then(work, work);
281 storeQueue = run.then(
282 () => undefined,
283 () => undefined
284 );
285 return run;
286}
287
288/** Reachability + credential check. Returns a line fit to show a human. */
289async function probeKey(
290 $: any,
291 key: string,
292 timeoutMs: number,
293 signal: AbortSignal | undefined
294): Promise<{ ok: boolean; line: string }> {
295 const stop = new AbortController();
296 try {
297 const probe = await Promise.race([
298 $.http.fetch(API_BASE + MODELS_PATH, {
299 method: "GET",
300 headers: { authorization: "Bearer " + key, accept: "application/json" },
301 }).finally(() => stop.abort()),
302 $.clock.sleep(timeoutMs, { signal: anySignal(stop.signal, signal) }).then(() => null),
303 ]);
304 if (probe === null) return { ok: false, line: `no answer from TypeSafe within ${timeoutMs}ms` };
305 const r = probe as { ok: boolean; status: number; text: string };
306 return r.ok ? { ok: true, line: "key accepted" } : { ok: false, line: describeHttpFailure(r.status, r.text) };
307 } catch (err) {
308 return { ok: false, line: `check failed — ${(err as Error).message}` };
309 } finally {
310 stop.abort();
311 }
312}
313
314/** One signal that fires when either input does; used to cancel a timeout race. */
315function anySignal(a: AbortSignal, b: AbortSignal | undefined): AbortSignal {
316 if (!b) return a;
317 const out = new AbortController();
318 const fire = () => out.abort();
319 if (a.aborted || b.aborted) out.abort();
320 else {
321 a.addEventListener("abort", fire, { once: true });
322 b.addEventListener("abort", fire, { once: true });
323 }
324 return out.signal;
325}
326
327/**
328 * One verification round trip. Throws with a message written for a human; the
329 * caller turns that into the configured onError action and counts it against
330 * gate health.
331 */
332async function askJev(
333 $: any,
334 config: ResolvedConfig,
335 apiKey: string,
336 plan: string | null,
337 history: HistoryEntry[],
338 tool: string,
339 args: Record<string, unknown>,
340 signal: AbortSignal | undefined
341): Promise<Verdict> {
342 const body = JSON.stringify(
343 buildRequest({ plan, recentCalls: history, current: { tool, args }, model: config.model })
344 );
345
346 const timedOut = Symbol("timeout");
347 const stop = new AbortController();
348 let response: unknown;
349 try {
350 response = await Promise.race([
351 $.http.fetch(API_BASE + SYSTEMONE_PATH, {
352 method: "POST",
353 headers: {
354 authorization: "Bearer " + apiKey,
355 "content-type": "application/json",
356 accept: "application/json",
357 },
358 body,
359 }).finally(() => stop.abort()),
360 // $.http.fetch takes no abort signal of its own, so the request is raced
361 // against the clock. The timer is cancelled as soon as either the fetch
362 // settles or the dispatch is abandoned, so no timer outlives the call.
363 $.clock.sleep(config.timeoutMs, { signal: anySignal(stop.signal, signal) }).then(() => timedOut),
364 ]);
365 } finally {
366 stop.abort();
367 }
368
369 if (response === timedOut) throw new Error(`Jev did not answer within ${config.timeoutMs}ms`);
370
371 const res = response as { ok: boolean; status: number; text: string };
372 if (!res.ok) throw new Error(describeHttpFailure(res.status, res.text));
373
374 let parsed: unknown;
375 try {
376 parsed = JSON.parse(res.text);
377 } catch {
378 throw new Error("TypeSafe returned a body that is not JSON");
379 }
380 return parseResponse(parsed);
381}
382
383/* --------------------------------------------------------------- the plugin */
384
385export const register: Register = (on, options) => {
386 // `options` holds this plugin's userConfig values, already type-checked by
387 // the engine (sensitive ones come from secure storage). Normalising once is
388 // safe: a change through /plugin configure reloads the module.
389 const config: ResolvedConfig = resolveConfig(options);
390 const raw = options as RawOptions;
391 // A fresh load starts from the stored config, never from the last session's
392 // /stepwarden switch.
393 modeOverride.clear();
394
395 // ------------------------------------------------------------ session.start
396
397 on("session.start", async ($: any, e: any, next: any) => {
398 try {
399 const sid = await sessionId($);
400 await $.store.set(sessionKey(sid, "interactive"), e?.isInteractive === true);
401
402 // Record this session and forget long-dead ones. Pruning by timestamp
403 // (rather than "delete everything that is not me") keeps concurrent
404 // sessions from deleting each other's plan mid-run.
405 try {
406 const now = await $.clock.now();
407 const seenRaw = await $.store.get(LAST_SEEN_KEY);
408 const seen: Record<string, unknown> =
409 seenRaw && typeof seenRaw === "object" ? { ...(seenRaw as Record<string, unknown>) } : {};
410 seen[sid] = now;
411 const allKeys = (await $.store.keys()) as string[];
412
413 // A session seen for the first time is recorded now and pruned only on
414 // a later start. Deleting it immediately would race a session that is
415 // starting concurrently and has not written its own timestamp yet.
416 for (const key of allKeys) {
417 const owner = parseKey(key);
418 if (owner && seen[owner.sessionId] === undefined) seen[owner.sessionId] = now;
419 }
420
421 const dead = staleSessions(allKeys, seen, now, SESSION_TTL_MS, sid);
422 for (const key of keysOf(allKeys, dead)) await $.store.delete(key);
423 for (const id of dead) delete seen[id];
424 await $.store.set(LAST_SEEN_KEY, seen);
425 } catch {
426 /* pruning is housekeeping; never let it cost the session */
427 }
428
429 // /stepwarden is registered here, not in engine.create: the command noun is not
430 // available that early, and a failure there would fail the whole load.
431 try {
432 await $.command.register({
433 name: "stepwarden",
434 description: "stepwarden: status, help, or switch mode (enforce / audit / off)",
435 argumentHint: "[enforce|audit|off|toggle|help]",
436 // Switching the gate on or off is most useful mid-turn — exactly when
437 // waiting for the turn to end would defeat the point.
438 immediate: true,
439 });
440 } catch {
441 /* no command surface here; the gate itself still works */
442 }
443
444 // A switch made with /stepwarden in an earlier session, unless the stored
445 // configuration has changed since — then the dialog wins.
446 try {
447 const stored = await $.store.get(MODE_KEY);
448 const remembered = overrideMode(stored ?? null, config.mode);
449 if (remembered !== null) modeOverride.set(sid, remembered);
450 else if (stored !== undefined && stored !== null) await $.store.delete(MODE_KEY);
451 } catch {
452 /* without it, the session simply starts from the stored config */
453 }
454
455 for (const w of config.warnings) $.ui.log(`[stepwarden] config: ${w}`);
456
457 if (currentMode(config, sid) === "off") {
458 $.ui.log("[stepwarden] mode is 'off' — no tool call is verified. Turn it back on with /stepwarden audit.");
459 return next(e);
460 }
461
462 const { key, source } = await readKey($, raw);
463 if (!key) {
464 // The most common setup failure by far. Say exactly what to do, and
465 // say plainly that nothing is being verified meanwhile.
466 $.ui.toast("stepwarden: no API key — nothing is being verified. Run /stepwarden.", { timeoutMs: 12000 });
467 $.ui.log(
468 "[stepwarden] No TypeSafe API key found. Set one with:\n" +
469 " /plugin configure stepwarden (stored in your OS keychain)\n" +
470 " or export TYPESAFE_API_KEY=... before starting Claude Code.\n" +
471 " Get a key at https://typesafe.ai — run /stepwarden any time to re-check."
472 );
473 await audit($, config, sid, { type: "no_api_key" });
474 return next(e);
475 }
476
477 // A cheap authenticated GET turns "your key is wrong" into a message at
478 // startup rather than a surprise on the first tool call.
479 const probe = await probeKey($, key, Math.min(config.timeoutMs, 3000), next.signal);
480 if (probe.ok) {
481 $.ui.log(
482 `[stepwarden] ready — mode=${currentMode(config, sid)}, model=${config.model}, key from ${source}; ` +
483 `block at or above ${config.denyAbove}, ask at or above ${config.flagAbove}.`
484 );
485 } else {
486 $.ui.toast(`stepwarden: ${probe.line}`, { timeoutMs: 12000 });
487 $.ui.log(`[stepwarden] ${probe.line}`);
488 await audit($, config, sid, { type: "key_check_failed", detail: probe.line });
489 }
490
491 if (currentMode(config, sid) === "audit") {
492 $.ui.log(
493 "[stepwarden] AUDIT mode: decisions are computed and logged, but nothing is ever blocked. " +
494 "Turn the gate on with /stepwarden enforce once the numbers look right."
495 );
496 }
497 } catch {
498 /* session.start must never break the session */
499 }
500 return next(e);
501 });
502
503 // ----------------------------------------------------------- prompt.submit
504
505 on("prompt.submit", async ($: any, e: any, next: any) => {
506 try {
507 const text = typeof e?.text === "string" ? e.text.trim() : "";
508 if (text.length > 0) {
509 const sid = await sessionId($);
510 const existing = await $.store.get(sessionKey(sid, "plan"));
511 if (typeof existing !== "string" || existing.length === 0) {
512 // Stored verbatim, never summarised: a summary could quietly
513 // misrepresent the intent that everything else is checked against.
514 await $.store.set(sessionKey(sid, "plan"), text);
515 $.ui.log(`[stepwarden] plan captured: "${text.slice(0, 120)}${text.length > 120 ? "…" : ""}"`);
516 }
517 }
518 } catch {
519 /* plan capture is best-effort */
520 }
521 return next(e);
522 });
523
524 // --------------------------------------------------------------- tool.call
525
526 on("tool.call", async ($: any, e: any, next: any) => {
527 try {
528 const sid = await sessionId($);
529 if (currentMode(config, sid) === "off") return next(e);
530
531 const { tool, toolUseId, args } = splitEvent(e ?? {});
532
533 if (config.skipTools.includes(tool)) {
534 await bumpStats($, sid, { skipped: 1 });
535 return next(e);
536 }
537
538 const { key } = await readKey($, raw);
539 if (!key) {
540 // The gate cannot run. `onError` decides, and it is recorded, so an
541 // unconfigured gate is never mistaken for a clean verdict.
542 await bumpStats($, sid, { errors: 1 });
543 await audit($, config, sid, {
544 type: "verification",
545 tool,
546 args: safeArgs(args),
547 failure: "no API key configured",
548 shadowAction: config.onError,
549 effectiveAction: currentMode(config, sid) === "enforce" ? config.onError : "allow",
550 });
551 if (currentMode(config, sid) === "enforce" && config.onError === "deny") {
552 return { deny: "stepwarden: no TypeSafe API key is configured and the failure policy is 'deny'. Run /stepwarden." };
553 }
554 return next(e);
555 }
556
557 const scrubbed = safeArgs(args);
558 let plan: string | null = null;
559 let history: HistoryEntry[] = [];
560 try {
561 const p = await $.store.get(sessionKey(sid, "plan"));
562 plan = typeof p === "string" ? p : null;
563 const h = await $.store.get(sessionKey(sid, "history"));
564 history = Array.isArray(h) ? (h as HistoryEntry[]) : [];
565 } catch {
566 /* an empty plan/history is a valid, lower-signal state */
567 }
568
569 let verdict: Verdict | null = null;
570 let failure: string | null = null;
571 try {
572 verdict = await askJev($, config, key, plan, history.slice(-HISTORY_KEEP), tool, scrubbed, next.signal);
573 } catch (err) {
574 failure = (err as Error).message;
575 }
576
577 // ---- gate health: is the gate itself still working?
578 // Every store touch in this hook is guarded. A throw here would propagate
579 // out of the hook, and a failed hook is skipped — so the tool would run
580 // unverified, with no deny, no prompt and no audit line: the gate would
581 // fail open at exactly the moment it is reporting that it is unhealthy.
582 const health = await serialize(async () => {
583 try {
584 let failures = 0;
585 try {
586 const rawFailures = await $.store.get(sessionKey(sid, "failures"));
587 failures = typeof rawFailures === "number" ? rawFailures : 0;
588 } catch {
589 failures = 0;
590 }
591 if (failure === null) {
592 if (failures > 0) await $.store.set(sessionKey(sid, "failures"), 0);
593 return { recovered: failures > 0, consecutive: 0 };
594 }
595 const now = failures + 1;
596 await $.store.set(sessionKey(sid, "failures"), now);
597 return { recovered: false, consecutive: now };
598 } catch {
599 // Health tracking is observability, never a reason to drop the gate.
600 return { recovered: false, consecutive: 0 };
601 }
602 });
603
604 if (failure === null) {
605 if (health.recovered) {
606 $.ui.log("[stepwarden] verification gate recovered.");
607 await audit($, config, sid, { type: "gate_recovered" });
608 }
609 } else {
610 const now = health.consecutive;
611 if (now === config.unhealthyAfter) {
612 // Fire once at the threshold, not on every later call: an outage must
613 // be visible without spamming every tool call.
614 $.ui.toast(
615 `stepwarden: ${now} verification failures in a row — calls are proceeding unverified (${config.onError}).`,
616 { timeoutMs: 12000 }
617 );
618 await audit($, config, sid, { type: "gate_unhealthy", consecutiveFailures: now, reason: failure });
619 }
620 }
621
622 // ---- decide
623 let decision =
624 verdict !== null
625 ? decide(verdict, config)
626 : { action: config.onError, reasons: [`verification failed: ${failure}`], maxProbability: 0 };
627
628 const shadowAction: Action = decision.action;
629 let effective: Action = shadowAction;
630 let askOutcome: string | null = null;
631
632 if (currentMode(config, sid) === "audit") {
633 effective = "allow"; // compute and record, change nothing
634 } else if (shadowAction === "flag") {
635 let interactive = false;
636 try {
637 interactive = (await $.store.get(sessionKey(sid, "interactive"))) === true;
638 } catch {
639 interactive = false;
640 }
641
642 if (!interactive) {
643 effective = config.onNoAnswer;
644 askOutcome = "non-interactive session";
645 // Say it. A flagged call nobody could be asked about must never read
646 // as a clean verdict. A headless run has no transcript of its own:
647 // the line reaches the SDK host as `ui_log` and the debug log, and
648 // the audit log carries `askOutcome` either way.
649 $.ui.toast(
650 `stepwarden: ${tool} was flagged and nobody could be asked — ` +
651 `${effective === "allow" ? "allowed" : "blocked"} under your 'if nobody answers' setting.`,
652 { timeoutMs: 10000 }
653 );
654 $.ui.log(
655 `[stepwarden] ${tool} was flagged and this run has nobody to ask, so 'if nobody answers' decided: ${effective}. ` +
656 "Run it in an interactive session to be asked, or change onNoAnswer with /plugin configure stepwarden."
657 );
658 } else {
659 try {
660 const answer = await $.ui.ask(
661 `${tool}: ${decision.reasons.join("; ")}. Allow this call?`,
662 { header: "stepwarden", options: ["Allow", "Block"] }
663 );
664 await bumpStats($, sid, { asked: 1 });
665 if (answer === "Allow") {
666 effective = "allow";
667 askOutcome = "allowed by a human";
668 } else {
669 effective = "deny";
670 askOutcome = "blocked by a human";
671 decision = { ...decision, reasons: [...decision.reasons, "a human chose to block it"] };
672 }
673 } catch (err) {
674 // Dismissed, rate-limited (one question per 2s per plugin), or past
675 // the session's question budget. None of those is an approval.
676 const why = (err as Error).message;
677 effective = config.onNoAnswer;
678 askOutcome = `no answer: ${why.slice(0, 120)}`;
679 // Say so. A flagged call quietly becoming an allow because the dialog
680 // was unavailable is exactly the failure this plugin exists to avoid.
681 if (effective === "allow") {
682 $.ui.toast(
683 `stepwarden: could not ask about this ${tool} call (${why.slice(0, 60)}) — it was allowed under your 'if nobody answers' setting.`,
684 { timeoutMs: 10000 }
685 );
686 }
687 }
688 }
689 }
690
691 await bumpStats($, sid, {
692 checked: 1,
693 allowed: effective === "allow" ? 1 : 0,
694 flagged: shadowAction === "flag" ? 1 : 0,
695 denied: effective === "deny" ? 1 : 0,
696 errors: failure === null ? 0 : 1,
697 inputTokens: verdict?.usage?.inputTokens ?? 0,
698 outputTokens: verdict?.usage?.outputTokens ?? 0,
699 });
700
701 await audit($, config, sid, {
702 type: "verification",
703 tool,
704 tool_use_id: toolUseId,
705 args: scrubbed,
706 shadowAction,
707 effectiveAction: effective,
708 askOutcome,
709 reasons: decision.reasons,
710 maxProbability: decision.maxProbability,
711 signals: verdict?.signals ?? null,
712 intent: verdict?.intent ?? null,
713 model: verdict?.model ?? null,
714 usage: verdict?.usage ?? null,
715 failure,
716 });
717
718 if (effective === "deny") return { deny: denyMessage(tool, decision, askOutcome) };
719
720 await serialize(async () => {
721 try {
722 const at = await $.clock.now();
723 const entry: HistoryEntry = { tool, args: scrubbed, at };
724 // Re-read inside the critical section: a concurrent call in the same
725 // batch may have appended since this hook read `history` above.
726 const current = await $.store.get(sessionKey(sid, "history"));
727 const base = Array.isArray(current) ? (current as HistoryEntry[]) : [];
728 await $.store.set(sessionKey(sid, "history"), [...base, entry].slice(-HISTORY_KEEP));
729 } catch {
730 /* history is an optimisation, not a correctness requirement */
731 }
732 });
733
734 return next(e);
735 } catch (err) {
736 // The gate itself crashed — a display call, a store write, a shape the
737 // engine changed. A hook that throws is skipped and the tool simply
738 // runs, so without this the gate would fail open at the one moment it
739 // most needs to be visible. Make it a configured, recorded outcome.
740 const why = (err as Error)?.message ?? String(err);
741 try {
742 $.ui.toast(`stepwarden crashed while checking a call: ${why.slice(0, 80)}`, { timeoutMs: 12000 });
743 } catch {
744 /* even the toast is best-effort here */
745 }
746 let sid = "unknown-session";
747 try {
748 sid = await sessionId($);
749 await audit($, config, sid, { type: "gate_crashed", tool: e?.tool, failure: why });
750 } catch {
751 /* nothing more we can do */
752 }
753 if (currentMode(config, sid) === "enforce" && config.onError === "deny") {
754 return { deny: `Blocked by stepwarden: the verifier itself failed (${why.slice(0, 120)}) and the failure policy is 'deny'.` };
755 }
756 return next(e);
757 }
758 });
759
760 // ------------------------------------------------------------ /stepwarden command
761
762 on("command.run", { command: "stepwarden" }, async ($: any, e: any, next: any) => {
763 const sid = await sessionId($);
764
765 // Bare, it reports. `help` explains the whole surface, and a mode switches
766 // the gate.
767 const asked = parseCommandArg(typeof e?.args === "string" ? e.args : "", currentMode(config, sid));
768 if (asked.kind === "help") return { text: helpText() };
769 if (asked.kind === "error") return { text: ` ${asked.message}` };
770 if (asked.kind === "mode") return { text: await applyMode($, config, sid, asked.mode) };
771
772 const { key, source } = await readKey($, raw);
773 const stats = await getStats($, sid);
774 let plan: unknown = null;
775 try {
776 plan = await $.store.get(sessionKey(sid, "plan"));
777 } catch {
778 plan = null;
779 }
780
781 // The host already draws the plugin's name above this output, so a leading
782 // blank line printed a bare "stepwarden:" row with nothing after it.
783 const lines: string[] = [];
784
785 if (!key) {
786 lines.push(
787 " API key: NOT SET — nothing is being verified.",
788 "",
789 " Set one, either way:",
790 " 1. /plugin configure stepwarden (stored in your OS keychain)",
791 " 2. export TYPESAFE_API_KEY=... (then restart Claude Code)",
792 "",
793 " Get a key at https://typesafe.ai"
794 );
795 } else {
796 const masked = key.length > 12 ? `${key.slice(0, 7)}…${key.slice(-4)}` : "(set)";
797 const probe = await probeKey($, key, Math.min(config.timeoutMs, 3000), next.signal);
798 lines.push(` API key: ${masked} from ${source}`, ` TypeSafe: ${probe.line}`);
799 }
800
801 const mode = currentMode(config, sid);
802 lines.push(
803 "",
804 ` Mode: ${mode} (${describeMode(mode)})` +
805 `${modeOverride.has(sid) ? ` — set with /stepwarden; the dialog says ${config.mode}` : ""}`,
806 ` Model: ${config.model}`,
807 ` Policy: block at or above ${config.denyAbove} · ask at or above ${config.flagAbove} · ask at intent ≤ ${config.lowIntentAtOrBelow} of 4`,
808 ` Fallbacks: on error ${config.onError} · on no answer ${config.onNoAnswer}`,
809 ` Skipping: ${config.skipTools.length > 0 ? config.skipTools.join(", ") : "(nothing)"}`,
810 "",
811 ` Plan: ${typeof plan === "string" && plan.length > 0 ? `"${plan.slice(0, 100)}${plan.length > 100 ? "…" : ""}"` : "(none captured yet)"}`,
812 "",
813 ` Session: ${stats.checked} verified · ${stats.allowed} allowed · ${stats.flagged} flagged · ` +
814 `${stats.denied} blocked · ${stats.skipped} skipped · ${stats.errors} errors · ${stats.asked} asked`,
815 ` Jev usage: ${stats.inputTokens} input tokens, ${stats.outputTokens} output tokens`
816 );
817
818 if (config.warnings.length > 0) {
819 lines.push("", " Config warnings:");
820 for (const w of config.warnings) lines.push(` - ${w}`);
821 }
822 lines.push("", " /stepwarden help for the settings, the commands and how to set your key.");
823
824 if (config.auditLog) {
825 // The real path, so the line can be pasted into a shell. `.claude/stepwarden`
826 // is the script's own default, read from the project you run it in.
827 // Spelled `$.plugin.root` exactly: the static scanner refuses any other
828 // shape, including optional chaining.
829 let root = "<path-to-stepwarden>";
830 try {
831 const here = $.plugin.root;
832 if (typeof here === "string" && here.length > 0) root = here;
833 } catch {
834 /* an older host without the field: the README still has the path */
835 }
836 lines.push(
837 "",
838 ` Audit log: ${AUDIT_DIR}/${sid}.jsonl`,
839 ` analyse it from this project: node ${root}/scripts/analyze-audit.mjs`
840 );
841 }
842
843 return { text: lines.join("\n") };
844 });
845};
846lib/config.ts 289 lines1/**
2 * Turns the raw `options` object that `register(on, options)` receives into a
3 * validated ResolvedConfig.
4 *
5 * The values come from plugin.json's `userConfig` (settings.json
6 * `pluginConfigs[<plugin>].options`, sensitive ones from secure storage). The
7 * engine checks each value's *type* before the module loads, but nothing
8 * checks that a number is in range or that the thresholds are ordered — so
9 * that happens here.
10 *
11 * Every bad value is repaired to a safe default and recorded in `warnings`
12 * rather than thrown. A typo in one field must not take the gate down; but it
13 * also must not be silent, so hooks/verify.ts prints the warnings once at
14 * session start.
15 */
16
17import type { Action, Mode, ResolvedConfig } from "./types";
18
19/** Raw options as the engine hands them over. */
20export type RawOptions = Readonly<
21 Record<string, string | number | boolean | readonly string[] | undefined>
22>;
23
24export const DEFAULTS: ResolvedConfig = {
25 // Audit, not enforce: an installed plugin must not block anyone's work on
26 // thresholds that have never seen their traffic. It verifies and records
27 // every call from the first minute, and the user turns the gate on with
28 // /stepwarden enforce when the numbers look right.
29 mode: "audit",
30 model: "jev-latest",
31 // Jev's own published accuracy sits below frontier models in several
32 // domains, so the shipped posture is "deny rarely, ask often": only a
33 // near-certain signal blocks outright, and the middle band asks a human
34 // rather than deciding for them.
35 denyAbove: 0.9,
36 flagAbove: 0.6,
37 lowIntentAtOrBelow: 1,
38 onError: "allow",
39 onNoAnswer: "allow",
40 skipTools: ["Read", "Glob", "Grep", "TodoWrite", "NotebookRead", "ListAgents", "TaskOutput"],
41 timeoutMs: 5000,
42 unhealthyAfter: 3,
43 auditLog: true,
44 warnings: [],
45};
46
47/**
48 * "flag" is deliberately absent: these two settings say what to do when a human
49 * CANNOT be asked, so "ask a human" is not an answer. A stored "flag" from an
50 * older config falls back to the default and warns.
51 */
52const FALLBACK_ACTIONS: Action[] = ["allow", "deny"];
53const MODES: Mode[] = ["enforce", "audit", "off"];
54
55function asString(v: unknown): string | undefined {
56 return typeof v === "string" && v.trim().length > 0 ? v.trim() : undefined;
57}
58
59function oneOf<T extends string>(v: unknown, allowed: T[], field: string, warnings: string[]): T | undefined {
60 const s = asString(v);
61 if (s === undefined) return undefined;
62 const hit = allowed.find((a) => a.toLowerCase() === s.toLowerCase());
63 if (hit) return hit;
64 warnings.push(`${field}: "${s}" is not one of ${allowed.join(", ")} — using the default`);
65 return undefined;
66}
67
68function probability(v: unknown, field: string, warnings: string[]): number | undefined {
69 // Number("") and Number(" ") are both 0, which would silently set a
70 // threshold of zero and make the gate block everything.
71 if (v === undefined || v === null || typeof v === "boolean") return undefined;
72 if (typeof v !== "number" && String(v).trim() === "") return undefined;
73 const n = typeof v === "number" ? v : Number(String(v).trim());
74 if (!Number.isFinite(n)) {
75 warnings.push(`${field}: "${String(v)}" is not a number — using the default`);
76 return undefined;
77 }
78 if (n < 0 || n > 1) {
79 warnings.push(`${field}: ${n} is outside 0..1 — using the default`);
80 return undefined;
81 }
82 return n;
83}
84
85function integer(v: unknown, field: string, min: number, max: number, warnings: string[]): number | undefined {
86 if (v === undefined || v === null || typeof v === "boolean") return undefined;
87 if (typeof v !== "number" && String(v).trim() === "") return undefined;
88 const n = typeof v === "number" ? v : Number(String(v).trim());
89 if (!Number.isFinite(n)) {
90 warnings.push(`${field}: "${String(v)}" is not a number — using the default`);
91 return undefined;
92 }
93 const r = Math.round(n);
94 if (r < min || r > max) {
95 warnings.push(`${field}: ${n} is outside ${min}..${max} — using the default`);
96 return undefined;
97 }
98 return r;
99}
100
101function toBool(v: unknown, fallback: boolean): boolean {
102 if (typeof v === "boolean") return v;
103 if (typeof v === "string") {
104 const s = v.trim().toLowerCase();
105 if (s === "false" || s === "no" || s === "off" || s === "0") return false;
106 if (s === "true" || s === "yes" || s === "on" || s === "1") return true;
107 }
108 return fallback;
109}
110
111/** Accepts a real list, or a comma/whitespace-separated string. */
112function toolList(v: unknown): string[] | undefined {
113 if (Array.isArray(v)) {
114 const out = v.map((x) => String(x).trim()).filter((x) => x.length > 0);
115 return out.length > 0 ? out : [];
116 }
117 const s = asString(v);
118 if (s === undefined) return undefined;
119 if (s.toLowerCase() === "none") return [];
120 const out = s
121 .split(/[,\s]+/)
122 .map((x) => x.trim())
123 .filter((x) => x.length > 0);
124 return out.length > 0 ? out : [];
125}
126
127export function resolveConfig(options: RawOptions | undefined): ResolvedConfig {
128 const o = options ?? {};
129 const warnings: string[] = [];
130
131 const denyAbove = probability(o.denyAbove, "denyAbove", warnings) ?? DEFAULTS.denyAbove;
132 let flagAbove = probability(o.flagAbove, "flagAbove", warnings) ?? DEFAULTS.flagAbove;
133
134 // An inverted pair would make `flag` unreachable and quietly turn the middle
135 // band into a deny. Repair it loudly instead of honouring a typo.
136 if (flagAbove > denyAbove) {
137 // Setting flagAbove == denyAbove would leave an ask band of zero width, so
138 // the "repair" would still mean every flagged call is really a block.
139 // Leave a real band below the block threshold instead.
140 const repaired = Math.max(0, Math.round((denyAbove - 0.1) * 100) / 100);
141 warnings.push(
142 `flagAbove (${flagAbove}) is above denyAbove (${denyAbove}), which would leave no band in which you are asked — lowering flagAbove to ${repaired}`
143 );
144 flagAbove = repaired;
145 }
146
147 const skip = toolList(o.skipTools);
148
149 return {
150 mode: oneOf(o.mode, MODES, "mode", warnings) ?? DEFAULTS.mode,
151 model: asString(o.model) ?? DEFAULTS.model,
152 denyAbove,
153 flagAbove,
154 lowIntentAtOrBelow:
155 integer(o.lowIntentAtOrBelow, "lowIntentAtOrBelow", 0, 4, warnings) ?? DEFAULTS.lowIntentAtOrBelow,
156 onError: oneOf(o.onError, FALLBACK_ACTIONS, "onError", warnings) ?? DEFAULTS.onError,
157 onNoAnswer: oneOf(o.onNoAnswer, FALLBACK_ACTIONS, "onNoAnswer", warnings) ?? DEFAULTS.onNoAnswer,
158 skipTools: skip ?? DEFAULTS.skipTools,
159 timeoutMs: integer(o.timeoutMs, "timeoutMs", 500, 30000, warnings) ?? DEFAULTS.timeoutMs,
160 unhealthyAfter: integer(o.unhealthyAfter, "unhealthyAfter", 1, 100, warnings) ?? DEFAULTS.unhealthyAfter,
161 auditLog: toBool(o.auditLog, DEFAULTS.auditLog),
162 warnings,
163 };
164}
165
166/** The API key, wherever the user chose to put it. */
167export function resolveApiKey(
168 options: RawOptions | undefined,
169 envKey: string | undefined
170): { key: string | null; source: "plugin-config" | "environment" | "none" } {
171 const fromOptions = asString(options?.TYPESAFE_API_KEY);
172 if (fromOptions) return { key: fromOptions, source: "plugin-config" };
173 const fromEnv = asString(envKey);
174 if (fromEnv) return { key: fromEnv, source: "environment" };
175 return { key: null, source: "none" };
176}
177
178/** What a mode means, in the one line the command and the status print. */
179export function describeMode(mode: Mode): string {
180 if (mode === "enforce") return "risky calls are blocked or put to you";
181 if (mode === "audit") return "every decision is logged, nothing is blocked";
182 return "nothing is verified";
183}
184
185/** What `/stepwarden <arg>` was asking for. */
186export type CommandArg =
187 | { kind: "status" }
188 | { kind: "help" }
189 | { kind: "mode"; mode: Mode }
190 | { kind: "error"; message: string };
191
192/**
193 * Reads the argument of `/stepwarden`.
194 *
195 * No argument reports status. An unknown word is an error rather than a silent
196 * no-op: a typo'd switch that quietly did nothing would leave the gate in
197 * exactly the state they meant to leave.
198 *
199 * `toggle` goes to enforce from audit, and to audit from anywhere else —
200 * turning a gate that is off all the way up to blocking in one word is not
201 * something a toggle should do.
202 */
203export function parseCommandArg(raw: string, current: Mode): CommandArg {
204 const word = raw.trim().toLowerCase();
205 if (word.length === 0) return { kind: "status" };
206 if (word === "help" || word === "?") return { kind: "help" };
207 if (word === "toggle") return { kind: "mode", mode: current === "audit" ? "enforce" : "audit" };
208 const hit = MODES.find((m) => m === word);
209 if (hit) return { kind: "mode", mode: hit };
210 return {
211 kind: "error",
212 message:
213 `"${raw.trim()}" is not something /stepwarden takes. Try enforce, audit, off, toggle, ` +
214 "or help — or /stepwarden on its own for status.",
215 };
216}
217
218/**
219 * `/stepwarden help`: the whole surface on one screen.
220 *
221 * Deliberately not the same text as the status command. Status answers "what is
222 * happening right now"; this answers "what can I do, and how do I set the key",
223 * which is what someone types `help` for. Defaults come from DEFAULTS so the
224 * two cannot drift apart.
225 */
226export function helpText(): string {
227 // Built as pairs so the second column lines up whatever the defaults are.
228 const settings: Array<[string, string]> = [
229 [`mode [${DEFAULTS.mode}]`, "enforce, audit or off"],
230 [`denyAbove [${DEFAULTS.denyAbove}]`, "block at or above this probability"],
231 [`flagAbove [${DEFAULTS.flagAbove}]`, "ask you at or above this probability"],
232 [`lowIntentAtOrBelow [${DEFAULTS.lowIntentAtOrBelow}]`, "ask when intent consistency is this low, of 0-4"],
233 [`onError [${DEFAULTS.onError}]`, "when Jev is unreachable, times out, or has no key"],
234 [`onNoAnswer [${DEFAULTS.onNoAnswer}]`, "when a flagged call cannot be put to a human"],
235 ["skipTools [Read, Glob, Grep, ...]", 'never verified; "none" verifies everything'],
236 [`model [${DEFAULTS.model}]`, "which TypeSafe model answers"],
237 [`timeoutMs [${DEFAULTS.timeoutMs}]`, "how long to wait before onError applies"],
238 [`unhealthyAfter [${DEFAULTS.unhealthyAfter}]`, "warn after this many failures in a row"],
239 [`auditLog [${DEFAULTS.auditLog}]`, "write .claude/stepwarden/<session>.jsonl"],
240 ];
241 const width = Math.max(...settings.map(([name]) => name.length)) + 3;
242
243 return [
244 " stepwarden checks every tool call before it runs. This is how you drive it.",
245 "",
246 " Commands",
247 " /stepwarden key, policy, and what it decided this session",
248 " /stepwarden enforce block or ask on risky calls",
249 ` /stepwarden audit verify and log everything, block nothing (${DEFAULTS.mode} is the default)`,
250 " /stepwarden off verify nothing",
251 " /stepwarden toggle flip between audit and enforce",
252 " /stepwarden help this text",
253 " /plugin configure stepwarden every setting below, each with an explanation",
254 "",
255 " Your TypeSafe API key",
256 " Without one, nothing is verified — the plugin says so at startup rather",
257 " than looking like a working gate. Get a key at https://typesafe.ai, then:",
258 " 1. /plugin configure stepwarden stored in your OS keychain (recommended)",
259 " 2. export TYPESAFE_API_KEY=... then restart Claude Code",
260 " A key in the plugin config wins over the environment. Claude Code does not",
261 " read .env files: source one into your shell first (set -a; . ./.env; set +a).",
262 "",
263 " What you can configure (defaults in brackets)",
264 ...settings.map(([name, what]) => ` ${name.padEnd(width)}${what}`),
265 "",
266 " Use it in audit mode on real work, read the log with",
267 " scripts/analyze-audit.mjs (/stepwarden prints the command), then switch on",
268 " enforcement with /stepwarden enforce.",
269 ].join("\n");
270}
271
272/**
273 * A mode switch the user made with `/stepwarden`, as it comes back from the
274 * plugin's store — or `null` when there is none to honour.
275 *
276 * The switch records the configured mode it was made against. If the stored
277 * configuration has changed since, whoever changed it in
278 * `/plugin configure stepwarden` meant it, and a switch made against the old
279 * value is stale: the dialog wins, and the override is dropped.
280 */
281export function overrideMode(stored: unknown, configMode: Mode): Mode | null {
282 if (stored === null || typeof stored !== "object") return null;
283 const o = stored as { mode?: unknown; basedOn?: unknown };
284 const mode = MODES.find((m) => m === o.mode);
285 const basedOn = MODES.find((m) => m === o.basedOn);
286 if (mode === undefined || basedOn === undefined) return null;
287 return basedOn === configMode ? mode : null;
288}
289lib/jev.ts 182 lines1/**
2 * The TypeSafe "System One" (Jev) wire protocol, as pure functions.
3 *
4 * This deliberately does NOT use @typesafe-ai/sdk. A function-hooks module may
5 * import only relative paths and "claude-code" — an npm import makes the whole
6 * plugin fail to load — so the request is built here and sent by the caller
7 * through `$.http.fetch`.
8 *
9 * Wire format (verified against api.typesafe.ai, 2026-09-18):
10 * POST https://api.typesafe.ai/v1/systemone
11 * Authorization: Bearer <key>
12 * { model, state, questions }
13 * -> { model, answers: { <name>: NoulAnswer | ScoreAnswer }, usage }
14 *
15 * NoulAnswer = { type: "noul", noul: number } // 0..1
16 * ScoreAnswer = { type: "score", score: number, confidence: number,
17 * legend: {...}, probabilities: {...} }
18 *
19 * Two shapes here are easy to get wrong and are the reason the previous
20 * version never produced a verdict at all:
21 * - a noul answer is `.noul`, NOT `.probability`;
22 * - a score question takes a RUBRIC ARRAY (>= 2 entries), not a level count,
23 * and its answer is `.score`, indexed FROM ZERO — a 5-entry rubric scores
24 * 0..4, and the score may be fractional.
25 */
26
27import type { HistoryEntry, Verdict } from "./types";
28
29export const API_BASE = "https://api.typesafe.ai";
30export const SYSTEMONE_PATH = "/v1/systemone";
31export const MODELS_PATH = "/v1/models";
32
33/**
34 * 0 = contradicts intent … 4 = fully consistent. Zero-indexed, so the highest
35 * level is `INTENT_RUBRIC.length - 1`.
36 */
37export const INTENT_RUBRIC = [
38 "Directly contradicts the stated plan or the user's intent.",
39 "Mostly inconsistent with it; hard to justify from the plan.",
40 "Unclear or unrelated; the plan does not speak to this.",
41 "Mostly consistent; a reasonable step toward the plan.",
42 "Fully consistent; plainly a step the plan calls for.",
43] as const;
44
45export const SIGNAL_LABELS: Record<string, string> = {
46 contradictsPlan: "contradicts the stated plan",
47 outOfScope: "touches an out-of-scope resource",
48 contradictsPrevious: "contradicts the previous tool call",
49 shouldFlag: "warrants independent human review",
50};
51
52/** The plan is the user's own words, so it is capped but never pattern-redacted. */
53export const MAX_PLAN_CHARS = 2000;
54
55export interface VerifyContext {
56 plan: string | null;
57 recentCalls: HistoryEntry[];
58 current: { tool: string; args: Record<string, unknown> };
59 model: string;
60}
61
62/** Builds the JSON body for one verification. Pure. */
63export function buildRequest(ctx: VerifyContext): Record<string, unknown> {
64 return {
65 model: ctx.model,
66 state: {
67 plan: ctx.plan
68 ? ctx.plan.length > MAX_PLAN_CHARS
69 ? `${ctx.plan.slice(0, MAX_PLAN_CHARS)}…[truncated from ${ctx.plan.length} chars]`
70 : ctx.plan
71 : "(no plan was declared for this session — treat plan-contradiction answers as low-signal)",
72 recent_tool_calls: ctx.recentCalls.map((h) => ({ tool: h.tool, args: h.args })),
73 current_tool_call: { tool: ctx.current.tool, args: ctx.current.args },
74 },
75 questions: {
76 // `criteria` on a noul describes what yes and no mean. It is optional,
77 // but supplying it measurably sharpens a yes/no with a vague subject.
78 contradictsPlan: {
79 type: "noul",
80 instructions: "Does current_tool_call contradict, undo, or work against the stated plan?",
81 criteria: {
82 true: "It reverses, discards, or works against something the plan asks for.",
83 false: "It advances the plan, or is unrelated but harmless to it.",
84 },
85 },
86 outOfScope: {
87 type: "noul",
88 instructions:
89 "Does current_tool_call read or modify a file, resource, host, or system outside what the plan describes?",
90 criteria: {
91 true: "It reaches outside the plan's stated subject matter.",
92 false: "It stays within what the plan describes.",
93 },
94 },
95 contradictsPrevious: {
96 type: "noul",
97 instructions:
98 "Does current_tool_call contradict, reverse, or undo the most recent entry in recent_tool_calls?",
99 criteria: {
100 true: "It undoes or reverses what the previous call just did.",
101 false: "It builds on the previous call, or is independent of it.",
102 },
103 },
104 shouldFlag: {
105 type: "noul",
106 instructions:
107 "Independent of the other answers: is this destructive, irreversible, or otherwise something a human should review before it runs?",
108 criteria: {
109 true: "Destructive, irreversible, touches production or credentials, or is otherwise high-consequence.",
110 false: "Routine and reversible.",
111 },
112 },
113 intentConsistency: {
114 type: "score",
115 instructions: "How consistent is current_tool_call with the stated plan?",
116 criteria: INTENT_RUBRIC,
117 },
118 },
119 };
120}
121
122function num(v: unknown): number | null {
123 return typeof v === "number" && Number.isFinite(v) ? v : null;
124}
125
126/**
127 * Parses a systemOne response into a Verdict.
128 *
129 * Throws when the response is not shaped as expected, so that a protocol
130 * change surfaces as a gate failure (visible, counted, handled by `onError`)
131 * rather than as a verdict full of zeros that reads like "nothing is wrong".
132 */
133export function parseResponse(body: unknown): Verdict {
134 if (!body || typeof body !== "object") throw new Error("response was not a JSON object");
135 const answers = (body as Record<string, unknown>).answers;
136 if (!answers || typeof answers !== "object") throw new Error("response had no `answers`");
137
138 const a = answers as Record<string, Record<string, unknown>>;
139 const signals = Object.keys(SIGNAL_LABELS).map((key) => {
140 const p = num(a[key]?.noul);
141 if (p === null) throw new Error(`answer "${key}" had no numeric \`noul\` field`);
142 return { key, label: SIGNAL_LABELS[key] ?? key, probability: p };
143 });
144
145 // Treated exactly like a missing noul: a silently absent intent score would
146 // disable the low-intent trigger and look like a clean verdict.
147 if (a.intentConsistency === undefined) throw new Error('answer "intentConsistency" was missing');
148 const scoreRaw = num(a.intentConsistency.score);
149 if (scoreRaw === null) throw new Error('answer "intentConsistency" had no numeric `score` field');
150 const intent = {
151 score: scoreRaw,
152 max: INTENT_RUBRIC.length - 1,
153 confidence: num(a.intentConsistency.confidence) ?? 0,
154 };
155
156 const usageRaw = (body as Record<string, unknown>).usage as Record<string, unknown> | undefined;
157 const usage = usageRaw
158 ? {
159 inputTokens: num(usageRaw.input_tokens) ?? 0,
160 outputTokens: num(usageRaw.output_tokens) ?? 0,
161 }
162 : null;
163
164 const model = (body as Record<string, unknown>).model;
165
166 return { signals, intent, usage, model: typeof model === "string" ? model : null };
167}
168
169/** Classifies an HTTP failure into something a human can act on. */
170export function describeHttpFailure(status: number, text: string): string {
171 const snippet = text.trim().slice(0, 200);
172 if (status === 401 || status === 403) {
173 return "TypeSafe rejected the API key (HTTP " + status + "). Run /stepwarden to check how the key is being supplied, then set a valid one.";
174 }
175 if (status === 429) return "TypeSafe rate-limited this session (HTTP 429).";
176 if (status >= 500) return `TypeSafe had a server error (HTTP ${status}).`;
177 if (status === 400 || status === 422) {
178 return `TypeSafe rejected the request (HTTP ${status})${snippet ? ": " + snippet : ""}. This usually means the wire format changed — regenerate types and check lib/jev.ts.`;
179 }
180 return `TypeSafe returned HTTP ${status}${snippet ? ": " + snippet : ""}`;
181}
182lib/keys.ts 75 lines1/**
2 * Store keys.
3 *
4 * `$.store` is the plugin's own key-value store and is "kept between sessions
5 * and hot reloads". Everything this plugin keeps — the plan, the recent-call
6 * history, gate health, the counters — is meaningful only within one session,
7 * so every key is scoped by the session id. Without that, a new session
8 * inherits the previous session's plan and silently verifies today's work
9 * against yesterday's intent.
10 *
11 * Scoping alone would grow the store forever, so old sessions are pruned. The
12 * prune is deliberately based on a per-session timestamp rather than "delete
13 * everything that is not me": concurrent sessions are normal, and they must
14 * not delete each other's state.
15 *
16 * There is no migration path for keys written under a previous plugin name:
17 * the engine keys each plugin's store by its manifest name, so a rename starts
18 * from an empty store and the old file is orphaned whole, where this code can
19 * never reach it.
20 */
21
22export const PREFIX = "stepwarden";
23
24/** Suffixes, all session-scoped. `seen` is the prune timestamp. */
25export const SUFFIXES = ["plan", "history", "failures", "stats", "interactive", "seen"] as const;
26export type Suffix = (typeof SUFFIXES)[number];
27
28export function sessionKey(sessionId: string, suffix: Suffix): string {
29 return `${PREFIX}:${sessionId}:${suffix}`;
30}
31
32/** The inverse of sessionKey, for keys this plugin owns; null for anything else. */
33export function parseKey(key: string): { sessionId: string; suffix: string } | null {
34 const parts = key.split(":");
35 if (parts.length !== 3 || parts[0] !== PREFIX) return null;
36 const sessionId = parts[1];
37 const suffix = parts[2];
38 if (!sessionId || !suffix) return null;
39 return { sessionId, suffix };
40}
41
42/** Sessions to forget: last seen too long ago, or carrying no timestamp at all. */
43export function staleSessions(
44 keys: readonly string[],
45 lastSeen: Readonly<Record<string, unknown>>,
46 now: number,
47 maxAgeMs: number,
48 currentSessionId: string
49): string[] {
50 const sessions = new Set<string>();
51 for (const key of keys) {
52 const parsed = parseKey(key);
53 if (parsed) sessions.add(parsed.sessionId);
54 }
55 sessions.delete(currentSessionId);
56
57 const stale: string[] = [];
58 for (const id of sessions) {
59 const seen = lastSeen[id];
60 // A session with no timestamp is from an older version of this plugin, or
61 // its session.start never completed: either way it is safe to forget.
62 if (typeof seen !== "number" || now - seen > maxAgeMs) stale.push(id);
63 }
64 return stale.sort();
65}
66
67/** Every key belonging to the given sessions. */
68export function keysOf(keys: readonly string[], sessionIds: readonly string[]): string[] {
69 const ids = new Set(sessionIds);
70 return keys.filter((k) => {
71 const parsed = parseKey(k);
72 return parsed !== null && ids.has(parsed.sessionId);
73 });
74}
75lib/policy.ts 79 lines1/**
2 * Turns a Verdict into allow / flag / deny.
3 *
4 * Shape of the decision, and why:
5 * - `deny` only on a near-certain single signal. TypeSafe publishes Jev
6 * accuracy below frontier models in several domains, so a hard block on a
7 * merely-probable signal produces false denials exactly where a gate is
8 * most wanted.
9 * - `flag` is the wide middle band, and it is a real tier now: it asks a
10 * human via $.ui.ask. A flag that resolves itself is not a gate.
11 * - low intent-consistency is its own flag trigger, independent of the
12 * probabilities, because "this is not what you said you were doing" is a
13 * different failure from "this looks dangerous".
14 */
15
16import type { Decision, ResolvedConfig, Verdict } from "./types";
17
18export function decide(verdict: Verdict, config: ResolvedConfig): Decision {
19 const reasons: string[] = [];
20 let maxProbability = 0;
21
22 for (const s of verdict.signals) {
23 if (s.probability > maxProbability) maxProbability = s.probability;
24 if (s.probability >= config.flagAbove) {
25 reasons.push(`${s.label} (p=${s.probability.toFixed(2)})`);
26 }
27 }
28
29 // Scores are zero-indexed: 0 is "contradicts", max is "fully consistent".
30 const intent = verdict.intent;
31 const lowIntent = intent !== null && intent.score <= config.lowIntentAtOrBelow;
32 if (lowIntent) {
33 reasons.push(
34 `low consistency with the stated plan (${intent.score.toFixed(1)} of ${intent.max}, confidence ${intent.confidence.toFixed(2)})`
35 );
36 }
37
38 if (maxProbability >= config.denyAbove) {
39 return { action: "deny", reasons, maxProbability };
40 }
41 if (reasons.length > 0) {
42 return { action: "flag", reasons, maxProbability };
43 }
44 return { action: "allow", reasons: [], maxProbability };
45}
46
47/** One line summarizing a decision, for the transcript and the audit log. */
48export function summarize(tool: string, decision: Decision): string {
49 if (decision.reasons.length === 0) {
50 return `${tool}: no signal above threshold (max p=${decision.maxProbability.toFixed(2)})`;
51 }
52 return `${tool}: ${decision.reasons.join("; ")}`;
53}
54
55/**
56 * What the agent is told when a call is blocked.
57 *
58 * Two different things block a call, and they are fixed with two different
59 * settings. A block that came from nobody being able to answer must not send
60 * the user to the thresholds: the thresholds were right, the question just
61 * never reached a human. A human's own Block keeps the threshold wording —
62 * there, the thresholds are why they were asked at all.
63 */
64export function denyMessage(tool: string, decision: Decision, askOutcome: string | null): string {
65 const nobodyAsked =
66 askOutcome === "non-interactive session" || (askOutcome !== null && askOutcome.startsWith("no answer:"));
67 if (nobodyAsked) {
68 return (
69 `Blocked by stepwarden: ${summarize(tool, decision)} — a human had to decide and nobody could be asked ` +
70 `(${askOutcome}). Run this in an interactive session to be asked, or set 'if nobody answers' to allow ` +
71 "with /plugin configure stepwarden."
72 );
73 }
74 return (
75 `Blocked by stepwarden: ${summarize(tool, decision)}. ` +
76 "If that is wrong, adjust the thresholds with /plugin configure stepwarden, or run /stepwarden to see the current policy."
77 );
78}
79lib/redact.ts 168 lines1/**
2 * Redaction + size-capping for anything that leaves the session.
3 *
4 * Two destinations make this necessary, and neither is obvious from the call
5 * site: every verified tool call is (a) sent to TypeSafe's API and (b) written
6 * to a local JSONL audit log. A `Bash` call can carry `AWS_SECRET=...` and a
7 * `Write` call can carry an entire file body, so raw tool arguments must not
8 * go out unfiltered — and a 2 MB file body would blow both the request latency
9 * and the log.
10 *
11 * This is deliberately conservative pattern matching, not a secret scanner: it
12 * catches the common shapes and truncates everything else. It reduces
13 * exposure; it does not eliminate it, which is why the README says plainly
14 * that tool arguments are sent to a third party.
15 */
16
17const MAX_STRING = 600;
18const MAX_ARRAY = 20;
19const MAX_DEPTH = 4;
20
21/**
22 * Terms that make a key's value a secret wherever they appear in the name.
23 * These essentially never occur in a benign tool argument.
24 */
25const SECRET_SUBSTRINGS = [
26 "secret",
27 "password",
28 "passwd",
29 "credential",
30 "apikey",
31 "api_key",
32 "api-key",
33 "private_key",
34 "privatekey",
35 "access_token",
36 "refresh_token",
37 "auth_token",
38 "bearer",
39];
40
41/**
42 * Terms that are secrets only as the WHOLE key. Matching these as substrings
43 * redacted ordinary arguments — `max_tokens` contains "token", `author`
44 * contains "auth" — which both destroys the signal Jev needs and makes the
45 * audit log unreadable.
46 */
47const SECRET_EXACT = new Set([
48 "token",
49 "auth",
50 "authorization",
51 "key",
52 "cookie",
53 "session_id",
54 "sessionid",
55 "pass",
56 "credentials",
57]);
58
59function isSecretKey(name: string): boolean {
60 const k = name.trim().toLowerCase();
61 if (SECRET_EXACT.has(k)) return true;
62 return SECRET_SUBSTRINGS.some((term) => k.includes(term));
63}
64
65/** Value shapes that look like credentials wherever they appear. */
66const SECRET_VALUE: RegExp[] = [
67 // KEY=value / KEY: value for a secret-ish name — a shell env assignment, and
68 // the quoted "api_key": "..." form that appears in JSON bodies and configs.
69 /(?:"|')?\b([A-Za-z0-9_]*(?:PASS(?:WD|WORD)?|SECRET|TOKEN|API[-_]?KEY|CREDENTIAL|PRIVATE[-_]?KEY)[A-Za-z0-9_]*)\b(?:"|')?\s*[=:]\s*("[^"]*"|'[^']*'|[^\s,;)}\]]+)/gi,
70 // Common vendor key prefixes
71 /\b((?:sk|pk|rk|apikey|ghp|gho|ghu|ghs|ghr|xox[baprs]|AKIA|ASIA)[-_][A-Za-z0-9_-]{8,})/g,
72 /\b(AKIA[0-9A-Z]{16})\b/g,
73 // Authorization headers
74 /\b(Bearer|Basic)\s+([A-Za-z0-9._~+/=-]{12,})/gi,
75 // PEM blocks
76 /-----BEGIN[A-Z ]*PRIVATE KEY-----[\s\S]*?-----END[A-Z ]*PRIVATE KEY-----/g,
77 // JWTs
78 /\beyJ[A-Za-z0-9_-]{8,}\.[A-Za-z0-9_-]{8,}\.[A-Za-z0-9_-]{8,}\b/g,
79];
80
81export const REDACTED = "«redacted»";
82
83/** Replaces credential-looking substrings inside a free-text value. */
84export function redactString(input: string): string {
85 let out = input;
86 for (const re of SECRET_VALUE) {
87 out = out.replace(re, (match: string, ...rest: unknown[]) => {
88 const groups = rest.filter((g) => typeof g === "string") as string[];
89 const name = groups[0];
90 // Keep the name and the real separator so the shape stays readable (and
91 // so a shell command is still parseable by Jev); replace only the value.
92 if (name !== undefined) {
93 const at = match.indexOf(name);
94 const after = match.slice(at + name.length);
95 const sep = after.match(/^["']?\s*[=:]\s*|^\s+/)?.[0];
96 if (sep !== undefined) return `${match.slice(0, at + name.length)}${sep}${REDACTED}`;
97 }
98 return REDACTED;
99 });
100 }
101 return out;
102}
103
104function truncate(s: string): string {
105 return s.length <= MAX_STRING ? s : `${s.slice(0, MAX_STRING)}…[${s.length} chars]`;
106}
107
108function scrub(value: unknown, depth: number, keyHint?: string): unknown {
109 if (keyHint && isSecretKey(keyHint)) return REDACTED;
110 if (value === null || value === undefined) return value ?? null;
111
112 const t = typeof value;
113 if (t === "string") return truncate(redactString(value as string));
114 if (t === "number" || t === "boolean") return value;
115 if (t !== "object") return String(value);
116
117 if (depth >= MAX_DEPTH) return "«depth»";
118
119 if (Array.isArray(value)) {
120 const head = value.slice(0, MAX_ARRAY).map((v) => scrub(v, depth + 1));
121 return value.length > MAX_ARRAY ? [...head, `…[${value.length} items]`] : head;
122 }
123
124 // A null-prototype object so an argument named __proto__ is kept as data
125 // rather than silently reassigning the result's prototype.
126 const out = Object.create(null) as Record<string, unknown>;
127 for (const [k, v] of Object.entries(value as Record<string, unknown>)) {
128 out[k] = scrub(v, depth + 1, k);
129 }
130 return { ...out };
131}
132
133/**
134 * The tool arguments as they may safely be sent to Jev and written to the
135 * audit log: secrets masked, long values truncated, deep structures clipped.
136 */
137export function safeArgs(args: Record<string, unknown>): Record<string, unknown> {
138 const scrubbed = scrub(args, 0);
139 return scrubbed && typeof scrubbed === "object" && !Array.isArray(scrubbed)
140 ? (scrubbed as Record<string, unknown>)
141 : {};
142}
143
144/**
145 * Splits a raw `tool.call` event into its identity and its arguments.
146 *
147 * A tool.call event is `{ tool, tool_use_id, ...the tool's own arguments }` —
148 * the arguments are spread at the top level, there is no `input` wrapper.
149 * These reserved keys are the engine's, not the tool's.
150 */
151const RESERVED = new Set(["tool", "tool_use_id", "agentId", "parentAgentId", "$shadowed"]);
152
153export function splitEvent(e: Record<string, unknown>): {
154 tool: string;
155 toolUseId: string | undefined;
156 args: Record<string, unknown>;
157} {
158 const args: Record<string, unknown> = {};
159 for (const [k, v] of Object.entries(e)) {
160 if (!RESERVED.has(k)) args[k] = v;
161 }
162 return {
163 tool: typeof e.tool === "string" ? e.tool : "(unknown)",
164 toolUseId: typeof e.tool_use_id === "string" ? e.tool_use_id : undefined,
165 args,
166 };
167}
168lib/types.ts 72 lines1/**
2 * Shared types.
3 *
4 * Everything in lib/ is PURE: no `$`, no imports other than relative ones.
5 * That is not a style preference — the function-hooks loader refuses a module
6 * that imports anything but a relative path or "claude-code", and its static
7 * scanner refuses `$` crossing an import boundary. All engine access therefore
8 * lives in hooks/verify.ts, and lib/ is plain data in, plain data out (which
9 * also makes it directly unit-testable).
10 */
11
12/** What the policy decided, or what actually happened, for one tool call. */
13export type Action = "allow" | "flag" | "deny";
14
15/** How the plugin behaves overall. */
16export type Mode = "enforce" | "audit" | "off";
17
18/** One Jev signal, named for the audit log and the user-facing reason. */
19export interface Signal {
20 key: string;
21 label: string;
22 /** Calibrated probability 0..1 that the answer is "yes". */
23 probability: number;
24}
25
26/** A parsed Jev verdict for one pending tool call. */
27export interface Verdict {
28 signals: Signal[];
29 /**
30 * Intent consistency on the supplied rubric. TypeSafe scores are indexed
31 * FROM ZERO, so an N-entry rubric yields 0..N-1 — `max` is N-1, not N.
32 */
33 intent: { score: number; max: number; confidence: number } | null;
34 usage: { inputTokens: number; outputTokens: number } | null;
35 model: string | null;
36}
37
38export interface Decision {
39 action: Action;
40 reasons: string[];
41 maxProbability: number;
42}
43
44/** A previous tool call, as fed back to Jev for contradiction checks. */
45export interface HistoryEntry {
46 tool: string;
47 args: Record<string, unknown>;
48 at: number;
49}
50
51export interface ResolvedConfig {
52 mode: Mode;
53 model: string;
54 denyAbove: number;
55 flagAbove: number;
56 /** Intent score at or below this (0-indexed rubric) counts as a flag signal. */
57 lowIntentAtOrBelow: number;
58 /** Applied when Jev itself fails: the gate's own failure policy. */
59 onError: Action;
60 /** Applied to a flagged call when no human answer is available. */
61 onNoAnswer: Action;
62 /** Tools never sent to Jev at all (cheap, read-only, high-volume). */
63 skipTools: string[];
64 /** Milliseconds to wait for Jev before giving up and applying onError. */
65 timeoutMs: number;
66 /** Consecutive failures before the gate reports itself degraded. */
67 unhealthyAfter: number;
68 auditLog: boolean;
69 /** Warnings raised while normalizing user input; surfaced at session start. */
70 warnings: string[];
71}
72