SLOPSHOPPER

spatz-mod

Recommends Claude models and effort with spatz and can apply them to subagents and the main session.

newguardcommandtoaststatusprocess
v0.1.6MITupdated 2026-10-07lorenzh/spatz/packages/claude-mod
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · spatz-mod
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /spatz ⎿ spatz-mod: spatz ⎿ spatz-mod: mode: show ⎿ spatz-mod: scope: subagent ⎿ spatz-mod: main: off ⎿ spatz-mod: record: auto (on) ⎿ spatz-mod: last: none yet ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts
README

<img src="docs/assets/banner-light.svg" alt="spatz" width="340" height="84">

spatz

Choose a model and effort for your coding task. Learn from the result.

Harness model defaults refresh from a daily catalog, without a CLI release.

License: MIT Version Bun 1.4 CI Harness catalog Nightly

spatz ranks the model and effort pairs that you can use for a coding task. It learns from task results to help choose cheaper pairs that succeed. You or your coding agent run the task with the recommended pair.

Quickstart · Documentation · Contributing · Releases

Why spatz?

Different tasks need different models and effort levels. spatz combines task classification with results from your previous tasks.

  • Use your available models. Start with harness defaults or set your own candidates.
  • Learn from results. Record outcomes through hooks or spatz report.
  • Compare performance. Use spatz stats to see results per task type or routing scope.
  • Connect to coding agents. Hooks support Claude Code and Codex CLI. Other agents can use the CLI directly.
  • Keep learning data local. spatz stores results in ~/.spatz/spatz.db. It does not store task text.

Quickstart

1. Install

Install the stable CLI with Node.js 18 or newer:

npm install -g @spatz/cli
spatz --version

The npm package includes the Bun runtime. Keep optional dependencies enabled. Linux needs glibc. spatz does not support Alpine Linux. For Windows, see Platform support.

For installation without Node.js, use a release archive. For the unstable nightly version, use npm install -g @spatz/cli@nightly. To test a release candidate, use npm install -g @spatz/cli@next.

2. Ask for a recommendation

From Claude Code or Codex, pass the task:

spatz "Fix the off-by-one error in src/list.ts"

spatz uses a preset for the detected harness. Override it with --models, SPATZ_MODELS, or models in configuration files. Use --family claude or --family gpt to filter candidates for a review. The output includes a suggestion_id and a ranked list of pairs. Add --json for machine-readable output.

spatz works without an API key. Without Jev or learned results, it recommends the most expensive candidate. Jev is the TypeSafe AI model that classifies the task and estimates candidate suitability. To enable Jev, set your TypeSafe AI key:

export TYPESAFE_AI_API_KEY="<your-key>"

When you enable Jev, spatz sends task text and candidate details to TypeSafe AI. Read Privacy before using sensitive task text.

3. Run the task and report the result

Run the task with the chosen model and effort in your coding agent. Then report the pair you actually used:

spatz report <suggestion_id> \
  --model claude-sonnet-5-5 --effort medium --result pass

spatz stats

Replace <suggestion_id> with the ID from the recommendation. Results can be pass, partial, or fail. An explicit report overrides hook signals for its selected attempt. A changed verdict creates a retry. Use --correct to fix a mistaken report. Use --attempt <id> --confirm to report on an existing attempt without creating a retry. Use spatz suggest "Retry the task" --retry-of <suggestion_id> to link a new suggestion to the same recovery chain. See report flags for explicit attempt selection.

For experiments, add --dry-run to the recommendation command. These suggestions never count toward learning or statistics.

How it works

  1. A local filter checks the task text for known secret patterns.
  2. Jev classifies the task. Without Jev, spatz uses local keyword rules.
  3. spatz ranks your candidates using classification and learned success estimates.
  4. You or your agent choose a pair and run the task.
  5. Hooks or spatz report record the outcome for future recommendations.

The CLI recommends pairs. The optional Claude Code mod can apply them automatically. See How it works and Recommendation rules for the decision process.

Agent integrations

IntegrationWhat it doesSetup
Claude Code spatz pluginRecord task signals and model usage.Hooks guide
Codex spatz pluginRecord shell results and model usage.Hooks guide
Claude Code spatz-mod plugin (mod)Show recommendations or apply model and effort choices.Mod guide
Other agentsRequest recommendations and report outcomes through the CLI.CLI reference

The Claude Code mod supports step, turn, subagent, session, and escalate routing scopes. Hooks and the mod can run together. With record: auto, the mod records step usage and hooks record signals.

Install for Claude Code

Use macOS or Linux with a POSIX shell (see Platform support). The plugins run the CLI through a bundled launcher: an installed spatz on PATH wins, otherwise it uses Bun or npx from the PATH Claude Code starts with. The mod needs Claude Code 2.1.287 or newer.

  1. Start Claude Code. Add the marketplace and install the hooks:
   /plugin marketplace add lorenzh/spatz
   /plugin install spatz@spatz

The hooks record test/build results and model usage. They do not switch models.

  1. Optional: install the mod for automatic recommendations:
   /plugin install spatz-mod@spatz

The mod defaults to show mode. To apply recommendations to subagents, run:

   /spatz mode apply
   /spatz status

When you install both plugins, keep record: auto. Hooks record signals. The mod owns usage for its registered execution segments. To apply recommendations to the main session too, enable /spatz main on.

  1. Remove any manual spatz hook entries from ~/.claude/settings.json and project settings. Keep unrelated hooks. This avoids duplicate records.

See the Claude Code mod guide for routing scopes and persistent configuration.

Install for Codex CLI

Use macOS or Linux with a POSIX shell and a Codex CLI version with plugin support. Install the spatz CLI first and check spatz --version in your terminal.

  1. Add the marketplace and install the hooks from your terminal:
   codex plugin marketplace add lorenzh/spatz
   codex plugin add spatz@spatz
   codex plugin list
  1. Start Codex. Run /hooks to review and trust the spatz hooks. Codex skips plugin hooks until you trust them. See OpenAI's hook documentation.
  1. Remove any manual spatz hook … --agent codex entries from ~/.codex/hooks.json. Keep unrelated hooks. This avoids duplicate records.

The plugin records shell results and model usage. Its routing skill guides the agent through recommendations and outcome reports. It does not automatically switch the Codex model. See the Codex hooks guide for recorded events and limits.

Check your setup

If you use Jev, set TYPESAFE_AI_API_KEY in the terminal before starting your coding agent. Ask your agent to request a recommendation for a real task using its available model and effort pairs. Keep the suggestion output unfiltered so the hooks can read its ID. After the task, ask the agent to report the actual pair and result with spatz report. Run spatz stats to see recorded outcomes.

Every plugin ships the routing skill (/spatz:routing, /spatz-mod:routing, or spatz:routing in Codex). The mod's /spatz command checks status and changes mode or scope. You do not need to edit your agent instructions. Every plugin can run without a global CLI through its bundled launcher using Bun or npx. The first run downloads about 60 MB. Codex hooks time out after 10 seconds. For this setup, warm the launcher before the first session with "<plugin root>/bin/spatz" --version. See Installation for launcher paths and setup without a global CLI.

Privacy

spatz stores learning data under ~/.spatz. It does not store task text or tool output. Hooks process task signals locally and make no network requests.

When you enable Jev, spatz sends task text and candidate details to TypeSafe AI. The secret filter catches known patterns. It cannot detect every secret. Do not put secrets in task text.

To disable Jev:

export SPATZ_NO_JEV=1

For a project, put {"jev": false} in .spatz.json in the working directory. Without TYPESAFE_AI_API_KEY, Jev is also disabled.

Disabling Jev still allows OpenRouter requests for model prices. These requests contain no task data. spatz caches prices for 24 hours. When the network is unavailable, spatz can still use cached prices.

See the Privacy guide for data flows and deletion instructions. See Configuration for environment variables and local files.

Releases

Download an archive and its matching .sha256 file from GitHub Releases. Archives include the runtime. You do not need Node.js or Bun installed.

Choose linux or darwin (macOS), then x64 or arm64. Apple Silicon uses darwin-arm64. Linux builds need glibc. For Windows archives, see Platform support.

Check the checksum before extraction. This Linux x64 example uses version 0.1.0:

version=0.1.0
archive="spatz-cli-$version-linux-x64.tar.gz"
sha256sum --check "$archive.sha256"

mkdir -p ~/.local/lib/spatz ~/.local/bin
tar -xzf "$archive" -C ~/.local/lib/spatz
ln -sfn "$HOME/.local/lib/spatz/${archive%.tar.gz}/spatz" ~/.local/bin/spatz
export PATH="$HOME/.local/bin:$PATH"
spatz --version

Use your downloaded version. On macOS, use shasum -a 256 --check "$archive.sha256". If your shell does not include ~/.local/bin, add the PATH line to your shell profile.

The nightly release is unstable. See Releasing spatz for the release process.

Install from source

Install Bun 1.4, then clone the repository:

git clone https://github.com/lorenzh/spatz.git
cd spatz
bun install

Put a wrapper on your PATH so hooks can find spatz in non-interactive shells:

mkdir -p ~/.local/bin
printf '#!/bin/sh\nexec bun "%s/packages/cli/src/cli.ts" "$@"\n' "$PWD" > ~/.local/bin/spatz
chmod +x ~/.local/bin/spatz
export PATH="$HOME/.local/bin:$PATH"
spatz --version

A shell alias does not work for hooks. See Contributing for the full development setup.

Documentation

GuideWhat you will find
CLI referenceCommands, flags, model IDs, output fields, and exit codes.
How it worksArchitecture and the flow from task to outcome.
Recommendation rulesRanking, exploration, and control groups.
MeasurementsBenchmark results for model and effort selection.
Bench rowsThe spatz-eval-row/1 format for benchmark results.
HooksClaude Code and Codex CLI setup and recorded signals.
Claude Code modModes, routing scopes, and /spatz commands.
ConfigurationEnvironment variables and local files.
PrivacyNetwork requests, stored data, and deletion.

Contributing

Bug reports and pull requests are welcome. For larger changes, open an issue first to agree on the scope. See CONTRIBUTING.md for setup and the test-first workflow. Changes to main must go through a pull request.

Run the checks before submitting a change:

bun test
bun run typecheck
bun run lint

For vulnerabilities, follow SECURITY.md. Do not open a public issue.

License

MIT © 2026 Lorenz Hilpert.

Source 3 files
hooks/register.ts 461 lines
1import type { EngineInterface, On, PluginOptions } from "claude-code";
2import {
3	aliasFor,
4	attemptCommand,
5	type Decision,
6	type Link,
7	linkAgent,
8	type Run,
9	recordUsage,
10	type Scope,
11	type StepUsage,
12	stronger,
13	suggest,
14	type Tokens,
15} from "./bridge.ts";
16import {
17	band,
18	change,
19	describeDecision,
20	readSettings,
21	USAGE,
22} from "./settings.ts";
23
24// Keyword match, as the CLI's hook parser does (packages/core/src/signals): a test or build
25// command that is the last step of a plain && chain, so its exit status is the one that counts.
26// ponytail: keyword regexes, not a shell parser.
27const PREFIX = String.raw`^(?:\w+=\S*\s+)*(?:(?:rtk(?:\s+proxy)?|time|npx|bunx|pnpx|uv\s+run|poetry\s+run|python3?\s+-m)\s+)*`;
28const TEST_OR_BUILD = new RegExp(
29	String.raw`${PREFIX}(?:(?:bun|npm|pnpm|yarn)\s+(?:run\s+)?(?:test|build)|pytest|(?:go|cargo)\s+(?:test|build)|vitest|jest|tsc|make(?:\s+-\S+)*(?:\s+(?:build|all|test|check))?(?:\s+-\S+)*\s*$)(?:\s|$)`,
30);
31const SETUP = /^(?:cd|pushd|export)(?:\s|$)/;
32// Failure text of a run that never reached the runner, as in packages/core/src/signals.
33const NOT_RUN =
34	/No such file or directory|[Pp]ermission denied|command not found/;
35function isTestOrBuild(command: string): boolean {
36	const plain = command.trim().replace(/'[^']*'|"(?:\\.|[^"\\])*"/g, "''");
37	if (/[|;&\n`]|\$\(/.test(plain.replaceAll("&&", " "))) return false;
38	const segs = plain.split("&&").map((x) => x.trim());
39	const last = segs.pop() ?? "";
40	if (!segs.every((x) => SETUP.test(x))) return false;
41	return (
42		TEST_OR_BUILD.test(last) &&
43		!/\s(?:--collect-only|--co|--help|-h|--version)(?:\s|$)/.test(last)
44	);
45}
46
47/** An interrupted run, or an error text of a run that never started, is no failure of the tests. */
48function neverRan(result: unknown): boolean {
49	if (typeof result === "string") return NOT_RUN.test(result);
50	const r = result as { interrupted?: boolean; stderr?: string } | null;
51	return !!r && (r.interrupted === true || NOT_RUN.test(r.stderr ?? ""));
52}
53
54/** What the hooks use of `$`; a hooks module may pass `$` only to a top-level function, so helpers take this. */
55interface Io {
56	run: Run;
57	sessionId(): Promise<string | undefined>;
58	status(text: string | undefined): void;
59	log(text: string): void;
60	toast(text: string): void;
61}
62
63function bind($: EngineInterface, spatz: string): Io {
64	const executable =
65		spatz === "spatz" ? ["sh", `${$.plugin.root}/bin/spatz`] : [spatz];
66	return {
67		run: (argv, init) =>
68			$.process.run(
69				argv[0] === spatz ? [...executable, ...argv.slice(1)] : [...argv],
70				// The mod always runs inside Claude Code; this keeps the CLI's harness preset working
71				// even when the engine environment lacks the markers Bash children get.
72				{ ...init, env: { CLAUDECODE: "1", ...init?.env } },
73			),
74		sessionId: () => $.session.id().catch(() => undefined),
75		status: (text) => $.ui.status(text),
76		toast: (text) => $.ui.toast(text),
77		log: (text) => $.ui.log(text, { to: "debug" }),
78	};
79}
80
81const ESCALATING: readonly Scope[] = ["escalate"];
82
83export function register(on: On, options: PluginOptions = {}) {
84	const s = readSettings(options);
85	/** Subagent and escalate decisions by agent id; kept across the agent's follow-up runs. */
86	const agents = new Map<string, Decision>();
87	/** turn and escalate decisions by turn id (main loop). */
88	const turns = new Map<string, Decision>();
89	/** Task text held only for the step scope, until the turn or agent ends. */
90	const prompts = new Map<string, string>();
91	const requested = new Map<
92		string,
93		Pick<Link, "requested" | "requestedAgent">
94	>();
95	type Segment = Omit<StepUsage, "attempt"> & {
96		attempt: Promise<string | undefined>;
97		turn: string;
98		model: string;
99		suggestionId: string;
100	};
101	const active = new Map<string, Segment>();
102	const failures = new Map<string, number>();
103	let lastTurn: Decision | undefined;
104	let sessionDecision: Decision | undefined;
105	let sessionTried = false;
106	let currentTurn: string | undefined;
107	let last: Decision | undefined;
108	let logged = false;
109	const logFailure = (io: Io, error: unknown) => {
110		if (logged) return;
111		logged = true;
112		try {
113			io.log(`spatz: decision failed, routing unchanged: ${error}`);
114		} catch {}
115		try {
116			io.toast(
117				"spatz: CLI call failed, routing unchanged (details: claude --debug)",
118			);
119		} catch {}
120	};
121
122	const show = (io: Io) => {
123		try {
124			io.status(band(s, last));
125		} catch {}
126	};
127
128	async function decide(
129		io: Io,
130		task: string,
131		link: Omit<Link, "session">,
132		withSession = true,
133	): Promise<Decision | undefined> {
134		const session = withSession ? await io.sessionId() : undefined;
135		const d = await suggest(
136			io.run,
137			task,
138			s.models,
139			{ ...link, session },
140			s.spatz,
141			(error) => logFailure(io, error),
142		);
143		if (!d) return undefined;
144		if (d.exploredRisky && !s.exploreHard) return undefined;
145		last = d;
146		show(io);
147		return d;
148	}
149
150	const applies = (agentId: string | undefined) =>
151		s.mode === "apply" && (agentId !== undefined || s.main);
152
153	async function pick(
154		io: Io,
155		e: { turnId: string; agentId?: string },
156	): Promise<Decision | undefined> {
157		const key = e.agentId ?? e.turnId;
158		if (s.scope === "step") {
159			const task = prompts.get(key);
160			return task
161				? decide(io, task, {
162						scope: "step",
163						turn: e.turnId,
164						agentId: e.agentId,
165						...requested.get(key),
166					})
167				: undefined;
168		}
169		if (e.agentId)
170			return s.scope === "subagent" || s.scope === "escalate"
171				? agents.get(e.agentId)
172				: undefined;
173		if (s.scope === "session") return sessionDecision;
174		return s.scope === "subagent" ? undefined : turns.get(e.turnId);
175	}
176
177	const tokens = (u: {
178		input_tokens: number;
179		output_tokens: number;
180		cache_read_input_tokens: number;
181		cache_creation_input_tokens: number;
182	}): Tokens => ({
183		input: u.input_tokens,
184		output: u.output_tokens,
185		cacheRead: u.cache_read_input_tokens,
186		cacheCreation: u.cache_creation_input_tokens,
187	});
188
189	on("session.start", async ($, e, next) => {
190		const result = await next(e);
191		try {
192			await $.command.register({
193				name: "spatz",
194				description: "Show or change the spatz model routing",
195				argumentHint: "[status|mode|scope|record|main]",
196			});
197		} catch {}
198		return result;
199	});
200
201	on("command.run", { command: "spatz" }, async ($, e) => {
202		const io = bind($, s.spatz);
203		const args = e.args.trim();
204		if (args === "" || args === "status") {
205			const record = s.record === "auto" ? "auto (on)" : s.record;
206			return {
207				text: `spatz\nmode: ${s.mode}\nscope: ${s.scope}\nmain: ${s.main ? "on" : "off"}\nrecord: ${record}\nlast: ${describeDecision(last)}`,
208			};
209		}
210		const answer = change(s, args);
211		show(io);
212		return { text: answer ?? USAGE };
213	});
214
215	on("agent.spawn", async ($, e, next) => {
216		const io = bind($, s.spatz);
217		if (s.mode === "off" || e.fork) return next(e);
218		// An explicit model or a named agent type is the caller's choice; so are all steps of that agent.
219		if (
220			s.respectPinned &&
221			(e.model || (e.subagentType && e.subagentType !== "general-purpose"))
222		)
223			return next(e);
224		if (s.scope === "step") {
225			const result = await next(e);
226			if (result.agentId && !result.deny) {
227				prompts.set(result.agentId, e.prompt);
228				requested.set(result.agentId, {
229					requested: e.model,
230					requestedAgent: e.subagentType,
231				});
232			}
233			return result;
234		}
235		if (s.scope !== "subagent" && s.scope !== "escalate") return next(e);
236		// The agent id exists only after the spawn: ask without session or agent, then link.
237		const d = await decide(
238			io,
239			e.prompt,
240			{ scope: s.scope, requested: e.model, requestedAgent: e.subagentType },
241			false,
242		);
243		if (!d) return next(e);
244		const alias = s.mode === "apply" ? aliasFor(d.model) : undefined;
245		const result = await next(alias ? { ...e, model: alias } : e);
246		if (result.agentId && !result.deny) {
247			agents.set(result.agentId, d);
248			try {
249				await linkAgent(
250					io.run,
251					d.suggestionId,
252					result.agentId,
253					await io.sessionId().catch(() => undefined),
254					s.spatz,
255				);
256			} catch {}
257		}
258		return result;
259	});
260
261	on("turn.start", async ($, e, next) => {
262		const io = bind($, s.spatz);
263		currentTurn = e.turnId;
264		if (s.mode !== "off") {
265			try {
266				if (s.scope === "step") prompts.set(e.turnId, e.text);
267				else if (s.scope === "session") {
268					if (!sessionTried && e.text) {
269						sessionTried = true;
270						sessionDecision = await decide(io, e.text, {
271							scope: "session",
272							turn: e.turnId,
273						});
274					}
275				} else if (s.scope !== "subagent") {
276					const long = e.text.length >= s.minPromptChars;
277					const d = long
278						? await decide(io, e.text, { scope: s.scope, turn: e.turnId })
279						: lastTurn;
280					if (long) lastTurn = d;
281					if (d) turns.set(e.turnId, d);
282				}
283			} catch {}
284		}
285		return next(e);
286	});
287
288	on("turn.step", async function* ($, e, next) {
289		const io = bind($, s.spatz);
290		let d: Decision | undefined;
291		try {
292			d = s.mode === "off" ? undefined : await pick(io, e);
293		} catch (error) {
294			logFailure(io, error);
295		}
296		const sent =
297			d && applies(e.agentId)
298				? {
299						...e,
300						model: d.model,
301						...(d.effort !== "none" && { effort: d.effort }),
302					}
303				: e;
304		let identity: Segment | undefined;
305		const agentKey = e.agentId ?? "";
306		if (!d || s.record === "off") active.delete(agentKey);
307		try {
308			if (d && s.record !== "off") {
309				const session = await io.sessionId();
310				const key = `${e.turnId}:${e.index}`;
311				const effort =
312					d && applies(e.agentId) && d.effort === "none" ? "none" : sent.effort;
313				if (session) {
314					const previous = active.get(agentKey);
315					const samePair =
316						previous?.suggestionId === d.suggestionId &&
317						previous.model === sent.model &&
318						previous.effort === effort;
319					if (!samePair) active.delete(agentKey);
320					const attempt = samePair
321						? previous.attempt
322						: attemptCommand(
323								io.run,
324								[
325									"start",
326									d.suggestionId,
327									"--key",
328									key,
329									"--model",
330									sent.model,
331									...(effort ? ["--effort", effort] : []),
332									"--session",
333									session,
334									"--turn",
335									e.turnId,
336									...(e.agentId ? ["--agent-id", e.agentId] : []),
337									"--owns-usage",
338									"--json",
339								],
340								s.spatz,
341							);
342					identity = {
343						attempt,
344						key,
345						session,
346						agentId: e.agentId,
347						effort,
348						turn: e.turnId,
349						model: sent.model,
350						suggestionId: d.suggestionId,
351					};
352					active.set(agentKey, identity);
353				}
354			}
355		} catch {}
356		const result = yield* next(sent);
357		try {
358			const attempt = await identity?.attempt;
359			if (identity && !attempt && active.get(agentKey) === identity)
360				active.delete(agentKey);
361			if (d && result.usage && identity && attempt) {
362				await recordUsage(
363					io.run,
364					d.suggestionId,
365					result.usage.model,
366					e.turnId,
367					tokens(result.usage),
368					s.spatz,
369					{ ...identity, attempt },
370				);
371			}
372		} catch {}
373		return result;
374	});
375
376	on("turn.complete", async ($, e, next) => {
377		const io = bind($, s.spatz);
378		const result = await next(e);
379		const key = e.agentId ?? "";
380		const identity = active.get(key);
381		if (identity?.turn === e.turnId && (await identity.attempt)) {
382			await attemptCommand(
383				io.run,
384				[
385					"finalize",
386					"--session",
387					identity.session,
388					...(e.agentId ? ["--agent-id", e.agentId] : []),
389				],
390				s.spatz,
391			);
392		}
393		prompts.delete(e.turnId);
394		failures.delete(e.agentId ?? e.turnId);
395		return result;
396	});
397
398	on("tool.call", async ($, e, next) => {
399		const io = bind($, s.spatz);
400		const identity = active.get(e.agentId ?? "");
401		if (identity && s.mode !== "off" && s.record !== "off")
402			void identity.attempt.then(
403				(attempt) =>
404					attempt &&
405					attemptCommand(
406						io.run,
407						[
408							"bind",
409							attempt,
410							"--call",
411							e.tool_use_id,
412							"--session",
413							identity.session,
414							...(e.agentId ? ["--agent-id", e.agentId] : []),
415						],
416						s.spatz,
417					),
418			);
419		const result = await next(e);
420		try {
421			if (
422				s.mode !== "off" &&
423				ESCALATING.includes(s.scope) &&
424				e.tool === "Bash" &&
425				result.isError &&
426				isTestOrBuild(e.command) &&
427				!neverRan(result.result)
428			)
429				escalate(io, e.agentId);
430		} catch {}
431		return result;
432	});
433
434	function escalate(io: Io, agentId: string | undefined) {
435		const key = agentId ?? currentTurn;
436		if (!key) return;
437		const count = (failures.get(key) ?? 0) + 1;
438		failures.set(key, count >= s.escalateAfter ? 0 : count);
439		if (count < s.escalateAfter) return;
440		const map = agentId ? agents : turns;
441		const current = map.get(key);
442		const pairs = current?.candidates ?? [];
443		const at = pairs.findIndex(
444			(p) => p.model === current?.model && p.effort === current?.effort,
445		);
446		const up =
447			current &&
448			(s.models.length
449				? stronger(s.models, current)
450				: at >= 0
451					? pairs[at + 1]
452					: undefined);
453		if (!current || !up) return;
454		const d = { ...current, ...up, escalated: true };
455		map.set(key, d);
456		last = d;
457		show(io);
458		io.toast(`spatz: escalating to ${d.model}:${d.effort}`);
459	}
460}
461
hooks/bridge.ts 302 lines
1export const SCOPES = [
2	"step",
3	"turn",
4	"subagent",
5	"session",
6	"escalate",
7] as const;
8export type Scope = (typeof SCOPES)[number];
9export const MODES = ["off", "show", "apply"] as const;
10export type Mode = (typeof MODES)[number];
11export const EFFORTS = [
12	"none",
13	"low",
14	"medium",
15	"high",
16	"xhigh",
17	"max",
18	"ultra",
19] as const;
20export type Effort = (typeof EFFORTS)[number];
21// ultra is a spatz effort, but Claude Code cannot dispatch it.
22const isClaudeEffort = (value: unknown): value is Exclude<Effort, "ultra"> =>
23	value !== "ultra" && (EFFORTS as readonly unknown[]).includes(value);
24
25export interface Decision {
26	suggestionId: string;
27	model: string;
28	effort: Effort;
29	scope: Scope;
30	escalated?: boolean;
31	/** An exploration pick on a hard or critical task. */
32	exploredRisky?: boolean;
33	candidates?: { model: string; effort: Effort }[];
34}
35
36/** What links a suggestion to the session: passed to the CLI as flags. */
37export interface Link {
38	requested?: string;
39	requestedAgent?: string;
40	scope: Scope;
41	session?: string;
42	turn?: string;
43	agentId?: string;
44}
45
46export interface Tokens {
47	input: number;
48	output: number;
49	cacheRead: number;
50	cacheCreation: number;
51}
52
53interface ProcessResult {
54	exitCode: number;
55	stdout: string;
56	stderr: string;
57}
58
59export type Run = (
60	argv: readonly string[],
61	init: { timeoutMs: number },
62) => Promise<ProcessResult>;
63
64export const SUGGEST_TIMEOUT_MS = 6000;
65const USAGE_TIMEOUT_MS = 2000;
66
67// ponytail: fixed table of what each Agent-tool alias resolves to today; the
68// engine offers no lookup. Update with the model catalog.
69const ALIASES: Record<string, string> = {
70	sonnet: "claude-sonnet-5-5",
71	opus: "claude-opus-5-5",
72	haiku: "claude-haiku-4-5-20251001",
73	fable: "claude-fable-5-1",
74};
75
76/** The Agent tool's model field takes aliases only: the alias that resolves to exactly this id, else undefined. */
77export function aliasFor(model: string): string | undefined {
78	return Object.keys(ALIASES).find((alias) => ALIASES[alias] === model);
79}
80
81function claudeModel(id: unknown): string | null {
82	if (typeof id !== "string" || !id.startsWith("anthropic/claude-"))
83		return null;
84	const model = id.slice("anthropic/".length).replaceAll(".", "-");
85	return (
86		Object.values(ALIASES).find((id) => id.replace(/-\d{8}$/, "") === model) ??
87		model
88	);
89}
90
91export async function suggest(
92	run: Run,
93	task: string,
94	models: string[],
95	link: Link,
96	spatz = "spatz",
97	onFailure?: (error: unknown) => void,
98): Promise<Decision | null> {
99	const fail = (error: unknown) => {
100		try {
101			onFailure?.(error);
102		} catch {}
103	};
104	try {
105		const { exitCode, stdout, stderr } = await run(
106			[
107				spatz,
108				task,
109				...(models.length ? ["--models", models.join(",")] : []),
110				"--json",
111				"--scope",
112				link.scope,
113				"--source",
114				"claude-code-mod",
115				"--requested",
116				link.requested ?? "-",
117				...(link.requestedAgent
118					? ["--requested-agent", link.requestedAgent]
119					: []),
120				...(link.session ? ["--session", link.session] : []),
121				...(link.turn ? ["--turn", link.turn] : []),
122				...(link.agentId ? ["--agent-id", link.agentId] : []),
123			],
124			{ timeoutMs: SUGGEST_TIMEOUT_MS },
125		);
126		if (exitCode !== 0) {
127			fail(
128				new Error(
129					`CLI exited ${exitCode}${stderr ? `: ${stderr.slice(0, 200)}` : ""}`,
130				),
131			);
132			return null;
133		}
134		const result = JSON.parse(stdout);
135		const first = result?.ranking?.[0];
136		const model = claudeModel(first?.model);
137		if (
138			typeof result?.suggestion_id !== "string" ||
139			!model ||
140			!isClaudeEffort(first?.effort)
141		) {
142			fail(new Error("invalid suggestion response"));
143			return null;
144		}
145		const candidates = Array.isArray(result.candidates)
146			? result.candidates.flatMap(
147					(pair: { model?: unknown; effort?: unknown }) => {
148						const model = claudeModel(pair?.model);
149						return model && isClaudeEffort(pair?.effort)
150							? [{ model, effort: pair.effort as Effort }]
151							: [];
152					},
153				)
154			: undefined;
155		const c = result.classification;
156		const exploredRisky =
157			result.explored === true &&
158			(c?.difficulty === "hard" || (c?.criticality ?? "none") !== "none");
159		return {
160			...(candidates && { candidates }),
161			...(exploredRisky && { exploredRisky }),
162			suggestionId: result.suggestion_id,
163			model,
164			effort: first.effort,
165			scope: link.scope,
166		};
167	} catch (error) {
168		fail(error);
169		return null;
170	}
171}
172
173/** `spatz link`: gives a spawn-time suggestion the real agent id and the session. Fails open: false on any error. */
174export async function linkAgent(
175	run: Run,
176	suggestionId: string,
177	agentId: string,
178	session: string | undefined,
179	spatz = "spatz",
180): Promise<boolean> {
181	if (!session) return false;
182	try {
183		const { exitCode } = await run(
184			[
185				spatz,
186				"link",
187				suggestionId,
188				"--agent-id",
189				agentId,
190				"--session",
191				session,
192			],
193			{ timeoutMs: USAGE_TIMEOUT_MS },
194		);
195		return exitCode === 0;
196	} catch {
197		return false;
198	}
199}
200
201/** Run an attempt lifecycle command. Only start returns an attempt id. */
202export async function attemptCommand(
203	run: Run,
204	args: string[],
205	spatz = "spatz",
206): Promise<string | undefined> {
207	try {
208		const result = await run([spatz, "attempt", ...args], {
209			timeoutMs: USAGE_TIMEOUT_MS,
210		});
211		if (result.exitCode !== 0 || args[0] !== "start") return undefined;
212		const value = JSON.parse(result.stdout);
213		return typeof value.id === "string" ? value.id : undefined;
214	} catch {
215		return undefined;
216	}
217}
218
219export interface StepUsage {
220	attempt: string;
221	key: string;
222	session: string;
223	agentId?: string;
224	effort?: string;
225}
226
227/** Disjoint step usage. Fails open: false on any error. */
228export async function recordUsage(
229	run: Run,
230	suggestionId: string,
231	model: string,
232	turn: string,
233	tokens: Tokens,
234	spatz = "spatz",
235	identity?: StepUsage,
236): Promise<boolean> {
237	try {
238		const { exitCode } = await run(
239			[
240				spatz,
241				"usage",
242				suggestionId,
243				...(identity
244					? [
245							"--attempt",
246							identity.attempt,
247							"--key",
248							identity.key,
249							"--session",
250							identity.session,
251							...(identity.agentId ? ["--agent-id", identity.agentId] : []),
252							...(identity.effort ? ["--effort", identity.effort] : []),
253						]
254					: []),
255				"--model",
256				model,
257				"--input",
258				String(tokens.input),
259				"--output",
260				String(tokens.output),
261				"--cache-read",
262				String(tokens.cacheRead),
263				"--cache-creation",
264				String(tokens.cacheCreation),
265				"--turn",
266				turn,
267				"--source",
268				"claude-code-mod",
269				"--json",
270			],
271			{ timeoutMs: USAGE_TIMEOUT_MS },
272		);
273		return exitCode === 0;
274	} catch {
275		return false;
276	}
277}
278
279/** Candidate pairs from "model:low+high,model2:..." weakest first: the list is strongest model first; efforts ascend. */
280export function ladder(models: string[]): { model: string; effort: Effort }[] {
281	return models.toReversed().flatMap((entry) => {
282		const [model = "", efforts = ""] = entry.split(":");
283		return efforts
284			.split("+")
285			.filter(isClaudeEffort)
286			.sort((a, b) => EFFORTS.indexOf(a) - EFFORTS.indexOf(b))
287			.map((effort) => ({ model, effort }));
288	});
289}
290
291/** The next stronger pair after `from`, or null when it is the top or not on the ladder. */
292export function stronger(
293	models: string[],
294	from: { model: string; effort: Effort },
295): { model: string; effort: Effort } | null {
296	const pairs = ladder(models);
297	const at = pairs.findIndex(
298		(p) => p.model === from.model && p.effort === from.effort,
299	);
300	return at === -1 ? null : (pairs[at + 1] ?? null);
301}
302
hooks/settings.ts 103 lines
1import {
2	type Decision,
3	MODES,
4	type Mode,
5	SCOPES,
6	type Scope,
7} from "./bridge.ts";
8
9export const RECORDS = ["auto", "on", "off"] as const;
10export type RecordSetting = (typeof RECORDS)[number];
11
12export interface Settings {
13	mode: Mode;
14	scope: Scope;
15	/** Main-session rewrites in apply mode need this explicit opt-in: a model switch drops the prompt cache. */
16	main: boolean;
17	record: RecordSetting;
18	/** turn and escalate skip prompts shorter than this and keep the last decision. */
19	minPromptChars: number;
20	/** escalate switches to the next stronger pair after this many failing test/build results. */
21	escalateAfter: number;
22	/** Leave a spawn alone when it names a model or an agent type other than general-purpose (the hook cannot see whether a type pins a model). */
23	respectPinned: boolean;
24	/** Allow exploration picks (a cheaper pair to collect data) on hard or critical tasks. */
25	exploreHard: boolean;
26	spatz: string;
27	models: string[];
28}
29
30type Options = Readonly<Record<string, unknown>>;
31
32const oneOf = <T extends string>(
33	list: readonly T[],
34	value: unknown,
35	fallback: T,
36): T => (list.includes(value as T) ? (value as T) : fallback);
37
38const count = (value: unknown, fallback: number) =>
39	typeof value === "number" && Number.isFinite(value) && value >= 0
40		? Math.floor(value)
41		: fallback;
42
43export function readSettings(options: Options): Settings {
44	const text = (key: string, fallback: string) =>
45		typeof options[key] === "string" && options[key]
46			? (options[key] as string)
47			: fallback;
48	return {
49		mode: oneOf(MODES, options.mode, "show"),
50		scope: oneOf(SCOPES, options.scope, "subagent"),
51		main: options.main === true || options.main === "true",
52		record: oneOf(RECORDS, options.record, "auto"),
53		minPromptChars: count(options.minPromptChars, 20),
54		escalateAfter: Math.max(1, count(options.escalateAfter, 2)),
55		respectPinned: !(
56			options.respectPinned === false || options.respectPinned === "false"
57		),
58		exploreHard: options.exploreHard === true || options.exploreHard === "true",
59		spatz: text("spatz", "spatz"),
60		models: text("models", "")
61			.split(",")
62			.map((model) => model.trim())
63			.filter(Boolean),
64	};
65}
66
67export const USAGE =
68	"usage: /spatz [status] | mode <off|show|apply> | scope <step|turn|subagent|session|escalate> | record <auto|on|off> | main <on|off>";
69
70export function describeDecision(d: Decision | undefined): string {
71	return d
72		? `${d.model}:${d.effort} (${d.scope}${d.escalated ? ", escalated" : ""})`
73		: "none yet";
74}
75
76/** The band text: the last recommendation with its scope; undefined clears it. */
77export function band(s: Settings, d: Decision | undefined): string | undefined {
78	if (s.mode === "off" || !d) return undefined;
79	return `spatz ${s.mode}: ${describeDecision(d)}`;
80}
81
82/** Runs one `/spatz` argument line against the settings and returns the answer. `status` is the caller's job. */
83export function change(s: Settings, args: string): string | null {
84	const [key, value, extra] = args.trim().split(/\s+/);
85	if (extra !== undefined) return null;
86	if (key === "mode" && (MODES as readonly string[]).includes(value ?? "")) {
87		s.mode = value as Mode;
88	} else if (
89		key === "scope" &&
90		(SCOPES as readonly string[]).includes(value ?? "")
91	) {
92		s.scope = value as Scope;
93	} else if (
94		key === "record" &&
95		(RECORDS as readonly string[]).includes(value ?? "")
96	) {
97		s.record = value as RecordSetting;
98	} else if (key === "main" && (value === "on" || value === "off")) {
99		s.main = value === "on";
100	} else return null;
101	return `spatz: ${key} is now ${value}`;
102}
103