SLOPSHOPPER

fleet

Wakes this session with a [fleet] turn when a watched container settles, as pi's fleet monitor does

newspinnerguardcommandtoasttool
★ 12v0.1.0MITupdated 2026-10-09jaqubowsky/fleet/claude/mods/fleet
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · fleet
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /fleet-watch ⎿ fleet: fleet: watching every container ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts
README

fleet

Give your coding agent a list of issues and get pull requests back. fleet runs each issue in its own sandbox with its own agent, all at once, and asks you only when a decision is yours.

You only talk to one agent, the host, on your Mac. It starts a sandbox per issue, keeps an eye on all of them and checks the finished work before anything lands. Questions it can't answer come to you. Every command any agent runs passes a tested guard first.

<img alt="Without fleet you watch three terminals, each waiting on you. With fleet you talk to one host agent, and it runs the three sandboxes." src="docs/before-after.svg">

license: MIT macOS: Apple Silicon works with: Claude Code works with: pi

How it works

<table> <tr> <td width="50%" valign="top"> <h4><code>1</code> You say what to ship</h4> <img alt="You tell the host agent: ship 12, 14 and 15" src="docs/step-1.svg" width="100%"> </td> <td width="50%" valign="top"> <h4><code>2</code> Each issue gets a sandbox</h4> <img alt="The host starts three sandboxes, one per issue" src="docs/step-2.svg" width="100%"> </td> </tr> <tr> <td width="50%" valign="top"> <h4><code>3</code> It asks only when it has to</h4> <img alt="Sandbox 14 asks which API to use; you answer v2" src="docs/step-3.svg" width="100%"> </td> <td width="50%" valign="top"> <h4><code>4</code> Pull requests come back</h4> <img alt="Three pull requests, each with tests passed and review done" src="docs/step-4.svg" width="100%"> </td> </tr> </table>

Quickstart

You need a Mac with Apple Silicon and zsh, Docker Sandboxes, herdr, Node, Claude Code and pi, and a Claude plan and a model provider for pi (setup step 1 matches the models to what you have).

git clone https://github.com/jaqubowsky/fleet ~/fleet
cd ~/fleet && claude

Then tell your agent: "read SETUP.md and set me up". It checks what you have, asks what it can't know, and ends by starting and stopping one test sandbox.

After setup you work from herdr: open a tab in your repository and run plain claude or pi. The host only wakes for the sandboxes its own herdr tab started.

Rather do it by hand? Follow the same steps yourself: SETUP.md lists them, one file each.

Inside one sandbox

Every sandbox agent follows the same run. It works out the task, cuts it into small tickets and builds each one test first, committing only after the checks pass.

Then a second agent reviews the change whenever it reaches beyond its own feature. It never saw the implementation, so it reads the diff the way a stranger would, for bugs and for code quality. When the change is something users see, the sandbox opens the running app in a real browser and takes one screenshot per acceptance criterion. It records a video walkthrough when you ask for one.

Once the pull request is open, tell the sandbox to babysit it. It answers review comments and fixes red checks, round after round.

<img alt="The sandbox analyzes, cuts tickets, writes tests and code per ticket, gets a review from a fresh agent and checks the running app. It hands the host screenshots, a summary and the diff; the host reads those, not the chat." src="docs/sandbox-to-host.svg">

You don't watch terminals

The host sleeps until a sandbox needs something, then wakes up with what changed.

<img alt="The host wakes when a sandbox finishes, asks a question, stalls, hits the usage limit or fills its context, and handles each" src="docs/host-watch.svg">

Every command passes a guard

Every tool call an agent makes goes through a policy first. This is what an agent gets back when it tries to rewrite history:

<img alt="An agent runs git push --force origin main and the guard stops it: a force, delete or mirror push rewrites what other people already hold" src="docs/guard-refusal.svg">

StoppedExample
Rewriting shared historygit push --force, delete and mirror pushes
Reading secretsSSH keys, the keychain, op read, gh auth token
Changing its own ruleswrites to ~/.claude and ~/.pi
Deleting your workrm -rf on home and project folders
Acting on GitHub for youmerging a PR in another repository

Commands

The host agent runs these for you. Each one is listed with what it changes, so nothing happens that you can't look up.

CommandWhat it doesWhat it changes
./sync.shshows what --apply would change for each agent whose CLI is on PATH, and names the backup archive it would writenothing
./sync.sh --applyinstalls the harness for each agent whose CLI is on PATH and builds its sandbox imagefirst packs every file it will change into ~/.fleet/backups/<UTC timestamp>.tar.gz; replaces ~/.claude/{rules,refs,skills,agents} and ~/.pi/{skills,agent/refs,agent/agents,agent/themes} of the agents it sets up; writes keys into ~/.claude/settings.json and links the guard hook into ~/.claude/hooks with claude; links fleet and a gh wrapper into ~/.local/bin and herdr's config
fleet up <label> --repo <path>starts a sandbox for a task, its agent waiting in a herdr tabcreates a sandbox with a private clone and your repository's ignored .env files; stores the profile's GitHub token as an sbx secret; adds a task folder under ~/.fleet/tasks/
fleet stop <sandbox>stops the sandbox, keeping its files and herdr tabends guest processes; the tab and watch status show stopped
fleet start <sandbox>restarts in the saved tab with the last saved session when one existsstarts the sandbox and agent; sends no prompt
fleet steer <sandbox> "<text>"sends the sandbox agent its next instructionnothing outside the sandbox
fleet watchwakes the host when a sandbox needs itnothing
fleet ls, fleet peek <sandbox>list the sandboxes, show what one is doingnothing
fleet diff [<sandbox>]toggles a live diff beside the current container's terminalopens or closes its Herdr pane; saves the selected file and scroll position in the host cache; reads Git without changing the checkout
fleet history, fleet artifactsreplay a task's status, list task foldersnothing
fleet exec <sandbox> -- <command>runs a command inside a sandboxwhatever that command changes there
fleet copy <src> <dst>copies a file in or out of a sandboxthe destination file
fleet handoff <sandbox>starts a fresh session in a sandbox that asked for onethe sandbox agent's session
fleet land <sandbox> [--push]brings the finished branch homemoves your local branch; signs commits per profile, which changes their SHAs; --push pushes to GitHub
fleet down <sandbox> [--force]closes the sandbox, keeping its task folderremoves the sandbox; refuses unlanded work, which --force throws away
fleet profile [<repo>] [--apply]shows who may push, open and merge pull requests, per repositorywith --apply: the checkout's git config (signing, HTTPS origin, credential helper) and its Linear MCP registration. Every host session start runs this on its own checkout
fleet tokens [set <name>]lists the GitHub tokens the profiles name and marks the ones missing from the macOS keychain; set asks for one and stores itwith set: one keychain item under fleet-gh. Run it yourself; the guard refuses it to agents
`fleet build [--pi\--claude]`rebuilds a sandbox imagethe local sbx image
fleet init <repo>lays out AGENTS.md and spec/vision.mdadds those files to the repository, never overwriting one
fleet renderrenders one seat's filesthe seat's home, or --out <dir>

fleet --help lists every flag.

In a container's Herdr tab, press Ctrl+B, then D to toggle its Changes pane. It works with both pi and Claude and leaves focus in the agent terminal. The host needs delta on PATH. Reload Herdr's configuration after installing the binding with herdr server reload-config.

The pane shows one continuous diff with a heading before each file. Long lines wrap. The mouse wheel scrolls through every file; clicking the file list jumps to that file without hiding the rest. With keyboard focus in the pane, [ and ] jump between files, j and k or the arrow keys scroll, Page Up and Page Down scroll by a page, and Home and End jump to the beginning and end. s switches between branch changes and uncommitted changes, and Tab selects a repository. The pane reads changes every two seconds after the previous read finishes, retaining the current file and its scroll offset. Branch changes compare the working tree to the merge base of HEAD and the container's origin base branch, so rebasing does not include unrelated upstream changes. The header names the base. No fetch runs while viewing. p pauses reads, r refreshes once, and q closes the pane. Closing stops its reads and reopening restores the file and scroll position. New untracked files are included; ignored files are excluded. A stopped or removed container leaves the last displayed snapshot and its status, without starting the container. Files larger than 4 MiB show a size notice instead of their contents.

Use fleet stop to release a sandbox's resources without deleting its work. fleet start restores the saved conversation in the same tab, not an interrupted tool call or app server. The task's status.md stays unchanged. fleet ls, fleet peek and the watch leave stopped sandboxes asleep; steer, handoff and exec require a start first. Raw sbx exec starts a stopped sandbox, even for a read.

Also in the box

  • Several repositories in one task. Repeat --repo, and one land brings all of them home.
  • Permissions per repository. fleet profile shows who may push, open and merge pull requests, for the host and for the sandbox.
  • Cost per task. fleet ls shows what each sandbox has spent so far.
  • A record of every task. Plan, review and logs stay in a task folder after the sandbox is gone, and fleet history replays how its status changed.
  • Session retrospectives. audit-harness reads a selected session, defaults to the current one, and proposes the smallest environment changes for observed friction and mistakes. Each proposal includes evidence and cost; nothing is edited.
  • Your phone as a remote. Drive pi sessions from your phone over Tailscale, set up as extensions/pi-remote describes.

Trust model

Sandboxes never hold your SSH or signing key, and their GitHub token reaches them only through the sandbox proxy. The host agent on your Mac gets a repository's token from the keychain only for the gh command that needs it. A repository's ignored .env files are copied into its sandbox. Three things never happen without you:

  • a force, delete or mirror push, from any seat;
  • a pull request opened or merged where the repository's profile doesn't give the host auto, since the guard refuses it;
  • a push or a signature with your key, since the key waits for your Touch ID on the Mac.

The guard matches patterns and doesn't understand the shell, so eval gets past it. A repository's own .claude/settings.json can also switch off user hooks. It stops mistakes. Someone who has read the rules can get around it.

Make it yours

Rules, skills and the guard are written once in this repo and rendered for both Claude Code and pi. To add a skill, drop a SKILL.md into skills/shared, skills/host or skills/container and run sync. It reaches each agent sync sets up, on your Mac, in the sandboxes or both. Edit or delete the bundled skills the same way. Keep them in the repo, because sync replaces ~/.claude/skills on every run.

This is my setup

It is opinionated and built around how I work. Fork it and let your agent bend it to yours.

License and warranty

MIT. The software comes as is, without warranty of any kind, and the authors are not liable for anything it does.

fleet runs AI agents that execute commands on your Mac and in sandboxes, with your GitHub tokens and, where a repository's profile allows, your permission to push and merge. The guard stops known mistakes, not every one (see the trust model above). Read fleet profile for each repository before its first task, scope every token to what that repository needs, and keep your own backups.

Third-party code and its licenses: THIRD_PARTY_NOTICES.md.

Source 2 files
hooks/register.tsx 154 lines
1import { atom, read, update } from "claude-code";
2import type { EngineInterface, Register } from "claude-code";
3
4type Line = { wake?: string; log?: string; watching?: string };
5
6const WATCH_DESCRIPTION =
7	"Watch containers beyond the ones this herdr pane put up or steered, which are watched by themselves. A watched container settling, stalling or failing wakes this session with a [fleet] <agent>: <change> turn. Pass the sandbox names fleet ls prints; an empty string watches every container.";
8
9const AGENT_STATUS: Record<string, string> = { working: "yellow", idle: "green", done: "green", blocked: "red", exited: "red", gone: "red" };
10
11const watching = atom({ plugin: "fleet", key: "watching" } as const, "");
12
13const watcher = {
14	stop: undefined as (() => void) | undefined,
15	running: false,
16	held: [] as string[],
17	said: new Set<string>(),
18};
19
20function send($: EngineInterface, text: string) {
21	$.ui.toast(text.split("\n", 1).join(""));
22	void $.prompt.submit({ text });
23}
24
25function heard($: EngineInterface, line: Line) {
26	const containers = line.watching;
27	if (containers !== undefined) void update($, watching, () => containers);
28	if (line.log !== undefined && !watcher.said.has(line.log)) {
29		watcher.said.add(line.log);
30		$.ui.toast(line.log);
31	}
32	if (line.wake === undefined) return;
33	if (watcher.running) watcher.held.push(line.wake);
34	else send($, line.wake);
35}
36
37function follow($: EngineInterface, args: string[]) {
38	watcher.stop?.();
39	watcher.said.clear();
40	const child = $.process.spawn({ argv: ["fleet", "watch", "--json", ...args] });
41	watcher.stop = () => void child.return(undefined as never);
42	void (async () => {
43		let buffer = "";
44		for await (const { stream, text } of child) {
45			if (stream === "stderr") {
46				$.ui.log(text, { to: "debug" });
47				continue;
48			}
49			buffer += text;
50			for (let nl = buffer.indexOf("\n"); nl >= 0; nl = buffer.indexOf("\n")) {
51				heard($, JSON.parse(buffer.slice(0, nl)) as Line);
52				buffer = buffer.slice(nl + 1);
53			}
54		}
55	})();
56}
57
58function watch($: EngineInterface, text: string) {
59	const names = text.trim().split(/[\s,]+/).filter(Boolean);
60	follow($, names.length ? names : ["--every"]);
61	return names.length ? `fleet: watching ${names.join(", ")}` : "fleet: watching every container";
62}
63
64function unwatch($: EngineInterface) {
65	follow($, []);
66	return "fleet: stopped";
67}
68
69export const register: Register = (on) => {
70	on("session.start", async ($, e, next) => {
71		const started = await next(e);
72		follow($, []);
73		await $.tool.register({
74			name: "watch",
75			description: WATCH_DESCRIPTION,
76			inputSchema: {
77				type: "object",
78				properties: {
79					agents: {
80						type: "string",
81						description:
82							"Space- or comma-separated sandbox names as fleet ls prints them. Empty string watches every container",
83					},
84				},
85				required: ["agents"],
86			},
87		});
88		await $.tool.register({
89			name: "unwatch",
90			description: "Stop the watch started by mcp__fleet__watch; the containers this pane put up or steered stay watched.",
91		});
92		await $.command.register({
93			name: "fleet-watch",
94			description: "Watch containers beyond the ones this pane put up or steered; no names watches every container",
95			argumentHint: "[sandbox...]",
96		});
97		await $.command.register({
98			name: "fleet-unwatch",
99			description: "Stop the watch started by /fleet-watch",
100		});
101		return started;
102	});
103
104	on("tool.call", { tool: "mcp__fleet__watch" }, ($, e) => {
105		const text = watch($, String((e as { agents?: unknown }).agents ?? ""));
106		return { result: text, text };
107	});
108	on("tool.call", { tool: "mcp__fleet__unwatch" }, ($) => {
109		const text = unwatch($);
110		return { result: text, text };
111	});
112	on("command.run", { command: "fleet-watch" }, ($, e) => ({ text: watch($, e.args) }));
113	on("command.run", { command: "fleet-unwatch" }, ($) => ({ text: unwatch($) }));
114
115	on("ui.render", { component: "SessionMode" }, async ($, e, next) => {
116		const drawn = await next(e);
117		const containers = await read($, watching);
118		if (!containers) return drawn;
119		const { Box, Text } = $.ui.resolve(e);
120		return (
121			<Box flexDirection="column">
122				{drawn}
123				<Text>
124					<Text color="yellow">◉ </Text>
125					<Text color="magenta">watching  </Text>
126					{containers.split(" · ").map((entry, i) => {
127						const cut = entry.lastIndexOf(" ");
128						const state = entry.slice(cut + 1);
129						return (
130							<Text key={entry}>
131								{i ? <Text dimColor> · </Text> : ""}
132								{`${entry.slice(0, cut)} `}
133								<Text color={AGENT_STATUS[state] ?? "magenta"}>{state}</Text>
134							</Text>
135						);
136					})}
137				</Text>
138			</Box>
139		);
140	});
141
142	on("turn.start", async ($, e, next) => {
143		watcher.running = true;
144		return next(e);
145	});
146	on("turn.complete", async ($, e, next) => {
147		const result = await next(e);
148		if (e.agentId) return result;
149		watcher.running = false;
150		if (watcher.held.length) send($, watcher.held.splice(0).join("\n\n"));
151		return result;
152	});
153};
154
types/index.d.ts 8 lines
1export type FleetWatching = string;
2
3declare module "claude-code" {
4	interface PluginState {
5		fleet: { watching: FleetWatching };
6	}
7}
8