SLOPSHOPPER

sandcastle

Shows a live sandcastle run in the Claude Code session: a row above the prompt, a notice when a ticket needs a person, and a prompt when the run ends; between…

newbandcommandtoaststatusprompt
★ 1v0.1.0MITupdated 2026-10-09henkisdabro/sandcastle-kit/mod
A shopper browsing a rack in a slop shop
README

<a href="https://henkisdabro.github.io/sandcastle-kit/"><img src="docs/img/logo.svg" alt="sandcastle-kit - turn your ticket backlog into a software factory" width="720"></a>

<img src="docs/img/status.gif" alt="sandcastle status playing an example night: thirteen tickets move through setup, implement, review and gates, land one by one as they go green, and end with nine merged and two held for a person" width="880">

Unattended coding agents burn down your queue of GitHub issues or ticket files in Docker sandboxes - implemented, reviewed, gated and merged while you are away from the keyboard.

Built on Sandcastle Licence: MIT Secret scan Release

Claude Code Codex Docker Node TypeScript pnpm

Website · Who it is for · Quick start · Why · Install details · Set up a project · Mark tickets ready · Trackers · Run · After a run · Herdr · Updating · Safety · Troubleshooting

[!NOTE] sandcastle-kit is an opinionated ticket-burndown kit built on Sandcastle by Matt Pocock. Sandcastle does the heavy lifting - sandboxing, worktrees and agent orchestration. This kit is one way to put it to work, learned from and improved on with gratitude.

Sandcastle runs coding agents in Docker sandboxes. This kit turns it into a ready-made loop for any repository whose tickets are GitHub Issues or Markdown files in the repo: it picks up tickets marked ready-for-agent, has one agent implement each ticket and a stronger one review it, runs your project's own lint/build/test gates itself (an agent never gets to say "tests pass"), merges what is green and closes the ticket. One install serves all your projects; each project adds a small config file.

🙋 Who is this for?

In one line: developers who already work with Claude Code or Codex, keep their to-do list as tickets, and want the well-specified ones done while they are away from the keyboard.

It is a good fit if you:

  • 🧑‍💻 are a solo developer, freelancer or small team with a backlog of clear, small-to-medium tickets - bug fixes, small features, refactors, test coverage, dependency and docs chores;
  • 🗂️ track work in GitHub Issues, or as Markdown files in the repo (including the .scratch/ layout from Matt Pocock's skills, if you use them - you do not have to);
  • ✅ have real gates - lint, typecheck, tests, a build - that a machine can run. They are how the kit knows work is good, so the better your tests, the more you can hand over;
  • 🖥️ have a Mac with OrbStack or Docker Desktop, or a Linux machine with rootful Docker Engine, and a Claude subscription or API key;
  • 🌙 are happy to come back to a report and review what merged, rather than watch it happen.

It is probably not for you if:

  • your tickets are vague ("make it faster", "improve onboarding") - an unattended agent with no chat context needs a closed spec, and the queue step helps you write one;
  • the repo has no automated checks - nothing then stands between an agent's guess and your base branch;
  • you want an agent pair-programming with you live (use Claude Code directly), or work that needs a production system, a device or a secret the sandbox must not have;
  • your tickets live only in Linear, Jira or another tool. Linear issues can be blockers today, and the tracker interface is built for more, but only GitHub and ticket files are queues so far.

What a day looks like: in the morning you write or triage a handful of tickets and mark them ready. You start a run and go about your day. When you are back there is a report: which tickets merged into your local base branch (nothing is ever pushed), which are red and why, and which need you. You review and push what you like, land a branch you fixed or requeue a ticket with a note, and repeat.

Do I need anyone else's tools? No, but one is recommended. The kit needs only Docker, Node, pnpm, git, jq and, for GitHub tickets, the gh CLI. It also supports the workflow from Matt Pocock's skills - the same author as Sandcastle, which this kit is built on - and we recommend installing them: /grill-with-docs, /to-spec and /to-tickets are a good way to turn an idea into tickets an unattended agent can finish, and /triage keeps incoming tickets in shape. Run his /setup-matt-pocock-skills in your repo once and the kit reads what it wrote (which tracker, which label), with no extra configuration. It stays optional: skip his skills and one line in config.ts does the same job.

📖 The words used here

New to GitHub or to agents? These are the only terms you need.

WordPlain meaning
TicketOne piece of work, written down: a title, a description of what should change, and how you will know it is done. On GitHub a ticket is an issue; in the files tracker it is a Markdown file.
GitHub issueGitHub's built-in ticket: a page in your repository (the Issues tab) with a title, a description and a comment thread.
LabelA coloured tag you stick on a GitHub issue, such as bug or ready-for-agent. Labels are how this kit knows which tickets to work on: it only touches the ones carrying its queue label.
ready-for-agentThe kit's default queue label. It means "an agent may do this without asking me anything". Putting it on a ticket is you saying yes. How to add it.
QueueEvery ticket currently marked ready-for-agent. A run works through it.
GateA command that must pass before anything merges: your linter, type checker, tests, build. You list them once in config.ts.
SandboxA throwaway Docker container in which one agent works on one ticket, on its own git branch, so it cannot touch your machine or other tickets.
Base branchThe branch (usually main) that green work is merged into, locally. Nothing is pushed.
ready-for-humanAdded by the kit when a change is risky or a ticket cannot be finished unattended. It takes the ticket out of the queue until you look. (Earlier versions called it needs-human; a ticket carrying that is still held.)
needs-triagePut by the kit on the follow-up tickets it files from the agents' <followup> lines (a label on GitHub, a Status: in ticket files). Never queued by itself: the closing summary lists them for you to triage.

[!TIP] AI coding agent? Start at the section written for you, then run sandcastle doctor.

⚡ Quick start

You need: macOS or Linux, Node 22.18+, pnpm, git, jq, ps (procps, on Linux), the GitHub CLI signed in (gh auth login; skippable if your tickets are files in the repo), and a container runtime - OrbStack on macOS (Podman is untested there), rootful Docker Engine on Linux, run as a normal user in the docker group (rootless Docker and Podman on Linux are not supported yet). Full requirements.

1. Install the kit (once per machine) - one line, paste it into your terminal:

git clone https://github.com/henkisdabro/sandcastle-kit.git ~/sandcastle-kit && cd ~/sandcastle-kit && pnpm install && ./bin/sandcastle setup

setup puts the sandcastle command on your PATH, installs the agent skill, walks you through your Claude and GitHub tokens, and finishes with a health check. Re-run it any time. (What it does, and doing it by hand.)

2. Point it at a project:

cd ~/code/your-project
sandcastle init              # or ask your agent: /sandcastle init
sandcastle build

init fills in the gate commands (lint, test, ...) from your stack; check them against CI in .sandcastle/config.ts - see Set up a project.

[!TIP] Recommended, optional: install Matt Pocock's skills too and run /setup-matt-pocock-skills in your project. The kit understands the tracker and labels it records, and his /grill-with-docs, /to-tickets and /triage are a good way to write the tickets this kit works through. Nothing here requires them.

3. Mark some tickets ready-for-agent - on GitHub that is a label you add to an issue (step by step, no experience needed); in ticket files it is a Status: line (Trackers) - and try a dry run:

DRY_RUN=1 sandcastle run     # implement, review, gate - never merge or close

💡 Why sandcastle-kit?

You already drive Claude Code or Codex every day - maybe with subagents, maybe several sessions side by side. That still needs you in the loop: approving prompts, watching terminals, copying context between chats. sandcastle-kit is the next step up.

  • 🎫 Your tickets stay where they are. GitHub Issues or Markdown files in the repo: write a ticket an agent could finish with no chat context, mark it ready, and it joins the queue. No new tool to learn.
  • 🏭 A software factory per repository. Each queued ticket gets its own git worktree and its own Docker sandbox, several at once, across several projects.
  • 🧪 Proof, not promises. The orchestrator - not the agent - runs your gates. Only green branches merge, and the merged branch is gated again.
  • 🔒 Safe to leave unattended. Agents run with permission prompts off, but inside a container with a token that cannot push, and risky changes wait for a human.
  • ☕ You come back to a report: what merged, what is red, what needs you. Nothing is pushed until you push it.
flowchart LR
    A["📝 You write tickets<br/>GitHub or repo files"] --> B["🏷️ Label<br/>ready-for-agent"]
    B --> C["🚀 sandcastle run"]
    C --> D1["🐳 Sandbox<br/>ticket #12"]
    C --> D2["🐳 Sandbox<br/>ticket #15"]
    C --> D3["🐳 Sandbox<br/>ticket #18"]
    D1 --> E{"🚦 Your gates<br/>green?"}
    D2 --> E
    D3 --> E
    E -- "yes" --> F["✅ Merged locally<br/>ticket closed"]
    E -- "no" --> G["🔴 Left red<br/>for you"]
    E -- "risky paths" --> H["🙋 ready-for-human"]
    F --> I["📊 Report<br/>you review and push"]
    G --> I
    H --> I

    classDef you fill:#dbeafe,stroke:#2563eb,color:#1e3a8a
    classDef kit fill:#fef3c7,stroke:#d97706,color:#78350f
    classDef ok fill:#dcfce7,stroke:#16a34a,color:#14532d
    classDef bad fill:#fee2e2,stroke:#dc2626,color:#7f1d1d
    class A,B,I you
    class C,D1,D2,D3,E kit
    class F ok
    class G,H bad

✨ What it adds to Sandcastle

FeatureWhat you get
🚦Gates run by the orchestratorYour lint/build/test, run after the agents, in the sandbox. Only green branches merge, and the merged base branch is gated once more.
🧱A green base firstEvery gate runs on the base commit in the image before any agent starts. A gate red there would be red on every branch, so the run stops before it spends anything.
🩹Repair on redA red gate gets one repair pass on the same warm sandbox, fed the gate's own output, then the gates run again - and up to two more while each pass turns up a failure the last one did not see. A branch a repair turned green gets the review pass again, on the repair commits, before it can land. A pass that commits nothing (the repairer judged the red a flake) is not counted: the run's per-ticket line says repair made no change, and repaired=N counts only passes that committed. A failure that is red on the base too, in a test file the branch did not change, gets no repair: the gates run once on the base's tip, and the closing summary names the test once under Needs you.
🗂️Your trackerTickets are GitHub Issues (the default) or Markdown files in the repo - in the layout Matt Pocock's setup skill uses, so a repo that ran it works with no extra config.
🔗Ticket dependenciesBlocked by #12 in a ticket body holds it back until #12 is closed. It can also wait on a Linear issue (ENG-42) or an in-repo task file - see Blockers.
🧑‍💻Implement, then reviewClaude Sonnet 5.5 implements and tests at the ticket's seams, Claude Opus 5.5 reviews the tests as well as the code, on the same warm sandbox; a failed review falls back to the implementer's model. A model: or effort: label gives one ticket a different implementer. Optional third review by an OpenAI model through Codex (CROSS_REVIEW=1).
🪶Lean sandboxesThe project's skills, subagents, commands and MCP servers are hidden from sandbox agents unless you keep them, and its plugins always, because each one costs context on every turn.
🪝Hooks enforcedThe project's Claude Code hooks are kept, and checked to be runnable in the image before any sandbox starts.
🐳One base image, a layer per projectRebuilt only when a Dockerfile or the Claude Code or Codex version changes - see The image's agent versions.
🛫PreflightOne short reply from every model before any sandbox starts, so an exhausted plan or a too-old CLI stops the run up front instead of halfway through.
♻️Picks up where it stoppedA ticket re-run on its old branch builds on it instead of starting again, and a branch that was already green costs only the gates - see Re-runs. A killed run's leftover sandboxes are stopped by the next run or sandcastle clean.
🛬Landing you controlEach ticket lands as soon as its gates are green, while the others still run, as a merge commit or squashed (land); a conflict or a red merge is sent back once for a second try in the same run (a conflict's resolve waits for the green branches that share its files to land first), and a ticket whose last blocker lands starts at once. Between runs, sandcastle preview shows which unlanded branches would conflict, sandcastle land <n> merges and gates one by hand, and sandcastle requeue <n> sends one back with a note. autonomy lets one run take further turns by itself, up to drain.
🔒Host safetyFine-grained tokens only, host git hooks off during a run, the shared .git fingerprinted, risky branches held for a human merge (see Safety model).
⚖️Machine-wide limitsSeveral projects can run at once without starving each other.
📺A live status viewsandcastle status in any terminal: every ticket of the run and where it is - working, ready to land, needing you, queued, blocked, merged - which gate is running, and when landing should start.
🖥️Best in HerdrThe run opens its own tab with the status view (the next run reuses that view wherever you put it), and rolls the run up in Herdr's sidebar (a pane per sandbox is opt-in) - see Works best in Herdr.
🔭A Claude Code modOptional: the session that started the run shows it above the prompt, says when a ticket needs you, and is told when the run ends - see The Claude Code mod.
🧩An agent skill/sandcastle in Claude Code, $sandcastle in Codex, also read by OpenCode - for setup, auditing a repo for work, ticket triage, starting and closing runs, and updating.

🔄 How it works

A run starts in a project, on its base branch, with a clean tree. It checks everything first, then fans out one sandbox per ticket, up to CONCURRENCY at once and within the machine-wide limit.

flowchart TD
    S(["🚀 sandcastle run"]) --> P["🔍 Checks<br/>token type · images (build if stale) · preflight<br/>prompts · hook check · git fingerprint"]
    P --> Q["📋 Queue: open tickets labelled ready-for-agent"]
    Q --> B{"🧱 Every gate on the base commit<br/>(skipped if green before at this commit)"}
    B -- "red" --> STOP(["⛔ Stop - no agent started"])
    B -- "green" --> W

    subgraph W ["🐳 Per ticket - git worktree + Docker sandbox on branch agent/issue-N"]
        direction TB
        L["🪶 Lean strip · lock worktree · setup install"] --> I["🧑‍💻 Implement agent<br/>commits on the branch"]
        I --> R["🔎 Review agent<br/>fixes on the branch<br/>(fallback model if it fails)"]
        R -.-> X["🤖 Codex review<br/>optional, non-blocking"]
        R --> G{"🚦 Gates<br/>run by the orchestrator"}
        X -.-> G
        G -- "red" --> FX["🩹 Repair agent<br/>fed the gate output<br/>(bounded; green again<br/>means a second review)"]
        FX --> G
    end

    G -- "green" --> PP{"🛡️ Touches hooks, CI,<br/>install scripts?"}
    G -- "still red" --> RED["🔴 Reported red"]
    PP -- "no" --> M["✅ Merge (or squash) into base<br/>close ticket with a comment"]
    PP -- "yes" --> NH["🙋 Labelled ready-for-human<br/>not merged"]
    M --> V["🔁 Verify: gates once more<br/>on the merged base branch"]
    V --> REP(["📊 Report: merged · red · held back<br/>Nothing is pushed"])
    RED --> REP
    NH --> REP

    classDef ok fill:#dcfce7,stroke:#16a34a,color:#14532d
    classDef bad fill:#fee2e2,stroke:#dc2626,color:#7f1d1d
    classDef agent fill:#ede9fe,stroke:#7c3aed,color:#3b0764
    classDef gate fill:#fef3c7,stroke:#d97706,color:#78350f
    class M,V ok
    class RED,NH bad
    class I,R,X,FX agent
    class G,PP,P gate

A green gate run proves only that the configured gate commands passed on that branch - no more than those commands check. Whether the change does what the ticket asked is checked by the review agent, not the gates, which is why every branch is reviewed before it is gated and a repaired branch is reviewed again. A live run rarely reaches the repair path, because agents run the gates themselves before they finish. SANDCASTLE_TEST_RED_GATE=1 (see the Configuration table) counts each ticket's first gate run as red to exercise it, at the cost of one repair pass per ticket.

The gates also run in the Linux sandbox, so a green run proves Linux only. A branch can still be red on macOS or Windows (BSD tools, bash 3.2, shell shims and terminal flags differ), and landing does not check that: it gates in the same Linux sandbox. A project that ships to macOS or Windows can add a CI job on that OS, or run its gates on the host before pushing.

🧱 Set up a project

From the project's root, on its base branch:

sandcastle init      # writes .sandcastle/config.ts, rules.md, .gitignore; prints the lean check

init reads the stack from the repo root - package.json (with its lockfile and scripts), pyproject.toml + uv.lock, go.mod or Cargo.toml - and fills in the gates and setup from it. Where the base image lacks the toolchain (uv, Go, Rust, Bun) it also writes .sandcastle/Dockerfile. Treat what it writes as a starting point. Anything else, poetry and pipenv included, gets a placeholder gate that fails until you copy in what CI runs.

Then edit, in this order:

  1. Where do your tickets live? Nothing to do for GitHub Issues. For Markdown tickets set tracker: "files" in config.ts (or let the kit read docs/agents/issue-tracker.md if you ran Matt Pocock's setup skill, recommended but optional). sandcastle queue shows which tracker the kit chose and why; see Trackers.
  2. .sandcastle/config.ts - gates (the commands CI runs: lint, typecheck, build, test), setup (dependency install in the sandbox), mounts (e.g. a host cache; pnpmStore: true mounts the host's pnpm store), lean. See Configuration. Gates run in the Linux sandbox, so green proves Linux only: if the project ships to macOS or Windows, add a CI job on that OS or run the gates on the host before pushing (see How it works).
  3. .sandcastle/rules.md - what an agent in this repo must read first, must never do (deploys, production databases), and how a visual or data change is proven. It is added to the implement, review and repair prompts, and reaches only those agents: the landing merge is the kit's own git merge, so rules do not reach it (see a branch conflicts at landing). init leaves three questions there for you to answer - which committed files a command writes (also generated in config.ts), which paths an agent must never touch, and which gate catches drift in generated files (see A gate for generated files) - and the /sandcastle skill's init action asks them for you.
  4. .sandcastle/Dockerfile - only if the gates need something the base image lacks (browsers, Python tooling, a pinned package manager). Start from templates/Dockerfile in the kit, then name it in .sandcastle/config.ts with dockerfile: ".sandcastle/Dockerfile": without the key the file is never built, and sandcastle build skips it with a "Not built" line (init adds the key itself when it writes the Dockerfile). A project that tests in Chromium (Playwright, Puppeteer) should launch it with --disable-dev-shm-usage: a sandbox gets the runtime's default /dev/shm, 64 MB on Docker (some runtimes give more), and a heavy page can crash there (Navigation failed because page crashed!) on the branch and on the base alike. The kit passes no --shm-size, since the docker run builder of @ai-hero/sandcastle has no such option.
  5. Lean and hooks - sandcastle lean lists what the repo would load into each sandbox and checks every kept hook. Keep nothing unless a run needs it; drop only host-only hooks. It also lists the project's permissions.ask rules (in the tracked .claude/settings.json), which sandboxes refuse - nobody there can answer - so a command matching one fails; sandcastle doctor warns of them too.
sandcastle build             # base image, then the project layer
sand
Source 6 files
hooks/register.tsx 587 lines
1// The sandcastle mod: shows a live run in the Claude Code session, in the status view's own
2// castle, glyphs and colours - a band above the prompt, a notice when a ticket needs a person -
3// and submits a prompt when the run ends, so the session that started it closes it. Between runs,
4// in a project `sandcastle init` has set up, it draws a quiet idle mark in sand above the prompt.
5//
6// It reads `.sandcastle/logs/run.json` under the session's root and asks whether the run's
7// process is alive. Once the session has used the sandcastle skill it also follows a run this
8// session started in another directory: every run records the id of the Claude Code session
9// that started it, and lists itself in a machine-wide directory while it lives. It writes
10// nothing, opens no pane and calls no model, and it does nothing at all in a project with no
11// `.sandcastle/` until the skill is used.
12
13import { atom, read, update } from "claude-code";
14import type { EngineInterface, Register } from "claude-code";
15
16import { band, building, CASTLE_FRAMES, endedHow, endPrompt, followable, HELD, line, needing, parse, parseRegistry, REGISTRY_SCRIPT, rows, type Run, SAND, startedBy, summarise } from "./run-state.ts";
17import { kitRunning } from "./run-live.ts";
18import { afterRead, type Choice, choiceAfter, dismissalEnded, due, MARK_USAGE, machineSwitch, markAction, markReport, markText, type MarkInput, parseChoice, parseEntry, readyIds, SETTINGS_SCRIPT, type Trigger } from "./idle.ts";
19
20const view = atom({ plugin: "sandcastle", key: "view" } as const, null);
21/** The castle frame the band draws: an index into CASTLE_FRAMES. */
22const castle = atom({ plugin: "sandcastle", key: "castle" } as const, HELD);
23/** The castle tower (a chess rook: one cell, text style, no emoji form) that leads the idle mark's row. */
24const MARK_ICON = "♜";
25/** The idle mark's line the band draws between runs, in sand; null for none. */
26const markLine = atom({ plugin: "sandcastle", key: "mark" } as const, null);
27
28const RECORD = ".sandcastle/logs/run.json";
29// What makes a project set up: `sandcastle init` writes it.
30const CONFIG = ".sandcastle/config.ts";
31const LIVE_MS = 3000;
32// A read of the queue is given up after this long, so one stuck tracker call never blocks the next.
33const READ_TIMEOUT_MS = 20000;
34// With no run alive the record is read this often: a run started by hand shows within it.
35const IDLE_MS = 15000;
36// Room for the control Claude Code draws at the band's right end.
37const BAND_MARGIN = 4;
38
39const NOTE =
40  "\n\n---\nThe sandcastle mod is loaded in this Claude Code session. When a run this session started ends, in this project or any other directory, the mod submits a prompt that says so: skip `sandcastle wait` in step 3 of the run action, and close the run (run.md) when that prompt arrives.";
41
42// Module variables on purpose: a reload starts the watch over, and the first look at the
43// record rebuilds all of it but `armed` and each project's `since`, which `$.store` keeps.
44let starting: Promise<boolean> | undefined;
45/** The project the watch reads. A `/cd` does not move it. */
46let watched: string | undefined;
47/** The sandcastle skill was used in this session, so this session closes the run. */
48let armed = false;
49/** The session root's own record and store entry are taken up (once `.sandcastle/` exists). */
50let adopted = false;
51/** What one watched project's looks remember. */
52type Seen = {
53  /**
54   * `startedAt` of the last record whose end this session has accounted for. A record that
55   * differs and whose process is gone is a run that ended - whether or not a look ever saw it
56   * alive: a run that dies on a missing Docker lasts seconds, inside one idle look.
57   */
58  since?: string;
59  /** False until the first look: what that finds was there before, so it announces nothing new. */
60  primed: boolean;
61  needs: string[];
62};
63const seen = new Map<string, Seen>();
64const seenAt = (root: string): Seen => seen.get(root) ?? (seen.set(root, { primed: false, needs: [] }).get(root) as Seen);
65/**
66 * Roots, resolved, of runs this session started outside its own root. Kept once seen: the
67 * registry file goes when the run exits (or later, once its Herdr tab has its report), and the end
68 * is still to be announced.
69 */
70const followed = new Set<string>();
71/**
72 * Every id this session has had in this process. `/clear` gives the session a new one, and a run
73 * started before it records the old: the terminal that started the run still closes it.
74 */
75const known = new Set<string>();
76/** What makes the next idle look read the queue ahead of the entry's age; one is enough, they all mean "now". */
77let trigger: Trigger | undefined;
78/** A queue read is under way in this session: one at a time. */
79let reading = false;
80/** A trigger fired during that read, which may have started before what the trigger is about. */
81let again = false;
82/** The sandcastle skill was used and its turn has not ended: the labelling it may do is still to come. */
83let skillTurn = false;
84/** The `sandcastle` to run, found once: the kit's own `bin/sandcastle` beside the mod, else the one on PATH. */
85let kitBin: string | undefined;
86/** The last round drew the idle mark (no needs-you text, no live run of the root): a choice made now redraws it at once. */
87let idling = false;
88let drawn = "";
89/** null: nothing pinned or cleared since this load, so the first call always reaches Claude Code. */
90let pinned: string | undefined | null = null;
91
92/**
93 * Whether `path` is a plain file. A plain file only: in a stranger's repository the path could be
94 * a link to a device that never ends.
95 */
96async function plain($: EngineInterface, path: string): Promise<boolean> {
97  const at = await stat($, path);
98  return at !== undefined && at.kind === "file" && !at.isLink;
99}
100
101/** The one place the mod stats a path; `resolve` also answers where the path lands. Undefined when it cannot be read. */
102async function stat($: EngineInterface, path: string, resolve = false) {
103  try {
104    return await $.fs.stat(path, resolve ? { resolve: true } : undefined);
105  } catch {
106    return undefined;
107  }
108}
109
110async function record($: EngineInterface, root: string): Promise<Run | undefined> {
111  const path = `${root}/${RECORD}`;
112  try {
113    if (!(await plain($, path))) return undefined;
114    return parse(await $.fs.read(path));
115  } catch {
116    return undefined;
117  }
118}
119
120const isProject = ($: EngineInterface, root: string) => $.fs.exists(`${root}/.sandcastle`);
121
122/** The one place the mod starts a process. */
123const exec = ($: EngineInterface, argv: string[], init?: { cwd?: string; timeoutMs?: number }) => $.process.run(argv, init);
124
125// Asks for the process's command line and sends nothing to it; that it is the kit's is
126// run-live.ts's rule, the same everywhere a run is asked about. The record's `finishedAt` is
127// left out on purpose: every turn of one run writes one, a killed run none, and between two
128// turns the run is still going.
129async function alive($: EngineInterface, pid: number): Promise<boolean> {
130  try {
131    // `-p` and `-o command=` are the flags BSD `ps` (macOS) and procps-ng share. BusyBox has no
132    // `-p`: the call fails, so there it is no live run and no false end.
133    const ps = await exec($, ["ps", "-p", String(pid), "-o", "command="]);
134    return kitRunning(pid, () => (ps.exitCode === 0 ? ps.stdout : undefined));
135  } catch {
136    return false;
137  }
138}
139
140type Kept = { session?: string; since?: string } | undefined;
141
142// Kept across a restart of Claude Code: a resumed session still hears that its run ended.
143// One entry per project, so of two sessions that used the skill there the later one owns it.
144async function remember($: EngineInterface, root: string) {
145  const since = seenAt(root).since;
146  await $.store.set(root, { session: await $.session.id(), ...(since === undefined ? {} : { since }) });
147}
148
149/** The session's id, kept in `known`. */
150async function me($: EngineInterface): Promise<string> {
151  const id = await $.session.id();
152  known.add(id);
153  return id;
154}
155
156/** Whether `run` records an id this session has had. */
157const ours = (run: Run | undefined) => [...known].some((id) => startedBy(run, id));
158
159async function owns($: EngineInterface, kept: Kept): Promise<boolean> {
160  return kept?.session === (await $.session.id());
161}
162
163/** The idle mark the band was last told to draw; undefined until the first look of this load. */
164let marked: string | null | undefined;
165
166/**
167 * The idle mark is drawn by the band, not pinned as a status line: Claude Code prefixes every
168 * pinned status line with its warning triangle and paints it in its notice colour, and a mod
169 * cannot change either. The pinned line stays for what needs a person.
170 */
171async function place($: EngineInterface, text: string | undefined) {
172  const line = text ?? null;
173  if (line === marked) return;
174  marked = line;
175  await update($, markLine, () => line);
176}
177
178function pin($: EngineInterface, text: string | undefined) {
179  if (text === pinned) return;
180  pinned = text;
181  $.ui.status(text);
182}
183
184/** The castle's timer while it builds; undefined while it stands still. */
185let tick: { cancel: () => void } | undefined;
186
187// A timer of its own, apart from the look at the record: a frame writes the one atom and reads
188// nothing, and a look never waits for a frame. One write per frame, about one a second while it
189// builds and none while it holds - far inside the band's redraw rate, so the prompt never flickers.
190function animate($: EngineInterface, on: boolean) {
191  if (on === (tick !== undefined)) return;
192  if (!on) {
193    tick?.cancel();
194    tick = undefined;
195    void update($, castle, () => HELD);
196    return;
197  }
198  const show = (i: number) => {
199    void update($, castle, () => i);
200    tick = $.clock.after(CASTLE_FRAMES[i]!.ms, () => show((i + 1) % CASTLE_FRAMES.length));
201  };
202  show(0);
203}
204
205async function draw($: EngineInterface, run: Run | undefined) {
206  const next = run ? summarise(run) : null;
207  animate($, next !== null && building(next));
208  if (JSON.stringify(next) === drawn) return;
209  drawn = JSON.stringify(next);
210  await update($, view, () => next);
211}
212
213/**
214 * One look at one project's record: the live run, if there is one. `follow` is set for a
215 * followed root: its run is this session's by the id it records, so it is never "old" and needs
216 * no store entry; a record that is another session's run, or none, ends the following.
217 */
218async function look($: EngineInterface, root: string, follow = false): Promise<Run | undefined> {
219  const run = await record($, root);
220  const w = seenAt(root);
221  const first = !follow && !w.primed;
222  w.primed = true;
223  if (follow && !ours(run)) {
224    followed.delete(root);
225    seen.delete(root);
226    return undefined;
227  }
228  if (run?.pid !== undefined && (await alive($, run.pid))) {
229    const now = needing(run);
230    const fresh = now.filter((n) => !w.needs.includes(n));
231    // What the first look finds needed a person already: the status line names it, no notice.
232    if (!first && fresh.length) $.ui.toast(`${fresh.join(", ")} ${fresh.length > 1 ? "need" : "needs"} you`, { timeoutMs: 10000 });
233    w.needs = now;
234    return run;
235  }
236  w.needs = [];
237  if (!run?.startedAt || run.startedAt === w.since) return undefined;
238  // Read before `since` moves: a store that cannot be read leaves the end to the next look.
239  // A run that records the session that started it is that session's, wherever it was started
240  // from: the store's "later session owns the project" would have a second session close it too.
241  // A run with no id (a plain terminal, Codex, OpenCode) is the store's.
242  const mine = armed && (follow || (run.session ? ours(run) : await owns($, (await $.store.get(root)) as Kept)));
243  w.since = run.startedAt;
244  // The end is accounted for: a later run of this session in that root is found again.
245  if (follow) {
246    followed.delete(root);
247    seen.delete(root);
248  }
249  // A session that was not waiting for a run meets an old record: nothing ended on its watch.
250  if (first && !armed) return undefined;
251  // The run closed or left tickets: the count is read again at the next idle look. A followed run
252  // too - it may be a second clone of this project, burning down the same tracker.
253  trigger = "run-ended";
254  const how = endedHow(run);
255  if (mine) {
256    // The store keeps the old `since` until the turn this starts has begun, which may be much
257    // later: a session quit with the prompt still queued hears it again when it is resumed.
258    const text = endPrompt(root, run);
259    void $.prompt.submit({ text }).then(() => (follow ? undefined : remember($, root))).catch(() => {});
260  }
261  $.ui.toast(`run ${how}`, { timeoutMs: 10000 });
262  return undefined;
263}
264
265/**
266 * Finds runs this session started outside its own root: the registered live runs whose record
267 * names an id this session has had. Fails closed - a shell or a record that cannot be read adds
268 * nothing.
269 */
270async function discover($: EngineInterface, root: string) {
271  let out: string;
272  try {
273    const ls = await exec($, ["sh", "-c", REGISTRY_SCRIPT, "sh", root]);
274    if (ls.exitCode !== 0) return;
275    out = ls.stdout;
276  } catch {
277    return;
278  }
279  for (const other of followable(parseRegistry(out))) {
280    if (followed.has(other)) continue;
281    const run = await record($, other);
282    if (ours(run) && run?.pid !== undefined && (await alive($, run.pid))) followed.add(other);
283  }
284}
285
286/** The session root's store entry, `/sandcastle-status` and `/sandcastle-mark`, once `.sandcastle/` exists. */
287async function adopt($: EngineInterface, root: string) {
288  if (adopted || !(await isProject($, root))) return;
289  adopted = true;
290  const kept = (await $.store.get(root)) as Kept;
291  if (await owns($, kept)) {
292    armed = true;
293    seenAt(root).since = kept?.since;
294  }
295  await $.command.register({ name: "sandcastle-status", description: "Show the sandcastle run in this project, with no model turn", immediate: true });
296  await $.command.register({ name: "sandcastle-mark", description: "Dismiss the idle mark's ready count, or hide or show the mark, with no model turn", argumentHint: "[dismiss|hide|show]", immediate: true });
297}
298
299/** The shared cache entry's key: one per project root, apart from the session entry under the bare root. */
300const readyKey = (root: string) => `ready:${root}`;
301
302/**
303 * The command that reads the queue. The mod is linked from the kit's checkout, so `bin/sandcastle`
304 * sits beside it; a mod copied elsewhere finds none and uses the `sandcastle` on PATH.
305 */
306async function sandcastle($: EngineInterface): Promise<string> {
307  if (kitBin !== undefined) return kitBin;
308  kitBin = "sandcastle";
309  // No stat of the mod's own folder, or no file beside it: the one on PATH.
310  const dir = ((await stat($, $.plugin.root, true))?.realPath ?? $.plugin.root).replace(/\/+$/, "");
311  const bin = `${dir.slice(0, Math.max(dir.lastIndexOf("/"), 0))}/bin/sandcastle`;
312  if (await plain($, bin)) kitBin = bin;
313  return kitBin;
314}
315
316/** The ready ids from `sandcastle queue --json`, or undefined when the read failed or timed out. */
317async function readQueue($: EngineInterface, root: string): Promise<string[] | undefined> {
318  try {
319    const out = await exec($, [await sandcastle($), "queue", "--json"], { cwd: root, timeoutMs: READ_TIMEOUT_MS });
320    return out.exitCode === 0 && !out.isStdoutTruncated ? readyIds(out.stdout) : undefined;
321  } catch {
322    return undefined;
323  }
324}
325
326/**
327 * Reads the queue into the shared entry, one read at a time in this session; never awaited by a
328 * look or a render, so a slow or hung tracker holds nothing up. A failed read keeps the last good
329 * ids and their time. Another session racing to the same stale entry may read too: accepted.
330 */
331async function refresh($: EngineInterface, root: string, forced: boolean) {
332  if (reading) {
333    // An age that is due again is the read under way; only a trigger asks for another after it.
334    if (forced) again = true;
335    return;
336  }
337  reading = true;
338  try {
339    do {
340      again = false;
341      const ids = await readQueue($, root);
342      const prior = parseEntry(await $.store.get(readyKey(root)));
343      await $.store.set(readyKey(root), afterRead(prior, ids, await $.clock.now()));
344    } while (again);
345  } catch {
346    // A store or clock that cannot be reached leaves the entry as it was: the next look reads again.
347  } finally {
348    reading = false;
349  }
350}
351
352/** The person's choices for a project's mark: the store's entry, none when it holds nothing usable. */
353const markKey = (root: string) => `mark:${root}`;
354
355/** The machine switch: one read of the personal settings, on unless they say `"idleMark": false`. */
356async function machineOn($: EngineInterface): Promise<boolean> {
357  try {
358    const out = await exec($, ["sh", "-c", SETTINGS_SCRIPT]);
359    if (out.exitCode === 0) return machineSwitch(out.stdout);
360  } catch {
361    // Settings that cannot be read leave the switch on.
362  }
363  return true;
364}
365
366/**
367 * Everything the idle mark is decided from that the mod already holds: set up is one `stat` of
368 * `.sandcastle/config.ts`, the machine switch one read of the personal settings, the choices and
369 * the count what the store holds. Nothing here reaches the tracker.
370 */
371async function facts($: EngineInterface, root: string): Promise<MarkInput> {
372  const setUp = await plain($, `${root}/${CONFIG}`);
373  const idleMark = setUp ? await machineOn($) : true;
374  const now = await $.clock.now();
375  const entry = parseEntry(await $.store.get(readyKey(root)));
376  const choice = parseChoice(await $.store.get(markKey(root)));
377  return { setUp, idleMark, hidden: choice?.hidden, dismissed: choice?.dismissed, entry, now };
378}
379
380/**
381 * The idle mark's text for the session root's project, which no followed run changes - a read it
382 * starts is for the next look to show.
383 */
384async function mark($: EngineInterface, root: string): Promise<string | undefined> {
385  const input = await facts($, root);
386  if (!input.setUp || !input.idleMark) return markText(input);
387  if (input.dismissed && dismissalEnded(input.dismissed, input.entry, input.now)) {
388    // A ticket not in the dismissal is ready: the dismissal is over for good, not only while it shows.
389    await $.store.set(markKey(root), { hidden: input.hidden === true });
390    input.dismissed = undefined;
391  }
392  if (!input.hidden && due(input.entry, input.now!, trigger)) {
393    void refresh($, root, trigger !== undefined);
394    trigger = undefined;
395  }
396  return markText(input);
397}
398
399/** `/sandcastle-mark`: the choice is kept, the line redrawn at once if the mark is what it shows, the reply is text. */
400async function choose($: EngineInterface, root: string, args: string): Promise<string> {
401  const action = markAction(args);
402  if (action === undefined) return MARK_USAGE;
403  if (action !== "report") {
404    const input = await facts($, root);
405    const prior: Choice = { hidden: input.hidden === true, ...(input.dismissed ? { dismissed: input.dismissed } : {}) };
406    await $.store.set(markKey(root), choiceAfter(action, prior, input.entry, input.now));
407    if (idling) await place($, await mark($, root));
408  }
409  const input = await facts($, root);
410  const done = { dismiss: "Dismissed: the count stays quiet until a ticket not ready now becomes ready.", hide: "Hidden in this project until /sandcastle-mark show.", show: "Shown." };
411  return action === "report" ? markReport(input) : `${done[action]}\n${markReport(input)}`;
412}
413
414/** One round: every watched project once; true while a run is alive. The newest live run is the one drawn. */
415async function round($: EngineInterface, root: string): Promise<boolean> {
416  await adopt($, root);
417  await me($);
418  if (armed) await discover($, root);
419  const live: Run[] = [];
420  const own = adopted ? await look($, root) : undefined;
421  if (own) live.push(own);
422  // oxlint-disable-next-line no-useless-spread -- a snapshot: `look` deletes from `followed`, and other rounds change it across the awaits
423  for (const other of [...followed]) {
424    const run = await look($, other, true);
425    if (run) live.push(run);
426  }
427  const shown = live.sort((a, b) => String(b.startedAt).localeCompare(String(a.startedAt)))[0];
428  const now = shown ? needing(shown) : [];
429  // The line is the needs-you text while a live run needs a person, and otherwise the idle mark -
430  // unless a run is alive, here or followed elsewhere: its castle and counts take over, and a
431  // mark above them would read as a second, stale queue. Last, so after the end notice: the mark
432  // returns once the run is over.
433  idling = !now.length && !shown && adopted;
434  pin($, now.length ? `${now.join(", ")} - /sandcastle-status` : undefined);
435  await place($, idling ? await mark($, root) : undefined);
436  await draw($, shown);
437  return live.length > 0;
438}
439
440async function watch($: EngineInterface, root: string) {
441  let live = false;
442  try {
443    live = await round($, root);
444  } catch {
445    // A look that failed is tried again at the next one.
446  }
447  $.clock.after(live ? LIVE_MS : IDLE_MS, () => void watch($, root));
448}
449
450/**
451 * Starts the loop. In a project with no `.sandcastle/` it starts only when the session has used
452 * the skill (`used`): the loop then watches runs started from here, and the project's own record
453 * once `sandcastle init` makes it.
454 */
455async function start($: EngineInterface, root: string, used: boolean): Promise<boolean> {
456  if (!used && !(await isProject($, root))) return false;
457  watched = root;
458  // Awaited: what the first look finds is what every later one is compared with.
459  await watch($, root);
460  return true;
461}
462
463/** Starts the watch, once however many callers ask at the same time; false in a project with no `.sandcastle/` unless `used`. */
464function begin($: EngineInterface, root: string, used = false): Promise<boolean> {
465  starting ??= start($, root, used).then(
466    (ok) => {
467      if (!ok) starting = undefined;
468      return ok;
469    },
470    (error) => {
471      starting = undefined;
472      throw error;
473    },
474  );
475  return starting;
476}
477
478export const register: Register = (on) => {
479  on("session.start", async ($, e, next) => {
480    const out = await next(e);
481    await begin($, await $.session.root());
482    return out;
483  });
484
485  // `/clear` and a `/resume` inside Claude Code change the session under the same process and
486  // start no session anew. The arming stays with this terminal, and the store learns the new
487  // session id, so it is still this one's after a restart.
488  on("classic.SessionStart", { source: ["clear", "resume", "fork"] }, async ($, e, next) => {
489    const out = await next(e);
490    if (watched === undefined) return out;
491    if (await owns($, (await $.store.get(watched)) as Kept)) armed = true;
492    else if (armed) await remember($, watched);
493    return out;
494  });
495
496  on("skill.prompt", { skill: "sandcastle" }, async ($, e, next) => {
497    const out = await next(e);
498    const root = await $.session.root();
499    // Whatever the root, the note is true: a run this session starts records the session's id and
500    // is followed wherever it lives, so the end prompt arrives. (`sandcastle init` may have made
501    // the project after this session started; a `/cd` leaves the watch on the first root, which
502    // the id-based follow does not depend on.)
503    await me($);
504    // The skill may just have labelled tickets: the next idle look reads the count. Triaging
505    // labels them well after this hook, so the end of the turn asks for one more read.
506    trigger = "skill";
507    skillTurn = true;
508    await begin($, root, true);
509    await remember($, root);
510    armed = true;
511    return { text: out.text + NOTE };
512  });
513
514  // The turn that used the skill is over: whatever it labelled is labelled, so the count is read
515  // again. A subagent's turn (one with an `agentId`) is not it: triage reads tickets through
516  // subagents before it labels, so their ends would spend the read too early.
517  on("turn.complete", async ($, e, next) => {
518    const out = await next(e);
519    if (skillTurn && e.agentId === undefined) {
520      skillTurn = false;
521      trigger = "skill";
522    }
523    return out;
524  });
525
526  on("command.run", { command: "sandcastle-status" }, async ($) => {
527    const roots = [await $.session.root(), ...followed];
528    const blocks: string[] = [];
529    for (const root of roots) {
530      const run = await record($, root);
531      if (!run) continue;
532      const live = run.pid !== undefined && (await alive($, run.pid));
533      const head = live ? "live" : run.finishedAt ? `ended (exit ${run.exitCode ?? "unknown"})` : "ended without a clean exit";
534      blocks.push(
535        [
536          ...(roots.length > 1 ? [root] : []),
537          `${head} · ${line(run, live)}`,
538          ...rows(run),
539          ...(live ? [] : ["`sandcastle report` prints the closing summary."]),
540        ].join("\n"),
541      );
542    }
543    return { text: blocks.join("\n\n") || "No sandcastle run on record in this project." };
544  });
545
546  on("command.run", { command: "sandcastle-mark" }, async ($, e) => ({ text: await choose($, await $.session.root(), e.args) }));
547
548  on("ui.render", { component: "AbovePrompt" }, async ($, e, next) => {
549    const now = await read($, view);
550    const line = await read($, markLine);
551    if ((now === null && line === null) || e.props.hasSurvey) return next(e);
552    const frame = CASTLE_FRAMES[await read($, castle)] ?? CASTLE_FRAMES[HELD];
553    const { Box, Text } = $.ui.resolve(e);
554    return (
555      <Box flexDirection="column">
556        {line === null ? null : (
557          <Box flexDirection="row" columnGap={1}>
558            <Text color={SAND.top} bold>
559              {MARK_ICON}
560            </Text>
561            <Text color={SAND.name}>{line}</Text>
562          </Box>
563        )}
564        {now === null
565          ? null
566          : band(now, e.props.bodyColumns - BAND_MARGIN, frame).map((row) => (
567              <Box flexDirection="row" columnGap={2}>
568                {row.map((seg) => (
569                  <Box flexDirection="row" columnGap={1}>
570                    <Text color={seg.colour} bold={seg.bold}>
571                      {seg.text}
572                    </Text>
573                    {seg.count === undefined ? null : (
574                      <Text color={seg.colour} bold>
575                        {String(seg.count)}
576                      </Text>
577                    )}
578                  </Box>
579                ))}
580              </Box>
581            ))}
582        {await next(e)}
583      </Box>
584    );
585  });
586};
587
hooks/run-state.ts 225 lines
1// What the mod says about a run record (.sandcastle/logs/run.json). Pure: nothing here
2// touches Claude Code, so the kit's own tests import this file and hold it to the status
3// view's vocabulary (status.sh, `style_of`). The record's types, the ticket states and the
4// tables from state to group and word live in run-record.ts.
5
6import { GROUPS, type Group, isTicketState, type RunRecord, sessionId, WORDS } from "./run-record.ts";
7
8/** A ticket as read from a file that may be a stranger's: its state is any short text until `isTicketState` says otherwise. */
9export type Ticket = { state?: string; note?: string; title?: string; order?: number };
10
11/** The run record as the mod reads it: the fields it shows, and tickets whose state is not yet trusted. */
12export type Run = Pick<RunRecord, "orchestrator" | "pid" | "session" | "startedAt" | "finishedAt" | "exitCode" | "stage" | "tokens"> & { tickets?: Record<string, Ticket> };
13
14// The status view's sand palette, as hex: a mod's Text takes no 256-colour index. Dry sand at
15// the castle's top, wet sand at its base.
16export const SAND = { top: "#e8d6b4", mid: "#cdb894", base: "#705c42", name: "#cdb894", stage: "#c0a478", muted: "#927c5e" };
17
18/** The castle's three rows, each five cells wide. */
19export type Castle = { top: string; mid: string; base: string };
20
21// The status view's three-row castle: the battlements stand a row above the text beside it.
22// A castle cut to one row (a bare slab) reads as no castle at all, so the view's one-row fold
23// draws none, only the wordmark.
24export const CASTLE: Castle = { top: "▄ ▄ ▄", mid: "█████", base: "██▀██" };
25
26// While a ticket is in work the castle builds from level sand, half a row at a time, and holds
27// complete before it starts again: a loop the eye follows as work going on, where a long hold
28// read as the run standing still. Every frame is a whole number of one 500 ms beat and the loop
29// is twelve of them (three bars of four), so the build keeps time and the restart lands on the
30// bar; the hold is never shorter than the build, so the finished castle is what the eye mostly sees. Every row of every frame is five cells, so the
31// text beside it never moves, and the band's height never changes.
32const AIR = "     ";
33export const CASTLE_FRAMES: (Castle & { ms: number })[] = [
34  { top: AIR, mid: AIR, base: "▁▁▁▁▁", ms: 500 },
35  { top: AIR, mid: AIR, base: "▄▄▄▄▄", ms: 500 },
36  { top: AIR, mid: AIR, base: CASTLE.base, ms: 500 },
37  { top: AIR, mid: "▄▄▄▄▄", base: CASTLE.base, ms: 500 },
38  { top: AIR, mid: CASTLE.mid, base: CASTLE.base, ms: 500 },
39  { ...CASTLE, ms: 3500 },
40];
41/** The held frame: the castle as the status view draws it. */
42export const HELD = CASTLE_FRAMES.length - 1;
43
44/** The status view's legend: its order, glyphs, words and colours. */
45export const LEGEND: { group: Group; glyph: string; label: string; colour: string }[] = [
46  { group: "working", glyph: "●", label: "working", colour: "#ffd75f" },
47  { group: "needs you", glyph: "!", label: "needs you", colour: "#ff5f5f" },
48  { group: "ready", glyph: ">", label: "ready to land", colour: "#5fd7d7" },
49  { group: "queued", glyph: "○", label: "queued", colour: "#87afff" },
50  { group: "blocked", glyph: "~", label: "blocked", colour: "#87afff" },
51  { group: "merged", glyph: "+", label: "merged", colour: "#5fd75f" },
52];
53
54// A state that is no ticket state (`constructor` included) is in no group and has no word of its own.
55const group = (t: Ticket): Group => (isTicketState(t.state) ? GROUPS[t.state] : "other");
56
57const word = (state: string | undefined) => (isTicketState(state) ? WORDS[state] : undefined) ?? state ?? "unknown";
58
59// The record is a file in a repository, which may be a stranger's: nothing in it is trusted to
60// be what src/run.ts writes. Text is cut to one short line with nothing in it that draws no
61// glyph - control, format (direction marks, tag characters), line separators, variation
62// selectors, private-use and half a surrogate pair - so it cannot pose as a line of its own in
63// what Claude reads, carry words a person cannot see, or redraw the terminal. Cut by code
64// point: a cut through an emoji would leave half of one. A number is a whole number or absent.
65const text = (v: unknown, max: number) =>
66  typeof v === "string"
67    ? Array.from(v.replace(/[\p{Cc}\p{Cf}\p{Cs}\p{Co}\p{Zl}\p{Zp}\p{Variation_Selector}]+/gu, " ").trim())
68        .slice(0, max)
69        .join("") || undefined
70    : undefined;
71const whole = (v: unknown) => (typeof v === "number" && Number.isSafeInteger(v) && v >= 0 ? v : undefined);
72const MAX_TICKETS = 200;
73
74export { RUN_COMMAND } from "./run-live.ts";
75
76// A ticket as the kit writes it (src/tracker.ts, `refOf`): `#12` for a number, the id itself
77// for a ticket file.
78const ref = (id: string) => (/^\d+$/.test(id) ? `#${id}` : id);
79
80/** How a run ended, from its exit code alone: a whole number or the word `unknown`, never record text. */
81export const endedHow = (run: Run): string => (run.finishedAt ? `ended (exit ${whole(run.exitCode) ?? "unknown"})` : "ended without a clean exit");
82
83/**
84 * The prompt that has the session close a run that ended. It names the run by numbers only - the
85 * pid when it is a whole number, the start time re-rendered as local `HH:MM` (as the closing
86 * summary's Run line prints it) when `startedAt` parses as a date - and leaves out anything else:
87 * the record is a file in a repository that may be a stranger's. The text is fixed when the end is
88 * seen and read when the session is idle, by which time a later run may be live in the same
89 * root; the numbers are how the agent tells the run that ended from that one.
90 */
91export const endPrompt = (root: string, run: Run): string => {
92  const pid = run.pid !== undefined && Number.isSafeInteger(run.pid) && run.pid > 0 ? `pid ${run.pid}` : undefined;
93  const at = run.startedAt === undefined ? NaN : Date.parse(run.startedAt);
94  const started = Number.isFinite(at) ? `started ${new Date(at).toTimeString().slice(0, 5)}` : undefined;
95  const named = [pid, started].filter(Boolean).join(", ");
96  return `The sandcastle run in ${root}${named ? ` (${named})` : ""} ${endedHow(run)}. Close that run now: read run.md in the sandcastle skill's directory and follow it.`;
97};
98
99/** A half-written or foreign file reads as no record. */
100export const parse = (raw: string): Run | undefined => {
101  let r: Record<string, unknown>;
102  try {
103    r = JSON.parse(raw);
104  } catch {
105    return undefined;
106  }
107  if (!r || typeof r !== "object") return undefined;
108  const tickets = r.tickets && typeof r.tickets === "object" ? Object.entries(r.tickets as Record<string, Record<string, unknown> | null>) : [];
109  return {
110    orchestrator: text(r.orchestrator, 40),
111    // 0 is no process.
112    pid: whole(r.pid) || undefined,
113    session: sessionId(r.session),
114    startedAt: text(r.startedAt, 40),
115    finishedAt: text(r.finishedAt, 40),
116    exitCode: whole(r.exitCode),
117    stage: text(r.stage, 40),
118    tokens: text(r.tokens, 40),
119    tickets: Object.fromEntries(
120      // A ticket file's id is a whole slug: cut short, two tickets of one feature would share a key.
121      tickets.slice(0, MAX_TICKETS).map(([id, t]) => [text(id, 100) ?? "?", { state: text(t?.state, 24), note: text(t?.note, 100), title: text(t?.title, 100), order: whole(t?.order) }]),
122    ),
123  };
124};
125
126const tickets = (run: Run) => Object.entries(run.tickets ?? {});
127
128/** The tickets a person has to act on: `#105 conflict`. */
129export const needing = (run: Run): string[] =>
130  tickets(run)
131    .filter(([, t]) => group(t) === "needs you")
132    .map(([id, t]) => `${ref(id)} ${word(t.state)}`);
133
134/** What the band above the prompt draws from: plain data, so it can live in `$.state`. */
135export type Summary = { name: string; stage: string; counts: number[]; tokens: string };
136
137/** `counts` runs parallel to LEGEND. */
138export const summarise = (run: Run): Summary => ({
139  name: run.orchestrator ?? "",
140  // One process can hold several turns: a finished record with a live process is between two.
141  stage: run.finishedAt ? "turn finished" : run.stage && run.stage !== "running" ? run.stage : "",
142  counts: LEGEND.map((g) => tickets(run).filter(([, t]) => group(t) === g.group).length),
143  tokens: run.tokens ?? "",
144});
145
146export type Segment = { text: string; colour: string; bold?: true; count?: number };
147
148/** Whether the run has a ticket in work: the castle builds only then. */
149export const building = (s: Summary): boolean => (s.counts[LEGEND.findIndex((g) => g.group === "working")] ?? 0) > 0;
150
151/**
152 * The band's three rows, each cut to the width it is drawn in: the castle's battlements alone,
153 * its walls with the run (name, stage, tokens), and its base with the legend - in words while
154 * they fit, then the glyphs alone. A row that wraps would push the prompt down every time a
155 * count gains a digit.
156 */
157export const band = (s: Summary, columns: number, castle: Castle = CASTLE): Segment[][] => {
158  // Two columns between segments, one between a segment's text and its count.
159  const width = (row: Segment[]) => row.reduce((n, seg) => n + seg.text.length + (seg.count === undefined ? 0 : 1 + String(seg.count).length), 2 * (row.length - 1));
160  const fit = (tries: Segment[][]) => tries.find((row) => width(row) <= columns) ?? tries.at(-1) ?? [];
161  const run = (name: boolean, tokens: boolean): Segment[] =>
162    [
163      { text: castle.mid, colour: SAND.mid },
164      { text: "sandcastle", colour: SAND.top, bold: true as const },
165      { text: name ? s.name : "", colour: SAND.name },
166      { text: s.stage, colour: SAND.stage },
167      { text: tokens ? s.tokens : "", colour: SAND.muted },
168    ].filter((seg) => seg.text);
169  const legend = (words: boolean): Segment[] => [
170    { text: castle.base, colour: SAND.base },
171    ...LEGEND.map((g, i) => ({ text: words ? `${g.glyph} ${g.label}` : g.glyph, count: s.counts[i] ?? 0, colour: g.colour })).filter((seg) => seg.count),
172  ];
173  return [[{ text: castle.top, colour: SAND.top }], fit([run(true, true), run(true, false), run(false, false)]), fit([legend(true), legend(false)])];
174};
175
176/** The same summary as one line of text, for `/sandcastle-status`. A run that is over has no stage. */
177export const line = (run: Run, live: boolean): string => {
178  const s = summarise(run);
179  return [live && s.stage, ...LEGEND.flatMap((g, i) => (s.counts[i] ? [`${g.glyph} ${g.label} ${s.counts[i]}`] : [])), s.tokens].filter(Boolean).join(" · ");
180};
181
182/** One line per ticket, in the status view's order: working first, then what needs a person. */
183export const rows = (run: Run): string[] => {
184  const place = (t: Ticket) => LEGEND.findIndex((g) => g.group === group(t));
185  const rank = (t: Ticket) => (place(t) < 0 ? LEGEND.length : place(t));
186  return tickets(run)
187    .sort(([, a], [, b]) => rank(a) - rank(b) || (a.order ?? 0) - (b.order ?? 0))
188    .map(([id, t]) => `${LEGEND[place(t)]?.glyph ?? "·"} ${ref(id)} ${word(t.state)}${t.note ? ` (${t.note})` : ""}${t.title ? ` - ${t.title}` : ""}`);
189};
190
191/**
192 * Lists the machine-wide live-runs registry (src/live-runs.ts): its directory is
193 * `$XDG_CACHE_HOME`, or `~/.cache` when that is unset or empty, on Linux and macOS alike. Prints
194 * the session root's resolved path (`pwd -P`: `/tmp` is `/private/tmp` on macOS, and one project
195 * must not read as two), then one resolved root per registered run, a line each. POSIX sh, and
196 * `cat`, `cd` and `pwd -P` only: BSD and GNU alike. A root that is gone prints nothing.
197 * `$1` is the session's root.
198 */
199export const REGISTRY_SCRIPT = [
200  'dir="${XDG_CACHE_HOME:-$HOME/.cache}/sandcastle-kit/runs"',
201  '(cd -- "$1" 2>/dev/null && pwd -P) || echo',
202  'for f in "$dir"/*; do',
203  '  [ -f "$f" ] || continue',
204  '  r=$(cat -- "$f" 2>/dev/null)',
205  '  [ -n "$r" ] && (cd -- "$r" 2>/dev/null && pwd -P)',
206  "done",
207  "exit 0",
208].join("\n");
209
210/** The script's output: the session's own root (empty when it could not be resolved) and the registered ones, as resolved strings compared as they are (no case folding). */
211export const parseRegistry = (stdout: string): { own: string; roots: string[] } => {
212  const [own = "", ...rest] = stdout.split("\n");
213  return { own, roots: [...new Set(rest.filter(Boolean))] };
214};
215
216/**
217 * The registered roots this session may follow: every one but its own, which the watch of the
218 * session's root already covers. With no resolved root of its own it follows none - failing closed
219 * (a shell that cannot resolve it, BusyBox `ps` and the like) beats watching one run twice.
220 */
221export const followable = ({ own, roots }: { own: string; roots: string[] }): string[] => (own ? roots.filter((r) => r !== own) : []);
222
223/** Whether `run` is the run of the session with this id. A run with no recorded id (a plain terminal, Codex, OpenCode) has no owner. */
224export const startedBy = (run: Run | undefined, session: string): boolean => !!run?.session && run.session === session;
225
hooks/run-live.ts 61 lines
1// Whether a run is live: the one rule behind the status view, the Herdr tab bar, the closing
2// summary, `sandcastle wait` and the Claude Code mod. A run is live when its record has not
3// finished and its pid is a process whose command line contains RUN_COMMAND: the pid alone
4// comes round again as some other process, and a run that was SIGKILLed would read as live
5// for as long as that process lasts. Pure and importing nothing, so the mod (which cannot
6// import the kit's source) and the kit's own tests can both read it; the process check is
7// passed in. status.sh keeps its own bash version of the same rule (`test/run-live-contract.test.ts`
8// holds the two together). A paused run (`sandcastle pause`; the record's `paused`) is live by this
9// rule, and rightly: its process runs and its record has not finished, so no view offers an end
10// prompt for it. Terms are GLOSSARY.md's.
11
12/** What the run's process shows in its command line (bin/sandcastle starts it): how a look tells it from a process that got its pid later. */
13export const RUN_COMMAND = "src/cli.ts";
14
15/**
16 * The process check: the command line of the process with this pid, or undefined when there is
17 * none. A process another user owns exists all the same (EPERM on a signal says so): the probe
18 * reads its command line like any other, as `ps -p <pid> -o command=` does.
19 */
20export type Probe = (pid: number) => string | undefined;
21
22/** The pid is a process of the kit, not one that has since taken the number. */
23export const isKit = (command: string | undefined): boolean => command !== undefined && command.includes(RUN_COMMAND);
24
25/** The pid names a process of the kit that is running now. */
26export const kitRunning = (pid: number | undefined, probe: Probe): boolean => Number.isInteger(pid) && (pid as number) > 0 && isKit(probe(pid as number));
27
28/** The fields of a run record that decide it; the rest of the record is not read. */
29export type LiveRecord = { pid?: unknown; finishedAt?: unknown; exitCode?: unknown };
30
31/**
32 * What a run is:
33 * - `live`: its process is running (the pid it holds the run lock under, or the one its
34 *   unfinished record names);
35 * - `finished`: its record says it ended, with the exit code it wrote (none for a run that was
36 *   not given one);
37 * - `own`: its record is unfinished and names the asking process itself - the run that is
38 *   summing itself up, between its own turns, which is not "another live run" and not a killed
39 *   one either;
40 * - `dead`: no record, or an unfinished one whose process is gone or is not the kit's: killed.
41 */
42export type Liveness = { state: "live"; pid: number } | { state: "finished"; exitCode?: number } | { state: "own" } | { state: "dead" };
43
44const finished = (record: LiveRecord) => (typeof record.finishedAt === "string" && record.finishedAt !== "") || Number.isInteger(record.exitCode);
45
46/**
47 * The lock goes before the process does (its exit handlers run in turn, and the record's end is
48 * written after the lock is released), so the lock's pid decides first: a process that still
49 * holds it is live whatever the record says, and a record that has not finished is live while its
50 * own pid is.
51 */
52export const liveness = ({ record, lockPid, self }: { record?: LiveRecord; lockPid?: number; self?: number }, probe: Probe): Liveness => {
53  if (kitRunning(lockPid, probe)) return { state: "live", pid: lockPid as number };
54  if (!record) return { state: "dead" };
55  if (finished(record)) return { state: "finished", ...(Number.isInteger(record.exitCode) ? { exitCode: record.exitCode as number } : {}) };
56  const pid = record.pid;
57  if (typeof pid !== "number" || !Number.isInteger(pid) || pid <= 0) return { state: "dead" };
58  if (pid === self) return { state: "own" };
59  return kitRunning(pid, probe) ? { state: "live", pid } : { state: "dead" };
60};
61
hooks/idle.ts 196 lines
1// The idle mark: the one row the mod draws in sand above the prompt (not in the status line) between runs, in a
2// project set up for sandcastle. Pure, and like run-state.ts it imports nothing: register.tsx
3// gathers the facts and this module decides the text. The input is one object on purpose - the
4// ready count, the project's hidden flag and a dismissal join it without a new signature.
5
6/** A project's cached ready-ticket read, one `$.store` entry shared by every session on the machine. */
7export type Entry = {
8  /** When `ids` were read, in ms since the epoch: the age of the count, kept through failed reads. */
9  at: number;
10  /** The ready ticket ids of the last good read. */
11  ids: string[];
12  /** Whether the latest read worked; a failure keeps `ids` and `at` and says so here. */
13  ok: boolean;
14  /**
15   * When the latest read ended, good or not: what `due` measures. Apart from `at` so that a
16   * tracker that cannot be reached is tried every ten minutes, not at every look.
17   */
18  tried: number;
19};
20
21/**
22 * A person's choices for one project, one `$.store` entry beside the ready entry: machine-local,
23 * kept across a restart, never in the repository.
24 */
25export type Choice = {
26  /** `/sandcastle-mark hide`: no mark in this project until `show`. */
27  hidden: boolean;
28  /** `/sandcastle-mark dismiss`: the ready ids at that moment; absent when nothing is dismissed. */
29  dismissed?: string[];
30};
31
32/** What the line is decided from. */
33export type MarkInput = {
34  /** The project's `.sandcastle/config.ts` is a plain file: `sandcastle init` ran there. */
35  setUp: boolean;
36  /** The machine switch: false when the personal settings say `"idleMark": false`. */
37  idleMark: boolean;
38  /** The project's hidden flag (`/sandcastle-mark hide`). */
39  hidden?: boolean;
40  /** The ids of a dismissal (`/sandcastle-mark dismiss`); the count stays quiet while the ready set is within them. */
41  dismissed?: string[];
42  /** The cached read; none yet, or with no `now`, gives the bare mark. */
43  entry?: Entry;
44  /** The time, ms since the epoch. */
45  now?: number;
46};
47
48/** A cached read is read again when it is this old. */
49export const READ_MS = 10 * 60 * 1000;
50/** A count older than this is not shown: the mark stands without one. */
51export const COUNT_MS = 60 * 60 * 1000;
52
53/** The ready ids of a read that is current - under `COUNT_MS` old - else none: what a count may be made of. */
54const shown = (entry: Entry | undefined, now: number | undefined): entry is Entry => entry !== undefined && now !== undefined && now - entry.at >= 0 && now - entry.at < COUNT_MS;
55const current = (entry: Entry | undefined, now: number | undefined): string[] => (shown(entry, now) ? entry.ids : []);
56
57/**
58 * Whether a dismissal has ended: the current ready set holds an id the dismissal does not. An id
59 * leaving the set ends nothing, and a count too old to show says nothing either way.
60 */
61export const dismissalEnded = (dismissed: string[], entry: Entry | undefined, now: number | undefined): boolean =>
62  current(entry, now).some((id) => !dismissed.includes(id));
63
64/**
65 * The line, or undefined to clear it. A count shows only while the last good read is under an
66 * hour old, so a tracker that cannot be reached keeps the number for a while and then goes quiet;
67 * the line never says why (`sandcastle queue` and the status view do). A dismissal takes the
68 * count off, never the mark; hiding takes both.
69 */
70export const markText = (input: MarkInput): string | undefined => {
71  if (!input.setUp || !input.idleMark || input.hidden) return undefined;
72  const { entry, now, dismissed } = input;
73  const ids = current(entry, now);
74  const quiet = dismissed !== undefined && !dismissalEnded(dismissed, entry, now);
75  return ids.length > 0 && !quiet ? `sandcastle · ${ids.length} ready - /sandcastle run` : "sandcastle";
76};
77
78/** What `/sandcastle-mark` takes. */
79export const MARK_USAGE = "Usage: /sandcastle-mark [dismiss|hide|show]";
80
81/** The command's argument, or undefined for one it does not know. Nothing is "no argument": the report. */
82export const markAction = (args: string): "dismiss" | "hide" | "show" | "report" | undefined => {
83  const word = args.trim().toLowerCase();
84  return word === "" ? "report" : word === "dismiss" || word === "hide" || word === "show" ? word : undefined;
85};
86
87/** The choice after `dismiss`, `hide` or `show`: the dismissal holds the ids of a count that is current, none otherwise. */
88export const choiceAfter = (action: "dismiss" | "hide" | "show", choice: Choice, entry: Entry | undefined, now: number | undefined): Choice =>
89  action === "show" ? { hidden: false } : action === "hide" ? { ...choice, hidden: true } : { ...choice, dismissed: [...current(entry, now)] };
90
91/** A stored value as a choice, or undefined when it is not one (the store is shared, so it is checked). */
92export const parseChoice = (value: unknown): Choice | undefined => {
93  if (typeof value !== "object" || value === null) return undefined;
94  const { hidden, dismissed } = value as Record<string, unknown>;
95  if (typeof hidden !== "boolean") return undefined;
96  if (dismissed === undefined) return { hidden };
97  if (!Array.isArray(dismissed) || !dismissed.every((id) => typeof id === "string")) return undefined;
98  return { hidden, dismissed: dismissed as string[] };
99};
100
101const ago = (ms: number): string => {
102  const min = Math.floor(Math.max(ms, 0) / 60000);
103  return min < 1 ? "under a minute ago" : min < 60 ? `${min} min ago` : `${Math.floor(min / 60)} h ${min % 60} min ago`;
104};
105
106/** The reply to a bare `/sandcastle-mark`: whether the mark is shown, and the cached count with its age. */
107export const markReport = (input: MarkInput): string => {
108  const { entry, now } = input;
109  const dismissed = input.dismissed !== undefined && !dismissalEnded(input.dismissed, entry, now);
110  const state = !input.setUp
111    ? "not shown: this project is not set up (no .sandcastle/config.ts)"
112    : !input.idleMark
113      ? `hidden on this machine ("idleMark": false in the personal settings)${input.hidden ? " and in this project" : ""}`
114      : input.hidden
115        ? "hidden in this project (/sandcastle-mark show brings it back)"
116        : dismissed
117          ? "shown without a count: dismissed until a new ticket becomes ready (/sandcastle-mark show ends it)"
118          : "shown";
119  const count = !entry
120    ? "No count read yet."
121    : `${entry.ids.length} ready, read ${now === undefined ? "earlier" : ago(now - entry.at)}${entry.ok ? "" : " (the latest read failed)"}${shown(entry, now) ? "" : "; too old to show"}.`;
122  return `Idle mark: ${state}.\nCached count: ${count}`;
123};
124
125/**
126 * The ready ticket ids in `sandcastle queue --json`'s stdout - the queued tickets with no open
127 * blocker - or undefined for a read that is no list of tickets. Never an empty list for output
128 * that is not one: a failed read must not count as "nothing ready".
129 */
130export const readyIds = (stdout: string): string[] | undefined => {
131  let rows: unknown;
132  try {
133    rows = JSON.parse(stdout);
134  } catch {
135    return undefined;
136  }
137  if (!Array.isArray(rows)) return undefined;
138  const ids: string[] = [];
139  for (const row of rows) {
140    if (typeof row !== "object" || row === null || Array.isArray(row)) return undefined;
141    const { id, blockedOn } = row as Record<string, unknown>;
142    if ((typeof id !== "string" && typeof id !== "number") || String(id) === "" || !Array.isArray(blockedOn)) return undefined;
143    if (blockedOn.length === 0 && !ids.includes(String(id))) ids.push(String(id));
144  }
145  return ids;
146};
147
148/**
149 * What makes a read due ahead of its age: a run ended - the project's own, or one this session
150 * followed elsewhere, perhaps a second clone of it - or the sandcastle skill was used (it may just
151 * have labelled tickets). A session's start needs none: its first look applies
152 * the age rule at once, so an entry that is missing or old is read then.
153 */
154export type Trigger = "run-ended" | "skill";
155
156/** Whether to read the tracker now: no entry, one `READ_MS` old or more, or a trigger fired. */
157export const due = (entry: Entry | undefined, now: number, trigger?: Trigger): boolean => {
158  if (trigger !== undefined || entry === undefined) return true;
159  const age = now - entry.tried;
160  // An entry from the future is a clock that moved: read it again rather than trust it.
161  return age < 0 || age >= READ_MS;
162};
163
164/** The entry after a read: `ids` is the read's result, or undefined for a failed one. */
165export const afterRead = (prior: Entry | undefined, ids: string[] | undefined, now: number): Entry =>
166  ids !== undefined ? { at: now, ids, ok: true, tried: now } : { at: prior?.at ?? now, ids: prior?.ids ?? [], ok: false, tried: now };
167
168/** A stored value as an entry, or undefined when it is not one: a store is shared, so it is checked. */
169export const parseEntry = (value: unknown): Entry | undefined => {
170  if (typeof value !== "object" || value === null) return undefined;
171  const { at, ids, ok, tried } = value as Record<string, unknown>;
172  if (typeof at !== "number" || !Number.isFinite(at) || typeof tried !== "number" || !Number.isFinite(tried) || typeof ok !== "boolean") return undefined;
173  if (!Array.isArray(ids) || !ids.every((id) => typeof id === "string")) return undefined;
174  return { at, ids: ids as string[], ok, tried };
175};
176
177/**
178 * Prints the personal machine settings, honouring `XDG_CONFIG_HOME` as the kit's `USER_CONFIG`
179 * does (src/sandbox.ts). `cat` only: BSD and GNU alike. A missing file prints nothing.
180 */
181export const SETTINGS_SCRIPT = ['cat -- "${XDG_CONFIG_HOME:-$HOME/.config}/sandcastle-kit/config.json" 2>/dev/null', "exit 0"].join("\n");
182
183/**
184 * The machine switch from the settings file's text. Off only for `"idleMark": false`: a file that
185 * is not JSON, or any other value, leaves the mark on - `sandcastle doctor` is where those are
186 * reported, and a typo never hides the mark silently the other way round.
187 */
188export const machineSwitch = (stdout: string): boolean => {
189  try {
190    const settings: unknown = JSON.parse(stdout);
191    return !(typeof settings === "object" && settings !== null && !Array.isArray(settings) && (settings as Record<string, unknown>).idleMark === false);
192  } catch {
193    return true;
194  }
195};
196
hooks/run-record.ts 336 lines
1// What a run record (.sandcastle/logs/run.json) holds and what the status view says about it:
2// the record's types, the closed set of ticket states, the derived states, and the tables from
3// ticket state to group and to word. Pure: it imports nothing, so the mod (which cannot import
4// the kit's source) and the kit's own tests can both read it. Terms are GLOSSARY.md's.
5
6/**
7 * Where one ticket of a run stands, as the run record holds it - every state src/run.ts
8 * writes. A phase (implement, review, gates ...) is one kind of ticket state; `requeued` is
9 * not one (it is a fact about a ticket's second attempt, `TicketRecord.requeued`). `paused` is a
10 * ticket parked at a juncture of a paused run: its sandbox is closed, its branch kept, and it
11 * resumes with the phase its note names.
12 */
13export const TICKET_STATES = [
14  "queued",
15  "blocked",
16  "setup",
17  "implement",
18  "resolve",
19  "review",
20  "cross-review",
21  "gates",
22  "repair",
23  "ready",
24  "landing",
25  "paused",
26  "merged",
27  "held",
28  "conflict",
29  "red",
30  "nochange",
31  "uncommitted",
32  "crashed",
33  "not landed",
34  "withdrawn",
35  "stopped",
36  "skipped",
37] as const;
38
39export type TicketState = (typeof TICKET_STATES)[number];
40
41/** Tells a ticket state from any other string: a record is a file in a repository, which may be a stranger's. */
42export const isTicketState = (s: unknown): s is TicketState => typeof s === "string" && (TICKET_STATES as readonly string[]).includes(s);
43
44/**
45 * A run record's tickets as read from a file, each one's state passed through the guard. A state
46 * outside the set (a record from an older kit, or edited by hand) is dropped, so the ticket
47 * falls to the "other" group and into no report section, never into one that asks for a person.
48 */
49export const readTickets = (record: unknown): Record<string, TicketRecord> => {
50  const tickets = (record as { tickets?: unknown } | null | undefined)?.tickets;
51  if (!tickets || typeof tickets !== "object" || Array.isArray(tickets)) return {};
52  return Object.fromEntries(
53    Object.entries(tickets as Record<string, unknown>).map(([id, t]) => {
54      const { state, ...rest } = (t && typeof t === "object" ? t : {}) as Record<string, unknown>;
55      return [id, (isTicketState(state) ? { ...rest, state } : rest) as TicketRecord];
56    }),
57  );
58};
59
60/**
61 * What a run says became of one ticket, as `.sandcastle/logs/outcomes.json` holds it beside the
62 * line: every reader (the status view, the report, the autonomy loop) decides on the kind, never
63 * on the line's words, so a new ending cannot read as "ready" for want of a prefix. `red` is red
64 * once merged with other tickets at landing; `gate red` is red in the ticket's own pipeline.
65 * `taken back` is a ticket a person marked for a human during the run; `green` is a branch gated
66 * green that this run has not (or, in a dry run, would have) landed.
67 */
68export const OUTCOME_KINDS = [
69  "green",
70  "merged",
71  "conflict",
72  "red",
73  "gate red",
74  "held",
75  "taken back",
76  "uncommitted",
77  "crashed",
78  "not landed",
79  "withdrawn",
80  "stopped",
81  "no change",
82] as const;
83
84export type OutcomeKind = (typeof OUTCOME_KINDS)[number];
85
86/** Tells an outcome kind from any other value: outcomes.json is a file in a repository, and an older run's entries carry none. */
87export const isOutcomeKind = (s: unknown): s is OutcomeKind => typeof s === "string" && (OUTCOME_KINDS as readonly string[]).includes(s);
88
89/** One ticket's outcome as a run writes it: the kind, the tickets it collided with, and the line a person reads. */
90export type Outcome = { kind: OutcomeKind; with?: string[]; text: string };
91
92/** One entry of outcomes.json as read: the run that wrote it, and no kind when an older kit did. */
93export type OutcomeEntry = Partial<Outcome> & { run?: string; at?: string };
94
95/**
96 * The states the status view works out for itself and no run record holds: a run that died
97 * (`stalled`, `orphaned`), a branch of an earlier run (`left over`), a branch of this run that
98 * waits for landing to decide it (`finished`), and `requeued`, the word it gives an older run's
99 * branch whose ticket was labelled again.
100 */
101export const DERIVED_STATES = ["stalled", "orphaned", "left over", "finished", "requeued"] as const;
102
103export type DerivedState = (typeof DERIVED_STATES)[number];
104
105/** What the status view sorts ticket states into, shared by every view so none disagrees about a ticket. */
106export type Group = "working" | "needs you" | "ready" | "queued" | "blocked" | "merged" | "other";
107
108/** The group each ticket state falls into. Keyed by the closed set: a missing or extra state fails the type check. */
109export const GROUPS: Record<TicketState, Group> = {
110  setup: "working",
111  implement: "working",
112  resolve: "working",
113  review: "working",
114  "cross-review": "working",
115  gates: "working",
116  repair: "working",
117  landing: "working",
118  red: "needs you",
119  conflict: "needs you",
120  held: "needs you",
121  uncommitted: "needs you",
122  crashed: "needs you",
123  "not landed": "needs you",
124  stopped: "needs you",
125  ready: "ready",
126  queued: "queued",
127  paused: "queued",
128  blocked: "blocked",
129  merged: "merged",
130  nochange: "other",
131  withdrawn: "other",
132  skipped: "other",
133};
134
135/** The status view's word for a ticket state, where it differs from the state's own name. */
136export const WORDS: Partial<Record<TicketState, string>> = { implement: "impl", "cross-review": "codex", red: "gate red", nochange: "no change" };
137
138/** One ticket of a run, as src/run.ts writes it. */
139export type TicketRecord = {
140  state?: TicketState;
141  /** Seconds since the epoch at which the state began. */
142  since?: number;
143  /** Seconds since the epoch at the ticket's first `setup`, kept through a requeue or a resume: TIME's start. */
144  started?: number;
145  /** Seconds since the epoch at the current attempt's `setup` (a requeued second attempt, a resume): the ETA's start. */
146  attemptStarted?: number;
147  order?: number;
148  note?: string | null;
149  title?: string;
150  commits?: number;
151  /** In and out as `tokenBrief` writes them (`3.1M in / 42k out`): the finished passes plus the one running, rewritten on the usage row's tick. */
152  tokens?: string;
153  minutes?: number;
154  /** Test ids a red gate named. */
155  failing?: string[];
156  /** Files a merge conflicted on, or protected paths a held branch changes. */
157  files?: string[];
158  /** The ticket's second attempt after a conflict or red at landing ("requeued after conflict with #3"); null once that attempt is not going to run. */
159  requeued?: string | null;
160  /** Merged, but the tracker refused the close: the error, short. */
161  closeFailed?: string;
162  /** What the reviewer said no gate exercises; a merged ticket with one needs a person. */
163  ungated?: string;
164  /** A sentence a reviewer left in prose naming a gap it filed neither as a `<followup>` nor as an `<unmet>` line; a merged ticket with one needs a person. */
165  gap?: string;
166  /** Changelog lines the implementer and reviewer asked for (`changelog: true`), each starting Added:, Changed:, Fixed: or Upgrading:. */
167  changelog?: string[];
168  /** How many `<changelog>` tags were no changelog line (too long, a list, a commit sha) and were left out of `changelog`. */
169  changelogDropped?: number;
170  /** The acceptance criterion an agent knowingly left undone: merged, the ticket still open; a merged ticket with one needs a person. */
171  unmet?: string;
172  /** Paths the branch changed beyond its ticket's `Touches:` line. */
173  overrun?: string[];
174};
175
176/**
177 * A Claude Code session id as an environment variable or a record holds it, or undefined when
178 * the value is not one. The record is a file in a repository, so only a short id of letters,
179 * digits, `-` and `_` passes.
180 */
181export const sessionId = (value: unknown): string | undefined => (typeof value === "string" && /^[\w-]{1,100}$/.test(value) ? value : undefined);
182
183/**
184 * The run's settings as one turn's record holds them (GLOSSARY.md: run setting): the autonomy
185 * level, the turn this record is, the level's cap, the repair attempts, the concurrency (asked
186 * and effective), whether cross-review runs, and the usage guard. Each field is optional and a
187 * reader shows only what is there - an older kit's record has no group at all, level 1 has no cap
188 * (it asks after every turn), and a record without the guard's fields shows nothing about it.
189 */
190export type RunSettings = {
191  /** The level, resolved once per run. */
192  autonomy?: 0 | 1 | 2 | 3 | "drain";
193  /** This record's turn, 1-based. */
194  turn?: number;
195  /** The most turns the level allows. */
196  cap?: number;
197  /** The repair attempts a ticket gets after a red gate; 0 is repair off. */
198  repair?: number;
199  /** The tickets the run takes at once, after the machine-wide sandbox cap. */
200  concurrency?: number;
201  /** The tickets at once the run asked for; the view shows it only when it differs from `concurrency`. */
202  asked?: number;
203  /** Whether cross-review runs, resolved once per run. */
204  crossReview?: boolean;
205  /** Cross-review's model: written only when it is on. */
206  crossReviewModel?: string;
207  /** Cross-review's effort: written only when it is on. */
208  crossReviewEffort?: string;
209  /** Whether the usage guard (`USAGE_CHECK=1`) was asked for. */
210  usageGuard?: boolean;
211  /** The guard's stop threshold in percent; only while it is on. */
212  usageStop?: number;
213  /**
214   * The guard's reading, a fact beside the setting and never a change to it: `unavailable` when it
215   * cannot get one (a 403 turns it off for the run, a rate limit or a missing OAuth token leaves it
216   * without one for now). The only settings field that may change during a turn.
217   */
218  usageReading?: "unavailable";
219  /** The plan usage in percent at which the run pauses itself (`USAGE_PAUSE`); only when it is on. */
220  usagePause?: number;
221  /** True when the sandboxes spend `ANTHROPIC_API_KEY`, billing API credits; absent otherwise. */
222  apiKey?: boolean;
223};
224
225/** One of a plan's usage windows: how much of it is spent (0 to 100) and when it resets (seconds since the epoch). */
226export type PlanWindow = { percent: number; resetsAt: number };
227
228/**
229 * One provider's plan usage, as a live run shows it (`usage` in the run record, read by the status view,
230 * the Herdr token and the closing summary), keyed by the provider that reports it. `windows` and `at` are
231 * absent until the first reading: the run is watching for one and none has come.
232 */
233export type PlanUsage = {
234  provider: "claude" | "codex";
235  /** The 5-hour and the weekly window: from a Claude agent's rate-limit event, or Codex's `rate_limits` (matched by their length, not their position). */
236  windows?: { fiveHour: PlanWindow; week: PlanWindow };
237  /** Seconds since the epoch at which the kit read the event: the reading's age is measured from it. */
238  at?: number;
239};
240
241/**
242 * Why a run is paused when no person asked for it: a plan window reached `USAGE_PAUSE` (or an agent hit the
243 * limit anyway), and the run resumes by itself at `resumesAt` - seconds since the epoch, a minute after the
244 * window's reset. `percent` is the window's usage when the run paused; `window` and `provider` say whose.
245 */
246export type UsagePaused = { cause: "usage"; provider: PlanUsage["provider"]; window: "fiveHour" | "week"; percent: number; resumesAt: number };
247
248/** The whole run record: the run's own fields and its tickets, by ticket id. Every field is optional - the file is read while the run is still filling it. */
249export type RunRecord = {
250  /** The project's name. */
251  orchestrator?: string;
252  pid?: number;
253  /** The Claude Code session that started the run (`CLAUDE_CODE_SESSION_ID`, nothing else of its environment); absent from a plain terminal, Codex or OpenCode. */
254  session?: string;
255  startedAt?: string;
256  /** Written on a clean exit; a pid that is gone without it is a run that was killed. */
257  finishedAt?: string;
258  exitCode?: number;
259  models?: string;
260  issues?: string[];
261  dryRun?: boolean;
262  versions?: { claude?: string; codex?: string };
263  /** Tickets held for another that is open: `on` names what each waits for. */
264  waiting?: { issue: string; on: string[] }[];
265  /** What the run line shows while the run is live. */
266  stage?: string;
267  concurrency?: number;
268  /**
269   * How loaded the run was, kept in the history line so the estimate prices a run from earlier runs of a similar load:
270   * the sandboxes it ran at once (its effective concurrency, after the machine-wide cap and its share) and its ticket
271   * count. An older kit's record has none, and the estimate counts that run as unknown.
272   */
273  load?: { concurrency: number; tickets: number };
274  /** The run settings: what the status view's settings row shows. */
275  settings?: RunSettings;
276  /**
277   * Present while a person has paused the run (`sandcastle pause`): no agent pass starts, the
278   * passes in flight finish and green branches still land. `since` is seconds since the epoch;
279   * `finishing` the tickets still doing something (a pass, a gate run, a landing). Absent when the
280   * run is not paused: a paused run is live all the same, its process is running. A pause the run
281   * took for its plan's usage (`USAGE_PAUSE`) says so: `cause: "usage"`, the window and when it resumes.
282   */
283  paused?: { since: number; finishing: string[] } & Partial<UsagePaused>;
284  /**
285   * The plan's usage, one entry per provider the run shows, each its newest reading across the run's agent logs
286   * (`src/usage.ts`): Claude's while the run spends a subscription on a Claude model, Codex's while cross-review
287   * runs on a ChatGPT plan. An older kit wrote the one entry as an object, and readers still take that.
288   */
289  usage?: PlanUsage[] | PlanUsage;
290  /** Live values, not settings: the sandbox slots the run could use now, and its share of the machine pool (src/pool.ts), rewritten as either changes. */
291  demand?: number;
292  share?: number;
293  /** A person's cap on the run's share (`sandcastle cap`); absent when there is none. */
294  cap?: number;
295  /**
296   * Present (true) while the run waits for a sandbox slot that its share of the machine pool (or its cap) holds back, or
297   * the slot kept for landing does, not only a full pool: the status view's next-to-start rows say so. The wait is the
298   * run's, not a ticket's - a worker leases its slot before it takes a ticket - and an older kit wrote it as each waiting
299   * ticket's note. Kept for older views: `waitsFor` says which.
300   */
301  waitsForShare?: boolean;
302  /**
303   * What holds that wait back, present with `waitsForShare`: `share` (the run's share or cap: `waits for the run's
304   * share`) or `landing` (the last slot of the share is kept for a landing: `waits: slot kept to land`).
305   * An older kit's record has only `waitsForShare`, which the view reads as `share`.
306   */
307  waitsFor?: "share" | "landing";
308  typical?: unknown;
309  tokens?: string;
310  /** Why the run stopped before the end of its queue. */
311  stopped?: string;
312  /** How a person ended the run: "sandcastle stop", "Ctrl-C", or the signal's name. Absent for a crash, a kill -9 and a run that ended by itself. */
313  stoppedBy?: string;
314  baseGates?: unknown;
315  /** Tests found red on the base mid-run, each once: a failure no branch caused, so none was repaired. */
316  baseRed?: string[];
317  /** Why a host merge check could not run for objects a partial clone lacks (`noteMissingObjects`): the closing summary says so, since those checks read the merge as clean. */
318  mergeUnchecked?: string;
319  /** Out-of-scope problems agents named in `<followup>` lines, recorded as each arrives: `id` is the ticket filed for triage, absent until it is filed (and for good in a dry run, or when filing `failed`, which a run that stopped on a `.git` change sets without trying). */
320  followUps?: { title: string; from: string; phase: string; id?: string; failed?: string }[];
321  /**
322   * The gates on the merged base. `image`: the tag they ran on, the run's own (built before any ticket landed).
323   * `failing`: the tests a red verify named, at most five (`failingMore`: it named others), read from the node:test, pytest, jest and similar output; absent when none were named.
324   * `dockerfiles`: the Dockerfiles the run's merges changed, which that image therefore lacks - absent when none.
325   * `gatedTree`: a red verify's tree is exactly one a fast-forward landing's ticket gates passed - the ticket's ref; the red is the sandbox's, not the merge's.
326   * `cleanTree`: the same for a landing merged and gated in a landing sandbox, a clean one - the red is a flaky or order-dependent test.
327   * `skipped`: the verify did not run, as the green-base record already named the merged tip: `by` is whose gates proved it and `kind` where they ran.
328   */
329  verify?: { green: boolean; line: string; image?: string; failing?: string[]; failingMore?: boolean; dockerfiles?: string[]; gatedTree?: string; cleanTree?: string; skipped?: { commit: string; by?: string; kind?: string } } | null;
330  keptWorktrees?: { issue: string; path: string }[];
331  /** Tracked files a gate rewrote and the kit put back, by path, once each. */
332  gateRewrites?: string[];
333  dryRunCheck?: string;
334  tickets?: Record<string, TicketRecord>;
335};
336
types/index.d.ts 13 lines
1/** What the band above the prompt draws from (run-state.ts, `Summary`); null while no run is live. */
2export type View = { name: string; stage: string; counts: number[]; tokens: string } | null;
3
4declare module "claude-code" {
5  interface PluginState {
6    /**
7     * `castle`: the frame of the castle the band draws, an index into run-state.ts's `CASTLE_FRAMES`.
8     * `mark`: the idle mark's line the band draws between runs (idle.ts's `markText`), null for none.
9     */
10    sandcastle: { view: View; castle: number; mark: string | null };
11  }
12}
13