XYZ-forge: /xyz-status prints recent hosted runs, tick claims and marathon/relay drivers with no model turn (read-only, GH-964)

A local-first engineering operations system for running AI coding agents as a workforce.
XYZ Forge coordinates several AI coding agents — Claude Code, Codex, agy (Google's Antigravity CLI), and others — working on the same repositories, and it carries the surrounding machinery an unattended agent run actually needs: specified work, collision-free lanes, verification gates, and a receipted record of what happened.
It started as a coordination kernel. It has become the loop around it:
capture → rate → plan → preflight → execute → gate → land → record
Every stage is a CLI verb or a skill, every artifact is on your disk, and no stage requires a server, an account, or an API key for the coordination layer itself.
Safety and warranty: XYZ Forge is provided "AS IS," without warranty, under the applicable license. Coding-agent automation is inherently risky: models may choose commands through their own runtimes and safety controls, outside the intended harness workflow. XYZ Forge cannot guarantee model behavior or data integrity; maintain tested, independent backups and follow industry-standard backup and recovery practices.
Claude consult and relay support explicit subscription checks and native reasoning effort. See Claude setup for restricted consult reads, supported CLI versions, and failure diagnostics.
Be clear-eyed about what you are adopting.
| Age | Current history rooted 2026-08-15 after a 2026-09-02 reset; the project predates it (archival tag bash-final-2026-07-28) |
| Release tags | No release tags — two archival tags only; versions are tracked in an internal ledger, not published (#452 tracks adopting them) |
| Primary operator | One, with a small collaborator set; most commits are agent-authored |
| Test suite | ~350 registered suites (count drifts weekly); npm run test:unit is a 14-case, sub-second kernel check |
| Gates | Local pre-push gate plus hosted CI on main and development |
| Runtime | Python by default since the XYZ_PYTHON flip; frozen Bash twins remain as a fallback |
| Known-weakest area | The frozen Bash fallback paths, and Python static analysis (there is none) |
This is software that has been used hard by its author against real repositories and has the scar tissue to show it — several guards in this repo exist because an earlier version destroyed work. That is a point in its favor and a warning at once. Run it on a branch. Keep independent backups.
Tested on Node 18+; requires git.
npm install # two parser deps used by the test tooling
npm run test:unit # ~1s — 14 kernel unit tests
Then, if you intend to contribute:
bash githooks/install.sh # ONCE PER CLONE — wires the pre-push gate
./validate.sh # the full suite (see the timing note)
./validate.sh runs the whole registered suite, no accounts or API keys required. Budget 5–10 minutes in parallel mode; below 4 cores it forces sequential and takes considerably longer. Run ./validate.sh --print-mode to see which mode your machine picks and why.
⚠️ Run
validate.shun-sandboxed. Under Claude Code's default Bash sandbox — or any sandboxed agent harness — the suite prints nothing for several minutes and then fails, because itsmktemp -dscratch directories are blocked. It looks like a hang, not a permissions error, and it is this repo's single most common false alarm. Turn the sandbox off for this command (/sandboxin Claude Code) before concluding anything is broken.
The pre-push hook is a correctness requirement, not optional setup. It lives in .git/hooks/, which does not travel with a clone, so a fresh clone is ungated until you install it. One install covers every branch and linked worktree of that clone. Check any time with bash githooks/install.sh --check.
tickA dependency-free Node CLI over an append-only event log at .tick/events/. Agents take path-scoped claims serialized by an O_EXCL lock, so two agents editing overlapping paths serialize instead of racing. Projection folds events into .tick/STATE.md. No server, no remote, no per-event network traffic.
Source: bin/tick, src/. Tests: test/.
Known limit: there is no stale-lock detection yet. A hard kill mid-claim leaves a lock that must be removed by hand — see the note at the top of src/lock.js.
Headless turns driven through each agent's own CLI, one turn at a time, with containment enforced by a shared core (relay-automation/relay-turn-lib.sh):
Four execution modes, in increasing order of autonomy:
| Mode | What it does | Needs |
|---|---|---|
| Consult | One question fans out to two models in parallel, isolated copies, answers reconciled. Advisory; nothing is modified. | This repo |
| Relay | Producer builds, Reviewer critiques, they hand off until the artifact converges. | This repo |
| Swarm | Multiple agents work concurrently on disjoint, path-scoped lanes. | Separate clones per automated runner |
| Marathon | A queue of preflighted work runs unattended, gated at every phase boundary. | The governance layer below |
Start at relay-automation/README.md. Live turns need each agent CLI installed and authenticated first.
Two marathon knobs worth knowing before the first unattended run: headless builders default to subscription-billed codex/agy; --builder claude remains an explicit operator choice. Claude also supports account-validated subscription mode for native consult, standalone relay review, and builds. CLAUDE_MAX_TURNS defaults to 12; CLAUDE_MAX_BUDGET (default $0.50) is an API budget, not a subscription quota guarantee. A target repo with known pre-existing test failures can pass --pre-advance-baseline <rc> (or MARATHON_GATE_BASELINE=<rc>) so the gate tolerates the existing exit code while still halting on regressions.
This is the half that grew, and the half worth explaining plainly.
Unattended agents fail for boring reasons: the work was under-specified, two lanes collided, a gate never ran, or nobody can tell afterward what actually landed. Each of those failures produced a durable surface here:
| Surface | Answers |
|---|---|
| Issue-first intake — GH issue → capture doc → parked ledger row | "Is this work written down anywhere?" |
Scored backlog — pri / sev / appeal / effort ratings | "What should run next, and why that?" |
| Wave planner — exact write-set intersection, zone caps, dependency gating | "Can these lanes run together without colliding?" |
| Preflight — freshness probes, already-landed detection, readiness verdict | "Is this specified well enough to run while I sleep?" |
| Verification gates — a gate must be able to start before turn 1, with CPU/wall/RSS caps | "Did anything actually prove this works?" |
| Release ledger, the Product Release System (PRS) — SQLite + a git-mergeable SQL dump, receipted writes | "What did we promise, and what shipped with evidence?" |
| Doc governance (PDDA) — frontmatter, status tables, ledger coverage | "Can an agent resume this work tomorrow from the docs alone?" |
| HQ — multi-repo resolution, capability tiers, previewed writes | "Do that, for project Acme, from wherever I am" |
| ~50 skills | Reusable procedures for all of the above |
Two design commitments hold this together:
MACHINE-CONTRACTS.md.The doc governance is not project-management paperwork. Agent work has to be stoppable, resumable, and handed off from PROJECT/ alone** — so the doc tree is serialized machine state, and drift between docs and code is a defect rather than untidiness. That is why the checks have teeth and why "done" means the suite is green and the doc contract holds.
The canonical product purpose is in Guiding Principles; this section explains that scope for operators.
It is: a local-first operations system for a small number of humans directing a larger number of agents across their own repositories. The operator is the sole decision authority; every write path of consequence previews first or requires an explicit gate.
It is not application lifecycle management. There is no multi-user identity, no role-based access control, no SSO, no cross-team capacity planning, no compliance certification, and no traceability matrix for regulatory submission. Those require multi-user, server-backed, permissioned state, and local-first is a deliberate constraint here, not a missing feature. If you need Polarion or Azure DevOps, you need Polarion or Azure DevOps.
It is not, yet, a spec-generation system. XYZ Forge has no structured intake → spec → PRD → project-plan generator of its own. The machinery underneath assumes specified work — PDDA's doc lifecycle expects a capture doc with a why, phases, and a verification contract, and preflight grades readiness against exactly that — but nothing here authors the spec. Writing it is left to the operator and the LLM working with them; /idea and /triage are thin front doors that scaffold the capture doc and synthesize a first-pass "why," not a rigid intake pipeline. If you want prompt-to-spec generation, adjacent tools cover it — GitHub Spec Kit for spec-driven flows, Task Master for PRD-to-task breakdown — and their output can land in PROJECT/1-INBOX/ as a capture doc.
XYZ Forge owns PDDA, its document-governance subsystem in utils/pdda/. Install governance into a target with bash utils/pdda/pdda-install.sh /path/to/repo. No separate PDDA clone is required; governance remains optional for Consult/Relay. The historical PDDA repository retains prior history. Migration and archive gates describe the cutover.
The relationship is not symmetrical, and earlier versions of this README stated it less directly:
hq fire refuses any repo that is not Tier A (PDDA and a vendored XYZ install), and the wave planner cannot rank a backlog it cannot read. PDDA is a prerequisite for the unattended path, not an optional enhancement.The dependency runs one way only: the harness reads governance structure; PDDA never calls the harness. PROJECT/PDDA.md is the shared document contract maintained here; distribution changes follow XYZ’s sync review policy. The constitution, anti-scope and mode guide are locally maintained PDDA-layer documents inherited from XYZ’s predecessor, not a blanket limit on XYZ’s product. The sync review policy is XYZ-owned and binding. ROUTER’s role split identifies their specific authority. XYZ’s canonical product purpose and principles live in GUIDING-PRINCIPLES.md; its behavioral rules live in AGENTS.md.
./install.sh ../my-app/xyz-tick --repo ../my-app
Copies the tick runtime and records the install in a machine-local registry at ~/.config/xyz/registry.tsv (never committed).
bash relay-automation/xyz-vendor.sh /path/to/target
Vendors the harness under .xyz/ in the target repo. Note: the harness alone leaves that repo at HQ Tier C, which cannot be dispatched to. For the unattended path you also need PDDA installed at the target's root — see skills/4-occasional/vendor-stack/SKILL.md for the two-step flow.
skills/ is grouped by how often a skill is used — 1-hourly/, 2-daily/, 3-weekly/, 4-occasional/ — so a directory listing reads as a usage map; the tier contract is in skills/README.md and the per-skill index in ARCHITECTURE.md → Skills Index.
Claude Code only scans ~/.claude/skills/, so skills must be symlinked in once per machine:
bash skills/1-hourly/relay-xyz/install.sh # the relay driver — start here
bash skills/2-daily/hq/install.sh # multi-repo command center
bash skills/2-daily/agent-chorus/install.sh # multi-session discussions
On a machine that runs the Skills Army HQ collection (skills/3-weekly/skills-army-hq), skip these: the collection owns those symlinks and deploys every skill at once. An installer now refuses to replace a live link it does not own (GH-678), so running one there is a no-op with a message.
Recommended minimum: 16 GB RAM for the serial marathon.sh --plan route. That covers one builder and its gate running serially with normal host reserve; it does not support per-lane parallel dispatch.
| Host RAM | Supported path |
|---|---|
| 16 GB | Serial only |
| 24 GB | Serial; small parallel wave after manual budgeting |
| 32 GB | Serial; per-lane parallel dispatch after manual budgeting |
| 64 GB | Wider parallel dispatch, still not an automatic width limit |
Measured: on a 32 GB M1 Max, 138 samples at 10-second intervals with a builder, a reviewer, and three pytest gates active showed a serial marathon at 2.19 GB average, 2.26 GB peak. Budget 1.5–2 GB per concurrent lane, then add your target repository's own test-suite memory — an unbounded term you must supply — plus host reserve.
Per-gate containment (wall clock, CPU, RSS) is enforced and kills an over-budget gate; a killed gate exits 108 and escalates distinctly, so a runaway is never triaged as a defect. Host-aware wave sizing remains the operator's responsibility — the guard does not inspect host RAM or clamp wave width.
bash relay-automation/marathon-recover.sh /path/to/target-repo
A read-only report over phase records, the tick log, and reachable commits. An UNGATED COMMIT means a phase landed a commit without an approval event: treat it as unverified, and re-run the gate or revert before trusting it. A stale driver lock (.git/relay-driver.lock left by a killed run) self-heals on the next marathon.sh / relay-drive.sh start; a LIVE lock means another driver is still running in this clone — do not delete it.
| You want to | Read |
|---|---|
| Understand the architecture | ARCHITECTURE.md |
| Work on this repo as an agent | ROUTER.md, then AGENTS.md |
| Know why it's built this way | GUIDING-PRINCIPLES.md |
| Run a relay or marathon | relay-automation/README.md |
| Operate it day to day, glossary, FAQ | HOW-TO-USE.md |
| Understand cross-subsystem contracts | MACHINE-CONTRACTS.md |
| Merge the release ledger safely | RELEASES-DB-FAQS.md |
| Use git worktrees with this | WORKTREE-SAFETY.md |
| Pick a skill for a job | ARCHITECTURE.md → Skills Index |
| See supported models and harnesses | HARNESS-MODELS-REGISTRY.md |
Project site: <https://hiqs-labs.github.io/XYZ-forge/>
AGPL-3.0-only. A commercial license is available — see LICENSE-COMMERCIAL.md.
hooks/register.ts 47 lines1// /xyz-status — read-only XYZ status with no model turn (GH-964).
2// Runs three existing readers and prints their output; it never approves,
3// blocks or rewrites anything, so it registers no tool or prompt hooks.
4import type { Register } from 'claude-code'
5
6const TIMEOUT_MS = 15_000
7
8export const register: Register = on => {
9 on('session.start', async ($, e, next) => {
10 await $.command.register({
11 name: 'xyz-status',
12 description: 'XYZ: recent hosted runs, tick claims, marathon/relay drivers (no model turn)',
13 immediate: true,
14 })
15 return next(e)
16 })
17
18 on('command.run', { command: 'xyz-status' }, async $ => {
19 const top = await $.process.run(['git', 'rev-parse', '--show-toplevel'], { timeoutMs: TIMEOUT_MS })
20 .catch(() => undefined)
21 const root = top?.exitCode === 0 ? top.stdout.trim() : ''
22 if (!root) return { text: 'xyz-status: not inside a git repository — start the session in an XYZ-forge clone.' }
23
24 // Repo-local readers run by absolute path; gh comes from PATH. All run in the root,
25 // with TICK_REPO_ROOT pinned so tick cannot follow an inherited root elsewhere.
26 const readers: [string, string[]][] = [
27 ['Hosted runs on development (recent 8)', ['gh', 'run', 'list', '--branch', 'development', '--limit', '8']],
28 ['tick claims', [`${root}/bin/tick`, 'claims']],
29 ['Marathon / relay drivers', ['bash', `${root}/relay-automation/marathon-ls.sh`]],
30 ]
31
32 const sections = await Promise.all(readers.map(async ([title, argv]) => {
33 try {
34 const r = await $.process.run(argv, { cwd: root, env: { TICK_REPO_ROOT: root }, timeoutMs: TIMEOUT_MS })
35 const out = r.stdout.trim()
36 if (r.exitCode === 0 && out) return `## ${title}\n${out}`
37 const why = r.stderr.trim().split('\n')[0] || (out ? out.split('\n')[0] : 'empty output')
38 return `## ${title}\nERROR (exit ${r.exitCode}): ${why}`
39 } catch (err) {
40 return `## ${title}\nERROR (${err instanceof Error ? err.message : String(err)})`
41 }
42 }))
43
44 return { text: [`root: ${root}`, ...sections].join('\n\n') }
45 })
46}
47