SLOPSHOPPER

supervisor

SUBAGENT SUPERVISOR: routing-and-budget Mod. Rewrites a spawn's model onto the capability graph instead of denying it, keeps a measured per-agent-type token…

newbandguardcommandstatustimer
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · supervisor
› fix the failing auth test and add an audit log call ● supervisor: ⟦supervisor⟧ armed · capability graph fable>opus>sonnet>haiku · maxSubagentCalls=40 (userConfig) · log → C:/Projects/claude-mods-rnd/prototypes/supervisor/logs/preview-session.jsonl ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /supervisor ⎿ supervisor: SUBAGENT SUPERVISOR · session preview-session ⎿ supervisor: spawns=0 rewritten=0 passed=0 denied=0 over-budget=0 live=0 ⎿ supervisor: reserved=0 tok charged=0 tok (an agent that never completes is charged its full reservation) ⎿ supervisor: main loop: 1 turns, 16,980 weighted tok · context 49% · five-hour 31% · cost $0.420 ⎿ supervisor: ROUTING LOG (0 decisions) ⎿ supervisor: none ╭──────────────────────────────────────────────────────────────────────────────────────────────────────────────────╮ │ SUBAGENT SUPERVISOR spawns=0 rewritten=0 denied=0 live=0 · reserved 0.0k / charged 0.0k · 5h │ │ ROUTE none yet · capability graph fable>opus>sonnet>haiku, advisor = +1 rung │ ╰──────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts ⚠ supervisor: supervisor · 0 spawns · 0 rewritten · 0 denied · 0 live · reserved 0.0k / charged 0.0k

Draws

Band
╭──────────────────────────────────────────────────────────────────────────────────────────────────╮ │ SUBAGENT SUPERVISOR spawns=0 rewritten=0 denied=0 live=0 · reserved 0.0k / charged 0.0k · 5h │ │ ROUTE none yet · capability graph fable>opus>sonnet>haiku, advisor = +1 rung │ ╰──────────────────────────────────────────────────────────────────────────────────────────────────╯
README

SUBAGENT SUPERVISOR

A Claude Code hooks module (function hooks, build 2.1.273) that does the routing-and-budget job that four separate classic PreToolUse guards do today in the production harness — agent-model-guard, subagent-budget-guard, capability-graph-guard, fable-delegate-guard — in one file, with three capabilities none of them has: it rewrites instead of denying, it measures instead of parsing a declaration, and it remembers across sessions.

None of those four guards is modified by this prototype. This is a parallel implementation in C:\Projects\claude-mods-rnd\prototypes\supervisor\, loaded only with --plugin-dir.


What it is

  1. Capability-graph routing on agent.spawn. Fable → Opus/Sonnet/Haiku, Opus → Sonnet/Haiku, Sonnet → Haiku, Haiku → nobody. Downward only; peers are not edges. The one upward edge is the advisor agent type, which is set to one rung above the parent (Sonnet → Opus, Opus → Fable, Fable → Fable) and is never capped. A spawn that violates the graph has its model field rewritten — next({ ...e, model: "haiku" }) — not refused. The only deny is a spawn from a Haiku loop, because Haiku is a leaf and no rewrite of model can make a leaf into a parent. The parent's model comes from e.parentModel, which the event carries; no registry, no transcript tail.
  2. A measured budget ledger in $.store under the key supervisor.ledger. Every spawn reserves tokens from a prior (17,000 for a lean agent type, 60,000 otherwise — the subagent-budget-guard constants). Every subagent turn.step adds its usage to that agent's live total; the subagent's turn.complete settles the row with the real number and pushes it into a running median for that subagent type, persisted across sessions. An agent that never completes is charged its full reservation. Where the Agent prompt carries the harness's # EST: <n> calls, <n> files marker, the ledger prints declared vs measured and the delta.
  3. A per-subagent tool-call cap. Every tool.call that carries an agentId is counted against that agent's row. Past maxSubagentCalls (a userConfig number field, default 40) the hook answers { deny: "supervisor: subagent over budget" } for that agent's further calls and leaves every other loop alone.
  4. A HUD and a command. An AbovePrompt band shows live subagents (type, model, calls, tokens so far), the last routing decision with its reason codes, the five-hour rate-limit percentage and the session cost. /supervisor [log|ledger|priors|all] (an immediate: true command) prints the routing log, the ledger and the learned priors.
  5. A JSONL routing log at prototypes/supervisor/logs/<sessionId>.jsonl, one object per session start, routing decision, settlement and budget denial.

Every decision carries an Agent-Task-Router style reasons[] array — e.g. ["peer-edge-not-in-graph", "downgraded-to-highest-child-rung"] — which is what appears in the HUD, the /supervisor log and the JSONL.


Architecture

                         Claude Code engine (build 2.1.273)
                                      |
   ┌──────────────────────────────────┴───────────────────────────────────┐
   │                        hooks worker, tier "user"                     │
   │                                                                      │
   │  session.start ──► $.session.id, $.store.get("supervisor.ledger"),   │
   │                    $.command.register("supervisor"),                 │
   │                    $.clock.every(2000) ──► flush + status + redraw   │
   │                                                                      │
   │  agent.spawn ──► decide(e, fableUsed)                                │
   │        │           rungOf(e.parentModel)  vs  rungOf(e.model)        │
   │        │           advisor?  → parentRung + 1                        │
   │        │           parent is haiku? → {deny}      ◄── the only deny  │
   │        │           else clamp to parentRung - 1                      │
   │        ├─ deny  ─► return {deny: reason}      (chain below never runs)│
   │        └─ else  ─► next({...e, model: target}) ─► core starts agent  │
   │                      │                                               │
   │                      └─ r.agentId ─► ledger row, reserve = prior(type)│
   │                                                                      │
   │  turn.step  (async generator, yield* next(e))                        │
   │        └─ e.agentId ─► row.tokens += weigh(r.usage)   [live]         │
   │                                                                      │
   │  turn.complete                                                       │
   │        ├─ e.agentId ─► settle: measured, median(type), $.store.set   │
   │        └─ main loop ─► session totals + $.session.usage()            │
   │                                                                      │
   │  tool.call                                                           │
   │        └─ e.agentId in ledger ─► row.calls++                         │
   │              calls > maxSubagentCalls ─► {deny: "...over budget"}    │
   │                                                                      │
   │  command.run{command:"supervisor"} ─► {text: log + ledger + priors}  │
   │  ui.render{AbovePrompt, terminal}  ─► HUD tree                       │
   └──────────────────────────────────────────────────────────────────────┘
                                      |
                       $.fs.write ──► logs/<sessionId>.jsonl

Middleware order. Every hook here is a wrapper: it calls next(e) and post-processes, except the two enforcement paths (agent.spawn deny, tool.call over-budget deny) which answer without calling next, so nothing beneath them runs. Load this plugin first (--plugin-dir supervisor before any other) if you also load a recorder: a plugin that answers without next is invisible to everything loaded after it. That lesson is the lab's own (lab/README.md: blackbox loaded after guardian reported "0 denied").


Run it

Validate:

claude plugin validate C:\Projects\claude-mods-rnd\prototypes\supervisor --json

Headless routing proof (a Sonnet loop asks for an Opus subagent; the supervisor rewrites it to Haiku):

claude --plugin-dir C:/Projects/claude-mods-rnd/prototypes/supervisor ^
  -p "Use the Agent tool exactly once. subagent_type haiku-scout, model opus, prompt: '# EST: 1 calls, 0 files # SPAWN_OK: prototype dogfood. Run the Bash command: echo scout-alive  # SEQ: probe . Then reply ALIVE.' Then reply PARENT-DONE." ^
  --model sonnet --allowedTools "Bash,Read,Agent,Task" --debug

(The production harness's own capability-graph-guard denies a Sonnet→Opus Agent call before the engine ever raises agent.spawn. Set its documented per-session off switch — CAPABILITY_GRAPH_GUARD=off — for the probe, or the prototype never gets the event to correct.)

Interactive, to see the HUD (the stock lab driver, unmodified):

python C:\Projects\claude-mods-rnd\tools\pty_drive.py out.txt 25 ^
  "C:/Projects/claude-mods-rnd/prototypes/supervisor" ^
  "<prompt that spawns one subagent>|||WAIT:25|||KEYS:/supervisor\r|||WAIT:8" 75 "" --model sonnet --debug

pty_drive.py hard-codes --model haiku into the command line, and a Haiku loop cannot spawn, so the HUD's interesting rows would stay empty. A second --model wins: verified directly — claude --model haiku --model sonnet -p "Reply with only your model id, nothing else." printed claude-sonnet-5. So --model sonnet is passed as an extra argument rather than editing the shared driver.


The four guards this supersedes

GuardWhat it didHow the supervisor does it nowWhat is lost
agent-model-guard.cjs (16 KB, PreToolUse on `Agent\Task\Workflow`)Denied any spawn whose prompt/model did not declare a model; capped Fable at 3 spawns/session; statically parsed Workflow JS source for agent( call sites, requiring an inline model: on each and Fable ones at module top level.agent.spawn carries model as a rewritable field. An undeclared model is not an error: reason model-undeclared is recorded and the model is set to the highest legal child rung. The Fable cap survives as a counter (FABLE_CAP = 3, reason fable-cap-3-per-session) applied to non-advisor spawns.The Workflow static parse. A Workflow({name}) runs its agent() calls through the engine, so each one does surface as an agent.spawn and gets corrected at runtime — but the guard's lint (a bare agent() blocks the whole script before it runs, Fable must be top-level and outside every fan-out) has no runtime equivalent. That half stays a source-lint job.
subagent-budget-guard.cjs (15 KB, PreToolUse + PostToolUse)Required # EST: <n> calls, <n> files in every Agent prompt; priced the declaration against inline work with fixed constants (LEAN_SPAWN=17000, FULL_SPAWN=60000, RESULT_TOKENS=2000, FILE_TOKENS=2000, CACHE_DISCOUNT=0.1); denied a spawn below break-even, showing the arithmetic; logged declared-vs-actual to .subagent-budget-log.jsonl for calibrate.cjs to refit offline. Its own header: "a PreToolUse hook sees only tool_input — never the files, never the result. So it does not guess."The same constants are the prior only. turn.step.usage and turn.complete.usage per agentId are the real cost; each settlement pushes into a running median per subagent type held in $.store under supervisor.ledger, so the constants refit themselves in-session with no separate calibrate.cjs pass. # EST: is still parsed, but only to print declared-vs-measured and the delta.The pre-spawn deny. The supervisor never refuses a spawn for being too small, because by the time it can measure, the spawn has happened. The economics become a visible ledger rather than a gate — which is the fable-delegate-guard lesson applied (see below). If a hard pre-spawn gate is wanted, the declared estimate is available in the same handler and {deny} is one line; it is deliberately not wired.
capability-graph-guard.cjs (18 KB, PreToolUse + SubagentStart + SubagentStop)Enforced the same graph, but needs the caller's model, which is in no PreToolUse payload. It maintained an agent_id → model registry across SubagentStart/SubagentStop — and the real SubagentStart payload carries subagent_config: null, so the registry alone resolves nothing. Denied a violating edge; rewrote the advisor's model through the updatedInput envelope.e.parentModel is the caller's model, pinned on the event. Three files of payload archaeology collapse into rungOf(e.parentModel). Every violating edge is rewritten rather than denied; only a Haiku parent is refused. advisor is min(4, parentRung + 1), uncapped, as the rule requires.Nothing functional — this is the clean win. The registry, the SubagentStart/SubagentStop wiring and the updatedInput envelope all disappear. One thing changes in kind: the guard's deny was final for a violating edge, the supervisor's rewrite is silent correction, so a model that intended Opus gets Haiku and is told so only through the HUD/ui.log line, not through a tool error.
fable-delegate-guard.cjs (16 KB, PreToolUse on edit/shell tools)Blocked nothing. Briefed a Fable main loop once per session on delegation economics via additionalContext, and logged hand-work for --report. Its retirement evidence: 1,046 events, 714 # FABLE_OK overrides, 266 shell denials (including npm test and a read-only grep), edit denials retried 5–7× on one file — "a model treats a deny like a transient error."The economics live in the AbovePrompt HUD and /supervisor: reserved vs charged tokens, live subagents with calls and tokens, the last routing decision. Advisory that stays visible costs no tokens per tool call and cannot be retried against.The once-per-session text briefing, which reached the model's context; the HUD reaches the human. prompt.section/prompt.context — the events that would put text into the model's system prompt — are withheld from the user tier on this machine (sec-default@builtin), so a Mod cannot reproduce the injected briefing at all. prompt.submit hidden context is the available substitute and is deliberately not used here (it would spend tokens every turn).

What it proves

  • agent.spawn is a sufficient seam for the whole routing problem: the parent's model, the requested model, the agent type, the prompt and the # EST: marker are all on one event, and model is rewritable. The four guards' shared workaround — reconstructing "who is calling, on what model" from registries and transcript tails — is unnecessary.
  • A rewrite is strictly better than a deny for a routing violation. The spawn still happens, on a legal model, and no retry loop starts. The contrast is in the same evidence set: two rewritten spawns produced zero retries, while one denied tool call produced four retries of the same call before the exit tool was exempted.
  • The upward edge works in the same handler as the downward one. advisor requested with model haiku from a Sonnet loop was rewritten up to Opus and core started claude-opus-5[1m] — one function, both directions, no separate envelope.
  • $.store makes the estimate self-correcting across sessions. Four consecutive sessions each opened with the previous session's median for haiku-scout and reserved against it (17,000 → 26,704 → 22,902 → 22,902), with no external calibration pass.
  • Subagent cost is measurable in-session: turn.step and turn.complete carry usage per agentId, so a declared estimate can be scored against the real number in the same session that made it, and the prior can refit itself in $.store.
  • tool.call scoped by agentId gives per-subagent enforcement that a PreToolUse hook cannot express, because a PreToolUse hook has no way to attribute a call to a particular live subagent.

What it does NOT prove

  • The Fable rung. Every observed spawn had a Sonnet parent. Fable → Opus/Sonnet/Haiku, Opus → Fable advisor escalation and the FABLE_CAP = 3 counter are implemented but never fired; there is no Fable spawn anywhere in the evidence.
  • The Haiku-parent deny. A Haiku loop cannot dispatch the Agent tool at all in this build, so the haiku-is-a-leaf branch — the only agent.spawn {deny} in the module — has never executed. Treat it as untested code.
  • That a tool-call deny actually stops an over-budget subagent. It does not. The denied agent keeps trying (see the retry finding in Evidence); what the cap buys is a hard ceiling on that agent's real work, not an orderly shutdown. The exit tool had to be exempted so the agent could finish at all.
  • Displacing the real guards. Nothing in the production harness was changed, and the probes had to turn one of the four off (CAPABILITY_GRAPH_GUARD=off) to get the event at all. A real migration means removing four settings.json entries, which this prototype does not do.
  • The agent-model-guard Workflow lint. No Workflow script was run through this Mod.
  • Cost in dollars. $.session.usage().cost.usd is whole-session, not per-subagent; the ledger is in weighted tokens, not money. The weigh() discount (0.1 × cache reads) is inherited from subagent-budget-guard, not re-derived here.
  • Concurrent subagents. Every probe spawned agents one at a time. Two live rows in the ledger, and two hooks racing on the same $.store key, were never exercised.

API assumptions

Every event and $ method used, with its status from MOD_CAPABILITY_MAP.md:

UsedStatus in the capability mapObserved here
on("session.start")CONFIRMED WORKING [RUN]yes — session.start settled in 46.8ms
on("agent.spawn") + next({...e, model})CONFIRMED (model forced, agentId returned) [RUN]yes — rewrite observed
agent.spawn {deny}§9 [DECL]no — coded, never fired
e.parentModel (pinned)§12 CONFIRMED [RUN]yes
on("turn.step") async generator, r.usage, e.agentIdCONFIRMED [RUN]yes
on("turn.complete"), e.usage, e.agentIdCONFIRMED incl. subagent turns [RUN]yes
on("tool.call") {deny} scoped by e.agentIdCONFIRMED [RUN] (deny), agentId pinned [RUN]counting yes, deny no
on("command.run", {command}) + $.command.register({immediate:true})CONFIRMED [RUN, registered]yes
on("ui.render", {component:"AbovePrompt", surface:"terminal"}) + $.ui.resolveCONFIRMED drawn [RUN]yes (interactive)
$.store.get / $.store.setPRESENT BUT EXPERIMENTAL [DECL]yes — round-tripped
$.clock.everyPRESENT BUT EXPERIMENTAL [DECL]yes
$.fs.writeCONFIRMED [RUN]yes
$.session.id / $.session.usageCONFIRMED [RUN]yes
$.ui.log / $.ui.status / $.ui.invalidateCONFIRMED [RUN]yes (ui.status is a no-op headless: "no status row in a headless session")
register(on, options) from plugin.json userConfig[DECL], not exercisedyes — newly confirmed, see Evidence

Assumptions that are policy, not API, and would be wrong in another harness:

  • Model rung is decided by substring: fable > opus > sonnet > haiku. A model id containing none of those is treated as Opus (reason parent-model-unrecognised-assumed-opus) — the conservative choice, since it permits Sonnet and Haiku children.
  • fork inherits the parent and its model is ignored by the engine, so a fork is passed through untouched with reason fork-inherits-parent.
  • advisor is the only upward edge and is never capped.
  • LEAN_TYPES is a hard-coded list of this harness's lean agent types.

Failure behaviour

  • A hook throws. The engine skips that hook and the chain continues beneath it; the transcript gets one dim line and ~/.claude/debug/<session>.txt gets the reason. Concretely: if decide() threw, the spawn would proceed unrouted — fail-open. Every $ call that can fail at runtime ($.fs.write, $.store.set, $.session.usage) is wrapped in try/catch so a full log directory or a store write error cannot take the routing with it.
  • The API is missing or renamed. The loader statically scans the module; an unknown event name in on("...") or a $ method that does not exist fails the load, and claude plugin validate reports it before a session ever starts. A field that vanishes (say parentModel) does not fail the load — rungOf(undefined) returns 0 and the parent-model-unrecognised-assumed-opus branch takes over, which routes everything as if the parent were Opus. That is the designed degradation.
  • Panes cannot draw. This uses AbovePrompt, not $.ui.open, precisely because an unasked pane waits unplaced below 144 terminal columns. The AbovePrompt band always draws. In headless (-p) mode there is no UI at all: ui.render never fires, $.ui.status logs "no status row in a headless session; kept for the next surface", and $.ui.log goes to the debug log. Enforcement and the JSONL are unaffected — the HUD is the only thing that is display-only.
  • The engine disables function hooks. After N worker crashes the engine turns function hooks off for the session. The supervisor then enforces nothing; there is no residual file-based fallback, by design. The production guards it supersedes would have to stay armed during any real migration.

Migration path if the API changes

Four functions are the adapter boundary; everything else is policy or presentation.

FunctionDepends onIf the API changes
rungOf(model)model-id stringsthe only place model naming is understood
decide(e, fableUsed)e.parentModel, e.model, e.subagentType, e.forkpure function of the event; the whole routing policy, testable without the engine
weigh(usage)TurnUsage field names (input_tokens, output_tokens, cache_read_input_tokens, cache_creation_input_tokens)one function to change if the usage shape moves
flush($) / persist($)$.fs.write, $.store.setthe only I/O; swap for $.http.fetch or a different store here

The event surface itself is five lines (on("agent.spawn"), on("turn.step"), on("turn.complete"), on("tool.call"), plus the two UI registrations). If agent.spawn loses its rewritability, the same decide() output feeds {deny} with the reasons as the message — i.e. it degrades exactly into what capability-graph-guard does today. The C:\Projects\claude-mods-rnd\adapter\claude-runtime shared adapter is not used: this prototype hooks engine events directly, so there is one less moving part between the policy and the engine.


Evidence

All of it from the real claude binary on this machine, 2026-09-16. Raw files in prototypes/supervisor/evidence/; the plugin's own JSONL in prototypes/supervisor/logs/.

0. Validation

claude plugin validate C:\Projects\claude-mods-rnd\prototypes\supervisor --json → exit 0:

"success": true
"./index.tsx hooks: session.start, agent.spawn, turn.step, turn.complete, tool.call,
                    command.run{command=supervisor}, ui.render{component=AbovePrompt, surface=terminal}"

First attempt failed on the manifest, not the module: userConfig.maxSubagentCalls.title — the field key is title, not label.

1. Load (headless smoke, session 5095533d) — evidence/01-*

[DEBUG] hooks module supervisor loaded (worker, environment 2, tier user); events: session.start,agent.spawn,turn.step,turn.complete,tool.call,command.run,ui.render
[DEBUG] hooks module supervisor session.start settled in 46.8ms (worker hop, next() included)
[DEBUG] plugin supervisor: no pluginConfigs["supervisor" or "supervisor@inline"]

and yet 01-smoke-log.jsonl line 1: "maxCalls":40,"maxCallsSource":"userConfig". With no stored config at all, register(on, options) still received the manifest's declared default. Zero hook failures in the log (grep -iE "fail|error|skip|refus|threw" over the 34 supervisor lines: empty).

2. Routing, both directions (session b49fdd17) — evidence/02-*

A Sonnet main loop asked for one Opus haiku-scout and one Haiku advisor. The engine's own debug lines are the proof — a hook, not the guard chain, changed the model:

[DEBUG] agent.spawn haiku-scout: model opus -> haiku by a hook
[DEBUG] agent.spawn advisor:     model haiku -> opus by a hook

02-routing-log.jsonl, trimmed:

{"kind":"route","route":{"type":"haiku-scout","parent":"claude-sonnet-5","parentRung":"sonnet",
 "requested":"opus","model":"haiku","action":"rewrite",
 "reasons":["upward-edge-not-in-graph","downgraded-to-highest-child-rung"],
 "resolvedModel":"claude-haiku-4-5-20251001"},"reserved":26704,"reservedFrom":"learned n=1"}

{"kind":"route","route":{"type":"advisor","parent":"claude-sonnet-5","parentRung":"sonnet",
 "requested":"haiku","model":"opus","action":"rewrite",
 "reasons":["advisor-upward-edge","advisor-uncapped","advisor-model-overridden","set-one-rung-above-parent"],
 "resolvedModel":"claude-opus-5[1m]"},"reserved":17000,"reservedFrom":"prior lean"}

resolvedModel is what core returned, so the rewrite was not merely accepted: an actual Haiku subagent and an actual Opus advisor ran. Neither spawn was denied; both happened, on legal models.

Settlements in the same file, declared vs measured:

{"kind":"settle","type":"advisor",    "calls":1,"declared":17000,"reserved":17000,"measured":21081,"delta
Source 1 files
hooks/index.tsx 498 lines
1// SUBAGENT SUPERVISOR — routing and budget for subagents, as one hooks module.
2//
3// Replaces four classic PreToolUse guards (agent-model-guard, subagent-budget-guard,
4// capability-graph-guard, fable-delegate-guard). All four reconstructed the same two invisible
5// facts — who is calling, on what model — from transcript tails and registry files, and could
6// only DENY. `agent.spawn` carries `parentModel` and a rewritable `model`, so this Mod corrects
7// instead of refusing; per-agent `turn.step`/`turn.complete` usage turns the declared `# EST:`
8// estimate into a measured one; `$.store` carries the fitted priors across sessions.
9//
10// Events hooked: session.start, agent.spawn, turn.step (streaming), turn.complete, tool.call,
11// command.run{supervisor}, ui.render{AbovePrompt,terminal}.
12
13const LOG_DIR = "C:/Projects/claude-mods-rnd/prototypes/supervisor/logs/";
14const STORE_KEY = "supervisor.ledger";
15
16// ---------------------------------------------------------------- capability graph
17// Fable -> Opus/Sonnet/Haiku, Opus -> Sonnet/Haiku, Sonnet -> Haiku, Haiku -> nobody.
18// Downward only; peers are not edges. The one upward edge is `advisor`, one rung ABOVE the parent.
19const RUNG_NAME = ["?", "haiku", "sonnet", "opus", "fable"];
20const ADVISOR_TYPES = ["advisor"];
21const FORK_TYPES = ["fork"];
22const FABLE_CAP = 3; // non-advisor fable spawns per session (advisor consultations are uncapped)
23
24// ---------------------------------------------------------------- budget priors
25// Constants inherited from subagent-budget-guard; they are only the PRIOR here — the ledger
26// learns a running median per subagent type from observed usage and overrides them.
27const LEAN_PRIOR = 17000;
28const FULL_PRIOR = 60000;
29const CALL_TOKENS = 2000;
30const FILE_TOKENS = 2000;
31const CACHE_DISCOUNT = 0.1;
32const MAX_SAMPLES = 20;
33// Never denied by the call cap: SubagentHandback is how a subagent returns its answer. Denying the
34// exit door does not stop an over-budget agent, it makes it retry the handback (observed: 4 retries
35// in evidence/03, the same "a deny reads as a transient error" pathology fable-delegate-guard logged).
36const EXEMPT_TOOLS = ["SubagentHandback"];
37const LEAN_TYPES = ["haiku-scout", "sonnet-implementer", "opus-owner", "advisor", "e2e-verifier", "Explore", "Plan", "statusline-setup", "security-reviewer", "fork"];
38
39const state = {
40  sessionId: "",
41  maxCalls: 40,
42  maxCallsSource: "default",
43  agents: {},        // agentId -> ledger row
44  order: [],         // agentIds in spawn order
45  routes: [],        // routing decisions (the routing log)
46  lastRoute: null,
47  priors: {},        // subagentType -> { samples: number[], median: number }
48  rows: [],          // JSONL rows
49  dirty: false,
50  t0: 0,
51  usage: null,
52  totals: { spawns: 0, rewrites: 0, denies: 0, passes: 0, fable: 0, overBudget: 0, mainTurns: 0, mainTokens: 0 },
53};
54
55// ---------------------------------------------------------------- pure helpers
56function rungOf(model) {
57  if (!model) return 0;
58  const m = String(model).toLowerCase();
59  if (m.indexOf("fable") >= 0) return 4;
60  if (m.indexOf("opus") >= 0) return 3;
61  if (m.indexOf("sonnet") >= 0) return 2;
62  if (m.indexOf("haiku") >= 0) return 1;
63  return 0;
64}
65function num(n) { return String(Math.round(n || 0)).replace(/\B(?=(\d{3})+(?!\d))/g, ","); }
66function k(n) { return (Math.round((n || 0) / 100) / 10).toFixed(1) + "k"; }
67function shortId(id) { return id ? String(id).slice(0, 8) : "-"; }
68function clip(v, n) {
69  let s;
70  try { s = typeof v === "string" ? v : JSON.stringify(v); } catch (err) { s = String(v); }
71  s = (s || "").replace(/\s+/g, " ");
72  return s.length > n ? s.slice(0, n - 1) + "\u2026" : s;
73}
74function parseEst(prompt) {
75  const m = /#\s*EST:\s*(\d+)\s*calls?\s*,\s*(\d+)\s*files?/i.exec(prompt || "");
76  if (!m) return null;
77  return { calls: Number(m[1]), files: Number(m[2]) };
78}
79function medianOf(arr) {
80  const s = (arr || []).slice().sort(function (a, b) { return a - b; });
81  const n = s.length;
82  if (!n) return 0;
83  return n % 2 ? s[(n - 1) / 2] : Math.round((s[n / 2 - 1] + s[n / 2]) / 2);
84}
85function isLean(type) { return LEAN_TYPES.indexOf(type) >= 0; }
86function priorFor(type) {
87  const p = state.priors[type];
88  if (p && p.median > 0) return p.median;
89  return isLean(type) ? LEAN_PRIOR : FULL_PRIOR;
90}
91function priorSource(type) {
92  const p = state.priors[type];
93  if (p && p.median > 0) return "learned n=" + p.samples.length;
94  return isLean(type) ? "prior lean" : "prior full";
95}
96function weigh(u) {
97  if (!u) return 0;
98  return (u.input_tokens || 0) + (u.output_tokens || 0) + (u.cache_creation_input_tokens || 0)
99    + Math.round(CACHE_DISCOUNT * (u.cache_read_input_tokens || 0));
100}
101// The subagent-budget-guard arithmetic, evaluated ONCE at spawn time and frozen on the ledger row:
102// recomputing it later would price the declaration against a prior the same agent just moved.
103function declaredCost(type, est) {
104  if (!est) return null;
105  return priorFor(type) + est.calls * CALL_TOKENS + est.files * FILE_TOKENS;
106}
107function charged(a) { return a.settled ? a.measured : a.reserved; }
108
109// The whole routing policy, as one pure function over the agent.spawn input.
110// Returns { action: "pass" | "rewrite" | "deny", model, reason, reasons[], parentRung, reqRung, targetRung }.
111function decide(e, fableUsed) {
112  const reasons = [];
113  const type = e.subagentType || "general-purpose";
114  const requested = e.model;
115  const reqRung = rungOf(requested);
116  let parentRung = rungOf(e.parentModel);
117
118  if (e.fork || FORK_TYPES.indexOf(type) >= 0) {
119    reasons.push("fork-inherits-parent");
120    reasons.push("model-field-ignored-by-engine");
121    return { action: "pass", model: requested, reason: "", reasons: reasons, parentRung: parentRung, reqRung: reqRung, targetRung: parentRung };
122  }
123  if (parentRung === 0) {
124    reasons.push("parent-model-unrecognised-assumed-opus");
125    parentRung = 3;
126  }
127  if (parentRung === 1) {
128    reasons.push("haiku-is-a-leaf");
129    reasons.push("no-rewrite-can-fix-a-leaf");
130    return {
131      action: "deny",
132      model: requested,
133      reason: "supervisor: haiku is a leaf in the capability graph (Haiku -> nobody); this loop may not spawn subagents.",
134      reasons: reasons, parentRung: parentRung, reqRung: reqRung, targetRung: 0,
135    };
136  }
137  if (ADVISOR_TYPES.indexOf(type) >= 0) {
138    const target = Math.min(4, parentRung + 1);
139    reasons.push("advisor-upward-edge");
140    reasons.push("advisor-uncapped");
141    if (!reqRung) reasons.push(requested ? "model-unrecognised" : "model-undeclared");
142    else if (reqRung !== target) reasons.push("advisor-model-overridden");
143    reasons.push("set-one-rung-above-parent");
144    return {
145      action: reqRung === target ? "pass" : "rewrite",
146      model: RUNG_NAME[target], reason: "", reasons: reasons,
147      parentRung: parentRung, reqRung: reqRung, targetRung: target,
148    };
149  }
150
151  let target = reqRung;
152  if (!reqRung) {
153    reasons.push(requested ? "model-unrecognised" : "model-undeclared");
154    target = parentRung - 1;
155  }
156  if (target === 4 && fableUsed >= FABLE_CAP) {
157    reasons.push("fable-cap-" + FABLE_CAP + "-per-session");
158    target = 3;
159  }
160  if (target >= parentRung) {
161    reasons.push(target === parentRung ? "peer-edge-not-in-graph" : "upward-edge-not-in-graph");
162    reasons.push("downgraded-to-highest-child-rung");
163    target = parentRung - 1;
164  }
165  if (target === reqRung && reqRung > 0) reasons.push("edge-in-graph");
166  return {
167    action: target === reqRung ? "pass" : "rewrite",
168    model: RUNG_NAME[target], reason: "", reasons: reasons,
169    parentRung: parentRung, reqRung: reqRung, targetRung: target,
170  };
171}
172
173function push(row) {
174  state.rows.push(Object.assign({ at: Date.now() - state.t0 }, row));
175  state.dirty = true;
176}
177function bumpAgent(agentId, usage) {
178  const a = state.agents[agentId];
179  if (!a) return;
180  a.steps++;
181  a.tokens += weigh(usage);
182  state.dirty = true;
183}
184function settleAgent(e) {
185  const a = state.agents[e.agentId];
186  if (!a) return null;
187  a.turns++;
188  a.measured += weigh(e.usage);
189  a.durationMs += e.durationMs || 0;
190  a.settled = true;
191  a.reason = e.reason;
192  const p = state.priors[a.type] || { samples: [], median: 0 };
193  p.samples.push(a.measured);
194  if (p.samples.length > MAX_SAMPLES) p.samples.shift();
195  p.median = medianOf(p.samples);
196  state.priors[a.type] = p;
197  state.dirty = true;
198  return a;
199}
200function liveAgents() {
201  return state.order.map(function (id) { return state.agents[id]; }).filter(function (a) { return a && !a.settled; });
202}
203function totalsLine() {
204  let reserved = 0, measured = 0;
205  state.order.forEach(function (id) {
206    const a = state.agents[id];
207    if (!a) return;
208    reserved += a.reserved;
209    measured += charged(a);
210  });
211  return { reserved: reserved, measured: measured };
212}
213function usageBits() {
214  const u = state.usage;
215  if (!u) return { five: "?", cost: "?", ctx: "?" };
216  const five = (u.rateLimits || []).filter(function (x) { return x.kind === "five_hour"; })[0];
217  return {
218    five: five ? Math.round(five.percentUsed) + "%" : "?",
219    cost: u.cost ? "$" + u.cost.usd.toFixed(3) : "?",
220    ctx: u.context && u.context.percent !== undefined ? u.context.percent + "%" : "?",
221  };
222}
223function routeLine(r) {
224  return (r.ms / 1000).toFixed(1).padStart(7) + "s  " + String(r.type).padEnd(20)
225    + " parent=" + String(r.parent).padEnd(7)
226    + " requested=" + String(r.requested === undefined ? "inherit" : r.requested).padEnd(8)
227    + " -> " + String(r.action === "deny" ? "DENIED" : r.model).padEnd(8)
228    + " [" + r.reasons.join(", ") + "]"
229    + (r.agentId ? "  agent=" + shortId(r.agentId) : "");
230}
231function ledgerLine(a) {
232  const dec = a.declared;
233  const ch = charged(a);
234  return "  " + shortId(a.agentId).padEnd(9) + String(a.type).padEnd(20) + String(a.model).padEnd(26)
235    + (a.calls + "/" + state.maxCalls).padEnd(8)
236    + (dec === null ? "no # EST:" : num(dec)).padStart(11)
237    + num(a.reserved).padStart(11)
238    + (a.settled ? num(a.measured) : num(a.tokens) + "*").padStart(11)
239    + (dec === null ? "" : "  delta=" + (ch - dec > 0 ? "+" : "") + num(ch - dec))
240    + (a.settled ? "" : "  LIVE")
241    + (a.overBudget ? "  OVER-BUDGET" : "");
242}
243function statusText() {
244  const t = totalsLine();
245  return "supervisor \u00b7 " + state.totals.spawns + " spawns \u00b7 " + state.totals.rewrites + " rewritten \u00b7 "
246    + state.totals.denies + " denied \u00b7 " + liveAgents().length + " live \u00b7 reserved " + k(t.reserved) + " / charged " + k(t.measured);
247}
248
249// ---------------------------------------------------------------- $-using helpers (top level)
250async function flush($) {
251  if (!state.dirty || !state.sessionId) return;
252  state.dirty = false;
253  try {
254    await $.fs.write(LOG_DIR + state.sessionId + ".jsonl", state.rows.map(function (r) { return JSON.stringify(r); }).join("\n") + "\n");
255  } catch (err) { /* a hooks module never blocks the engine */ }
256}
257async function persist($) {
258  try {
259    await $.store.set(STORE_KEY, { priors: state.priors, updatedAt: Date.now(), sessionId: state.sessionId });
260  } catch (err) { /* store is best-effort */ }
261}
262function tick($) {
263  void flush($);
264  $.ui.status(statusText());
265  $.ui.invalidate("ui.render");
266}
267async function refreshUsage($) {
268  try { state.usage = await $.session.usage(); } catch (err) { /* usage is advisory */ }
269}
270
271// ---------------------------------------------------------------- module
272export const register = (on, options) => {
273  const configured = options && options.maxSubagentCalls;
274  if (typeof configured === "number" && configured > 0) { state.maxCalls = configured; state.maxCallsSource = "userConfig"; }
275
276  on("session.start", async ($, e, next) => {
277    state.sessionId = await $.session.id();
278    state.t0 = Date.now();
279    const stored = await $.store.get(STORE_KEY);
280    if (stored && typeof stored === "object" && stored.priors) state.priors = stored.priors;
281    await $.command.register({
282      name: "supervisor",
283      description: "SUBAGENT SUPERVISOR: routing log + measured budget ledger (log | ledger | priors | all)",
284      argumentHint: "[log|ledger|priors|all]",
285      immediate: true,
286    });
287    push({ kind: "session", sessionId: state.sessionId, maxCalls: state.maxCalls, maxCallsSource: state.maxCallsSource, priorTypes: Object.keys(state.priors) });
288    $.ui.status(statusText());
289    $.clock.every(2000, () => tick($));
290    $.ui.log("\u27e6supervisor\u27e7 armed \u00b7 capability graph fable>opus>sonnet>haiku \u00b7 maxSubagentCalls=" + state.maxCalls
291      + " (" + state.maxCallsSource + ") \u00b7 log \u2192 " + LOG_DIR + state.sessionId + ".jsonl");
292    await refreshUsage($);
293    return next(e);
294  });
295
296  // ---- 1. routing: rewrite onto the capability graph, deny only a leaf ----
297  on("agent.spawn", async ($, e, next) => {
298    const d = decide(e, state.totals.fable);
299    const est = parseEst(e.prompt);
300    const type = e.subagentType || "general-purpose";
301    const route = {
302      ms: Date.now() - state.t0,
303      type: type,
304      parent: e.parentModel,
305      parentRung: RUNG_NAME[d.parentRung] || String(e.parentModel),
306      requested: e.model,
307      model: d.model,
308      action: d.action,
309      reasons: d.reasons,
310      est: est,
311      agentId: null,
312    };
313    state.totals.spawns++;
314
315    if (d.action === "deny") {
316      state.totals.denies++;
317      state.routes.push(route);
318      state.lastRoute = route;
319      push({ kind: "route", route: route, denied: d.reason });
320      $.ui.log("\u27e6supervisor\u27e7 DENY " + type + " parent=" + e.parentModel + " [" + d.reasons.join(", ") + "]");
321      await flush($);
322      return { deny: d.reason };
323    }
324
325    const sent = d.action === "rewrite" ? Object.assign({}, e, { model: d.model }) : e;
326    if (d.action === "rewrite") state.totals.rewrites++; else state.totals.passes++;
327    // Advisor consultations land on fable but are explicitly uncapped, so they do not feed the counter.
328    if (d.targetRung === 4 && ADVISOR_TYPES.indexOf(type) < 0) state.totals.fable++;
329
330    // Priced before the spawn, against the prior in force now.
331    const reserved = priorFor(type);
332    const reservedFrom = priorSource(type);
333    const declared = declaredCost(type, est);
334    route.reserved = reserved;
335    route.declared = declared;
336
337    const r = await next(sent);
338    const agentId = r && r.agentId ? r.agentId : null;
339    route.agentId = agentId;
340    route.resolvedModel = r ? r.model : null;
341    state.routes.push(route);
342    state.lastRoute = route;
343
344    if (agentId) {
345      state.agents[agentId] = {
346        agentId: agentId, type: type, model: r.model, requested: e.model, rewritten: d.action === "rewrite",
347        reasons: d.reasons, est: est, declared: declared, reserved: reserved, reservedFrom: reservedFrom,
348        calls: 0, steps: 0, turns: 0, tokens: 0, measured: 0, durationMs: 0,
349        settled: false, overBudget: false, reason: "", startedAt: Date.now() - state.t0,
350        description: clip(e.description, 60),
351      };
352      state.order.push(agentId);
353    }
354    push({ kind: "route", route: route, reserved: reserved, reservedFrom: reservedFrom, declared: declared });
355    $.ui.log("\u27e6supervisor\u27e7 " + (d.action === "rewrite" ? "REWRITE" : "PASS") + " " + type
356      + " " + (e.model === undefined ? "inherit" : e.model) + " \u2192 " + d.model
357      + " (parent " + e.parentModel + ") [" + d.reasons.join(", ") + "]"
358      + (agentId ? " agent=" + shortId(agentId) + " reserve=" + num(state.agents[agentId].reserved) : ""));
359    await flush($);
360    return r;
361  });
362
363  // ---- 2. live token accounting per subagent ----
364  on("turn.step", async function* ($, e, next) {
365    const r = yield* next(e);
366    if (e.agentId) bumpAgent(e.agentId, r ? r.usage : null);
367    return r;
368  });
369
370  // ---- 3. settle the ledger; maintain session totals from the main loop ----
371  on("turn.complete", async ($, e, next) => {
372    if (e.agentId) {
373      const a = settleAgent(e);
374      if (a) {
375        const dec = a.declared;
376        push({
377          kind: "settle", agentId: a.agentId, type: a.type, model: a.model, calls: a.calls,
378          declared: dec, reserved: a.reserved, measured: a.measured, delta: dec === null ? null : a.measured - dec,
379          durationMs: a.durationMs, reason: a.reason, newMedian: state.priors[a.type].median,
380        });
381        $.ui.log("\u27e6supervisor\u27e7 SETTLE " + shortId(a.agentId) + " " + a.type + " " + a.model
382          + " calls=" + a.calls + " declared=" + (dec === null ? "n/a" : num(dec))
383          + " reserved=" + num(a.reserved) + " measured=" + num(a.measured)
384          + " \u00b7 median(" + a.type + ")=" + num(state.priors[a.type].median));
385        await persist($);
386      }
387    } else {
388      state.totals.mainTurns++;
389      state.totals.mainTokens += weigh(e.usage);
390      await refreshUsage($);
391    }
392    await flush($);
393    return next(e);
394  });
395
396  // ---- 4. per-subagent tool-call budget ----
397  on("tool.call", async ($, e, next) => {
398    const a = e && e.agentId ? state.agents[e.agentId] : null;
399    if (!a) return next(e);
400    a.calls++;
401    state.dirty = true;
402    if (EXEMPT_TOOLS.indexOf(e.tool) >= 0) return next(e);
403    if (a.calls > state.maxCalls) {
404      if (!a.overBudget) {
405        a.overBudget = true;
406        state.totals.overBudget++;
407        $.ui.log("\u27e6supervisor\u27e7 OVER BUDGET " + shortId(a.agentId) + " " + a.type + " at call " + a.calls + "/" + state.maxCalls);
408      }
409      push({ kind: "budget-deny", agentId: a.agentId, type: a.type, tool: e.tool, call: a.calls, max: state.maxCalls });
410      await flush($);
411      return { deny: "supervisor: subagent over budget" };
412    }
413    return next(e);
414  });
415
416  // ---- 5. /supervisor ----
417  on("command.run", { command: "supervisor" }, async ($, e, next) => {
418    await flush($);
419    await refreshUsage($);
420    const arg = (e.args || "").trim().toLowerCase() || "all";
421    const t = totalsLine();
422    const u = usageBits();
423    const head = "SUBAGENT SUPERVISOR \u00b7 session " + state.sessionId
424      + "\n  spawns=" + state.totals.spawns + "  rewritten=" + state.totals.rewrites + "  passed=" + state.totals.passes
425      + "  denied=" + state.totals.denies + "  over-budget=" + state.totals.overBudget
426      + "  live=" + liveAgents().length
427      + "\n  reserved=" + num(t.reserved) + " tok  charged=" + num(t.measured) + " tok"
428      + "  (an agent that never completes is charged its full reservation)"
429      + "\n  main loop: " + state.totals.mainTurns + " turns, " + num(state.totals.mainTokens) + " weighted tok"
430      + "  \u00b7 context " + u.ctx + " \u00b7 five-hour " + u.five + " \u00b7 cost " + u.cost;
431    const log = "\nROUTING LOG (" + state.routes.length + " decisions)\n"
432      + (state.routes.length ? state.routes.slice(-25).map(routeLine).join("\n") : "  none");
433    const cols = "  " + "agent".padEnd(9) + "type".padEnd(20) + "model".padEnd(26) + "calls".padEnd(8)
434      + "declared".padStart(11) + "reserved".padStart(11) + "measured".padStart(11);
435    const ledger = "\nLEDGER (* = live, still accruing)\n" + cols + "\n"
436      + (state.order.length ? state.order.map(function (id) { return ledgerLine(state.agents[id]); }).join("\n") : "  none");
437    const priors = "\nPRIORS ($.store " + STORE_KEY + ", learned medians across sessions)\n"
438      + (Object.keys(state.priors).length
439        ? Object.keys(state.priors).map(function (ty) {
440          const p = state.priors[ty];
441          return "  " + ty.padEnd(22) + "n=" + String(p.samples.length).padEnd(4) + "median=" + num(p.median).padStart(9)
442            + "   default=" + num(isLean(ty) ? LEAN_PRIOR : FULL_PRIOR);
443        }).join("\n")
444        : "  none yet \u2014 defaults: lean " + num(LEAN_PRIOR) + ", full " + num(FULL_PRIOR));
445    const tail = "\nlog file: " + LOG_DIR + state.sessionId + ".jsonl";
446    if (arg === "log") return { text: head + log + tail };
447    if (arg === "ledger") return { text: head + ledger + tail };
448    if (arg === "priors") return { text: head + priors + tail };
449    return { text: head + log + ledger + priors + tail };
450  });
451
452  // ---- 6. HUD ----
453  on("ui.render", { component: "AbovePrompt", surface: "terminal" }, async ($, e, next) => {
454    const { Box, Text } = $.ui.resolve(e);
455    const width = Math.max(40, (e.props && e.props.bodyColumns ? e.props.bodyColumns : 100) - 4);
456    const u = usageBits();
457    const t = totalsLine();
458    const r = state.lastRoute;
459    const routeText = r
460      ? "ROUTE " + r.type + "  " + (r.requested === undefined ? "inherit" : r.requested) + " \u2192 "
461        + (r.action === "deny" ? "DENIED" : r.model) + "  [" + r.reasons.join(", ") + "]"
462      : "ROUTE none yet \u00b7 capability graph fable>opus>sonnet>haiku, advisor = +1 rung";
463    const live = liveAgents().slice(-4).map(function (a, i) {
464      return (
465        <Text key={"l" + i} color="green" wrap="truncate-end">
466          {("LIVE  " + shortId(a.agentId) + " " + a.type + " " + a.model + " calls=" + a.calls + "/" + state.maxCalls
467            + " tok=" + k(a.tokens) + " (reserved " + k(a.reserved) + ")").slice(0, width)}
468        </Text>
469      );
470    });
471    const done = state.order.map(function (id) { return state.agents[id]; })
472      .filter(function (a) { return a && a.settled; }).slice(-3).map(function (a, i) {
473        const dec = a.declared;
474        return (
475          <Text key={"d" + i} dimColor wrap="truncate-end">
476            {("DONE  " + shortId(a.agentId) + " " + a.type + " " + a.model + " calls=" + a.calls
477              + " measured=" + k(a.measured) + (dec === null ? " (no # EST:)" : " vs declared " + k(dec))).slice(0, width)}
478          </Text>
479        );
480      });
481    return (
482      <Box flexDirection="column" borderStyle="round" borderColor="cyan" paddingX={1}>
483        <Text bold color="cyan">
484          {("SUBAGENT SUPERVISOR  spawns=" + state.totals.spawns + " rewritten=" + state.totals.rewrites
485            + " denied=" + state.totals.denies + " live=" + liveAgents().length
486            + "  \u00b7 reserved " + k(t.reserved) + " / charged " + k(t.measured)
487            + "  \u00b7 5h " + u.five + "  \u00b7 " + u.cost).slice(0, width)}
488        </Text>
489        <Text color={r && r.action === "deny" ? "red" : r && r.action === "rewrite" ? "yellow" : "gray"} wrap="truncate-end">
490          {routeText.slice(0, width)}
491        </Text>
492        {live}
493        {done}
494      </Box>
495    );
496  });
497};
498