A Claude Code mod that holds subagent fan-out, heavy-context prompts and retry loops before they spend your usage window. It reads your plan's usage, counts…

Stops Claude Code from spending your five-hour usage window in ten minutes.
agent-usage-guard is a Claude Code mod: it runs inside Claude Code and holds runaway spend before it happens. It reads your plan's live usage, and it holds subagent fan-out, heavy-context churn and retry loops.
<img src="https://github.com/rbartoli/agent-usage-guard/releases/download/agent-usage-guard--v1.0.0/agent-usage-guard.gif" alt="Four scenes in Claude Code's interactive UI. The guard asks before a prompt re-writes the expired cache of a 486k-token session resumed after two hours, before a fifth parallel agent starts against a limit of four, before a twenty-first tool call at 433k context, and before a new agent once subagents have processed 10.9M tokens." width="900">
<sub>Recorded from Claude Code's interactive UI with the guard loaded. A <a href="demo/scripted-api.ts">scripted local API</a> stands in for Claude, so the sessions and token counts are staged and no usage is spent. Every question on screen is the guard's own. <a href="demo/render.sh">Rebuild it</a> from the <a href="demo/fixture/northstar-api">Northstar fixture</a>.</sub>
It acts before the spend, not after it: before an agent starts, a prompt is sent, or a tool call runs. Agent limits hold across every Claude Code session on your machine, and fan-out and heavy context are capped at any usage level, not only near the limit. When you are at the keyboard it asks you, in Claude Code's own question dialog. When nobody can answer, it refuses and tells Claude why, so the model can wait, batch the work, or stop. /agent-guard report shows what it did.
Run these commands inside Claude Code v2.1.287 or later:
/plugin marketplace add rbartoli/agent-usage-guard
/plugin install agent-usage-guard@agent-usage-guard
/reload-plugins
Then run /agent-guard to see what it sees. Mods are on by default; the overview lists the settings that turn them off.
One afternoon I gave Claude Code a research prompt at max effort. It spawned 394 subagent calls, 341 of them running at once, and pushed 84M context tokens through the API in ten minutes. Session over: "Claude usage limit reached." Locked out mid-workday.
That wasn't a one-off. Sixty days of my own logs held 12 lockouts and 166 ten-minute windows with agent bursts. The failure mode multiplies: a large context × rapid tool loops × parallel agents × blind retries. Each protection below removes one factor.
And it's not just me: 339 subagents from a single prompt, half a weekly Max plan consumed by recursive spawning, and multi-agent runs using ~15× the tokens of a chat session.
| Area | Rule | Default | What happens |
|---|---|---|---|
| Plan windows | New agents need approval once a window is this full | 80% | asks; approval lasts until that window resets |
| New agents are refused once a window is this full | 95% | refuses | |
| New agents need approval when the 5-hour window rises this fast | 20 points / 10 min | asks | |
| Agents | Agents running at once, across every session on this machine | 4 | asks |
| Agents started or resumed per 10 minutes, across this machine | 12 | asks | |
| Tokens processed by subagents per 10 minutes, across this machine | 10M | asks | |
| Subagents may not start their own agents | depth 1 | refuses | |
| Context | Claude is told once to keep the turn bounded | 300k tokens | adds a note to the prompt |
| A prompt that would re-read this much context | 500k tokens | asks: compact then send, send anyway, or cancel | |
| Tool calls per 10 minutes while a conversation or agent carries this much | 20 calls at 400k | asks: keep going, compact after this turn, or stop | |
| A heavy session idle long enough for its prompt cache to expire | 150k tokens, 1 h idle | asks before re-writing the cache | |
| Loops | The same tool call failing again and again, with no other call in between | 3 identical failures | refuses the identical retry |
| Refused agent calls within 10 minutes | 5 | pauses agent spawns for 10 min |
Agents covers subagents, agent-team teammates, workflow agents, and finished subagents resumed with SendMessage. Each tool call is judged on the context of the loop that makes it, so a small subagent never inherits the main conversation's size.
When it asks. The question appears in Claude Code's own dialog, the one Claude uses to ask you something, and the call waits for your answer. Agents started in parallel while it is open are held back, and Claude starts them again once you answer. If you allow it, that condition stays allowed for a while: a plan window until it resets, a heavy context until it shrinks (for tool calls, an hour at most), the agent counts for 10 minutes. If you decline or dismiss the question, Claude is told not to start agents for the next 10 minutes (AGENT_GUARD_WINDOW_SECONDS), and is not asked again in that time. Typed words reach Claude, so you can answer "use haiku agents instead".
When nobody can answer. In claude -p, the Agent SDK and dontAsk mode, an agent or tool gate refuses instead of asking. A heavy prompt is held with the reason, and so is a dormant heavy resume once the machine-wide cap of one per 10 minutes is reached. Claude reads each refusal as the tool's result. Repeat refusals of the same condition are worded differently each time and escalate: the second says nothing has changed and to try something else, the third says to end the turn. If a Stop hook such as /goal keeps reopening the turn, the refusal says the user has to decide.
What it leaves alone. It never holds a background task's notification or another session's message, because that would lose the result. It no longer blocks prompts after a usage limit: since v2.1.234, Claude Code waits at a limit and continues at reset on its own (autoContinueAtUsageLimit). The guard's job is to keep you from reaching the limit, and its journal counts the lockouts it did not prevent.
/agent-guard [status] what the guard sees right now
/agent-guard allow [agents|context] [minutes] lift its limits for this session (default: all, 10 min)
/agent-guard pause [minutes] the same as allow, for every limit
/agent-guard resume end an allow early, lift the agent pause, and let declined questions ask again
/agent-guard report [days] what it did on this machine (default: 7 days)
/agent-guard help these commands
/agent-guard runs at once, even mid-turn, and never starts a model turn, so Claude never sees an override and cannot take one as permission. Sending [allow-usage-guard] or [allow-agent-burst] alone as a prompt, the markers of earlier versions, does the same as /agent-guard allow or allow agents.
Checked against Claude Code v2.1.295:
| Path | Coverage |
|---|---|
| Subagents, teammates and workflow agents | Held before they start (agent.spawn) |
A finished subagent sent new work with SendMessage | Held before it resumes, and counted as a start |
/subtask forks, and anything else that starts without agent.spawn | Counted when they start, never held: Claude Code offers no event before they do. Forks Claude starts through the Agent tool are held like any agent |
| Claude Code's internal agents (prompt suggestions, compaction) | Not counted |
| Plan-window percentages | From the last API response on a subscription, in claude -p too. API-key sessions report none, so only the token and count budgets apply |
| Teammates running in their own terminal panes | Not visible: their loops run in other processes |
| Sessions on other machines | Not counted: the cross-session records live in this machine's mod store |
| Tools the API runs itself, such as the advisor | Not holdable: no tool event fires for them |
Mods run in claude in a terminal, in the Desktop app's Code tab, in claude -p and the Agent SDK. They don't run in a Desktop-app WSL session (where mods run).
Native limits still matter. Claude Code caps a session at 20 concurrent subagents (CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS) and three layers of nesting (CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH). It has no limit on the total number of agents a session starts. The guard's defaults are stricter and span sessions; the native caps remain the outer wall for the paths the guard only counts.
/agent-guard report 30
The report counts usage-limit lockouts per week, plan-window threshold crossings, and every question and refusal by rule, with what you answered. It reads every session's journal on this machine and never leaves it. Set AGENT_GUARD_JOURNAL=0 to stop recording.
A mod runs inside Claude Code with your permissions; read what a mod can reach before you install one. This one is about 2,400 lines of TypeScript with no dependencies, and claude plugin validate . lists every API it calls and every environment variable it reads:
SessionStart hook changes nothing.~/.claude/plugins/store/. Each session writes its own counts (running agents, recent starts and dormant resumes, subagent tokens per minute, its latest 5-hour reading), when it was last active (kept 30 days, so a resumed session knows how long it sat idle), and its journal rows (time, a session-id prefix, rule, tool name, context bucket such as 400-500k, plan window and percentage, refusal number, and your answer).AGENT_GUARD_JOURNAL_DAYS).demo/, which builds the GIF above, is not part of what runs. It starts Claude Code in a temporary home directory with a made-up API key for a scripted local API, and reads none of your credentials.Set AGENT_GUARD=0 to turn it off without uninstalling it.
Set these as environment variables, or under env in ~/.claude/settings.json. Times are in seconds; a value that isn't a number in range keeps its default, and /agent-guard says so.
| Variable | Meaning (default) |
|---|---|
AGENT_GUARD | 0, false, off or no turns the guard off (1) |
AGENT_GUARD_JOURNAL | 0 stops the journal (1) |
AGENT_GUARD_JOURNAL_DAYS | Days of journal kept (90) |
AGENT_GUARD_WINDOW_SECONDS | Rolling window for budgets, burn and refusals (600) |
AGENT_GUARD_AGENT_MAX | Agents running at once on this machine (4) |
AGENT_GUARD_ROLLING_MAX | Agent starts and resumes per window on this machine (12) |
AGENT_GUARD_AGENT_TOKENS_MAX | Subagent tokens per window on this machine (10000000) |
AGENT_GUARD_DEPTH_MAX | Agent layers below the main conversation (1) |
AGENT_GUARD_LIMIT_ASK_PERCENT | Plan-window percentage from which new agents ask (80) |
AGENT_GUARD_LIMIT_DENY_PERCENT | Plan-window percentage from which new agents are refused (95) |
AGENT_GUARD_BURN_PERCENT | 5-hour-window points per window that make new agents ask (20) |
AGENT_GUARD_CONTEXT_WARN | Context at which Claude is told once to keep the turn bounded (300000) |
AGENT_GUARD_CONTEXT_HARD | Context at which a prompt asks before it is sent (500000) |
AGENT_GUARD_TOOL_CONTEXT | Context above which tool calls count against the tool budget (400000) |
AGENT_GUARD_TOOL_MAX | Tool calls per window above that context (20) |
AGENT_GUARD_DORMANT_SECONDS | Idle time after which a heavy session's cache counts as expired (3600) |
AGENT_GUARD_DORMANT_CONTEXT | Context at which an idle session counts as heavy (150000) |
AGENT_GUARD_DORMANT_MAX | Heavy dormant resumes per window on this machine when nobody can be asked (1) |
AGENT_GUARD_TOOL_FAILURE_MAX | Identical failures before the identical retry is refused (3) |
AGENT_GUARD_FUSE_MAX | Refused agent calls per fuse period that pause spawns (5) |
AGENT_GUARD_FUSE_SECONDS | The fuse period, and how long the pause lasts (600) |
AGENT_GUARD_PEER_TTL_SECONDS | How long another session's counts, or an agent with no sign of activity, still count (900) |
The dormant default assumes the one-hour prompt cache that subscriptions get within plan usage. With an API key or extra usage, the cache lasts five minutes, so set AGENT_GUARD_DORMANT_SECONDS=300.
The rules live in a harness-agnostic core (core/) that performs no I/O. The Claude Code mod is its first adapter. docs/architecture.md describes the contract an adapter keeps and what another agent harness must offer to carry each rule.
claude plugin validate . --strict
claude plugin test
claude plugin test runs the suite inside Claude Code's own test kit: unit tests of the core, and tests that fire real Claude Code events through the hooks module with no session, model or network. TESTING.md covers the layers and what each test proves.
MIT
hooks/register.ts 742 lines1// The Claude Code adapter. It turns Claude Code's events into the core's facts
2// and requests, asks the user when a verdict needs an answer, and keeps the
3// cross-session records in $.store. Every decision lives in ../core.
4
5import type { EngineInterface, On } from 'claude-code'
6
7import {
8 type AgentRequest,
9 type Ask,
10 type Config,
11 DEFAULT_ALLOW_MINUTES,
12 type Deny,
13 HEADER,
14 type JournalRow,
15 MAIN,
16 type PeerRecord,
17 type PromptSource,
18 SIGNATURE,
19 type SessionState,
20 type ToolRequest,
21 type Verdict,
22 agentGate,
23 applyOverride,
24 canonical,
25 contextBucket,
26 contextNote,
27 defaultConfig,
28 expireAgents,
29 fingerprint,
30 formatReport,
31 heldForQuestion,
32 helpText,
33 isJournalRow,
34 journalKey,
35 livePeers,
36 lockoutKind,
37 markerScope,
38 newSession,
39 observeAgentActive,
40 observeAgentRunEnd,
41 observeAgentStart,
42 observeCompaction,
43 observeLimits,
44 observeMainContext,
45 observeResponse,
46 observeSpawn,
47 observeToolResult,
48 observeToolStart,
49 observeTurnEnd,
50 parseCommand,
51 parseConfig,
52 peerRecord,
53 promptGate,
54 prune,
55 releaseStart,
56 reserveStart,
57 resolveAgentAsk,
58 resolveHeavyAsk,
59 resolvePromptAsk,
60 resumeGuard,
61 splitJournalKeys,
62 statusLine,
63 statusReport,
64 toolGate,
65} from '../core/index.ts'
66
67type Api = EngineInterface
68
69const COMMAND = 'agent-guard'
70const PEER_PREFIX = 'peer:'
71const LAST_ACTIVE_PREFIX = 'last-active:'
72const HEARTBEAT_MS = 60_000
73/** Keeps one session's day of journal rows well inside the store's 4 MiB. */
74const MAX_ROWS_PER_DAY = 2000
75/** All journal days together stay below half the store, oldest dropped first. */
76const JOURNAL_BUDGET_BYTES = 2_000_000
77/** A peer record nobody refreshed for a day is from a session that is gone. */
78const STALE_PEER_MS = 86_400_000
79/** When a session was last active outlives the session by a month, for a later resume. */
80const LAST_ACTIVE_KEEP_MS = 30 * 86_400_000
81/** Keys of the tool call's envelope, not of the tool's own arguments. */
82const ENVELOPE_KEYS = new Set(['tool', 'tool_use_id', 'agentId', 'consent'])
83
84let config: Config = defaultConfig()
85let problems: string[] = []
86let session: SessionState = newSession('', false)
87let interactive = false
88/** The session has no AskUserQuestion tool (`--tools` left it out), so nobody can be asked. */
89let askUnavailable = false
90let permissionMode: string | undefined
91let peers: PeerRecord[] = []
92let lastPublished = 0
93/** The last-active time this process last stored, so a heartbeat does not rewrite it. */
94let lastActiveKept: number | undefined
95let lastStatus: string | undefined
96let compactAfterTurn = false
97/** A Stop event happened and no prompt has started a new turn since. */
98let stopSeen = false
99let journalCache: { key: string; rows: JournalRow[] } | undefined
100/** One question per family at a time: calls that trip while it is open are held back. */
101const pending = new Map<string, Promise<unknown>>()
102
103export function register(on: On): void {
104 on('session.start', async ($, e, next) => {
105 await loadConfig($)
106 interactive = e.isInteractive
107 await startSession($)
108 $.clock.after(0, () => housekeeping($))
109 try {
110 await $.command.register({
111 name: COMMAND,
112 description: 'agent-usage-guard: status, allow, resume, report',
113 argumentHint: '[status | allow [agents|context] [minutes] | pause [minutes] | resume | report [days] | help]',
114 immediate: true,
115 })
116 } catch (error) {
117 $.ui.log(`agent-usage-guard could not add /${COMMAND}: ${String(error)}`, { to: 'debug' })
118 }
119 return next(e)
120 }).catch(($, e, next) => failOpen($, 'session.start', next.error, () => next(e)))
121
122 on('classic.SessionStart', async ($, e, next) => {
123 if (e.source === 'compact') observeCompaction(session)
124 if (e.source === 'clear' || e.source === 'resume' || e.source === 'fork') await startSession($)
125 return next(e)
126 }).catch(($, e, next) => failOpen($, 'classic.SessionStart', next.error, () => next(e)))
127
128 on('classic.UserPromptSubmit', async ($, e, next) => {
129 trackMode(e.permission_mode, e.agent_id)
130 return next(e)
131 }).catch(($, e, next) => failOpen($, 'classic.UserPromptSubmit', next.error, () => next(e)))
132
133 on('classic.PostToolUse', async ($, e, next) => {
134 trackMode(e.permission_mode, e.agent_id)
135 return next(e)
136 }).catch(($, e, next) => failOpen($, 'classic.PostToolUse', next.error, () => next(e)))
137
138 on('classic.PostCompact', async ($, e, next) => {
139 observeCompaction(session)
140 await refreshStatus($)
141 return next(e)
142 }).catch(($, e, next) => failOpen($, 'classic.PostCompact', next.error, () => next(e)))
143
144 on('classic.Stop', async ($, e, next) => {
145 trackMode(e.permission_mode, e.agent_id)
146 stopSeen = true
147 if (e.stop_hook_active) session.reopened = true
148 return next(e)
149 }).catch(($, e, next) => failOpen($, 'classic.Stop', next.error, () => next(e)))
150
151 on('classic.SubagentStart', async ($, e, next) => {
152 const now = await $.clock.now()
153 if (observeAgentStart(session, now, e.agent_id, e.agent_type)) await publish($, now)
154 return next(e)
155 }).catch(($, e, next) => failOpen($, 'classic.SubagentStart', next.error, () => next(e)))
156
157 on('classic.StopFailure', async ($, e, next) => {
158 if (config.enabled) {
159 // On StopFailure, last_assistant_message is the failed request's own
160 // error message, the text Claude Code shows, not Claude's earlier prose.
161 const text = `${e.error_details ?? ''} ${e.last_assistant_message ?? ''}`
162 const kind = lockoutKind(e.error, text, Object.values(session.limits))
163 if (kind) {
164 const reading = session.limits[kind]
165 await journal($, await $.clock.now(), { ev: 'lockout', kind, ...(reading ? { pct: reading.percent } : {}) })
166 }
167 }
168 return next(e)
169 }).catch(($, e, next) => failOpen($, 'classic.StopFailure', next.error, () => next(e)))
170
171 on('session.end', async ($, e, next) => {
172 await keepLastActive($)
173 try {
174 if (session.id) await $.store.delete(PEER_PREFIX + session.id)
175 } catch {
176 // The record expires on its own once its heartbeat is old.
177 }
178 return next(e)
179 }).catch(($, e, next) => failOpen($, 'session.end', next.error, () => next(e)))
180
181 on('session.measure', async ($, e, next) => {
182 if (config.enabled) {
183 const now = await $.clock.now()
184 if (e.changed.includes('rateLimits')) {
185 const readings = e.rateLimits.map((r) => ({
186 at: now,
187 kind: r.kind,
188 percent: r.percentUsed,
189 ...(r.resetsAt ? { resetsAt: Date.parse(r.resetsAt) } : {}),
190 }))
191 for (const { reading, threshold } of observeLimits(session, config, readings)) {
192 await journal($, now, { ev: 'limit', kind: reading.kind, pct: reading.percent, rule: `${threshold}%` })
193 }
194 await publish($, now)
195 }
196 if (e.context.tokens !== undefined) observeMainContext(session, config, e.context.tokens)
197 await refreshStatus($)
198 }
199 return next(e)
200 }).catch(($, e, next) => failOpen($, 'session.measure', next.error, () => next(e)))
201
202 on('turn.step', async function* ($, e, next) {
203 const result = yield* next(e)
204 // A streaming hook fails open by hand: recording a response must never fail the request.
205 try {
206 if (config.enabled && result.usage) {
207 const now = await $.clock.now()
208 const loop = e.agentId ?? MAIN
209 if (loop === MAIN && stopSeen) session.reopened = true
210 observeResponse(session, config, now, loop, {
211 input: result.usage.input_tokens,
212 output: result.usage.output_tokens,
213 cacheRead: result.usage.cache_read_input_tokens,
214 cacheWrite: result.usage.cache_creation_input_tokens,
215 })
216 if (now - lastPublished >= HEARTBEAT_MS) await publish($, now)
217 }
218 } catch (error) {
219 $.ui.log(`agent-usage-guard: could not record a response (${String(error)})`, { to: 'debug' })
220 }
221 return result
222 })
223
224 on('turn.complete', async ($, e, next) => {
225 const result = await next(e)
226 if (!config.enabled) return result
227 const now = await $.clock.now()
228 if (e.agentId) {
229 observeAgentRunEnd(session, e.agentId)
230 await publish($, now)
231 return result
232 }
233 observeTurnEnd(session)
234 if (compactAfterTurn) {
235 compactAfterTurn = false
236 $.clock.after(0, async () => {
237 const failed = await compact($)
238 if (failed !== undefined) $.ui.log(`agent-usage-guard could not compact (${failed}).`)
239 })
240 }
241 await refreshStatus($)
242 return result
243 }).catch(($, e, next) => failOpen($, 'turn.complete', next.error, () => next(e)))
244
245 on('prompt.submit', async ($, e, next) => {
246 if (!config.enabled) return next(e)
247 const now = await $.clock.now()
248 const source = sourceOf(e.origin.kind)
249
250 const scope = source === 'other' ? undefined : markerScope(e.text)
251 if (scope) {
252 const text = applyOverride(session, now, scope, DEFAULT_ALLOW_MINUTES)
253 await journal($, now, { ev: 'override', rule: scope })
254 await refreshStatus($)
255 return { drop: `${SIGNATURE}: ${text}` }
256 }
257
258 await readMainContext($)
259 const context = session.loops[MAIN]?.context
260 if (context !== undefined && context >= Math.min(config.dormantContext, config.contextHard)) {
261 await loadPeers($, now)
262 }
263 const gate = async (at: number): Promise<Verdict> =>
264 promptGate(session, peers, config, at, { source, text: e.text, idleMs: await idleTime($, at) })
265 const send = (note: string | undefined) => (note ? next({ ...e, context: [...(e.context ?? []), note] }) : next(e))
266 const verdict = await gate(now)
267 if (session.dormantResumes.at(-1) === now) await publish($, now)
268 stopSeen = false
269 session.reopened = false
270
271 if (verdict.kind === 'allow') return send(verdict.note)
272 if (verdict.kind === 'drop') {
273 await journal($, now, { ev: 'drop', rule: verdict.rule, ctx: contextBucket(context) })
274 return { drop: verdict.message }
275 }
276 if (verdict.kind !== 'ask') return next(e)
277
278 const asked = await askUser($, verdict, contextBucket(context))
279 if (asked.unavailable) {
280 await journalAnswer($, verdict, asked)
281 const alone = await gate(asked.at)
282 if (alone.kind === 'drop') return { drop: alone.message }
283 return send(alone.kind === 'allow' ? alone.note : undefined)
284 }
285 const outcome = resolvePromptAsk(session, verdict, asked.answer)
286 await journalAnswer($, verdict, asked, outcome.kind === 'compact-then-send' ? 'compact' : outcome.kind === 'allow' ? 'allow' : 'cancel')
287 if (outcome.kind === 'allow') return send(contextNote(session, config))
288 if (outcome.kind === 'compact-then-send') {
289 const resend = { text: e.text, ...(e.attachments ? { attachments: e.attachments } : {}) }
290 $.clock.after(0, async () => {
291 // Claude Code puts a dropped prompt back in the input box; this one is sent for the user.
292 if ((await $.prompt.read()).text === resend.text) await fillInput($, '')
293 // Sending at full size after a failed compaction is the one thing the user chose against.
294 const failed = await compact($)
295 if (failed !== undefined) {
296 $.ui.log(`agent-usage-guard could not compact (${failed}), so your prompt is back in the input box.`)
297 await fillInput($, resend.text)
298 return
299 }
300 try {
301 await $.prompt.submit({ ...resend, asUser: true })
302 } catch (error) {
303 $.ui.log(`agent-usage-guard could not resend your prompt: ${String(error)}`)
304 await fillInput($, resend.text)
305 }
306 })
307 return { drop: outcome.message }
308 }
309 await fillInput($, e.text)
310 return { drop: outcome.message }
311 }).catch(($, e, next) => failOpen($, 'prompt.submit', next.error, () => next(e)))
312
313 on('agent.spawn', async ($, e, next) => {
314 if (!config.enabled) return next(e)
315 trackMode(e.permissionMode, e.parentAgentId)
316 const now = await $.clock.now()
317 const loop = e.parentAgentId ?? MAIN
318 const parentDepth = e.parentAgentId ? (session.agents[e.parentAgentId]?.depth ?? 1) : 0
319 const request = { loop, depth: parentDepth + 1, resume: false }
320 const { verdict, reservedAt } = await agentVerdict($, now, request)
321 if (verdict.kind === 'deny') {
322 await journalDenial($, verdict, 'Agent')
323 return { deny: verdict.message }
324 }
325 let started: string | undefined
326 try {
327 const result = await next(e)
328 if (!result.deny && result.agentId) {
329 started = result.agentId
330 const kind = e.isTeammate ? 'teammate' : e.workflow ? 'workflow' : e.fork ? 'fork' : 'subagent'
331 observeSpawn(session, await $.clock.now(), { agentId: started, kind, depth: request.depth, ...(e.name ? { name: e.name } : {}) })
332 }
333 return result
334 } finally {
335 if (reservedAt !== undefined) releaseStart(session, reservedAt, started)
336 await publish($, await $.clock.now())
337 }
338 }).catch(($, e, next) =>
339 // An overrun before the agent started leaves it unjudged; starting it anyway
340 // would let a stuck guard wave through every agent of a parallel batch.
341 next.error?.kind === 'timeout' && !next.called
342 ? { deny: `agent-usage-guard could not judge this agent in time. Start it again. [${SIGNATURE}]` }
343 : failOpen($, 'agent.spawn', next.error, () => next(e)),
344 )
345
346 on('tool.call', async ($, e, next) => {
347 if (!config.enabled) return next(e)
348 const now = await $.clock.now()
349 const loop = e.agentId ?? MAIN
350
351 let resumed: { id: string; reservedAt: number | undefined } | undefined
352 if (e.tool === 'SendMessage') {
353 const target = resumeTarget(e.to, e.message)
354 if (target) {
355 const { verdict, reservedAt } = await agentVerdict($, now, { loop, depth: session.agents[target]!.depth, resume: true })
356 if (verdict.kind === 'deny') {
357 await journalDenial($, verdict, 'SendMessage')
358 return { deny: verdict.message }
359 }
360 resumed = { id: target, reservedAt }
361 }
362 }
363
364 const print = fingerprint(canonical(argumentsOf(e)))
365 if (!resumed) {
366 const label = loop === MAIN ? 'this conversation' : `agent ${session.agents[loop]?.name ?? loop.slice(0, 8)}`
367 const request = { loop, tool: e.tool, fingerprint: print, label }
368 let verdict = toolGate(session, config, now, request)
369 if (verdict.kind === 'ask') verdict = await settleHeavy($, verdict, request)
370 if (verdict.kind === 'deny') {
371 await journalDenial($, verdict, e.tool, contextBucket(session.loops[loop]?.context))
372 return { deny: verdict.message }
373 }
374 }
375
376 observeToolStart(session, loop)
377 let result: Awaited<ReturnType<typeof next>> | undefined
378 try {
379 result = await next(e)
380 return result
381 } finally {
382 observeToolResult(session, await $.clock.now(), loop, print, Boolean(result?.isError) && !result?.deny)
383 if (resumed) {
384 const delivered = result !== undefined && !result.deny && !result.isError
385 if (resumed.reservedAt !== undefined) releaseStart(session, resumed.reservedAt, delivered ? resumed.id : undefined)
386 if (delivered) observeAgentActive(session, resumed.id)
387 await publish($, await $.clock.now())
388 }
389 }
390 }).catch(($, e, next) => failOpen($, 'tool.call', next.error, () => next(e)))
391
392 on('command.run', { command: COMMAND }, async ($, e) => {
393 const now = await $.clock.now()
394 const command = parseCommand(e.args)
395 if (command.action === 'help') return { text: helpText(command.error) }
396 if (command.action === 'report') return { text: await report($, now, command.days) }
397 if (command.action === 'allow') {
398 const text = applyOverride(session, now, command.scope, command.minutes)
399 await journal($, now, { ev: 'override', rule: command.scope })
400 await refreshStatus($)
401 return { text }
402 }
403 if (command.action === 'resume') {
404 const text = resumeGuard(session)
405 await refreshStatus($)
406 return { text }
407 }
408 await loadPeers($, now)
409 return { text: statusReport(session, peers, config, now, problems) }
410 }).catch(($, e, next) => ({ text: `The command failed: ${next.error?.message ?? 'unknown error'}` }))
411}
412
413/** Each variable is named in full, so `claude plugin validate` can list every one the mod reads. */
414async function loadConfig($: Api): Promise<void> {
415 const parsed = parseConfig({
416 AGENT_GUARD: await $.env.get('AGENT_GUARD'),
417 AGENT_GUARD_JOURNAL: await $.env.get('AGENT_GUARD_JOURNAL'),
418 AGENT_GUARD_WINDOW_SECONDS: await $.env.get('AGENT_GUARD_WINDOW_SECONDS'),
419 AGENT_GUARD_AGENT_MAX: await $.env.get('AGENT_GUARD_AGENT_MAX'),
420 AGENT_GUARD_ROLLING_MAX: await $.env.get('AGENT_GUARD_ROLLING_MAX'),
421 AGENT_GUARD_DEPTH_MAX: await $.env.get('AGENT_GUARD_DEPTH_MAX'),
422 AGENT_GUARD_AGENT_TOKENS_MAX: await $.env.get('AGENT_GUARD_AGENT_TOKENS_MAX'),
423 AGENT_GUARD_LIMIT_ASK_PERCENT: await $.env.get('AGENT_GUARD_LIMIT_ASK_PERCENT'),
424 AGENT_GUARD_LIMIT_DENY_PERCENT: await $.env.get('AGENT_GUARD_LIMIT_DENY_PERCENT'),
425 AGENT_GUARD_BURN_PERCENT: await $.env.get('AGENT_GUARD_BURN_PERCENT'),
426 AGENT_GUARD_CONTEXT_WARN: await $.env.get('AGENT_GUARD_CONTEXT_WARN'),
427 AGENT_GUARD_CONTEXT_HARD: await $.env.get('AGENT_GUARD_CONTEXT_HARD'),
428 AGENT_GUARD_TOOL_CONTEXT: await $.env.get('AGENT_GUARD_TOOL_CONTEXT'),
429 AGENT_GUARD_TOOL_MAX: await $.env.get('AGENT_GUARD_TOOL_MAX'),
430 AGENT_GUARD_DORMANT_SECONDS: await $.env.get('AGENT_GUARD_DORMANT_SECONDS'),
431 AGENT_GUARD_DORMANT_CONTEXT: await $.env.get('AGENT_GUARD_DORMANT_CONTEXT'),
432 AGENT_GUARD_DORMANT_MAX: await $.env.get('AGENT_GUARD_DORMANT_MAX'),
433 AGENT_GUARD_TOOL_FAILURE_MAX: await $.env.get('AGENT_GUARD_TOOL_FAILURE_MAX'),
434 AGENT_GUARD_FUSE_MAX: await $.env.get('AGENT_GUARD_FUSE_MAX'),
435 AGENT_GUARD_FUSE_SECONDS: await $.env.get('AGENT_GUARD_FUSE_SECONDS'),
436 AGENT_GUARD_PEER_TTL_SECONDS: await $.env.get('AGENT_GUARD_PEER_TTL_SECONDS'),
437 AGENT_GUARD_JOURNAL_DAYS: await $.env.get('AGENT_GUARD_JOURNAL_DAYS'),
438 })
439 config = parsed.config
440 problems = parsed.problems
441 for (const problem of problems) $.ui.log(`agent-usage-guard: ${problem}`, { to: 'debug' })
442}
443
444/** Starts fresh state for the session id Claude Code reports now. */
445async function startSession($: Api): Promise<void> {
446 const previous = session.id
447 const id = await $.session.id()
448 if (previous && previous !== id) {
449 await keepLastActive($)
450 await $.store.delete(PEER_PREFIX + previous).catch(() => undefined)
451 }
452 session = newSession(id, canAsk())
453 lastActiveKept = undefined
454 journalCache = undefined
455 stopSeen = false
456 compactAfterTurn = false
457}
458
459/** Drops journal days past their retention and records of sessions long gone. */
460async function housekeeping($: Api): Promise<void> {
461 try {
462 const now = await $.clock.now()
463 const keys = await $.store.keys()
464 const { recent, expired } = splitJournalKeys(keys, now, config.journalDays)
465 for (const key of expired) await $.store.delete(key)
466 const days = recent.sort()
467 const sizes = await Promise.all(days.map(async (k) => JSON.stringify((await $.store.get(k)) ?? null).length))
468 let total = sizes.reduce((sum, n) => sum + n, 0)
469 for (let i = 0; total > JOURNAL_BUDGET_BYTES && i < days.length; i++) {
470 await $.store.delete(days[i]!)
471 total -= sizes[i]!
472 }
473 for (const key of keys.filter((k) => k.startsWith(PEER_PREFIX))) {
474 const record = (await $.store.get(key)) as { at?: unknown } | undefined
475 if (typeof record?.at !== 'number' || now - record.at > STALE_PEER_MS) await $.store.delete(key)
476 }
477 for (const key of keys.filter((k) => k.startsWith(LAST_ACTIVE_PREFIX))) {
478 const at = await $.store.get(key)
479 if (typeof at !== 'number' || now - at > LAST_ACTIVE_KEEP_MS) await $.store.delete(key)
480 }
481 } catch (error) {
482 $.ui.log(`agent-usage-guard housekeeping failed: ${String(error)}`, { to: 'debug' })
483 }
484}
485
486/** A person can answer a question: an interactive session not set to never ask. */
487function canAsk(): boolean {
488 return interactive && !askUnavailable && permissionMode !== 'dontAsk'
489}
490
491/** Follows the main loop's permission mode; a subagent's own mode says nothing about who can answer. */
492function trackMode(mode: string | undefined, agentId: string | undefined): void {
493 if (!mode || agentId) return
494 permissionMode = mode
495 session.canAsk = canAsk()
496}
497
498function sourceOf(kind: string): PromptSource {
499 if (kind === 'composer' || kind === 'bridge') return 'user'
500 if (kind === 'sdk') return 'headless'
501 return 'other'
502}
503
504/** Publishes this session's counts for the other sessions on this machine. */
505async function publish($: Api, now: number): Promise<void> {
506 if (!session.id) return
507 lastPublished = now
508 expireAgents(session, now, config.peerTtlMs)
509 prune(session, now, config.windowMs)
510 try {
511 await $.store.set(PEER_PREFIX + session.id, peerRecord(session, now, config.windowMs))
512 } catch (error) {
513 $.ui.log(`agent-usage-guard could not publish its counts: ${String(error)}`, { to: 'debug' })
514 }
515 await keepLastActive($)
516}
517
518/**
519 * Stores when the main conversation was last active, so a later process
520 * that resumes the session knows how long it sat idle. The transcript cannot
521 * say: Claude Code rewrites its modification time on resume, and hourly while
522 * the session is open.
523 */
524async function keepLastActive($: Api): Promise<void> {
525 const last = session.loops[MAIN]?.activeAt
526 if (!session.id || last === undefined || last === lastActiveKept) return
527 try {
528 await $.store.set(LAST_ACTIVE_PREFIX + session.id, last)
529 lastActiveKept = last
530 } catch (error) {
531 $.ui.log(`agent-usage-guard could not store when the session was last active: ${String(error)}`, { to: 'debug' })
532 }
533}
534
535/** Reads the other sessions' records; a failure leaves the last ones read. */
536async function loadPeers($: Api, now: number): Promise<void> {
537 try {
538 const own = PEER_PREFIX + session.id
539 const keys = (await $.store.keys()).filter((k) => k.startsWith(PEER_PREFIX) && k !== own)
540 const records = await Promise.all(keys.map((k) => $.store.get(k)))
541 peers = livePeers(records, now, config.peerTtlMs)
542 } catch (error) {
543 $.ui.log(`agent-usage-guard could not read other sessions: ${String(error)}`, { to: 'debug' })
544 }
545}
546
547/** The main context before a prompt, from Claude Code's own figure. */
548async function readMainContext($: Api): Promise<void> {
549 try {
550 const usage = await $.session.usage()
551 if (usage.context.tokens !== undefined) observeMainContext(session, config, usage.context.tokens)
552 } catch {
553 // Keep the last figure a response reported.
554 }
555}
556
557/** Time since the main conversation was last active: in this process, or, after a resume, in the one before. */
558async function idleTime($: Api, now: number): Promise<number | undefined> {
559 const last = session.loops[MAIN]?.activeAt
560 if (last !== undefined) return now - last
561 try {
562 const kept = await $.store.get(LAST_ACTIVE_PREFIX + session.id)
563 return typeof kept === 'number' ? now - kept : undefined
564 } catch {
565 return undefined
566 }
567}
568
569async function refreshStatus($: Api): Promise<void> {
570 if (!interactive) return
571 try {
572 const now = await $.clock.now()
573 const text = statusLine(session, peers, config, now)
574 if (text === lastStatus) return
575 lastStatus = text
576 $.ui.status(text)
577 } catch (error) {
578 $.ui.log(`agent-usage-guard could not update its status line: ${String(error)}`, { to: 'debug' })
579 }
580}
581
582type Asked = { answer: string | undefined; unavailable: boolean; at: number }
583
584/**
585 * Journals a question and asks it in Claude Code's own question dialog.
586 * `answer` is undefined when the dialog came back without one: the user
587 * dismissed it. `unavailable` means the question could not be asked at all,
588 * because the session has no AskUserQuestion tool: from then on the session
589 * counts as one nobody can answer, and the gate decides alone. Any other
590 * failure, such as the turn being interrupted while the dialog was open, is
591 * not the user's answer: it is rethrown, and the hook fails open. `at` is when
592 * the question ended.
593 */
594async function askUser($: Api, ask: Ask, ctx?: string): Promise<Asked> {
595 await journal($, await $.clock.now(), { ev: 'ask', rule: ask.trips[0]?.rule, ctx })
596 let answer: string | undefined
597 let unavailable = false
598 try {
599 answer = await $.ui.ask(ask.question, { options: ask.options, header: HEADER })
600 } catch (error) {
601 const text = String(error)
602 if (/no tool named "AskUserQuestion"/.test(text)) {
603 askUnavailable = true
604 session.canAsk = false
605 unavailable = true
606 } else if (!/\$\.ui\.ask: no answer\b/.test(text)) {
607 throw error
608 }
609 }
610 return { answer, unavailable, at: await $.clock.now() }
611}
612
613/** Journals how a question ended: the user's `choice`, a dismissal, or no dialog at all. */
614async function journalAnswer($: Api, ask: Ask, asked: Asked, choice?: string): Promise<void> {
615 const answer = asked.unavailable ? 'unavailable' : asked.answer === undefined ? 'dismiss' : choice
616 await journal($, asked.at, { ev: 'answer', rule: ask.trips[0]?.rule, answer })
617}
618
619/**
620 * One question per key at a time: a call that trips while its question is
621 * open is held back at once, and the model makes it again once it is
622 * answered. Waiting instead could outlast the hook's time limit.
623 */
624async function oneAtATime<T>(key: string, held: () => T, run: () => Promise<T>): Promise<T> {
625 // The check for an open question and the claim on it must not straddle an
626 // await, or two parallel calls would each ask.
627 if (pending.has(key)) return held()
628 const asking = run()
629 pending.set(key, asking)
630 try {
631 return await asking
632 } finally {
633 pending.delete(key)
634 }
635}
636
637/** Compacts the conversation. Returns why it did not, or undefined once it has. */
638async function compact($: Api): Promise<string | undefined> {
639 try {
640 return (await $.session.compact()).skip
641 } catch (error) {
642 return error instanceof Error ? error.message : String(error)
643 }
644}
645
646/** A verdict on an agent, and the slot it holds when allowed, for `releaseStart`. */
647type AgentVerdict = { verdict: Verdict; reservedAt?: number }
648
649/** The agent gate, asking at most one question about agents at a time. */
650async function agentVerdict($: Api, now: number, request: AgentRequest): Promise<AgentVerdict> {
651 await loadPeers($, now)
652 expireAgents(session, now, config.peerTtlMs)
653 prune(session, now, config.windowMs)
654 const verdict = agentGate(session, peers, config, now, request)
655 if (verdict.kind === 'allow') return { verdict, reservedAt: reserveStart(session, now) }
656 if (verdict.kind !== 'ask') return { verdict }
657 return oneAtATime<AgentVerdict>('agents', () => ({ verdict: heldForQuestion('agents') }), async () => {
658 const asked = await askUser($, verdict)
659 const resolved = asked.unavailable
660 ? agentGate(session, peers, config, asked.at, request)
661 : resolveAgentAsk(session, config, asked.at, verdict, asked.answer)
662 const reservedAt = resolved.kind === 'allow' ? reserveStart(session, asked.at) : undefined
663 await journalAnswer($, verdict, asked, resolved.kind === 'allow' ? 'allow' : 'refuse')
664 await refreshStatus($)
665 return { verdict: resolved, ...(reservedAt === undefined ? {} : { reservedAt }) }
666 })
667}
668
669/** The heavy-context question, one per loop at a time. */
670async function settleHeavy($: Api, ask: Ask, request: ToolRequest): Promise<Verdict> {
671 return oneAtATime(`heavy:${ask.loop}`, () => heldForQuestion('heavy'), async () => {
672 const asked = await askUser($, ask, contextBucket(session.loops[ask.loop]?.context))
673 if (asked.unavailable) {
674 await journalAnswer($, ask, asked)
675 return toolGate(session, config, asked.at, request)
676 }
677 const outcome = resolveHeavyAsk(session, config, asked.at, ask, asked.answer)
678 if (outcome.compactAfterTurn) compactAfterTurn = true
679 await journalAnswer($, ask, asked, outcome.compactAfterTurn ? 'compact' : outcome.verdict.kind === 'allow' ? 'allow' : 'refuse')
680 return outcome.verdict
681 })
682}
683
684/** The agent a plain-text SendMessage would resume: a finished subagent of this session. */
685function resumeTarget(to: unknown, message: unknown): string | undefined {
686 if (typeof to !== 'string' || typeof message !== 'string') return undefined
687 const agent = session.agents[to] ?? Object.values(session.agents).find((a) => a.name === to)
688 if (!agent || agent.running || agent.kind === 'teammate') return undefined
689 return agent.id
690}
691
692/** The tool's own arguments, for its fingerprint. */
693function argumentsOf(e: Readonly<Record<string, unknown>>): Record<string, unknown> {
694 const args: Record<string, unknown> = {}
695 for (const [key, value] of Object.entries(e)) if (!ENVELOPE_KEYS.has(key)) args[key] = value
696 return { tool: e.tool, args }
697}
698
699/** Puts `text` in the input box; a failure leaves the box as it was. */
700async function fillInput($: Api, text: string): Promise<void> {
701 await $.prompt.fill({ text }).catch(() => undefined)
702}
703
704/** Journals a refusal the model reads. */
705async function journalDenial($: Api, verdict: Deny, tool: string, ctx?: string): Promise<void> {
706 await journal($, await $.clock.now(), { ev: 'deny', rule: verdict.rule, tool, ctx, n: verdict.n })
707}
708
709/** Appends a row to this session's journal key for the day. Never throws. */
710async function journal($: Api, now: number, row: Omit<JournalRow, 't' | 's'>): Promise<void> {
711 if (!config.journal || !session.id) return
712 try {
713 const key = journalKey(session.id, now)
714 if (journalCache?.key !== key) {
715 const stored = await $.store.get(key)
716 journalCache = { key, rows: Array.isArray(stored) ? stored.filter(isJournalRow) : [] }
717 }
718 const clean = Object.fromEntries(Object.entries(row).filter(([, v]) => v !== undefined)) as Omit<JournalRow, 't' | 's'>
719 journalCache.rows.push({ t: now, s: session.id.slice(0, 8), ...clean })
720 if (journalCache.rows.length > MAX_ROWS_PER_DAY) journalCache.rows.splice(0, journalCache.rows.length - MAX_ROWS_PER_DAY)
721 await $.store.set(key, journalCache.rows)
722 } catch (error) {
723 $.ui.log(`agent-usage-guard could not write its journal: ${String(error)}`, { to: 'debug' })
724 }
725}
726
727async function report($: Api, now: number, days: number): Promise<string> {
728 try {
729 const { recent } = splitJournalKeys(await $.store.keys(), now, days)
730 const rows = (await Promise.all(recent.map((k) => $.store.get(k)))).flatMap((v) => (Array.isArray(v) ? v.filter(isJournalRow) : []))
731 return formatReport(rows, now, days)
732 } catch (error) {
733 return `agent-usage-guard could not read its journal: ${String(error)}`
734 }
735}
736
737/** A hook that throws or overruns lets the event go ahead: a broken guard must not break the session. */
738function failOpen<T>($: Api, event: string, error: { kind?: string; message?: string } | undefined, proceed: () => T): T {
739 $.ui.log(`agent-usage-guard: the ${event} hook failed, so the event went ahead (${error?.kind ?? 'error'}: ${error?.message ?? 'unknown'})`, { to: 'debug' })
740 return proceed()
741}
742core/index.ts 55 lines1// The harness-agnostic core of agent-usage-guard. It performs no I/O and reads
2// no clock: an adapter reports facts, asks for verdicts, and carries them out.
3// See docs/architecture.md for the contract an adapter keeps.
4
5export { type AgentRequest, agentGate, heldForQuestion, highestLimit, resolveAgentAsk } from './agents.ts'
6export { type Command, DEFAULT_ALLOW_MINUTES, applyOverride, helpText, parseCommand, resumeGuard } from './commands.ts'
7export { type Config, SPECS, SWITCHES, defaultConfig, parseConfig } from './config.ts'
8export { type JournalEvent, type JournalRow, formatReport, isJournalRow, journalKey, splitJournalKeys } from './journal.ts'
9export { lockoutKind } from './lockout.ts'
10export {
11 type SpawnFacts,
12 type TokenUsage,
13 expireAgents,
14 observeAgentActive,
15 observeAgentRunEnd,
16 observeAgentStart,
17 observeCompaction,
18 observeLimits,
19 observeMainContext,
20 observeResponse,
21 observeSpawn,
22 observeToolResult,
23 observeToolStart,
24 observeTurnEnd,
25 releaseStart,
26 reserveStart,
27} from './observe.ts'
28export {
29 type PromptOutcome,
30 type PromptRequest,
31 type PromptSource,
32 contextNote,
33 markerScope,
34 promptGate,
35 resolvePromptAsk,
36} from './prompts.ts'
37export {
38 type AgentKind,
39 type AgentRecord,
40 type LimitReading,
41 type LoopId,
42 MAIN,
43 type PeerRecord,
44 type Scope,
45 type SessionState,
46 livePeers,
47 newSession,
48 peerRecord,
49 prune,
50} from './state.ts'
51export { USAGE, statusLine, statusReport } from './status.ts'
52export { canonical, contextBucket, fingerprint } from './text.ts'
53export { type HeavyOutcome, type ToolRequest, resolveHeavyAsk, toolGate } from './tools.ts'
54export { type Ask, type Deny, type Family, HEADER, type Rule, SIGNATURE, type Verdict } from './verdict.ts'
55core/agents.ts 229 lines1// The agent gate: every subagent, teammate, workflow agent and resumed agent
2// passes here before it starts.
3
4import type { Config } from './config.ts'
5import { type LimitReading, type LoopId, type PeerRecord, type SessionState, countOnMachine, overrideCovers, runningAgents, tokensWithin } from './state.ts'
6import { count, duration, join, sentence, shortClock, tokens, windowName } from './text.ts'
7import { ALLOW, type Ask, type Deny, SIGNATURE, type Trip, type Verdict, refuse } from './verdict.ts'
8
9export type AgentRequest = {
10 /** The loop asking for the agent. */
11 loop: LoopId
12 /** Layers below the main conversation the new agent would run at. */
13 depth: number
14 /** True when a finished agent is being sent a new task. */
15 resume: boolean
16}
17
18const AGENT_ALLOW = 'Allow'
19const AGENT_REFUSE = "Don't start it"
20
21export function agentGate(
22 state: SessionState,
23 peers: readonly PeerRecord[],
24 config: Config,
25 now: number,
26 request: AgentRequest,
27): Verdict {
28 if (!config.enabled || overrideCovers(state, now, 'agents')) return ALLOW
29 const what = request.resume ? 'resuming this agent' : 'this agent'
30
31 if (request.depth > config.depthMax) {
32 const base = `${request.resume ? 'This agent cannot be resumed' : 'This agent was not started'}: subagents may not start agents of their own here (depth limit ${config.depthMax}). Do the work yourself and report back.`
33 return refuse(state, config, now, 'depth', 'depth', base, true)
34 }
35 if (state.fuseUntil !== undefined && now < state.fuseUntil) {
36 const base = `Agent spawns in this session are paused until ${shortClock(state.fuseUntil)} after repeated refused agent calls. Continue without agents; the user can lift this with /agent-guard allow.`
37 return refuse(state, config, now, 'fuse', 'fuse', base, false)
38 }
39
40 const limit = highestLimit(state, peers, now)
41 if (limit && limit.percent >= config.limitDenyPercent) {
42 const base = `${sentence(what)} was refused: ${limitDetail(limit)}, and the guard keeps the rest of the window for the user. Continue without agents; the user can lift this with /agent-guard allow.`
43 return refuse(state, config, now, 'limit-deny', 'limit-deny', base, true)
44 }
45
46 const trips = agentTrips(state, peers, config, now, limit).filter((t) => !leased(state, t, now))
47 if (trips.length === 0) return ALLOW
48
49 const details = join(trips.map((t) => t.detail))
50 const declinedAt = state.refusals.agents
51 if (declinedAt !== undefined && now < declinedAt) {
52 const base = `The user declined new agents for now (${details}). Continue without starting agents; the user can lift this with /agent-guard allow.`
53 return refuse(state, config, now, trips[0]!.rule, `agents:${trips[0]!.rule}`, base, true)
54 }
55 if (!state.canAsk) return unattended(state, config, now, trips, details)
56 return {
57 kind: 'ask',
58 family: 'agents',
59 loop: request.loop,
60 trips,
61 question: `${sentence(details)}. Allowing lifts ${trips.length > 1 ? 'these limits' : 'this limit'} until ${shortClock(Math.max(...trips.map((t) => t.leaseUntil)))}. ${request.resume ? 'Resume this agent' : 'Start this agent'}?`,
62 options: [AGENT_ALLOW, AGENT_REFUSE],
63 }
64}
65
66/**
67 * Applies the user's answer to an agents question. Any answer other than the
68 * allow label refuses, and typed words reach the model so it can act on them.
69 * `answer` is undefined when the question was dismissed.
70 */
71export function resolveAgentAsk(
72 state: SessionState,
73 config: Config,
74 now: number,
75 ask: Ask,
76 answer: string | undefined,
77): Verdict {
78 if (answer === AGENT_ALLOW) {
79 for (const trip of ask.trips) state.leases[trip.leaseKey] = Math.max(state.leases[trip.leaseKey] ?? 0, trip.leaseUntil)
80 return ALLOW
81 }
82 state.refusals.agents = now + config.windowMs
83 const details = join(ask.trips.map((t) => t.detail))
84 const said =
85 answer === undefined
86 ? 'The user dismissed the question about starting it'
87 : answer === AGENT_REFUSE
88 ? 'The user chose not to start it'
89 : `The user answered "${answer}" instead of allowing it`
90 const base = `This agent was not started: ${details}. ${said}. Continue without starting agents unless the user says otherwise.`
91 return refuse(state, config, now, ask.trips[0]!.rule, `agents:${ask.trips[0]!.rule}`, base, true)
92}
93
94/**
95 * Holds back a call that tripped while the user is already being asked the
96 * same question. It does not wait for the answer, which could outlast the
97 * hook's time limit, and it is not a refusal of the condition: no ladder rung,
98 * no fuse.
99 */
100export function heldForQuestion(subject: 'agents' | 'heavy'): Deny {
101 const message =
102 subject === 'agents'
103 ? 'The user is being asked about starting agents. Wait for that answer, then start this agent again.'
104 : 'The user is being asked whether to keep working at this context size. Wait for that answer, then make this call again.'
105 return { kind: 'deny', rule: 'asking', message: `${message} [${SIGNATURE}]` }
106}
107
108function agentTrips(
109 state: SessionState,
110 peers: readonly PeerRecord[],
111 config: Config,
112 now: number,
113 limit: LimitReading | undefined,
114): Trip[] {
115 const window = config.windowMs
116 const span = duration(window)
117 const until = now + window
118 const trips: Trip[] = []
119
120 const running = runningAgents(state) + peers.reduce((sum, p) => sum + p.running, 0)
121 if (running >= config.agentMax) {
122 trips.push({
123 rule: 'concurrency',
124 detail: `${count(running, 'agent')} ${running === 1 ? 'is' : 'are'} running on this machine (limit ${config.agentMax})`,
125 leaseKey: 'concurrency',
126 leaseUntil: until,
127 })
128 }
129
130 const starts = countOnMachine(state, peers, 'starts', now, window)
131 if (starts >= config.rollingMax) {
132 trips.push({
133 rule: 'starts',
134 detail: `${count(starts, 'agent')} started on this machine in the last ${span} (limit ${config.rollingMax})`,
135 leaseKey: 'starts',
136 leaseUntil: until,
137 })
138 }
139
140 const processed = tokensWithin(state.agentTokens, now, window) + peers.reduce((sum, p) => sum + tokensWithin(p.agentTokens, now, window), 0)
141 if (processed >= config.agentTokensMax) {
142 trips.push({
143 rule: 'agent-tokens',
144 detail: `subagents processed ${tokens(processed)} tokens in the last ${span} (limit ${tokens(config.agentTokensMax)})`,
145 leaseKey: 'agent-tokens',
146 leaseUntil: until,
147 })
148 }
149
150 const burned = burn(state, peers, now, window)
151 if (burned && burned.points >= config.burnPercent) {
152 trips.push({
153 rule: 'burn',
154 detail: `the 5-hour window rose ${burned.points} points in the last ${span} (now ${burned.percent}%)`,
155 leaseKey: 'burn',
156 leaseUntil: until,
157 })
158 }
159
160 if (limit && limit.percent >= config.limitAskPercent) {
161 trips.push({
162 rule: 'limit',
163 detail: limitDetail(limit),
164 leaseKey: `limit:${limit.kind}`,
165 leaseUntil: limit.resetsAt ?? until,
166 })
167 }
168 return trips
169}
170
171function unattended(state: SessionState, config: Config, now: number, trips: Trip[], details: string): Deny {
172 const first = trips[0]!
173 const advice: Record<string, string> = {
174 concurrency: 'Wait for a running agent to finish before starting another',
175 starts: 'Starts free up as they age out of the window; batch the work or do it yourself',
176 'agent-tokens': 'Do the remaining work yourself instead of in agents',
177 burn: 'Continue without agents until the burn slows',
178 limit: 'Continue without agents',
179 }
180 const base = `This agent was not started: ${details}. Nobody can approve it in this session. ${advice[first.rule] ?? 'Continue without agents'}.`
181 return refuse(state, config, now, first.rule, `agents:${first.rule}`, base, true)
182}
183
184function leased(state: SessionState, trip: Trip, now: number): boolean {
185 const until = state.leases[trip.leaseKey]
186 return until !== undefined && now < until
187}
188
189/** The most-used plan window with a current reading, from this session or its peers. */
190export function highestLimit(state: SessionState, peers: readonly PeerRecord[], now: number): LimitReading | undefined {
191 const readings = [...Object.values(state.limits), ...peers.flatMap((p) => (p.limit ? [p.limit] : []))]
192 const freshest = new Map<string, LimitReading>()
193 for (const r of readings) {
194 if (r.resetsAt !== undefined && r.resetsAt <= now) continue
195 const seen = freshest.get(r.kind)
196 if (!seen || r.at > seen.at) freshest.set(r.kind, r)
197 }
198 let highest: LimitReading | undefined
199 for (const r of freshest.values()) if (!highest || r.percent > highest.percent) highest = r
200 return highest
201}
202
203function limitDetail(limit: LimitReading): string {
204 const reset = limit.resetsAt === undefined ? '' : ` (resets ${shortClock(limit.resetsAt)})`
205 return `the ${windowName(limit.kind)} window is ${limit.percent}% used${reset}`
206}
207
208/** Points the 5-hour window rose within the rolling window, across this machine. */
209function burn(
210 state: SessionState,
211 peers: readonly PeerRecord[],
212 now: number,
213 windowMs: number,
214): { points: number; percent: number } | undefined {
215 const own = state.limits.five_hour
216 const peerReadings = peers.flatMap((p) => (p.limit?.kind === 'five_hour' ? [p.limit] : []))
217 const current = [own, ...peerReadings].reduce<LimitReading | undefined>(
218 (best, r) => (r && (!best || r.at > best.at) ? r : best),
219 undefined,
220 )
221 if (!current) return undefined
222 const sameWindow = [...state.limitHistory, ...peerReadings].filter(
223 (r) => r.kind === 'five_hour' && r.resetsAt === current.resetsAt && now - r.at < windowMs,
224 )
225 if (sameWindow.length === 0) return undefined
226 const lowest = Math.min(...sameWindow.map((r) => r.percent))
227 return { points: Math.round((current.percent - lowest) * 10) / 10, percent: current.percent }
228}
229core/commands.ts 82 lines1// `/agent-guard` and its arguments. The command runs without a model turn, so
2// the model never sees an override and cannot take one as permission.
3
4import { MAIN, type Scope, type SessionState, loopOf } from './state.ts'
5import { duration, sentence, shortClock } from './text.ts'
6import { USAGE, scopeName } from './status.ts'
7
8export type Command =
9 | { action: 'status' }
10 | { action: 'allow'; scope: Scope; minutes: number }
11 | { action: 'resume' }
12 | { action: 'report'; days: number }
13 | { action: 'help'; error?: string }
14
15export const DEFAULT_ALLOW_MINUTES = 10
16const MAX_ALLOW_MINUTES = 8 * 60
17const DEFAULT_REPORT_DAYS = 7
18const MAX_REPORT_DAYS = 3650
19
20export function parseCommand(args: string): Command {
21 const words = args.trim().toLowerCase().split(/\s+/).filter(Boolean)
22 const [verb, ...rest] = words
23 if (verb === undefined || verb === 'status') return { action: 'status' }
24 if (verb === 'help') return { action: 'help' }
25 if (verb === 'resume') return { action: 'resume' }
26 if (verb === 'report') {
27 const days = rest[0] === undefined ? DEFAULT_REPORT_DAYS : Number(rest[0])
28 if (!Number.isInteger(days) || days < 1 || days > MAX_REPORT_DAYS) return { action: 'help', error: `"${rest[0]}" is not a number of days` }
29 return { action: 'report', days }
30 }
31 if (verb === 'allow' || verb === 'pause') {
32 let scope: Scope = 'all'
33 let minutes = DEFAULT_ALLOW_MINUTES
34 for (const word of rest) {
35 if (verb === 'allow' && (word === 'agents' || word === 'context' || word === 'all')) {
36 scope = word
37 continue
38 }
39 const n = Number(word.replace(/m(in(utes?)?)?$/, ''))
40 if (!Number.isInteger(n) || n < 1 || n > MAX_ALLOW_MINUTES) {
41 return { action: 'help', error: `"${word}" is not ${verb === 'allow' ? 'agents, context or ' : ''}a number of minutes from 1 to ${MAX_ALLOW_MINUTES}` }
42 }
43 minutes = n
44 }
45 return { action: 'allow', scope, minutes }
46 }
47 return { action: 'help', error: `unknown subcommand "${verb}"` }
48}
49
50/** Lifts limits for this session, and forgets the declines and the fuse they cover. */
51export function applyOverride(state: SessionState, now: number, scope: Scope, minutes: number): string {
52 const until = now + minutes * 60_000
53 state.override = { until, scope }
54 if (scope !== 'context') {
55 delete state.fuseUntil
56 delete state.refusals.agents
57 state.denials = state.denials.filter((d) => !d.agent)
58 }
59 if (scope !== 'agents') {
60 for (const loop of Object.values(state.loops)) loop.stopped = false
61 state.failures = {}
62 }
63 return `${sentence(scopeName(scope))} lifted for ${duration(minutes * 60_000)}, until ${shortClock(until)}. /agent-guard resume ends it early.`
64}
65
66/** Ends an override and clears declined questions, so the guard asks again. */
67export function resumeGuard(state: SessionState): string {
68 const had = state.override !== undefined
69 delete state.override
70 state.refusals = {}
71 delete state.fuseUntil
72 state.denials = state.denials.filter((d) => !d.agent)
73 loopOf(state, MAIN).stopped = false
74 return had
75 ? 'Limits apply again.'
76 : 'No override was active. Declined questions are cleared, so the guard asks again.'
77}
78
79export function helpText(error?: string): string {
80 return error ? `${sentence(error)}.\n\n${USAGE}` : USAGE
81}
82core/config.ts 127 lines1// Thresholds and switches. The core reads them from a plain record of strings
2// so that any harness can supply them: environment variables, a settings file,
3// or a test's literal object.
4
5export type Config = {
6 enabled: boolean
7 /** Rolling window for agent starts, token burn, tool budgets and denials. */
8 windowMs: number
9 /** Agents running at once across every session on this machine. */
10 agentMax: number
11 /** Agent starts and resumes per window across every session on this machine. */
12 rollingMax: number
13 /** Subagent layers below the main conversation; 1 means subagents cannot spawn. */
14 depthMax: number
15 /** Tokens processed by subagent requests per window across this machine. */
16 agentTokensMax: number
17 /** Plan-window percentage at which new agents need the user's approval. */
18 limitAskPercent: number
19 /** Plan-window percentage at which new agents are refused. */
20 limitDenyPercent: number
21 /** Plan-window points burned within one window that make new agents ask. */
22 burnPercent: number
23 /** Context at which Claude is told, once, to keep the turn bounded. */
24 contextWarn: number
25 /** Context at which a prompt needs the user's approval before it is sent. */
26 contextHard: number
27 /** Context above which a loop's tool calls count against the tool budget. */
28 toolContext: number
29 /** Tool calls per window a loop may make above `toolContext`. */
30 toolMax: number
31 /** Idle time after which a heavy session's prompt cache has likely expired. */
32 dormantMs: number
33 /** Context at which a dormant session counts as heavy. */
34 dormantContext: number
35 /** Dormant heavy resumes per window on this machine when nobody can be asked. */
36 dormantMax: number
37 /** Identical consecutive tool failures before the identical retry is refused. */
38 failureMax: number
39 /** Refused agent calls per `fuseMs` that pause agent spawns for `fuseMs`. */
40 fuseMax: number
41 fuseMs: number
42 /** How long another session's record, or an agent with no sign of activity, still counts. */
43 peerTtlMs: number
44 journal: boolean
45 journalDays: number
46}
47
48type Field = Exclude<keyof Config, 'enabled' | 'journal'>
49
50type Spec = {
51 field: Field
52 env: string
53 default: number
54 min: number
55 max: number
56 /** Multiplier from the variable's unit to the field's (seconds to milliseconds). */
57 scale: number
58}
59
60const SECONDS = 1000
61
62export const SPECS: readonly Spec[] = [
63 { field: 'windowMs', env: 'AGENT_GUARD_WINDOW_SECONDS', default: 600, min: 60, max: 86_400, scale: SECONDS },
64 { field: 'agentMax', env: 'AGENT_GUARD_AGENT_MAX', default: 4, min: 1, max: 1000, scale: 1 },
65 { field: 'rollingMax', env: 'AGENT_GUARD_ROLLING_MAX', default: 12, min: 1, max: 10_000, scale: 1 },
66 { field: 'depthMax', env: 'AGENT_GUARD_DEPTH_MAX', default: 1, min: 1, max: 10, scale: 1 },
67 { field: 'agentTokensMax', env: 'AGENT_GUARD_AGENT_TOKENS_MAX', default: 10_000_000, min: 1, max: 1e12, scale: 1 },
68 { field: 'limitAskPercent', env: 'AGENT_GUARD_LIMIT_ASK_PERCENT', default: 80, min: 1, max: 100, scale: 1 },
69 { field: 'limitDenyPercent', env: 'AGENT_GUARD_LIMIT_DENY_PERCENT', default: 95, min: 1, max: 101, scale: 1 },
70 { field: 'burnPercent', env: 'AGENT_GUARD_BURN_PERCENT', default: 20, min: 1, max: 101, scale: 1 },
71 { field: 'contextWarn', env: 'AGENT_GUARD_CONTEXT_WARN', default: 300_000, min: 1, max: 1e9, scale: 1 },
72 { field: 'contextHard', env: 'AGENT_GUARD_CONTEXT_HARD', default: 500_000, min: 1, max: 1e9, scale: 1 },
73 { field: 'toolContext', env: 'AGENT_GUARD_TOOL_CONTEXT', default: 400_000, min: 1, max: 1e9, scale: 1 },
74 { field: 'toolMax', env: 'AGENT_GUARD_TOOL_MAX', default: 20, min: 1, max: 100_000, scale: 1 },
75 { field: 'dormantMs', env: 'AGENT_GUARD_DORMANT_SECONDS', default: 3600, min: 60, max: 30 * 86_400, scale: SECONDS },
76 { field: 'dormantContext', env: 'AGENT_GUARD_DORMANT_CONTEXT', default: 150_000, min: 1, max: 1e9, scale: 1 },
77 { field: 'dormantMax', env: 'AGENT_GUARD_DORMANT_MAX', default: 1, min: 0, max: 1000, scale: 1 },
78 { field: 'failureMax', env: 'AGENT_GUARD_TOOL_FAILURE_MAX', default: 3, min: 1, max: 1000, scale: 1 },
79 { field: 'fuseMax', env: 'AGENT_GUARD_FUSE_MAX', default: 5, min: 1, max: 1000, scale: 1 },
80 { field: 'fuseMs', env: 'AGENT_GUARD_FUSE_SECONDS', default: 600, min: 10, max: 86_400, scale: SECONDS },
81 { field: 'peerTtlMs', env: 'AGENT_GUARD_PEER_TTL_SECONDS', default: 900, min: 60, max: 86_400, scale: SECONDS },
82 { field: 'journalDays', env: 'AGENT_GUARD_JOURNAL_DAYS', default: 90, min: 1, max: 3650, scale: 1 },
83]
84
85export const SWITCHES = { enabled: 'AGENT_GUARD', journal: 'AGENT_GUARD_JOURNAL' } as const
86
87const OFF = new Set(['0', 'false', 'off', 'no'])
88
89export function defaultConfig(): Config {
90 const config = { enabled: true, journal: true } as Config
91 for (const spec of SPECS) config[spec.field] = spec.default * spec.scale
92 return config
93}
94
95/**
96 * Builds the config from variables by name. Unset variables take their
97 * defaults; a value that is not a number in range is reported in `problems`
98 * and falls back to the default, so a typo never disables a protection.
99 */
100export function parseConfig(vars: Readonly<Record<string, string | undefined>>): {
101 config: Config
102 problems: string[]
103} {
104 const config = defaultConfig()
105 const problems: string[] = []
106 for (const [field, name] of Object.entries(SWITCHES) as [keyof typeof SWITCHES, string][]) {
107 const raw = vars[name]?.trim().toLowerCase()
108 if (raw) config[field] = !OFF.has(raw)
109 }
110 for (const spec of SPECS) {
111 const raw = vars[spec.env]?.trim()
112 if (!raw) continue
113 const value = Number(raw.replaceAll('_', ''))
114 if (!Number.isFinite(value) || value < spec.min || value > spec.max) {
115 problems.push(`${spec.env}=${raw} is not a number from ${spec.min} to ${spec.max}; using ${spec.default}`)
116 continue
117 }
118 config[spec.field] = Math.round(value * spec.scale)
119 }
120 if (config.limitDenyPercent < config.limitAskPercent) {
121 problems.push(
122 `AGENT_GUARD_LIMIT_DENY_PERCENT is below AGENT_GUARD_LIMIT_ASK_PERCENT; new agents are refused from ${config.limitDenyPercent}%`,
123 )
124 }
125 return { config, problems }
126}
127core/journal.ts 149 lines1// The local journal: one row per intervention, answer, override, limit
2// crossing and lockout. Rows hold rule names, tool names, coarse buckets and
3// times, never prompt text, tool inputs or model output.
4
5import { count, duration, windowName } from './text.ts'
6
7export type JournalEvent =
8 | 'ask' // the guard asked the user
9 | 'answer' // the user answered (or dismissed) a question
10 | 'deny' // a model-facing refusal
11 | 'drop' // a prompt was held
12 | 'override' // /agent-guard allow or pause
13 | 'limit' // a plan window crossed an ask or deny threshold
14 | 'lockout' // a request failed on a usage limit
15
16export type JournalRow = {
17 t: number
18 /** First 8 characters of the session id. */
19 s: string
20 ev: JournalEvent
21 rule?: string
22 /** For `answer`: allow, compact, refuse, cancel or dismiss, or unavailable when the question could not be shown. */
23 answer?: string
24 tool?: string
25 ctx?: string
26 /** Plan-window percentage, for `limit` and `lockout`. */
27 pct?: number
28 kind?: string
29 /** For `deny`: which refusal of the condition this was. */
30 n?: number
31}
32
33const KEY_PREFIX = 'journal:'
34const DAY_MS = 86_400_000
35
36/** Each session appends only to its own key for the day, so sessions never overwrite each other. */
37export function journalKey(sessionId: string, now: number): string {
38 return `${KEY_PREFIX}${new Date(now).toISOString().slice(0, 10)}:${sessionId}`
39}
40
41/**
42 * Picks the journal's keys out of the store's and splits them by age: `recent`
43 * days can hold rows from the last `days` days, `expired` days ended before
44 * those began.
45 */
46export function splitJournalKeys(keys: readonly string[], now: number, days: number): { recent: string[]; expired: string[] } {
47 const start = now - days * DAY_MS
48 const split = { recent: [] as string[], expired: [] as string[] }
49 for (const key of keys) {
50 if (!key.startsWith(KEY_PREFIX)) continue
51 const day = Date.parse(key.slice(KEY_PREFIX.length, KEY_PREFIX.length + 10))
52 if (Number.isNaN(day)) continue
53 split[day + DAY_MS < start ? 'expired' : 'recent'].push(key)
54 }
55 return split
56}
57
58export function isJournalRow(value: unknown): value is JournalRow {
59 if (typeof value !== 'object' || value === null) return false
60 const r = value as Record<string, unknown>
61 return typeof r.t === 'number' && typeof r.ev === 'string' && typeof r.s === 'string'
62}
63
64type RuleCount = { asks: number; allowed: number; declined: number; denials: number; drops: number }
65
66/** The text `/agent-guard report` prints. */
67export function formatReport(rows: readonly JournalRow[], now: number, days: number): string {
68 const recent = rows.filter((r) => now - r.t < days * DAY_MS).sort((a, b) => a.t - b.t)
69 const lines = [`agent-usage-guard report: last ${count(days, 'day')}`]
70 if (recent.length === 0) {
71 lines.push('', 'Nothing recorded yet. The journal fills as the guard asks, refuses, or sees a limit.')
72 return lines.join('\n')
73 }
74 const sessions = new Set(recent.map((r) => r.s)).size
75 lines.push(`${count(recent.length, 'event')} from ${count(sessions, 'session')}`)
76
77 const lockouts = recent.filter((r) => r.ev === 'lockout' && r.kind !== 'model')
78 const modelScoped = recent.filter((r) => r.ev === 'lockout' && r.kind === 'model').length
79 const weeks = Math.max(days / 7, 1)
80 lines.push('', `Usage-limit lockouts: ${lockouts.length} (${(lockouts.length / weeks).toFixed(1)} per week)`)
81 for (const row of lockouts.slice(-5)) {
82 lines.push(` ${stamp(row.t)}${row.kind && row.kind !== 'unknown' ? ` · ${windowName(row.kind)} window` : ''}`)
83 }
84 if (modelScoped > 0) lines.push(`Limits on one model, where /model switched away: ${modelScoped}`)
85
86 const crossings = recent.filter((r) => r.ev === 'limit')
87 if (crossings.length > 0) {
88 const byThreshold = countBy(crossings, (r) => `${windowName(r.kind ?? '?')} ${r.rule ?? ''}`.trim())
89 lines.push(`Limit crossings: ${[...byThreshold].map(([k, n]) => `${k} ×${n}`).join(', ')}`)
90 }
91
92 const perRule = new Map<string, RuleCount>()
93 const bump = (rule: string | undefined, field: keyof RuleCount): void => {
94 const key = rule ?? 'unknown'
95 const entry = perRule.get(key) ?? { asks: 0, allowed: 0, declined: 0, denials: 0, drops: 0 }
96 entry[field] += 1
97 perRule.set(key, entry)
98 }
99 for (const row of recent) {
100 if (row.ev === 'ask') bump(row.rule, 'asks')
101 else if (row.ev === 'deny') bump(row.rule, 'denials')
102 else if (row.ev === 'drop') bump(row.rule, 'drops')
103 // A question that could not be shown was neither allowed nor declined.
104 else if (row.ev === 'answer' && row.answer !== 'unavailable') {
105 bump(row.rule, row.answer === 'allow' || row.answer === 'compact' ? 'allowed' : 'declined')
106 }
107 }
108 if (perRule.size > 0) {
109 lines.push('', 'By rule (asks: allowed / declined · refusals to the model · prompts held)')
110 const width = Math.max(...[...perRule.keys()].map((k) => k.length))
111 for (const [rule, c] of [...perRule].sort((a, b) => total(b[1]) - total(a[1]))) {
112 lines.push(` ${rule.padEnd(width)} asks ${c.asks}: ${c.allowed} / ${c.declined} · refusals ${c.denials} · held ${c.drops}`)
113 }
114 }
115
116 const denials = recent.filter((r) => r.ev === 'deny')
117 const late = denials.filter((r) => (r.n ?? 1) >= 3).length
118 const fuses = denials.filter((r) => r.rule === 'fuse').length
119 if (denials.length > 0) lines.push(`Refusals at rung 3 or later: ${late} · refused while the agent fuse was burning: ${fuses}`)
120
121 const tools = countBy(denials.filter((r) => r.tool), (r) => r.tool ?? '')
122 if (tools.size > 0) lines.push(`Refusals by tool: ${[...tools].map(([k, n]) => `${k} ${n}`).join(', ')}`)
123
124 const contexts = countBy(recent.filter((r) => r.ctx), (r) => r.ctx ?? '')
125 if (contexts.size > 0) lines.push(`Context when it acted: ${[...contexts].map(([k, n]) => `${k} ${n}`).join(', ')}`)
126
127 const overrides = recent.filter((r) => r.ev === 'override')
128 if (overrides.length > 0) lines.push(`Overrides: ${overrides.length} (${[...countBy(overrides, (r) => r.rule ?? 'all')].map(([k, n]) => `${k} ${n}`).join(', ')})`)
129
130 lines.push('', `Span: ${stamp(recent[0]!.t)} to ${stamp(recent.at(-1)!.t)} (${duration(recent.at(-1)!.t - recent[0]!.t)})`)
131 return lines.join('\n')
132}
133
134function total(c: RuleCount): number {
135 return c.asks + c.denials + c.drops
136}
137
138function countBy<T>(items: readonly T[], key: (item: T) => string): Map<string, number> {
139 const counts = new Map<string, number>()
140 for (const item of items) counts.set(key(item), (counts.get(key(item)) ?? 0) + 1)
141 return new Map([...counts].sort((a, b) => b[1] - a[1]))
142}
143
144function stamp(ms: number): string {
145 const d = new Date(ms)
146 const pad = (n: number): string => String(n).padStart(2, '0')
147 return `${d.getFullYear()}-${pad(d.getMonth() + 1)}-${pad(d.getDate())} ${pad(d.getHours())}:${pad(d.getMinutes())}`
148}
149core/lockout.ts 29 lines1// Which failed requests were usage-limit lockouts: the outcome the guard
2// exists to prevent, and the figure its report is judged by.
3
4import type { LimitReading } from './state.ts'
5
6/** Claude Code offers `/model` when only one model's limit was reached. */
7const MODEL_SWITCH = /\/model\b[^.\n]{0,40}\bswitch\b|\bswitch(?:ing)?\s+(?:to\s+)?(?:another\s+|a\s+different\s+)?models?\b|\bopus\b[^.\n]{0,24}\blimit\b/i
8
9const PHRASES: ReadonlyArray<[RegExp, string]> = [
10 [/\b(?:5-hour|five-hour|session) limit\b/i, 'five_hour'],
11 [/\bweekly limit\b/i, 'seven_day'],
12 [/\bmonthly spend limit\b|\bspend limit\b|\busage credits\b/i, 'spend_limit'],
13 [/\busage limit\b|\blimit reached\b|\bhit your limit\b/i, 'unknown'],
14]
15
16/**
17 * The window a failed request ran into, or undefined for a failure that is
18 * not a usage limit, such as an overloaded or transient 429. A limit scoped to
19 * one model is reported as `model`: switching models continues the work.
20 */
21export function lockoutKind(error: string, text: string, readings: readonly LimitReading[]): string | undefined {
22 if (error !== 'rate_limit' && error !== 'billing_error') return undefined
23 if (MODEL_SWITCH.test(text)) return 'model'
24 const named = PHRASES.find(([pattern]) => pattern.test(text))?.[1]
25 if (named !== undefined && named !== 'unknown') return named
26 // A generic phrase, or none, is pinned to the window a reading shows full.
27 return readings.find((r) => r.percent >= 100)?.kind ?? named
28}
29core/observe.ts 224 lines1// Facts the harness reports after they happen. None of these decide anything;
2// they keep the state the gates read accurate.
3
4import type { Config } from './config.ts'
5import { type AgentKind, type LimitReading, type LoopId, MAIN, type SessionState, loopOf } from './state.ts'
6
7export type TokenUsage = {
8 input: number
9 output: number
10 cacheRead: number
11 cacheWrite: number
12}
13
14/** Tokens a request carried in: what the loop's next request re-reads, roughly. */
15function contextOf(usage: TokenUsage): number {
16 return usage.input + usage.cacheRead + usage.cacheWrite
17}
18
19/**
20 * One model request finished in `loop`, with the usage the API reported. An
21 * agent that makes one is running, even if it had gone quiet for long enough
22 * to stop counting.
23 */
24export function observeResponse(state: SessionState, config: Config, now: number, loop: LoopId, usage: TokenUsage): void {
25 const context = contextOf(usage)
26 setContext(state, config, loop, context)
27 loopOf(state, loop).activeAt = now
28 const agent = state.agents[loop]
29 if (!agent) return
30 agent.running = true
31 state.agentTokens.push([now, context + usage.output])
32}
33
34/** The main conversation's context as the harness reports it, without a request. */
35export function observeMainContext(state: SessionState, config: Config, tokens: number): void {
36 setContext(state, config, MAIN, tokens)
37}
38
39/** Records a loop's context. Below a threshold, what was tied to it ends: the heavy approval and tool budget, the warning, the prompt approval. */
40function setContext(state: SessionState, config: Config, loop: LoopId, context: number): void {
41 const l = loopOf(state, loop)
42 l.context = context
43 if (context < config.toolContext) {
44 delete l.heavyLeaseUntil
45 l.heavyCalls = []
46 }
47 if (loop !== MAIN) return
48 if (context < config.contextWarn) state.warned = false
49 if (context < config.contextHard) delete state.leases['prompt-context']
50}
51
52/** The conversation was compacted: the old context and approvals tied to it are gone. */
53export function observeCompaction(state: SessionState, loop: LoopId = MAIN): void {
54 const l = loopOf(state, loop)
55 delete l.context
56 delete l.heavyLeaseUntil
57 l.heavyCalls = []
58 if (loop === MAIN) {
59 state.warned = false
60 delete state.leases['prompt-context']
61 }
62}
63
64/**
65 * The plan windows the latest response reported. Returns the windows whose
66 * percentage crossed one of the guard's thresholds since the previous reading,
67 * so the adapter can journal the crossing once.
68 */
69export function observeLimits(
70 state: SessionState,
71 config: Config,
72 readings: readonly LimitReading[],
73): Array<{ reading: LimitReading; threshold: number }> {
74 const crossings: Array<{ reading: LimitReading; threshold: number }> = []
75 for (const reading of readings) {
76 const previous = state.limits[reading.kind]
77 const sameWindow = previous?.resetsAt === reading.resetsAt
78 for (const threshold of [config.limitAskPercent, config.limitDenyPercent]) {
79 const before = sameWindow && previous ? previous.percent : 0
80 if (before < threshold && reading.percent >= threshold) crossings.push({ reading, threshold })
81 }
82 state.limits[reading.kind] = reading
83 state.limitHistory.push(reading)
84 }
85 return crossings
86}
87
88export type SpawnFacts = {
89 agentId: string
90 kind: AgentKind
91 depth: number
92 name?: string
93}
94
95/**
96 * Holds a slot for an agent the gate just allowed. Call it in the same
97 * synchronous step as the verdict: parallel spawns are judged concurrently, and
98 * each must see the slots the others already took. Returns the start's time,
99 * which `releaseStart` takes back if the agent never starts.
100 */
101export function reserveStart(state: SessionState, now: number): number {
102 state.pendingStarts += 1
103 state.starts.push(now)
104 return now
105}
106
107/**
108 * Ends a reservation: `agentId` started, or, when it is undefined, the start
109 * failed and is not counted. Once no gated agent is starting, the starts none
110 * of them claimed are counted.
111 */
112export function releaseStart(state: SessionState, reservedAt: number, agentId: string | undefined): void {
113 state.pendingStarts = Math.max(0, state.pendingStarts - 1)
114 if (agentId === undefined) {
115 const index = state.starts.indexOf(reservedAt)
116 if (index >= 0) state.starts.splice(index, 1)
117 } else {
118 delete state.unclaimed[agentId]
119 }
120 countUnclaimed(state)
121}
122
123/** An agent the gate allowed has started; its start was counted by `reserveStart`. */
124export function observeSpawn(state: SessionState, now: number, facts: SpawnFacts): void {
125 state.agents[facts.agentId] = {
126 id: facts.agentId,
127 kind: facts.kind,
128 depth: facts.depth,
129 running: true,
130 startedAt: now,
131 ...(facts.name ? { name: facts.name } : {}),
132 }
133}
134
135/** A known agent's loop ran again: a resumed subagent or a teammate that woke. */
136export function observeAgentActive(state: SessionState, agentId: string): void {
137 const agent = state.agents[agentId]
138 if (agent) agent.running = true
139}
140
141/**
142 * The harness reported an agent's loop starting. A known agent is running
143 * again. An unknown one started without passing the gate, such as a /subtask
144 * fork the user ran: it is counted, never blocked. While a gated agent is
145 * still starting, its own report looks the same, so an unknown start waits
146 * until every gated agent has claimed its own. An agent with no type is the
147 * harness's own and is left out, which is the one case that returns false.
148 */
149export function observeAgentStart(state: SessionState, now: number, agentId: string, type: string): boolean {
150 if (state.agents[agentId]) {
151 observeAgentActive(state, agentId)
152 return true
153 }
154 if (!type) return false
155 state.unclaimed[agentId] = { at: now, kind: type === 'fork' ? 'fork' : 'subagent' }
156 countUnclaimed(state)
157 return true
158}
159
160/** Once no gated agent is starting, counts the starts none of them claimed. */
161function countUnclaimed(state: SessionState): void {
162 if (state.pendingStarts > 0) return
163 for (const [agentId, { at, kind }] of Object.entries(state.unclaimed)) {
164 observeSpawn(state, at, { agentId, kind, depth: 1 })
165 state.starts.push(at)
166 }
167 state.unclaimed = {}
168}
169
170/** A known agent finished a run. Its context leaves with it. */
171export function observeAgentRunEnd(state: SessionState, agentId: string): void {
172 const agent = state.agents[agentId]
173 if (agent) agent.running = false
174 const loop = state.loops[agentId]
175 if (loop) loop.stopped = false
176}
177
178/**
179 * Agents with no sign of life for `ttlMs` stop counting as running: no model
180 * request, and no tool call running or finished. That catches an agent whose
181 * end the harness never reported; one that is only slow comes back with its
182 * next request.
183 */
184export function expireAgents(state: SessionState, now: number, ttlMs: number): void {
185 for (const agent of Object.values(state.agents)) {
186 const loop = state.loops[agent.id]
187 if (!agent.running || (loop?.toolsRunning ?? 0) > 0) continue
188 if (now - Math.max(agent.startedAt, loop?.activeAt ?? 0) > ttlMs) agent.running = false
189 }
190}
191
192/** A tool call is starting to run in `loop`. */
193export function observeToolStart(state: SessionState, loop: LoopId): void {
194 const l = loopOf(state, loop)
195 l.toolsRunning = (l.toolsRunning ?? 0) + 1
196}
197
198/**
199 * A tool call finished; identical failures in a row feed the retry fuse. Any
200 * other call by the same loop breaks the row, so test, edit, re-run is never
201 * a retry loop.
202 */
203export function observeToolResult(state: SessionState, now: number, loop: LoopId, fingerprint: string, failed: boolean): void {
204 const l = loopOf(state, loop)
205 l.toolsRunning = Math.max(0, (l.toolsRunning ?? 0) - 1)
206 l.activeAt = now
207 for (const [key, failure] of Object.entries(state.failures)) {
208 if (failure.loop === loop && key !== fingerprint) delete state.failures[key]
209 }
210 if (!failed) {
211 delete state.failures[fingerprint]
212 return
213 }
214 const previous = state.failures[fingerprint]
215 state.failures[fingerprint] = { count: (previous?.count ?? 0) + 1, last: now, loop }
216}
217
218/** The main turn ended: per-turn conditions reset. */
219export function observeTurnEnd(state: SessionState): void {
220 state.reopened = false
221 const main = state.loops[MAIN]
222 if (main) main.stopped = false
223}
224core/prompts.ts 126 lines1// The prompt gate. It never touches a prompt the user did not type, such as a
2// background task's notification or another session's message: holding those
3// loses results, and the context warning waits for the user's own next prompt.
4// It never holds a prompt over plan limits either; Claude Code waits at a limit
5// and resumes on its own.
6
7import type { Config } from './config.ts'
8import { MAIN, type PeerRecord, type Scope, type SessionState, countOnMachine, overrideCovers } from './state.ts'
9import { count, duration, sentence, tokens } from './text.ts'
10import { ALLOW, type Allow, type Ask, type Drop, type Rule, SIGNATURE, type Verdict } from './verdict.ts'
11
12/** Who a prompt came from, as far as the guard cares. */
13export type PromptSource = 'user' | 'headless' | 'other'
14
15export type PromptRequest = {
16 source: PromptSource
17 text: string
18 /** Time since the session's last model response, when known. */
19 idleMs?: number
20}
21
22const PROMPT_COMPACT = 'Compact, then send'
23const PROMPT_SEND = 'Send anyway'
24const PROMPT_CANCEL = 'Cancel'
25
26const MARKERS: Readonly<Record<string, Scope>> = {
27 '[allow-usage-guard]': 'all',
28 '[allow-agent-burst]': 'agents',
29}
30
31/**
32 * The bracket markers of earlier versions, when one is the whole prompt. They
33 * act like `/agent-guard allow`: the prompt is not sent, so no model turn runs
34 * and the model never reads the marker.
35 */
36export function markerScope(text: string): Scope | undefined {
37 return MARKERS[text.trim().toLowerCase()]
38}
39
40export function promptGate(
41 state: SessionState,
42 peers: readonly PeerRecord[],
43 config: Config,
44 now: number,
45 request: PromptRequest,
46): Verdict {
47 if (!config.enabled || request.source === 'other') return ALLOW
48 const context = state.loops[MAIN]?.context
49 if (context === undefined || overrideCovers(state, now, 'context')) return withWarning(state, config, context)
50
51 const hardLeased = state.leases['prompt-context'] !== undefined
52 if (context >= config.contextHard && !hardLeased) {
53 const detail = `this session holds ${tokens(context)} context tokens (limit ${tokens(config.contextHard)}), and sending re-reads all of them`
54 if (state.canAsk && request.source === 'user') {
55 return promptAsk('prompt-context', `This session holds ${tokens(context)} context tokens, and sending re-reads all of them; compacting first shrinks this request and every later one. Send this prompt?`, detail)
56 }
57 return drop('prompt-context', `${detail}. Compact the session first, or raise AGENT_GUARD_CONTEXT_HARD.`)
58 }
59
60 const idle = request.idleMs
61 if (idle !== undefined && idle >= config.dormantMs && context >= config.dormantContext) {
62 const detail = `this session was idle for ${duration(idle)}, so its prompt cache has likely expired and this request re-writes about ${tokens(context)} tokens`
63 if (state.canAsk && request.source === 'user') {
64 return promptAsk('dormant', `${sentence(detail)}. Send this prompt?`, detail)
65 }
66 const recent = countOnMachine(state, peers, 'dormantResumes', now, config.windowMs)
67 if (recent >= config.dormantMax) {
68 return drop(
69 'dormant',
70 `${detail}, and ${count(recent, 'heavy idle session')} already resumed on this machine in the last ${duration(config.windowMs)} (limit ${config.dormantMax}). Try again later, or compact the session first.`,
71 )
72 }
73 state.dormantResumes.push(now)
74 }
75 return withWarning(state, config, context)
76}
77
78/** The user chose to compact first: drop the prompt, compact, then send it. */
79export type CompactThenSend = { kind: 'compact-then-send'; rule: Rule; message: string }
80
81export type PromptOutcome = Allow | Drop | CompactThenSend
82
83/** Applies the user's answer to a prompt question. Dismissing it cancels. */
84export function resolvePromptAsk(state: SessionState, ask: Ask, answer: string | undefined): PromptOutcome {
85 const trip = ask.trips[0]!
86 if (answer === PROMPT_COMPACT) {
87 return { kind: 'compact-then-send', rule: trip.rule, message: `${SIGNATURE}: compacting first, because ${trip.detail}; your prompt is sent once that finishes.` }
88 }
89 if (answer === PROMPT_SEND) {
90 if (trip.rule === 'prompt-context') state.leases['prompt-context'] = Number.MAX_SAFE_INTEGER
91 return ALLOW
92 }
93 return drop(trip.rule, 'cancelled; your prompt is back in the input box.')
94}
95
96function promptAsk(rule: 'prompt-context' | 'dormant', question: string, detail: string): Ask {
97 return {
98 kind: 'ask',
99 family: 'prompt',
100 loop: MAIN,
101 trips: [{ rule, detail, leaseKey: rule, leaseUntil: Number.MAX_SAFE_INTEGER }],
102 question,
103 options: [PROMPT_COMPACT, PROMPT_SEND, PROMPT_CANCEL],
104 }
105}
106
107function drop(rule: Rule, reason: string): Drop {
108 return { kind: 'drop', rule, message: `${SIGNATURE}: ${reason}` }
109}
110
111/** The warning for a prompt the user chose to send, if Claude has not had it yet. */
112export function contextNote(state: SessionState, config: Config): string | undefined {
113 const verdict = withWarning(state, config, state.loops[MAIN]?.context)
114 return verdict.kind === 'allow' ? verdict.note : undefined
115}
116
117/** Tells Claude once, when the context first crosses the warning threshold. */
118function withWarning(state: SessionState, config: Config, context: number | undefined): Verdict {
119 if (context === undefined || context < config.contextWarn || state.warned) return ALLOW
120 state.warned = true
121 return {
122 kind: 'allow',
123 note: `${SIGNATURE}: this session's context is ${tokens(context)} tokens. Keep this turn bounded, and recommend /compact to the user at the next natural break.`,
124 }
125}
126core/state.ts 237 lines1// The guard's state for one session, plus the small record each session
2// publishes for the others on the same machine. Everything here is plain JSON,
3// so an adapter can keep it in memory, in a key-value store, or in a file
4// between hook processes.
5//
6// The core updates a SessionState in place, inside synchronous functions only:
7// one adapter owns one state, and no update spans an await, so concurrent hooks
8// in the same process cannot interleave halfway through a change.
9
10/** `main` for the main conversation, otherwise the agent's id. */
11export type LoopId = string
12
13export const MAIN: LoopId = 'main'
14
15export type AgentKind = 'subagent' | 'teammate' | 'workflow' | 'fork'
16
17export type AgentRecord = {
18 id: string
19 kind: AgentKind
20 /** Layers below the main conversation: 1 for an agent the main loop started. */
21 depth: number
22 running: boolean
23 startedAt: number
24 name?: string
25}
26
27export type LoopState = {
28 /** Input tokens of the loop's latest model request: what the next one re-reads. */
29 context?: number
30 /** When the loop last finished a model request or a tool call. */
31 activeAt?: number
32 /** Tool calls the loop has running: an agent in one is running, however long the tool takes. */
33 toolsRunning?: number
34 /** Tool calls made while the loop was above the tool-budget threshold. */
35 heavyCalls: number[]
36 /** The user approved heavy work in this loop until then, or until it compacts. */
37 heavyLeaseUntil?: number
38 /** The user chose to stop this turn: refuse its tool calls until it ends. */
39 stopped?: boolean
40}
41
42export type LimitReading = {
43 at: number
44 /** `five_hour`, `seven_day`, a gateway's `spend_limit`, ... */
45 kind: string
46 percent: number
47 resetsAt?: number
48}
49
50export type Scope = 'all' | 'agents' | 'context'
51
52export type Denial = { at: number; key: string; agent: boolean }
53
54export type SessionState = {
55 id: string
56 /** A person can answer questions in this session. */
57 canAsk: boolean
58 loops: Record<LoopId, LoopState>
59 agents: Record<string, AgentRecord>
60 /** Agents the gate allowed that Claude Code has not reported started yet. */
61 pendingStarts: number
62 /** Agents that started outside the gate while a gated agent was starting, by id: counted once it has. */
63 unclaimed: Record<string, { at: number; kind: AgentKind }>
64 /** This session's agent starts and resumes, newest last. */
65 starts: number[]
66 /** Tokens each subagent request processed, as [time, tokens]. */
67 agentTokens: Array<[number, number]>
68 /** Latest reading per plan window, and the readings inside the current window. */
69 limits: Record<string, LimitReading>
70 limitHistory: LimitReading[]
71 /** An approved condition stays approved until the time stored under its key. */
72 leases: Record<string, number>
73 /** A family the user declined is refused without asking until then. */
74 refusals: Record<string, number>
75 denials: Denial[]
76 fuseUntil?: number
77 /** Consecutive identical failures per tool-call fingerprint, and the loop that made them. */
78 failures: Record<string, { count: number; last: number; loop: LoopId }>
79 override?: { until: number; scope: Scope }
80 /** The context warning was given and still applies. */
81 warned: boolean
82 /** A Stop hook reopened the current turn, so "end the turn" cannot clear a block. */
83 reopened: boolean
84 dormantResumes: number[]
85}
86
87/** What one session publishes for the others: counts, never content. */
88export type PeerRecord = {
89 v: 2
90 /** When the session last made a model request or changed this record. */
91 at: number
92 running: number
93 starts: number[]
94 /** Subagent tokens per minute, as [the minute's latest time, tokens], so they age out like starts. */
95 agentTokens: Array<[number, number]>
96 dormantResumes: number[]
97 limit?: LimitReading
98}
99
100export function newSession(id: string, canAsk: boolean): SessionState {
101 return {
102 id,
103 canAsk,
104 loops: {},
105 agents: {},
106 pendingStarts: 0,
107 unclaimed: {},
108 starts: [],
109 agentTokens: [],
110 limits: {},
111 limitHistory: [],
112 leases: {},
113 refusals: {},
114 denials: [],
115 failures: {},
116 warned: false,
117 reopened: false,
118 dormantResumes: [],
119 }
120}
121
122export function loopOf(state: SessionState, id: LoopId): LoopState {
123 const existing = state.loops[id]
124 if (existing) return existing
125 const created: LoopState = { heavyCalls: [] }
126 state.loops[id] = created
127 return created
128}
129
130export function within(times: readonly number[], now: number, windowMs: number): number[] {
131 return times.filter((t) => now - t < windowMs)
132}
133
134/** Tokens of the [time, tokens] pairs inside the window. */
135export function tokensWithin(pairs: ReadonlyArray<readonly [number, number]>, now: number, windowMs: number): number {
136 return pairs.reduce((sum, [t, n]) => (now - t < windowMs ? sum + n : sum), 0)
137}
138
139/** Entries of one rolling list inside the window, in this session and its peers together. */
140export function countOnMachine(
141 state: SessionState,
142 peers: readonly PeerRecord[],
143 list: 'starts' | 'dormantResumes',
144 now: number,
145 windowMs: number,
146): number {
147 return [state[list], ...peers.map((p) => p[list])].reduce((sum, times) => sum + within(times, now, windowMs).length, 0)
148}
149
150/** Running agents, counting those allowed a moment ago that are still starting. */
151export function runningAgents(state: SessionState): number {
152 return Object.values(state.agents).filter((a) => a.running).length + state.pendingStarts
153}
154
155export function overrideCovers(state: SessionState, now: number, scope: Exclude<Scope, 'all'>): boolean {
156 const override = state.override
157 if (!override || now >= override.until) return false
158 return override.scope === 'all' || override.scope === scope
159}
160
161/** Drops what has aged out of every rolling list, so the state stays small. */
162export function prune(state: SessionState, now: number, windowMs: number): void {
163 state.starts = within(state.starts, now, windowMs)
164 state.agentTokens = state.agentTokens.filter(([t]) => now - t < windowMs)
165 state.limitHistory = state.limitHistory.filter((r) => now - r.at < windowMs)
166 state.denials = state.denials.filter((d) => now - d.at < windowMs)
167 state.dormantResumes = within(state.dormantResumes, now, windowMs)
168 for (const [key, until] of Object.entries(state.leases)) if (now >= until) delete state.leases[key]
169 for (const [key, until] of Object.entries(state.refusals)) if (now >= until) delete state.refusals[key]
170 for (const [key, failure] of Object.entries(state.failures)) {
171 if (now - failure.last >= windowMs) delete state.failures[key]
172 }
173 for (const loop of Object.values(state.loops)) loop.heavyCalls = within(loop.heavyCalls, now, windowMs)
174 if (state.fuseUntil !== undefined && now >= state.fuseUntil) delete state.fuseUntil
175 if (state.override && now >= state.override.until) delete state.override
176}
177
178/** The record this session publishes for its peers. */
179export function peerRecord(state: SessionState, now: number, windowMs: number): PeerRecord {
180 const fiveHour = state.limits.five_hour
181 return {
182 v: 2,
183 at: now,
184 running: runningAgents(state),
185 starts: within(state.starts, now, windowMs),
186 agentTokens: perMinute(state.agentTokens, now, windowMs),
187 dormantResumes: within(state.dormantResumes, now, windowMs),
188 ...(fiveHour ? { limit: fiveHour } : {}),
189 }
190}
191
192/** Sums [time, tokens] pairs inside the window per minute, so a record stays small however busy its session. */
193function perMinute(pairs: ReadonlyArray<readonly [number, number]>, now: number, windowMs: number): Array<[number, number]> {
194 const minutes = new Map<number, [number, number]>()
195 for (const [t, n] of pairs) {
196 if (now - t >= windowMs) continue
197 const key = Math.floor(t / 60_000)
198 const minute = minutes.get(key)
199 if (minute) {
200 minute[0] = Math.max(minute[0], t)
201 minute[1] += n
202 } else {
203 minutes.set(key, [t, n])
204 }
205 }
206 return [...minutes.values()]
207}
208
209/**
210 * Accepts only records that look like ours and are fresh: a session that has
211 * made no model request for `ttlMs` is not spending, so its agents stop
212 * counting even if it crashed without clearing its record.
213 */
214export function livePeers(records: readonly unknown[], now: number, ttlMs: number): PeerRecord[] {
215 const peers: PeerRecord[] = []
216 for (const record of records) {
217 if (!isPeerRecord(record)) continue
218 if (now - record.at > ttlMs) continue
219 peers.push(record)
220 }
221 return peers
222}
223
224function isPeerRecord(value: unknown): value is PeerRecord {
225 if (typeof value !== 'object' || value === null) return false
226 const r = value as Record<string, unknown>
227 return (
228 r.v === 2 &&
229 typeof r.at === 'number' &&
230 typeof r.running === 'number' &&
231 Array.isArray(r.starts) &&
232 Array.isArray(r.agentTokens) &&
233 r.agentTokens.every((p) => Array.isArray(p) && typeof p[0] === 'number' && typeof p[1] === 'number') &&
234 Array.isArray(r.dormantResumes)
235 )
236}
237core/status.ts 98 lines1// What the user sees without asking: one status line while something needs
2// attention, and the `/agent-guard` command's text.
3
4import { highestLimit } from './agents.ts'
5import type { Config } from './config.ts'
6import { MAIN, type PeerRecord, type Scope, type SessionState, countOnMachine, runningAgents } from './state.ts'
7import { count, duration, shortClock, tokens, windowName } from './text.ts'
8
9/** The line under the prompt, or undefined when there is nothing to say. */
10export function statusLine(state: SessionState, peers: readonly PeerRecord[], config: Config, now: number): string | undefined {
11 if (!config.enabled) return undefined
12 const parts = liftedOrPaused(state, now)
13 const limit = highestLimit(state, peers, now)
14 if (limit && limit.percent >= config.limitDenyPercent) {
15 parts.push(`${windowName(limit.kind)} window ${limit.percent}%: new agents refused`)
16 } else if (limit && limit.percent >= config.limitAskPercent) {
17 parts.push(`${windowName(limit.kind)} window ${limit.percent}%: new agents need approval`)
18 }
19 const context = state.loops[MAIN]?.context
20 if (context !== undefined && context >= config.contextHard) {
21 parts.push(`context ${tokens(context)}: /compact recommended`)
22 } else if (context !== undefined && context >= config.contextWarn) {
23 parts.push(`context ${tokens(context)}`)
24 }
25 return parts.length > 0 ? parts.join(' · ') : undefined
26}
27
28/** What is lifted or paused now, worded alike in the status line and the report. */
29function liftedOrPaused(state: SessionState, now: number): string[] {
30 const parts: string[] = []
31 if (state.override && now < state.override.until) {
32 parts.push(`${scopeName(state.override.scope)} lifted until ${shortClock(state.override.until)}`)
33 }
34 if (state.fuseUntil !== undefined && now < state.fuseUntil) {
35 parts.push(`agent spawns paused until ${shortClock(state.fuseUntil)}`)
36 }
37 return parts
38}
39
40export function scopeName(scope: Scope): string {
41 if (scope === 'agents') return 'agent limits'
42 if (scope === 'context') return 'context limits'
43 return 'all limits'
44}
45
46/** The text `/agent-guard` and `/agent-guard status` print. */
47export function statusReport(
48 state: SessionState,
49 peers: readonly PeerRecord[],
50 config: Config,
51 now: number,
52 problems: readonly string[],
53): string {
54 if (!config.enabled) return 'Off: AGENT_GUARD is set to turn it off.'
55 const window = config.windowMs
56 const lines = ['On.']
57
58 const peerRunning = peers.reduce((sum, p) => sum + p.running, 0)
59 const starts = countOnMachine(state, peers, 'starts', now, window)
60 lines.push(
61 `Agents: ${runningAgents(state)} running here, ${peerRunning} in ${count(peers.length, 'other session')} (limit ${config.agentMax}); ${starts} started in the last ${duration(window)} (limit ${config.rollingMax}).`,
62 )
63
64 const limits = Object.values(state.limits).filter((r) => r.resetsAt === undefined || r.resetsAt > now)
65 if (limits.length > 0) {
66 const text = limits
67 .map((r) => `${windowName(r.kind)} ${r.percent}%${r.resetsAt === undefined ? '' : ` (resets ${shortClock(r.resetsAt)})`}`)
68 .join(', ')
69 lines.push(`Plan windows: ${text}. New agents ask from ${config.limitAskPercent}% and are refused from ${config.limitDenyPercent}%.`)
70 } else {
71 lines.push('Plan windows: no reading yet (one arrives with the next response on a subscription).')
72 }
73
74 const context = state.loops[MAIN]?.context
75 lines.push(
76 `Context: ${context === undefined ? 'unknown until the next response' : `${tokens(context)} tokens`} (warn ${tokens(config.contextWarn)}, ask before a prompt from ${tokens(config.contextHard)}).`,
77 )
78
79 const held = liftedOrPaused(state, now)
80 for (const [key, until] of Object.entries(state.leases)) {
81 if (now < until) held.push(`${key} approved${until < Number.MAX_SAFE_INTEGER ? ` until ${shortClock(until)}` : ' until the context shrinks'}`)
82 }
83 for (const [key, until] of Object.entries(state.refusals)) if (now < until) held.push(`${key} declined until ${shortClock(until)}`)
84 if (held.length > 0) lines.push(`Now: ${held.join('; ')}.`)
85 for (const problem of problems) lines.push(`Config: ${problem}.`)
86 lines.push('', USAGE)
87 return lines.join('\n')
88}
89
90export const USAGE = [
91 'Commands:',
92 ' /agent-guard allow [agents|context] [minutes] lift the guard\'s limits for this session (default: all, 10 min)',
93 ' /agent-guard pause [minutes] same as allow, for every limit',
94 ' /agent-guard resume end an allow or pause early, and clear declined questions',
95 ' /agent-guard report [days] what the guard did on this machine (default 7 days)',
96 ' /agent-guard status this summary',
97].join('\n')
98core/text.ts 94 lines1// Formatting shared by every message the guard writes.
2
3/** `1 agent`, `4 agents`. */
4export function count(n: number, noun: string): string {
5 return `${n} ${noun}${n === 1 ? '' : 's'}`
6}
7
8/** Lists details as prose: "a", "a and b", "a, b and c". */
9export function join(details: readonly string[]): string {
10 if (details.length <= 1) return details.join('')
11 return `${details.slice(0, -1).join(', ')} and ${details.at(-1)}`
12}
13
14/** Capitalises the first letter of a sentence assembled from parts. */
15export function sentence(text: string): string {
16 return text.charAt(0).toUpperCase() + text.slice(1)
17}
18
19export function tokens(n: number): string {
20 if (n >= 1_000_000) return `${trim(n / 1_000_000)}M`
21 if (n >= 1000) return `${Math.round(n / 1000)}k`
22 return String(n)
23}
24
25function trim(value: number): string {
26 return value >= 100 ? String(Math.round(value)) : value.toFixed(1).replace(/\.0$/, '')
27}
28
29/** Local wall-clock time, `14:32:07`. */
30export function clock(ms: number): string {
31 const d = new Date(ms)
32 return [d.getHours(), d.getMinutes(), d.getSeconds()].map((n) => String(n).padStart(2, '0')).join(':')
33}
34
35/** Local time without seconds, `14:32`. */
36export function shortClock(ms: number): string {
37 return clock(ms).slice(0, 5)
38}
39
40/** `45 s`, `12 min`, `2 h 5 min`. */
41export function duration(ms: number): string {
42 const seconds = Math.round(ms / 1000)
43 if (seconds < 60) return `${seconds} s`
44 const minutes = Math.round(seconds / 60)
45 if (minutes < 60) return `${minutes} min`
46 const hours = Math.floor(minutes / 60)
47 const rest = minutes % 60
48 return rest ? `${hours} h ${rest} min` : `${hours} h`
49}
50
51const WINDOW_NAMES: Record<string, string> = {
52 five_hour: '5-hour',
53 seven_day: 'weekly',
54 spend_limit: 'spend',
55}
56
57export function windowName(kind: string): string {
58 return WINDOW_NAMES[kind] ?? kind.replaceAll('_', ' ')
59}
60
61/** Coarse context bucket for the journal, never an exact figure. */
62export function contextBucket(n: number | undefined): string | undefined {
63 if (n === undefined) return undefined
64 if (n < 150_000) return '<150k'
65 if (n < 300_000) return '150-300k'
66 if (n < 400_000) return '300-400k'
67 if (n < 500_000) return '400-500k'
68 return '>=500k'
69}
70
71/** FNV-1a over the text, as 13 base-36 digits. Identifies a call without storing it. */
72export function fingerprint(text: string): string {
73 let hash = 0xcbf29ce484222325n
74 const prime = 0x100000001b3n
75 const mask = 0xffffffffffffffffn
76 for (let i = 0; i < text.length; i++) {
77 hash ^= BigInt(text.charCodeAt(i))
78 hash = (hash * prime) & mask
79 }
80 return hash.toString(36).padStart(13, '0')
81}
82
83/** JSON with sorted keys, so equal inputs give equal fingerprints. */
84export function canonical(value: unknown): string {
85 if (Array.isArray(value)) return `[${value.map(canonical).join(',')}]`
86 if (value && typeof value === 'object') {
87 const entries = Object.entries(value as Record<string, unknown>)
88 .filter(([, v]) => v !== undefined)
89 .sort(([a], [b]) => (a < b ? -1 : a > b ? 1 : 0))
90 return `{${entries.map(([k, v]) => `${JSON.stringify(k)}:${canonical(v)}`).join(',')}}`
91 }
92 return JSON.stringify(value) ?? 'null'
93}
94