SLOPSHOPPER

subagent-limits

Keeps subagents small: tells the main agent how to size and split work, asks a growing subagent to checkpoint, and refuses its tools past a hard limit so a…

newguardtoastpromptagents
A shopper browsing a rack in a slop shop
README

subagent-limits

Keeps subagents small. The main agent still decides how work is split. This mod tells it how to split well, watches each subagent's context, and steps in when one grows too large.

  • It adds a short section to the main agent's system prompt with the rules for sizing and splitting work.
  • It adds a context budget to every subagent's brief: read files by line range, keep output short, and write a handoff note when asked to checkpoint.
  • When a subagent's requests carry nudgeAtK thousand tokens of context, it asks the agent to finish its current item and checkpoint.
  • Past stopAtK, the agent gets 8 more tool calls to commit and write its note, and then every tool call is refused.
  • When a subagent waits over 4 minutes on one of its own tool calls and its next request rebuilds most of a large context (100K tokens or more), its prompt cache expired. The mod shows you a toast and tells that agent once, with the measured size and pause, so it checks on long jobs sooner. Agents whose caches never expire never see it.
  • It ships the subagent-limits:split-work skill, with templates for briefs and handoff notes.
SettingDefaultWhat it does
nudgeAtK300Asks a subagent to checkpoint when its requests carry this many thousand tokens of context.
stopAtK450Gives a subagent 8 tool calls to checkpoint past this many thousand tokens, then refuses its tools. Always at least 50 above nudgeAtK.
cacheNotestrueTells a subagent once when its prompt cache expired during a long pause. Off: only you get the toast.

The main README covers installing, why the limits sit where they do, and what it can't see.

Check the limits against your own agents

The defaults come from one person's agents. scripts/ctx_study.py reads your subagent transcripts and tells you whether 300K and 450K suit how you work, or which values would cost less.

It reads ~/.claude/projects/<project>/<session>/subagents/agent-*.jsonl and never writes there. Forks (isFork in the .meta.json) are left out of every number. It measures:

  • context at each agent's first request, at its first Edit/Write (R_edit), and at its peak
  • how many agents pass each limit, and what share of weighted cost (input + 1.25 cache_write + 0.1 cache_read + 5 output) went to requests above it
  • cache-expiry rewrites: requests after an agent's first whose cache write is at least half its context
  • this mod's activity: briefs carrying the contract, nudges, wind-down notes, refusals, final reports starting with CHECKPOINT:, and predecessor-to-successor pairs
  • the re-orientation fraction: a successor's growth before its first edit, divided by the median for other agents
  • a what-if simulation of the cost change at limits of 200K to 450K, and a verdict on the current limits

Run it

Python 3.8 or newer, standard library only. Installed from the marketplace, the script is at ~/.claude/plugins/cache/claude-ops/subagent-limits/<version>/scripts/ctx_study.py. In a clone of the repo it's at plugins/subagent-limits/scripts/ctx_study.py.

py -3 <path>/ctx_study.py              # Windows
python3 <path>/ctx_study.py            # macOS, Linux

Options: --since DAYS (default 7, 0 for all time), --limits 300,450 (in K tokens; pass your own if you changed nudgeAtK or stopAtK), --projects DIR (default ~/.claude/projects), --out DIR (default ~/.claude/ctx-study).

To run it every Monday at 09:00, point the schedule at a clone. The <version> folder in the plugin cache changes on every update, which would break the schedule.

schtasks /Create /TN ctx-study /SC WEEKLY /D MON /ST 09:00 /TR "C:\Windows\pyw.exe -3 C:\path\to\claude-ops\plugins\subagent-limits\scripts\ctx_study.py"
0 9 * * 1  python3 /path/to/claude-ops/plugins/subagent-limits/scripts/ctx_study.py >/dev/null   # crontab

Output

Everything goes to --out:

  • report-YYYY-MM-DD.md: the report. A run with a window other than 7 days adds a suffix (-all, -30d), so it doesn't overwrite the weekly one.
  • agents-YYYY-MM-DD.csv: one row per agent, forks included and flagged.
  • history.csv: one line per run, appended, to show trends across weeks. Filter by window_days when comparing.

The reports contain project names and agent descriptions, so they stay in your home directory. Don't commit them anywhere.

How the mod's activity is detected

The marker strings come from hooks/register.ts. If they change there, update them in the script. Before 0.3.0 this mod was called pit-stop and tagged its notes [pit-stop]; both tags are matched, so older transcripts still count.

  • Contract: the agent's brief contains [subagent-limits] Context budget.
  • Nudge, wind-down and refusal: the note text with a concrete token count, found in anything except the agent's own output. A note whose count is more than 10% above the agent's own peak is treated as a quotation and ignored. That stops an agent that read the mod's source or tests from being counted.
  • CHECKPOINT: the agent's last text message starts with CHECKPOINT:.
  • Successor: a predecessor is an agent that ended with CHECKPOINT or was refused. Its successor is the first non-fork agent in the same parent session that started after the predecessor's last event and whose brief (minus the appended contract) either names the predecessor's agent id, names a note file from its report (a path containing handoff, checkpoint, note or relay), or shares at least three 8-word runs with the report. Each agent can succeed only one predecessor.

Limitations

  • The simulation uses 0.6 as the re-orientation fraction until at least 3 measured pairs exist. Its cost model ignores output tokens.
  • The verdict flags the limits when the cheapest simulated limit is more than 50K from the nudge limit. It also prints the rule of thumb (2 x median R_edit), which ignores the cost of handing off and so usually comes out lower.
  • When CLAUDE_CODE_SESSION_ID is set (the script is run from inside Claude Code), that session's transcripts are skipped. A scheduled run doesn't skip anything, so it includes partly written transcripts from any session still running at the time.
  • Only edits made through Edit, Write, MultiEdit and NotebookEdit count as an agent's first edit. Edits made through Bash don't.
Source 1 files
hooks/register.ts 269 lines
1import type { EngineInterface, Register, ToolCallResult, TurnUsage } from 'claude-code'
2
3// what the mod has seen of one subagent or in-process teammate
4export type Budget = {
5  // tokens its latest request carried: input plus cache reads and writes
6  context: number
7  // it has been asked to checkpoint
8  isNudged: boolean
9  // tool calls it has made since passing the stop limit
10  windDownCalls: number
11  // its tool calls are being refused
12  isRefused: boolean
13  // its previous request, to tell a cache that expired from a first request, a model switch or a compaction
14  last?: Step
15  // a cache drop it has not yet been told about
16  pendingDrop?: Drop
17  // it has been told about a cache drop, and the person has been shown one: each happens once per agent
18  isCacheNoted: boolean
19  isDropToasted: boolean
20}
21
22// what the mod keeps of one request: its turn, when its response arrived, the model that answered, how many
23// messages it carried, and whether the response called tools (so the pause after it was the agent's own wait)
24export type Step = { turnId: string; endedAt: number; model: string; messageCount: number; calledTools: boolean }
25
26// the request being made now, as far as the drop rule reads it
27export type Request = { turnId: string; startedAt: number; messageCount: number }
28
29// a request that re-wrote most of a large context to the prompt cache after a long pause
30export type Drop = { rebuilt: number; pauseMs: number }
31
32export type Limits = { nudge: number; stop: number }
33
34// what an agent's next tool call gets
35export type Verdict =
36  | { kind: 'run' }
37  | { kind: 'nudge' }
38  | { kind: 'wind-down'; left: number }
39  | { kind: 'refuse' }
40
41// tool calls an agent keeps after passing the stop limit, to commit and write its handoff note
42export const WIND_DOWN_CALLS = 8
43// marks text this mod wrote
44export const MARK = '[subagent-limits]'
45// a brief that already carries the contract (a successor's, copied from its predecessor) gets no second copy;
46// one that merely quotes a note still gets it
47export const CONTRACT_HEAD = `${MARK} Context budget.`
48
49export const CONTRACT = `${CONTRACT_HEAD} Every request you make re-sends your whole context, so keep it small.
50- Read files by line range (grep -n first, then read only the lines you need), not whole.
51- Keep command and test output short. Send long output to a file and read only the summary lines.
52- Stay inside the scope above. If the job is bigger than the brief suggests, say so in your report rather than expanding it.
53- Run the narrowest command that proves your change, and don't wait in sleep loops: a long pause between your requests lets the prompt cache expire, and your next request pays to rebuild your whole context.
54If you are asked to checkpoint, finish or back out the change in progress, commit if you work in git, and write a handoff note. The note says what is done and verified, what is left as a numbered list naming files and functions, any traps you found, and exactly which files and line ranges the next agent should read first. Put the note where your brief says, or in your final report if the brief names no place. Start your final report with "CHECKPOINT:" so whoever started you knows the work is unfinished.`
55
56// 955492 -> 955K, 1200000 -> 1.2M
57export function tokens(n: number): string {
58  return n >= 1_000_000 ? `${(n / 1_000_000).toFixed(1)}M` : `${Math.round(n / 1000)}K`
59}
60
61// a configured number, or the default when unset or not a number, held within bounds
62export function setting(value: unknown, fallback: number, min: number, max: number): number {
63  const n = Number(value ?? fallback)
64  return Number.isFinite(n) ? Math.min(max, Math.max(min, n)) : fallback
65}
66
67// the stop limit always leaves room above the nudge to finish an item and checkpoint
68export function limitsFrom(options: { nudgeAtK?: unknown; stopAtK?: unknown }): Limits {
69  const nudge = setting(options.nudgeAtK, 300, 50, 10_000) * 1000
70  const stop = Math.max(setting(options.stopAtK, 450, 50, 10_000) * 1000, nudge + 50_000)
71  return { nudge, stop }
72}
73
74export function fresh(): Budget {
75  return { context: 0, isNudged: false, windDownCalls: 0, isRefused: false, isCacheNoted: false, isDropToasted: false }
76}
77
78// A cache drop worth mentioning. A subagent's prompt cache lives 5 minutes unless the person set it longer, so a
79// pause under 4 minutes (measured from the previous response, a little after the cache was last used) cannot have
80// expired it; a rebuild under 100K costs too little to talk about; and a request that wrote less than half its
81// context had most of it served from the cache.
82export const DROP = { minContext: 100_000, minShare: 0.5, minPauseMs: 4 * 60_000 }
83
84// The drop this request shows, if any. Only a pause the agent spent waiting on its own tool call, in the same run,
85// counts: an agent resumed after it finished did not choose its pause. Not an agent's first request, not one
86// after a model switch (a new model has no cache), and not one after a compaction (fewer messages, a new prefix).
87export function cacheDrop(last: Step | undefined, request: Request, usage: TurnUsage): Drop | undefined {
88  if (last === undefined || !last.calledTools || last.turnId !== request.turnId) return undefined
89  if (usage.model !== last.model || request.messageCount < last.messageCount) return undefined
90  const context = usage.input_tokens + usage.cache_read_input_tokens + usage.cache_creation_input_tokens
91  const rebuilt = usage.cache_creation_input_tokens
92  const pauseMs = request.startedAt - last.endedAt
93  if (context < DROP.minContext || rebuilt < context * DROP.minShare || pauseMs <= DROP.minPauseMs) return undefined
94  return { rebuilt, pauseMs }
95}
96
97export function minutes(ms: number): number {
98  return Math.round(ms / 60_000)
99}
100
101export function cacheNote(drop: Drop): string {
102  return `${MARK} Your last request rebuilt ${tokens(drop.rebuilt)} tokens of context because the prompt cache expired during a ${minutes(drop.pauseMs)}-minute pause. If you need to wait on a long job again, check on it before it runs that long, or run something shorter.`
103}
104
105export function dropToast(name: string, drop: Drop): string {
106  return `${name} rebuilt ${tokens(drop.rebuilt)} of context: its prompt cache expired during a ${minutes(drop.pauseMs)}-minute pause`
107}
108
109// past the stop limit an agent keeps a few calls to checkpoint, then everything is refused;
110// below it, the first call past the nudge limit carries the request to checkpoint
111export function verdict(budget: Budget, limits: Limits): Verdict {
112  if (budget.context >= limits.stop) {
113    const left = WIND_DOWN_CALLS - budget.windDownCalls
114    return left > 0 ? { kind: 'wind-down', left: left - 1 } : { kind: 'refuse' }
115  }
116  if (budget.context >= limits.nudge && !budget.isNudged) return { kind: 'nudge' }
117  return { kind: 'run' }
118}
119
120export function nudgeNote(context: number, limits: Limits): string {
121  return `${MARK} Your context has reached ${tokens(context)} tokens. Finish the item you are on, or back it out if it is far from done, then checkpoint: commit, write your handoff note, and end with a final report that starts with "CHECKPOINT:". A fresh agent will continue from your note. At ${tokens(limits.stop)} your tools start being refused.`
122}
123
124export function windDownNote(context: number, left: number, limits: Limits): string {
125  const budget =
126    left > 0 ? `You have ${left} tool calls left before every tool is refused.` : 'That was your last tool call.'
127  return `${MARK} Your context is ${tokens(context)} tokens, past the ${tokens(limits.stop)} limit. ${budget} Use what is left only to commit and write your handoff note, then end with a final report that starts with "CHECKPOINT:".`
128}
129
130export function refusal(context: number, limits: Limits): string {
131  return `${MARK} Refused: your context (${tokens(context)} tokens) is past the ${tokens(limits.stop)} limit. End now with a final report that starts with "CHECKPOINT:" and says what is done and what is left. Resuming this agent will not help; a fresh agent should continue from your report.`
132}
133
134// the main agent decides the split; this tells it what the mod enforces and how to work with it
135export function planningRules(limits: Limits): string {
136  return `# Sizing delegated work
137The subagent-limits plugin is installed. You decide how work is split; the plugin only measures each subagent's context and enforces two limits. When you hand work to subagents:
138- Give each agent one phase it can finish well under ${tokens(limits.nudge)} tokens of context. At ${tokens(limits.nudge)} the plugin asks the agent to checkpoint, and at ${tokens(limits.stop)} it starts refusing the agent's tools.
139- Run agents in parallel only when they edit different files. When they would share files, run a relay instead: one fresh agent per phase, each starting from the previous one's handoff note.
140- Fresh agents often read 100K tokens or more before their first edit. Cut that with a brief that names the exact files, functions and line ranges to read, the test command, and what done means.
141- Set \`model\` on every Agent call.
142- A final report that starts with "CHECKPOINT:" means the agent stopped on purpose with work left. Start a fresh agent from its handoff note rather than resuming the old one.
143For brief and handoff templates, load the subagent-limits:split-work skill.`
144}
145
146export function withContract(prompt: string): string {
147  return prompt.includes(CONTRACT_HEAD) ? prompt : `${prompt}\n\n${CONTRACT}`
148}
149
150export function agentLabel(id: string, description: string | undefined): string {
151  return description !== undefined && description !== '' ? description : `Agent ${id.slice(0, 8)}`
152}
153
154// the call's result with a note the model reads after it; a refusal stays as it is
155function withNote(result: ToolCallResult, note: string): ToolCallResult {
156  if (result.deny !== undefined) return result
157  return { ...result, context: [...(result.context ?? []), note] }
158}
159
160async function label($: EngineInterface, id: string): Promise<string> {
161  const agent = (await $.agent.list()).find(a => a.id === id)
162  return agentLabel(id, agent?.description)
163}
164
165export const register: Register = (on, options) => {
166  const limits = limitsFrom(options)
167  // off: an agent whose cache expired is not told; the person still gets the toast
168  const cacheNotes = options.cacheNotes !== false
169  // module state: a reload starts these over, so a nudge may repeat once after one
170  const budgets = new Map<string, Budget>()
171  // forks inherit the parent's context and its prompt cache: never limited
172  const forks = new Set<string>()
173
174  on('session.end', { reason: 'clear' }, async ($, e, next) => {
175    budgets.clear()
176    forks.clear()
177    return next(e)
178  })
179
180  // only a model that can start agents needs the planning rules
181  on('prompt.compose', async ($, e, next) => {
182    const composed = await next(e)
183    if (!e.tools.includes('Agent') || e.traits.includes('bare')) return composed
184    const rules = { id: 'subagent-limits:planning', text: planningRules(limits), scope: 'session' } as const
185    return { sections: [...composed.sections, rules] }
186  })
187
188  on('agent.spawn', async ($, e, next) => {
189    if (e.fork) {
190      const started = await next(e)
191      if (started.agentId !== undefined) forks.add(started.agentId)
192      return started
193    }
194    return next({ ...e, prompt: withContract(e.prompt) })
195  })
196
197  // every model request of a subagent or an in-process teammate; the main thread is never limited
198  on('turn.step', async function* ($, e, next) {
199    const id = e.agentId
200    if (id === undefined || forks.has(id)) return yield* next(e)
201    const startedAt = await $.clock.now()
202    const result = yield* next(e)
203    if (result.usage === null) return result
204    const usage = result.usage
205    const context = usage.input_tokens + usage.cache_read_input_tokens + usage.cache_creation_input_tokens
206    const budget = budgets.get(id) ?? fresh()
207    budgets.set(id, budget)
208    const drop = cacheDrop(budget.last, { turnId: e.turnId, startedAt, messageCount: e.messageCount }, usage)
209    budget.context = context
210    budget.last = {
211      turnId: e.turnId,
212      endedAt: await $.clock.now(),
213      model: usage.model,
214      messageCount: e.messageCount,
215      calledTools: result.toolUses.length > 0,
216    }
217    if (drop !== undefined) {
218      if (cacheNotes && !budget.isCacheNoted) budget.pendingDrop = drop
219      if (!budget.isDropToasted) {
220        budget.isDropToasted = true
221        $.ui.toast(dropToast(await label($, id), drop))
222      }
223    }
224    return result
225  })
226
227  on('tool.call', async ($, e, next) => {
228    const id = e.agentId
229    const budget = id === undefined ? undefined : budgets.get(id)
230    if (id === undefined || budget === undefined) return next(e)
231    // taken before any await, so of calls made in parallel only the first carries it
232    const drop = budget.pendingDrop
233    if (drop !== undefined) {
234      budget.pendingDrop = undefined
235      budget.isCacheNoted = true
236    }
237    // the tool's result, with the cache note when one is due
238    const ran = async () => {
239      const result = await next(e)
240      return drop === undefined ? result : withNote(result, cacheNote(drop))
241    }
242    const decided = verdict(budget, limits)
243    switch (decided.kind) {
244      case 'run':
245        return ran()
246      case 'nudge': {
247        budget.isNudged = true
248        $.ui.toast(`Asked ${await label($, id)} to checkpoint at ${tokens(budget.context)} of context`)
249        return withNote(await ran(), nudgeNote(budget.context, limits))
250      }
251      case 'wind-down': {
252        // counted before any await, so calls made in parallel each see the ones before them
253        budget.windDownCalls += 1
254        if (budget.windDownCalls === 1) {
255          $.ui.toast(`${await label($, id)} passed ${tokens(limits.stop)}: ${WIND_DOWN_CALLS} tool calls left to checkpoint`)
256        }
257        return withNote(await ran(), windDownNote(budget.context, decided.left, limits))
258      }
259      case 'refuse': {
260        if (!budget.isRefused) {
261          budget.isRefused = true
262          $.ui.toast(`Refusing tools for ${await label($, id)} at ${tokens(budget.context)} of context`)
263        }
264        return { deny: refusal(budget.context, limits) }
265      }
266    }
267  })
268}
269