Keeps subagents small: tells the main agent how to size and split work, asks a growing subagent to checkpoint, and refuses its tools past a hard limit so a…

Keeps subagents small. The main agent still decides how work is split. This mod tells it how to split well, watches each subagent's context, and steps in when one grows too large.
nudgeAtK thousand tokens of context, it asks the agent to finish its current item and checkpoint.stopAtK, the agent gets 8 more tool calls to commit and write its note, and then every tool call is refused.subagent-limits:split-work skill, with templates for briefs and handoff notes.| Setting | Default | What it does |
|---|---|---|
nudgeAtK | 300 | Asks a subagent to checkpoint when its requests carry this many thousand tokens of context. |
stopAtK | 450 | Gives a subagent 8 tool calls to checkpoint past this many thousand tokens, then refuses its tools. Always at least 50 above nudgeAtK. |
cacheNotes | true | Tells a subagent once when its prompt cache expired during a long pause. Off: only you get the toast. |
The main README covers installing, why the limits sit where they do, and what it can't see.
The defaults come from one person's agents. scripts/ctx_study.py reads your subagent transcripts and tells you whether 300K and 450K suit how you work, or which values would cost less.
It reads ~/.claude/projects/<project>/<session>/subagents/agent-*.jsonl and never writes there. Forks (isFork in the .meta.json) are left out of every number. It measures:
input + 1.25 cache_write + 0.1 cache_read + 5 output) went to requests above itCHECKPOINT:, and predecessor-to-successor pairsPython 3.8 or newer, standard library only. Installed from the marketplace, the script is at ~/.claude/plugins/cache/claude-ops/subagent-limits/<version>/scripts/ctx_study.py. In a clone of the repo it's at plugins/subagent-limits/scripts/ctx_study.py.
py -3 <path>/ctx_study.py # Windows
python3 <path>/ctx_study.py # macOS, Linux
Options: --since DAYS (default 7, 0 for all time), --limits 300,450 (in K tokens; pass your own if you changed nudgeAtK or stopAtK), --projects DIR (default ~/.claude/projects), --out DIR (default ~/.claude/ctx-study).
To run it every Monday at 09:00, point the schedule at a clone. The <version> folder in the plugin cache changes on every update, which would break the schedule.
schtasks /Create /TN ctx-study /SC WEEKLY /D MON /ST 09:00 /TR "C:\Windows\pyw.exe -3 C:\path\to\claude-ops\plugins\subagent-limits\scripts\ctx_study.py"
0 9 * * 1 python3 /path/to/claude-ops/plugins/subagent-limits/scripts/ctx_study.py >/dev/null # crontab
Everything goes to --out:
report-YYYY-MM-DD.md: the report. A run with a window other than 7 days adds a suffix (-all, -30d), so it doesn't overwrite the weekly one.agents-YYYY-MM-DD.csv: one row per agent, forks included and flagged.history.csv: one line per run, appended, to show trends across weeks. Filter by window_days when comparing.The reports contain project names and agent descriptions, so they stay in your home directory. Don't commit them anywhere.
The marker strings come from hooks/register.ts. If they change there, update them in the script. Before 0.3.0 this mod was called pit-stop and tagged its notes [pit-stop]; both tags are matched, so older transcripts still count.
[subagent-limits] Context budget.CHECKPOINT:.CLAUDE_CODE_SESSION_ID is set (the script is run from inside Claude Code), that session's transcripts are skipped. A scheduled run doesn't skip anything, so it includes partly written transcripts from any session still running at the time.hooks/register.ts 269 lines1import type { EngineInterface, Register, ToolCallResult, TurnUsage } from 'claude-code'
2
3// what the mod has seen of one subagent or in-process teammate
4export type Budget = {
5 // tokens its latest request carried: input plus cache reads and writes
6 context: number
7 // it has been asked to checkpoint
8 isNudged: boolean
9 // tool calls it has made since passing the stop limit
10 windDownCalls: number
11 // its tool calls are being refused
12 isRefused: boolean
13 // its previous request, to tell a cache that expired from a first request, a model switch or a compaction
14 last?: Step
15 // a cache drop it has not yet been told about
16 pendingDrop?: Drop
17 // it has been told about a cache drop, and the person has been shown one: each happens once per agent
18 isCacheNoted: boolean
19 isDropToasted: boolean
20}
21
22// what the mod keeps of one request: its turn, when its response arrived, the model that answered, how many
23// messages it carried, and whether the response called tools (so the pause after it was the agent's own wait)
24export type Step = { turnId: string; endedAt: number; model: string; messageCount: number; calledTools: boolean }
25
26// the request being made now, as far as the drop rule reads it
27export type Request = { turnId: string; startedAt: number; messageCount: number }
28
29// a request that re-wrote most of a large context to the prompt cache after a long pause
30export type Drop = { rebuilt: number; pauseMs: number }
31
32export type Limits = { nudge: number; stop: number }
33
34// what an agent's next tool call gets
35export type Verdict =
36 | { kind: 'run' }
37 | { kind: 'nudge' }
38 | { kind: 'wind-down'; left: number }
39 | { kind: 'refuse' }
40
41// tool calls an agent keeps after passing the stop limit, to commit and write its handoff note
42export const WIND_DOWN_CALLS = 8
43// marks text this mod wrote
44export const MARK = '[subagent-limits]'
45// a brief that already carries the contract (a successor's, copied from its predecessor) gets no second copy;
46// one that merely quotes a note still gets it
47export const CONTRACT_HEAD = `${MARK} Context budget.`
48
49export const CONTRACT = `${CONTRACT_HEAD} Every request you make re-sends your whole context, so keep it small.
50- Read files by line range (grep -n first, then read only the lines you need), not whole.
51- Keep command and test output short. Send long output to a file and read only the summary lines.
52- Stay inside the scope above. If the job is bigger than the brief suggests, say so in your report rather than expanding it.
53- Run the narrowest command that proves your change, and don't wait in sleep loops: a long pause between your requests lets the prompt cache expire, and your next request pays to rebuild your whole context.
54If you are asked to checkpoint, finish or back out the change in progress, commit if you work in git, and write a handoff note. The note says what is done and verified, what is left as a numbered list naming files and functions, any traps you found, and exactly which files and line ranges the next agent should read first. Put the note where your brief says, or in your final report if the brief names no place. Start your final report with "CHECKPOINT:" so whoever started you knows the work is unfinished.`
55
56// 955492 -> 955K, 1200000 -> 1.2M
57export function tokens(n: number): string {
58 return n >= 1_000_000 ? `${(n / 1_000_000).toFixed(1)}M` : `${Math.round(n / 1000)}K`
59}
60
61// a configured number, or the default when unset or not a number, held within bounds
62export function setting(value: unknown, fallback: number, min: number, max: number): number {
63 const n = Number(value ?? fallback)
64 return Number.isFinite(n) ? Math.min(max, Math.max(min, n)) : fallback
65}
66
67// the stop limit always leaves room above the nudge to finish an item and checkpoint
68export function limitsFrom(options: { nudgeAtK?: unknown; stopAtK?: unknown }): Limits {
69 const nudge = setting(options.nudgeAtK, 300, 50, 10_000) * 1000
70 const stop = Math.max(setting(options.stopAtK, 450, 50, 10_000) * 1000, nudge + 50_000)
71 return { nudge, stop }
72}
73
74export function fresh(): Budget {
75 return { context: 0, isNudged: false, windDownCalls: 0, isRefused: false, isCacheNoted: false, isDropToasted: false }
76}
77
78// A cache drop worth mentioning. A subagent's prompt cache lives 5 minutes unless the person set it longer, so a
79// pause under 4 minutes (measured from the previous response, a little after the cache was last used) cannot have
80// expired it; a rebuild under 100K costs too little to talk about; and a request that wrote less than half its
81// context had most of it served from the cache.
82export const DROP = { minContext: 100_000, minShare: 0.5, minPauseMs: 4 * 60_000 }
83
84// The drop this request shows, if any. Only a pause the agent spent waiting on its own tool call, in the same run,
85// counts: an agent resumed after it finished did not choose its pause. Not an agent's first request, not one
86// after a model switch (a new model has no cache), and not one after a compaction (fewer messages, a new prefix).
87export function cacheDrop(last: Step | undefined, request: Request, usage: TurnUsage): Drop | undefined {
88 if (last === undefined || !last.calledTools || last.turnId !== request.turnId) return undefined
89 if (usage.model !== last.model || request.messageCount < last.messageCount) return undefined
90 const context = usage.input_tokens + usage.cache_read_input_tokens + usage.cache_creation_input_tokens
91 const rebuilt = usage.cache_creation_input_tokens
92 const pauseMs = request.startedAt - last.endedAt
93 if (context < DROP.minContext || rebuilt < context * DROP.minShare || pauseMs <= DROP.minPauseMs) return undefined
94 return { rebuilt, pauseMs }
95}
96
97export function minutes(ms: number): number {
98 return Math.round(ms / 60_000)
99}
100
101export function cacheNote(drop: Drop): string {
102 return `${MARK} Your last request rebuilt ${tokens(drop.rebuilt)} tokens of context because the prompt cache expired during a ${minutes(drop.pauseMs)}-minute pause. If you need to wait on a long job again, check on it before it runs that long, or run something shorter.`
103}
104
105export function dropToast(name: string, drop: Drop): string {
106 return `${name} rebuilt ${tokens(drop.rebuilt)} of context: its prompt cache expired during a ${minutes(drop.pauseMs)}-minute pause`
107}
108
109// past the stop limit an agent keeps a few calls to checkpoint, then everything is refused;
110// below it, the first call past the nudge limit carries the request to checkpoint
111export function verdict(budget: Budget, limits: Limits): Verdict {
112 if (budget.context >= limits.stop) {
113 const left = WIND_DOWN_CALLS - budget.windDownCalls
114 return left > 0 ? { kind: 'wind-down', left: left - 1 } : { kind: 'refuse' }
115 }
116 if (budget.context >= limits.nudge && !budget.isNudged) return { kind: 'nudge' }
117 return { kind: 'run' }
118}
119
120export function nudgeNote(context: number, limits: Limits): string {
121 return `${MARK} Your context has reached ${tokens(context)} tokens. Finish the item you are on, or back it out if it is far from done, then checkpoint: commit, write your handoff note, and end with a final report that starts with "CHECKPOINT:". A fresh agent will continue from your note. At ${tokens(limits.stop)} your tools start being refused.`
122}
123
124export function windDownNote(context: number, left: number, limits: Limits): string {
125 const budget =
126 left > 0 ? `You have ${left} tool calls left before every tool is refused.` : 'That was your last tool call.'
127 return `${MARK} Your context is ${tokens(context)} tokens, past the ${tokens(limits.stop)} limit. ${budget} Use what is left only to commit and write your handoff note, then end with a final report that starts with "CHECKPOINT:".`
128}
129
130export function refusal(context: number, limits: Limits): string {
131 return `${MARK} Refused: your context (${tokens(context)} tokens) is past the ${tokens(limits.stop)} limit. End now with a final report that starts with "CHECKPOINT:" and says what is done and what is left. Resuming this agent will not help; a fresh agent should continue from your report.`
132}
133
134// the main agent decides the split; this tells it what the mod enforces and how to work with it
135export function planningRules(limits: Limits): string {
136 return `# Sizing delegated work
137The subagent-limits plugin is installed. You decide how work is split; the plugin only measures each subagent's context and enforces two limits. When you hand work to subagents:
138- Give each agent one phase it can finish well under ${tokens(limits.nudge)} tokens of context. At ${tokens(limits.nudge)} the plugin asks the agent to checkpoint, and at ${tokens(limits.stop)} it starts refusing the agent's tools.
139- Run agents in parallel only when they edit different files. When they would share files, run a relay instead: one fresh agent per phase, each starting from the previous one's handoff note.
140- Fresh agents often read 100K tokens or more before their first edit. Cut that with a brief that names the exact files, functions and line ranges to read, the test command, and what done means.
141- Set \`model\` on every Agent call.
142- A final report that starts with "CHECKPOINT:" means the agent stopped on purpose with work left. Start a fresh agent from its handoff note rather than resuming the old one.
143For brief and handoff templates, load the subagent-limits:split-work skill.`
144}
145
146export function withContract(prompt: string): string {
147 return prompt.includes(CONTRACT_HEAD) ? prompt : `${prompt}\n\n${CONTRACT}`
148}
149
150export function agentLabel(id: string, description: string | undefined): string {
151 return description !== undefined && description !== '' ? description : `Agent ${id.slice(0, 8)}`
152}
153
154// the call's result with a note the model reads after it; a refusal stays as it is
155function withNote(result: ToolCallResult, note: string): ToolCallResult {
156 if (result.deny !== undefined) return result
157 return { ...result, context: [...(result.context ?? []), note] }
158}
159
160async function label($: EngineInterface, id: string): Promise<string> {
161 const agent = (await $.agent.list()).find(a => a.id === id)
162 return agentLabel(id, agent?.description)
163}
164
165export const register: Register = (on, options) => {
166 const limits = limitsFrom(options)
167 // off: an agent whose cache expired is not told; the person still gets the toast
168 const cacheNotes = options.cacheNotes !== false
169 // module state: a reload starts these over, so a nudge may repeat once after one
170 const budgets = new Map<string, Budget>()
171 // forks inherit the parent's context and its prompt cache: never limited
172 const forks = new Set<string>()
173
174 on('session.end', { reason: 'clear' }, async ($, e, next) => {
175 budgets.clear()
176 forks.clear()
177 return next(e)
178 })
179
180 // only a model that can start agents needs the planning rules
181 on('prompt.compose', async ($, e, next) => {
182 const composed = await next(e)
183 if (!e.tools.includes('Agent') || e.traits.includes('bare')) return composed
184 const rules = { id: 'subagent-limits:planning', text: planningRules(limits), scope: 'session' } as const
185 return { sections: [...composed.sections, rules] }
186 })
187
188 on('agent.spawn', async ($, e, next) => {
189 if (e.fork) {
190 const started = await next(e)
191 if (started.agentId !== undefined) forks.add(started.agentId)
192 return started
193 }
194 return next({ ...e, prompt: withContract(e.prompt) })
195 })
196
197 // every model request of a subagent or an in-process teammate; the main thread is never limited
198 on('turn.step', async function* ($, e, next) {
199 const id = e.agentId
200 if (id === undefined || forks.has(id)) return yield* next(e)
201 const startedAt = await $.clock.now()
202 const result = yield* next(e)
203 if (result.usage === null) return result
204 const usage = result.usage
205 const context = usage.input_tokens + usage.cache_read_input_tokens + usage.cache_creation_input_tokens
206 const budget = budgets.get(id) ?? fresh()
207 budgets.set(id, budget)
208 const drop = cacheDrop(budget.last, { turnId: e.turnId, startedAt, messageCount: e.messageCount }, usage)
209 budget.context = context
210 budget.last = {
211 turnId: e.turnId,
212 endedAt: await $.clock.now(),
213 model: usage.model,
214 messageCount: e.messageCount,
215 calledTools: result.toolUses.length > 0,
216 }
217 if (drop !== undefined) {
218 if (cacheNotes && !budget.isCacheNoted) budget.pendingDrop = drop
219 if (!budget.isDropToasted) {
220 budget.isDropToasted = true
221 $.ui.toast(dropToast(await label($, id), drop))
222 }
223 }
224 return result
225 })
226
227 on('tool.call', async ($, e, next) => {
228 const id = e.agentId
229 const budget = id === undefined ? undefined : budgets.get(id)
230 if (id === undefined || budget === undefined) return next(e)
231 // taken before any await, so of calls made in parallel only the first carries it
232 const drop = budget.pendingDrop
233 if (drop !== undefined) {
234 budget.pendingDrop = undefined
235 budget.isCacheNoted = true
236 }
237 // the tool's result, with the cache note when one is due
238 const ran = async () => {
239 const result = await next(e)
240 return drop === undefined ? result : withNote(result, cacheNote(drop))
241 }
242 const decided = verdict(budget, limits)
243 switch (decided.kind) {
244 case 'run':
245 return ran()
246 case 'nudge': {
247 budget.isNudged = true
248 $.ui.toast(`Asked ${await label($, id)} to checkpoint at ${tokens(budget.context)} of context`)
249 return withNote(await ran(), nudgeNote(budget.context, limits))
250 }
251 case 'wind-down': {
252 // counted before any await, so calls made in parallel each see the ones before them
253 budget.windDownCalls += 1
254 if (budget.windDownCalls === 1) {
255 $.ui.toast(`${await label($, id)} passed ${tokens(limits.stop)}: ${WIND_DOWN_CALLS} tool calls left to checkpoint`)
256 }
257 return withNote(await ran(), windDownNote(budget.context, decided.left, limits))
258 }
259 case 'refuse': {
260 if (!budget.isRefused) {
261 budget.isRefused = true
262 $.ui.toast(`Refusing tools for ${await label($, id)} at ${tokens(budget.context)} of context`)
263 }
264 return { deny: refusal(budget.context, limits) }
265 }
266 }
267 })
268}
269