When one turn uses more than +5 session points or 500,000 uncached input tokens, writes a handoff from exact session data and continues in a fresh session by…

Make the $20 Claude Pro plan last longer in Claude Code.
Twenty-two small mods that show you exactly where your usage goes and cut the waste. One mod makes a model call: prompt-polish, once each time you press Improve, and never on its own. The others make none. None adds anything to the system prompt (read-cap adds one tool, which Claude Code lists by name only until Claude first uses it; write-guard adds one line to a tool result at most once per conversation; compact-keeper adds a note of at most 4,000 characters after a compaction): every figure on screen is one Claude Code already reports, or a time the mod measured.
| Mod | What it does | Where |
|---|---|---|
| tool-diet | Loads tools you have not used lately on demand instead of with every request | Everywhere |
| skill-diet | Lists skills you have not used lately in this project by name only, without their descriptions | Everywhere |
| agent-diet | Runs Explore subagents on Haiku instead of your main model | Everywhere |
| context-xray | /xray opens the exact breakdown of what fills your context window | Everywhere |
| pro-hud | Live meters above the prompt for your 5-hour session, your week and the context window, plus a per-turn receipt of tokens in, cached and out | Claude desktop app |
| output-diet | Trims long shell output, search results and subagent reports before Claude reads them, keeping the head, the tail and the error lines; the untrimmed text is saved to a file Claude can open without a permission prompt | Everywhere |
| reread-guard | Skips Claude re-reading a file it already read when the file has not changed, with Read or a plain cat, sed -n, head, tail or Get-Content; a deliberate retry still goes through | Everywhere |
| write-guard | Steers Claude to Edit instead of rewriting an existing file in full with Write; a deliberate rewrite still goes through | Everywhere |
| loop-guard | Holds back a shell command that already failed twice in a row, so Claude changes approach | Everywhere |
| cmd-diet | Adds quiet flags to noisy shell commands before they run, so Claude reads short output from the start; errors, failures and warnings stay in full | Everywhere |
| gh-account | Runs each git push, pull, fetch, clone and gh command as the logged-in GitHub account that can see the repository, without switching the active account | Everywhere |
| cache-clock | Counts down until the prompt cache expires; once it has, shows exactly how many tokens your next message will re-send uncached | Everywhere |
| read-cap | Stops Claude reading a file over 1,000 lines whole (with Read, cat or Get-Content) and gives it an outline tool, so it reads only the lines it needs; a retry still reads the whole file | Everywhere |
| session-receipt | /receipt opens a pane with the exact tokens every turn of the session spent, and the costliest turns | Everywhere |
| peek | Typing status while background tasks run opens a pane with each one's exact elapsed time and last output lines, instead of sending the prompt to Claude | Everywhere |
| budget-guard | Holds a prompt back once your 5-hour or weekly usage reaches your limit (90% by default); sending it again goes through | Everywhere |
| turn-budget | When one turn uses more than +5 session points, writes a handoff and continues in a fresh session by itself; /handoff any time | Everywhere |
| compact-keeper | After a compaction, adds a note of exact facts from before it: your latest prompt in full, the todo list, the last failed command, files edited | Everywhere |
| collision-guard | Asks before Claude edits a file another chat on this machine changed in the last 30 minutes | Everywhere |
| answer-pane | Explain, plan and ELI5 pages drawn natively in a side pane; plans have decision buttons and Respond fills the prompt box | Desktop app (no diagrams in the terminal) |
| prompt-polish | An Improve button beside Send (above the prompt in the terminal) rewrites your draft with Opus at low effort and puts it back in the box; Undo restores it. One model call per press | Everywhere |
| kit-updates | Tells you when an installed mod from this kit has a newer version or a new mod joins the kit; /kit-update installs updates and new mods | Everywhere |

In Claude Code:
/plugin marketplace add VedantAndhale/claude-pro-kit
/plugin install tool-diet@claude-pro-kit
/plugin install skill-diet@claude-pro-kit
/plugin install agent-diet@claude-pro-kit
/plugin install context-xray@claude-pro-kit
/plugin install pro-hud@claude-pro-kit
/plugin install output-diet@claude-pro-kit
/plugin install reread-guard@claude-pro-kit
/plugin install write-guard@claude-pro-kit
/plugin install loop-guard@claude-pro-kit
/plugin install cmd-diet@claude-pro-kit
/plugin install gh-account@claude-pro-kit
/plugin install cache-clock@claude-pro-kit
/plugin install read-cap@claude-pro-kit
/plugin install session-receipt@claude-pro-kit
/plugin install peek@claude-pro-kit
/plugin install budget-guard@claude-pro-kit
/plugin install turn-budget@claude-pro-kit
/plugin install compact-keeper@claude-pro-kit
/plugin install collision-guard@claude-pro-kit
/plugin install answer-pane@claude-pro-kit
/plugin install prompt-polish@claude-pro-kit
/plugin install kit-updates@claude-pro-kit
Updates are off by default for marketplaces you add yourself. With kit-updates installed you are told when a fix ships and /kit-update installs it, from the desktop app or the terminal. Without it, turn on auto-update once: in a terminal, run claude, then /plugin → Marketplaces → claude-pro-kit → Enable auto-update; the desktop app has no toggle for it.
Install any one on its own; they do not depend on each other. Mods are not sandboxed, so read the code before installing: each mod is a single file under plugins/<name>/hooks/.
Every tool listed in front sends its whole description and schema with every request. A deferred tool is listed by name only, and Claude loads it through ToolSearch when it needs it; Claude Code already does this for most MCP tools. tool-diet does it for the rest of the tools you are not using: anything outside the core set (Bash, PowerShell, Read, Edit, Write, Glob, Grep, Agent, Skill, ToolSearch, TodoWrite, AskUserQuestion) that you have not used in your last five sessions.
Measured with one prompt, "Reply with just OK.", in a fresh session, as the API reported each request:
| Prompt tokens per request | |
|---|---|
| Without tool-diet | 43,859 |
| With tool-diet | 27,905 |
| Change | −15,954 (−36%) |
On that setup, the largest tool moved was Artifact, whose description alone is 19,870 characters. What moves depends on your tools and your habits; /xray shows yours.
/tool-diet lists what is on demand this session, grouped by where each tool comes from; /tool-diet keep <tool> always loads one, unkeep undoes it, and /tool-diet off|on switches it. Answers are toasts, so they add nothing to the conversation. The status line keeps the count: 38 tools on demand.
Every request carries the skill listing: each installed skill's name and its whole description. With a few plugins installed that is dozens of skills, most of which a given project never uses (video skills in a backend repo, document skills in a game). skill-diet keeps the skills you used lately in this project listed in full and lists the rest on one line by name only. The Skill tool still loads any of them, and typing /name still works.
Measured with one prompt, "Reply with just OK.", in a fresh session with 45 skills installed and the other mods on, as the API reported the request:
| Prompt tokens per request | |
|---|---|
| Without skill-diet | 27,423 |
| With skill-diet | 19,197 |
| Change | −8,226 (−30%) |
In the same setup, asked to fill in a PDF form, Claude found anthropic-skills:pdf from its name alone and loaded it with the Skill tool. What moves depends on your skills; /xray shows yours.
/name or called by Claude through the Skill tool./skill-diet shows what is listed by name only this session and how many characters left the listing; /skill-diet keep <skill> always lists one in full, unkeep undoes it, and /skill-diet off|on switches it. Answers are toasts, so they add nothing to the conversation. The status line keeps the count, in the form <n> skills by name only, <n> characters off./xray shows the skills line before and after, in tokens.A subagent runs on your main model unless its definition or Claude's Agent call names another. Explore only searches and reads, yet in the author's own 85 subagent transcripts, every one of the 2,043 Explore requests ran on Opus or Sonnet, reading 147,298,992 input tokens. agent-diet starts the agent types you list (Explore by default) on Haiku.
/agent-diet shows the setting and how many subagents it moved this session; /agent-diet model haiku|sonnet|opus picks the model, /agent-diet add|remove <agent type> changes the list, and /agent-diet off|on switches it. Answers are toasts, so they add nothing to the conversation. The status line keeps the count, in the form 2 subagents on haiku.Measured with claude -p --model sonnet --output-format json on a task that sends one Explore subagent to find a file, three runs each in alternating order. Every run found the file. Figures are the cost and tokens Claude Code reported per model:
| Run | Without: cost | Without: Sonnet tokens in | With: cost | With: Sonnet tokens in | With: Haiku tokens in |
|---|---|---|---|---|---|
| 1 | $0.05692 | 68,395 | $0.03724 | 40,213 | 27,848 |
| 2 | $0.05577 | 68,111 | $0.03764 | 39,593 | 27,872 |
| 3 | $0.05441 | 68,123 | $0.03736 | 40,184 | 42,770 |
| Average | $0.05570 | $0.03741 |

/xray opens a pane with the exact breakdown /context computes: what is sent with every request (system prompt, tools, MCP tools, memory files, skills, messages), what is loaded on demand, which MCP tools load every time, and each memory file's size. It measures when the pane opens and when you press Refresh (or r), never in the background, because the exact count sends one token-count request per tool and memory file.
Opens by itself when the context crosses 60% and again at 80%, with a toast pointing at /compact and /handoff. /xray auto off keeps it to /xray only.
The recording at the top of this page is the band. In text:
Session ━━━━━━━━━━━━━━────── 69% resets in 32m
Week ━━━━━━━━━━━━━━━───── 74% resets in 4d 17h
Context ━━━━━━━━──────────── 44% 436,034 tokens
This turn 1m 35s · 1 tool · 0 files edited · 436,034 in · 431,260 cached · 580 out
On a wide window the three meters sit side by side on one row; on a narrow one each figure on the turn line wraps whole rather than being cut off.
/compact when it climbs.· session 20%, and finished tool calls draw as one line: status dot, tool, target, time./hud shows what is on; /hud all on|off, or /hud band|spinner|cards on|off. The answer is a toast, so toggling adds nothing to the conversation. It draws in the desktop app only and leaves the terminal as it is. Rows other mods put above the prompt still draw beneath the band.
Weekly toasts at 50%, 75% and 90% of the week, once each, because the weekly limit drains quietly across many sessions.
When a Bash or PowerShell result runs past 120 lines or 8,000 characters, Claude reads:
[output-diet: 82/500 lines shown; all 500 in ~/.claude/projects/<project>/<session>/tool-results/output-diet-<call>.txt]
<first 30 lines>
… [lines 31–450 omitted; error/warning lines from them:]
250│ ERROR: build failed in src/app.ts
300│ Warning: deprecated API
…
<last 50 lines>
3 trimmed · 41,200 chars saved.tool-results folder, beside the transcript, where Claude Code keeps the outputs it saves itself, so Claude can open it without a permission prompt. When that folder cannot be found, it goes under ~/.claude/output-diet/ (or CLAUDE_CONFIG_DIR).Grep result past the same limits keeps its first 100 lines, with a note to narrow the search: [output-diet: first 100/252 lines shown; narrow the search, or read all in …]. In the author's 1,048 past Grep results, 26 went past the limits, and keeping the first 100 lines would have cut 104,052 characters. Glob is left alone: none of 128 went past.When Claude asks to Read the same range of the same file again in the same conversation, and the file's size and modification time have not changed, the read is skipped and Claude is told to use the copy it has. Any edit, a different range, or a subagent (which has its own context) reads freely. Claude Code can clear old tool results from context, so retrying the identical read straight after a skip always goes through. The record resets on /compact and /clear.
What Claude is told in place of the repeat read, kept to one line because the model reads it:
api.ts unchanged since you read it; use that copy. If it's gone from context, retry the same Read.
Shell reads too. Claude often reads a file with the shell instead of Read: in the author's transcripts, 318 sed -n runs went past the guard. A shell command that only prints one file is now treated the same way: cat FILE, sed -n 'A,Bp' FILE, head -n N / tail -n N, and in PowerShell Get-Content, gc, cat or type (with -TotalCount, -Head, -First, -Tail, -Last or -Raw). The same range of the same unchanged file is skipped once, and retrying the same command goes through. A whole cat after a whole Read that showed every line counts as a repeat. A pipe, a chain, a variable, a glob or any other flag runs untouched. A shell read never counts as a Read, since Edit needs a real one first.
The status line counts them: 2 re-reads skipped.
Write sends the whole file as Claude's output, the most expensive kind of token; Edit sends only the lines that change. Claude sometimes rewrites a file it has already read in full to change a few lines. In the author's own 272 session transcripts, 141 writes rewrote a file Claude had already read or written (999,273 characters), and 112 of them followed an earlier rewrite in the same session. Of the 512,847 characters in the rewrites whose previous version was in the transcript, 171,325 had changed.
A hook runs after Claude has written the content, so write-guard cannot save the rewrite it sees; it stops the ones after it:
You rewrote all of api.ts. For changes to an existing file use Edit: it sends only the changed lines.api.ts exists; change it with Edit, not a full Write. If a full rewrite is intended, retry the same Write. Retrying the same Write goes through./clear starts over.1 full rewrite held back.Each retry of a failing command re-sends the whole conversation, and the same command usually fails the same way. In the author's 272 session transcripts, 10 commands failed 3 or more times, 38 runs between them.
Once the exact same Bash or PowerShell command has failed twice in a row, the next try is held back once, and Claude is told:
This exact command failed 2 times in a row. Change the approach instead of rerunning it. If a rerun is intended, retry the same command.
/clear starts over.1 failing retry held back.output-diet trims long output after a command has run; cmd-diet keeps the noise from being printed at all. Before a Bash or PowerShell command runs, a known noisy command gets its own quiet flags, which drop progress lines and keep errors, test failures and warnings:
| Command | Runs as | Output, measured |
|---|---|---|
git status | git status --short --branch | 415 to 58 characters |
pytest | pytest -q | 1,063 to 495 characters, failure details unchanged |
cargo build / test / check / clippy / run | cargo build -q | 197 to 0 characters on success, warnings unchanged |
npm install / ci | npm install --no-audit --no-fund | 62 to 18 characters |
curl (Bash only) | curl -sS | 1,049 to 577 characters, the progress meter removed |
mvn | mvn -B -ntp | not measured: drops download progress |
wget (Bash only) | wget -nv | not measured: one line per file |
docker pull | docker pull -q | not measured: drops layer progress |
In a live headless run (git status && python -m pytest on 41 test files, one failing), the request cost 41,535 tokens without cmd-diet and 40,571 with it, by the API's usage, and both runs named the failing test and its reason correctly. The shorter output stays in context, so every later request in the session sends less too.
&&, || or ; chain is handled on its own. A step that pipes, redirects or substitutes (| grep, > file, $(...)) is left alone, since something else reads its output; 2>&1 is fine.git status --porcelain, pytest -v, curl -fsSL, npm install --silent) is left alone, and a flag already present is not added twice.3 commands quieted.With two GitHub accounts logged in to gh, a push to a repository the other account owns fails, and Claude spends turns on gh auth status and gh auth switch. Before a Bash or PowerShell command runs, gh-account matches each git push, pull, fetch, clone, ls-remote and gh step to the repository's owner, and the owner to a logged-in account. When that account is not gh's active one, that step alone runs with its token:
GH_TOKEN="$(gh auth token --user ACCOUNT)" git push
The token itself never appears in the command or its output, and the active account is not switched. For git it also asks gh for the credential.
-R owner/repo, a gh api repos/<owner>/... path, or the folder's remotes. With no remote named, every remote must point at the same owner.gh auth status, and organizations from gh api user/orgs. What it learns is kept across sessions.gh auth and other gh commands that reach no repository, a command that already sets GH_TOKEN, a token from the environment, git over ssh, hosts other than github.com, and subshells or substitutions./gh-account shows which account each owner uses, and /gh-account forget clears it. Answers are toasts. The status line counts them: 2 commands sent as <account>.Claude's prompt cache keeps your conversation for a fixed time after each request. Reply within it and the context is read from cache; reply after it and the whole context is sent again at full price. That message pays the cache-write price (1.25 times the input price for a 5-minute cache, 2 times for a 1-hour one) on every token instead of the cache-read price (0.1 times): 12.5 to 20 times more for the same context, and nothing on screen says so.
cache-clock puts the countdown in the status line, restarted by every response from the main conversation:
cache warm · 3m left
cache cold · next message re-sends 61,204 tokens
When the cache runs out, a toast says so once (after the lifetime is known, see below). The token figure is the previous response
hooks/register.ts 255 lines1import type { EngineInterface, Register } from 'claude-code'
2
3import {
4 CONTINUE,
5 CONTINUE_QUIETLY,
6 DEFAULTS,
7 STOP,
8 choiceOf,
9 describeLimits,
10 firstMarks,
11 isOver,
12 nextMarks,
13 parse,
14 question,
15 status,
16 stopNote,
17 uncached,
18 HANDOFF,
19} from './budget'
20import type { Limits, Marks, Spend } from './budget'
21import { continueMessage, handoffDoc } from './handoff'
22import type { Todo } from './handoff'
23
24// Before each model request of a turn (subagents' included), checks what the
25// turn has spent; past the limit it asks, before the request is sent, whether
26// to go on. Only between requests, so a running command is never cut.
27
28async function limitsOf($: EngineInterface): Promise<Limits> {
29 return { ...DEFAULTS, ...((await $.store.get('limits')) as Partial<Limits> | undefined) }
30}
31
32type Turn = {
33 id: string
34 start?: number
35 spend: Spend
36 marks: Marks
37 isQuiet: boolean
38 isStopped: boolean
39 /** One question at a time, however many subagents step at once. */
40 asking?: Promise<boolean>
41}
42
43const EDITS = new Set(['Edit', 'MultiEdit', 'Write', 'NotebookEdit'])
44
45// What this session has done, for the handoff: kept for the whole session.
46const tracked = { files: {} as Record<string, number>, todos: [] as Todo[] }
47
48const join = (...parts: string[]) => parts.map((p, i) => (i === 0 ? p.replace(/[\\/]+$/, '') : p)).join('/')
49
50async function gitOut($: EngineInterface, cwd: string, args: string[]) {
51 const ran = await $.process.run(['git', ...args], { cwd, timeoutMs: 10_000 }).catch(() => undefined)
52 return ran && ran.exitCode === 0 ? ran.stdout : undefined
53}
54
55/**
56 * Writes the handoff from exact session data (no model call), stops the turn,
57 * clears the conversation and starts the fresh one with the handoff as its
58 * first message. If the clear is refused, the handoff waits in the prompt box.
59 */
60async function handOff($: EngineInterface, turnId: string | undefined, reason: string) {
61 const cwd = await $.session.cwd()
62 const messages = await $.session.messages().catch(() => [])
63 const prompts = messages.filter(m => m.role === 'user' && m.text.trim() && !m.text.trimStart().startsWith('<')).map(m => m.text)
64 const lastAnswer = [...messages].reverse().find(m => m.role === 'assistant' && m.text.trim())?.text
65 const now = await $.clock.now()
66 const doc = handoffDoc({
67 reason,
68 cwd,
69 prompts,
70 lastAnswer,
71 files: tracked.files,
72 todos: tracked.todos,
73 gitStatus: await gitOut($, cwd, ['status', '--short']),
74 gitDiffStat: await gitOut($, cwd, ['diff', '--stat']),
75 at: new Date(now).toISOString().replace('T', ' ').slice(0, 16),
76 })
77
78 const config = await $.env.get('CLAUDE_CONFIG_DIR')
79 const home = (await $.env.get('USERPROFILE')) ?? (await $.env.get('HOME'))
80 const base = config ?? (home === undefined ? undefined : join(home, '.claude'))
81 const path = base === undefined ? 'not saved' : join(base, 'handoffs', `${new Date(now).toISOString().replace(/[:.]/g, '-')}.md`)
82 if (base !== undefined) await $.fs.write(path, doc).catch(() => undefined)
83
84 if (turnId) await $.turn.abort({ turnId }).catch(() => undefined)
85 $.ui.status(undefined)
86
87 // After this hook returns: a clear and a new prompt cannot start inside it.
88 $.clock.after(300, async () => {
89 const message = continueMessage(doc, path)
90 const cleared = await $.command
91 .run({ command: 'clear' })
92 .then(() => true)
93 .catch(() => false)
94 if (cleared) {
95 tracked.files = {}
96 tracked.todos = []
97 await $.prompt.submit({ text: message })
98 $.ui.toast(`Fresh session started from a handoff. The old conversation is in /resume. Handoff: ${path}`, { timeoutMs: 15_000 })
99 } else {
100 await $.prompt.fill({ text: message, mode: 'replace' }).catch(() => undefined)
101 $.ui.toast(`Handoff written (${path}) and placed in the prompt box. Run /clear, then send it.`, { timeoutMs: 15_000 })
102 }
103 })
104}
105
106// Resolves true to go on, false when the person chose to stop.
107async function check($: EngineInterface, t: Turn, latest: number | undefined): Promise<boolean> {
108 if (t.isStopped) return false
109 if (t.isQuiet || !isOver(t.spend, t.marks)) return true
110 if (t.asking) return t.asking
111
112 if ((await limitsOf($)).onLimit === 'handoff') {
113 t.isStopped = true
114 const used = question(t.spend, t.marks, t.start, latest).replace(/^This turn has used /, '').replace(/ Keep going\?$/, '')
115 await handOff($, t.id, `The previous session was handed off by turn-budget: one turn used ${used}`)
116 return false
117 }
118
119 t.asking = (async () => {
120 let answer: string
121 try {
122 answer = await $.ui.ask(question(t.spend, t.marks, t.start, latest), {
123 options: [HANDOFF, CONTINUE, CONTINUE_QUIETLY, STOP],
124 header: 'Turn budget',
125 })
126 } catch {
127 // Nobody to ask (a -p run, or the dialog was dismissed): never stop a
128 // turn on a question nobody saw; note it and go on.
129 $.ui.status(`${status(t.spend, t.marks)} · over, not asked`)
130 t.isQuiet = true
131 return true
132 }
133 const choice = choiceOf(answer)
134 if (choice === 'handoff') {
135 t.isStopped = true
136 await handOff($, t.id, "The previous session was handed off at the user's request when a turn crossed its budget.")
137 return false
138 }
139 if (choice === 'stop') {
140 t.isStopped = true
141 await $.session
142 .append({ message: { type: 'user', content: [{ type: 'text', text: stopNote(t.spend, t.marks) }] } })
143 .catch(() => undefined)
144 await $.turn.abort({ turnId: t.id }).catch(() => undefined)
145 $.ui.status(undefined)
146 return false
147 }
148 if (choice === 'quiet') t.isQuiet = true
149 else t.marks = nextMarks(t.spend, t.marks, await limitsOf($))
150 return true
151 })()
152
153 try {
154 return await t.asking
155 } finally {
156 t.asking = undefined
157 }
158}
159
160export const register: Register = on => {
161 let latest: number | undefined
162 let turn: Turn | undefined
163
164 const fivePoints = (start: number | undefined) =>
165 start !== undefined && latest !== undefined ? Math.max(0, Math.round((latest - start) * 10) / 10) : undefined
166
167 on('session.start', async ($, e, next) => {
168 await $.command.register({ name: 'turn-budget', description: 'Turn budget: /turn-budget · <points> · tokens <n> · handoff|ask · off|on' })
169 await $.command.register({ name: 'handoff', description: 'Write a handoff from this session (no model call) and continue in a fresh session' })
170 return next(e)
171 })
172
173 // What the handoff needs: the files edited and the latest todo list.
174 on('tool.call', async ($, e, next) => {
175 if (e.agentId === undefined) {
176 const path = (e as { file_path?: unknown }).file_path ?? (e as { notebook_path?: unknown }).notebook_path
177 if (EDITS.has(e.tool) && typeof path === 'string') tracked.files[path] = (tracked.files[path] ?? 0) + 1
178 const todos = (e as { todos?: unknown }).todos
179 if (e.tool === 'TodoWrite' && Array.isArray(todos)) tracked.todos = todos as Todo[]
180 }
181 return next(e)
182 })
183
184 on('command.run', { command: 'handoff' }, async $ => {
185 await handOff($, turn?.id, 'The previous session was handed off with /handoff.')
186 return {}
187 })
188
189 on('session.measure', async ($, e, next) => {
190 const five = e.rateLimits.find(l => l.kind === 'five_hour')
191 if (five) {
192 // The window reset mid-turn: the new reading is the new baseline.
193 if (turn?.start !== undefined && five.percentUsed < turn.start) turn.start = five.percentUsed
194 latest = five.percentUsed
195 if (turn && !turn.isStopped) {
196 turn.spend.points = fivePoints(turn.start)
197 $.ui.status(status(turn.spend, turn.marks))
198 }
199 }
200 return next(e)
201 })
202
203 on('turn.start', async ($, e, next) => {
204 const limits = await limitsOf($)
205 turn = limits.isOff
206 ? undefined
207 : {
208 id: e.turnId,
209 start: latest,
210 spend: { points: latest === undefined ? undefined : 0, tokens: 0, requests: 0 },
211 marks: firstMarks(limits),
212 isQuiet: false,
213 isStopped: false,
214 }
215 $.ui.status(undefined)
216 return next(e)
217 })
218
219 on('turn.complete', async ($, e, next) => {
220 if (e.agentId === undefined) {
221 turn = undefined
222 $.ui.status(undefined)
223 }
224 return next(e)
225 })
226
227 on('turn.step', async function* ($, e, next) {
228 const t = turn
229 if (!t) return yield* next(e)
230
231 if (!(await check($, t, latest))) {
232 // Stopped: answer the step without sending the request.
233 return { turnId: e.turnId, index: e.index, answer: '', toolUses: [], stopReason: null, usage: null }
234 }
235
236 const step = yield* next(e)
237 if (step.usage) {
238 t.spend.tokens += uncached(step.usage)
239 t.spend.requests += 1
240 t.spend.points = fivePoints(t.start)
241 $.ui.status(status(t.spend, t.marks))
242 }
243 return step
244 })
245
246 // Answered with a toast and no text, so the command adds nothing to the conversation.
247 on('command.run', { command: 'turn-budget' }, async ($, e) => {
248 const current = await limitsOf($)
249 const changed = parse(e.args, current)
250 if (changed) await $.store.set('limits', changed)
251 $.ui.toast(describeLimits(changed ?? current), { timeoutMs: 10_000 })
252 return {}
253 })
254}
255hooks/budget.ts 92 lines1// The budget arithmetic: pure, so the tests exercise it directly.
2
3export type Limits = {
4 /** Session points (5-hour window) one turn may use before it asks. */
5 points: number
6 /** Uncached input tokens (uncached + cache-written) one turn may send before it asks. */
7 tokens: number
8 isOff: boolean
9 /** At the limit: hand off to a fresh session by itself (default), or ask first. */
10 onLimit: 'handoff' | 'ask'
11}
12
13export const DEFAULTS: Limits = { points: 5, tokens: 500_000, isOff: false, onLimit: 'handoff' }
14
15export type Spend = {
16 /** Session points this turn has used, from Claude Code's readings; undefined off a subscription. */
17 points?: number
18 tokens: number
19 requests: number
20}
21
22/** The next marks to ask at; each Continue raises the crossed one by its own step. */
23export type Marks = { points: number; tokens: number }
24
25export const firstMarks = (limits: Limits): Marks => ({ points: limits.points, tokens: limits.tokens })
26
27export const isOver = (spend: Spend, marks: Marks): boolean =>
28 (spend.points !== undefined && spend.points >= marks.points) || spend.tokens >= marks.tokens
29
30/** Which limit the turn crossed: the one to name when asking and in the note. Points first when both are. */
31export type Crossed = 'points' | 'tokens'
32
33export const crossed = (spend: Spend, marks: Marks): Crossed =>
34 spend.points !== undefined && spend.points >= marks.points ? 'points' : 'tokens'
35
36export const nextMarks = (spend: Spend, marks: Marks, limits: Limits): Marks => {
37 let { points, tokens } = marks
38 while (spend.points !== undefined && spend.points >= points) points += limits.points
39 while (spend.tokens >= tokens) tokens += limits.tokens
40 return { points, tokens }
41}
42
43export type ModelUsageLike = {
44 input_tokens: number
45 cache_creation_input_tokens?: number
46}
47
48/** Cache reads are left out: they cost a fraction of new input. */
49export const uncached = (usage: ModelUsageLike) => usage.input_tokens + (usage.cache_creation_input_tokens ?? 0)
50
51const n = (value: number) => value.toLocaleString('en-US')
52
53export const question = (spend: Spend, marks: Marks, start: number | undefined, now: number | undefined) => {
54 const used =
55 crossed(spend, marks) === 'points' && start !== undefined && now !== undefined
56 ? `${spend.points} ${spend.points === 1 ? 'point' : 'points'} of your session (${start}% → ${now}%, limit +${marks.points})`
57 : `${n(spend.tokens)} uncached input tokens (limit ${n(marks.tokens)})`
58 return `This turn has used ${used} over ${spend.requests} ${spend.requests === 1 ? 'request' : 'requests'}. Keep going?`
59}
60
61export const HANDOFF = 'Hand off to a fresh session'
62export const CONTINUE = 'Continue'
63export const CONTINUE_QUIETLY = "Don't ask again"
64export const STOP = 'Stop here'
65
66export type Choice = 'handoff' | 'continue' | 'quiet' | 'stop'
67
68/** Anything but a Continue label stops: an answer that is not a clear yes protects the session. */
69export const choiceOf = (answer: string): Choice =>
70 answer === HANDOFF ? 'handoff' : answer === CONTINUE ? 'continue' : answer === CONTINUE_QUIETLY ? 'quiet' : 'stop'
71
72export const stopNote = (spend: Spend, marks: Marks) =>
73 `[turn-budget] The user stopped the previous turn at ${crossed(spend, marks) === 'points' ? `+${spend.points} session points` : `${n(spend.tokens)} uncached input tokens`}. Ask before continuing that work.`
74
75export const status = (spend: Spend, marks: Marks) =>
76 spend.points !== undefined ? `turn budget ${spend.points}/${marks.points} pts` : `turn budget ${n(spend.tokens)}/${n(marks.tokens)} tok`
77
78/** `/turn-budget 8`, `/turn-budget tokens 1000000`, `/turn-budget off|on`; undefined for anything else. */
79export const parse = (args: string, limits: Limits): Limits | undefined => {
80 const [a, b] = args.trim().toLowerCase().split(/\s+/)
81 if (a === 'off' || a === 'on') return { ...limits, isOff: a === 'off' }
82 if (a === 'handoff' || a === 'ask') return { ...limits, onLimit: a }
83 if (a === 'tokens' && b && /^\d+$/.test(b) && Number(b) > 0) return { ...limits, tokens: Number(b) }
84 if (a && /^\d+(\.\d+)?$/.test(a) && Number(a) > 0) return { ...limits, points: Number(a) }
85 return undefined
86}
87
88export const describeLimits = (limits: Limits) =>
89 limits.isOff
90 ? 'Turn budget is off. /turn-budget on to turn it back on.'
91 : `Turn budget: at +${limits.points} session points or ${n(limits.tokens)} uncached input tokens in one turn, it ${limits.onLimit === 'handoff' ? 'hands off to a fresh session by itself' : 'asks first'}. /turn-budget <points> · tokens <n> · handoff|ask · off`
92hooks/handoff.ts 76 lines1// The handoff document: built from exact session data, with no model call, so a
2// fresh session can pick the work up without re-sending the old context.
3
4export type Todo = { content: string; status: string }
5
6export type HandoffInput = {
7 /** Why the old session stopped, in one line. */
8 reason: string
9 cwd: string
10 /** The person's prompts, oldest first. */
11 prompts: string[]
12 /** The last thing Claude said before the stop. */
13 lastAnswer?: string
14 /** Files Claude edited in the old session, with how many edits each. */
15 files: Record<string, number>
16 todos: Todo[]
17 gitStatus?: string
18 gitDiffStat?: string
19 at: string
20}
21
22const PROMPTS = 5
23const PROMPT_CHARS = 800
24const ANSWER_CHARS = 1_200
25const GIT_LINES = 40
26const MAX_CHARS = 12_000
27
28const clip = (text: string, n: number) => {
29 const t = text.trim()
30 return t.length > n ? `${t.slice(0, n - 1)}…` : t
31}
32
33const lines = (text: string | undefined, n: number) => {
34 // Leading spaces matter: in `git status --short`, " M" is not "M ".
35 const all = (text ?? '').split('\n').filter(l => l.trim() !== '')
36 return all.length > n ? [...all.slice(0, n), `… ${all.length - n} more`] : all
37}
38
39const MARK: Record<string, string> = { completed: '[x]', in_progress: '[~]', pending: '[ ]' }
40
41export const handoffDoc = (h: HandoffInput) => {
42 const out: string[] = [
43 `# Handoff from a previous session`,
44 '',
45 `${h.reason} Written ${h.at} in \`${h.cwd}\`, without a model call: everything below is taken from the session as it was.`,
46 '',
47 '## What the user asked (latest last, in full; earlier ones shortened)',
48 // The latest prompt is what to continue: never cut, its newlines kept.
49 ...h.prompts
50 .slice(-PROMPTS)
51 .map((p, i, all) => (i === all.length - 1 ? `${i + 1} (latest). ${p.trim()}` : `${i + 1}. ${clip(p, PROMPT_CHARS).replace(/\n+/g, ' ')}`)),
52 ]
53 if (h.todos.length) out.push('', '## Todo list when it stopped', ...h.todos.map(t => `- ${MARK[t.status] ?? '[ ]'} ${t.content}`))
54 const files = Object.entries(h.files).sort((a, b) => b[1] - a[1])
55 if (files.length) out.push('', '## Files edited', ...files.map(([f, n]) => `- \`${f}\` (${n} ${n === 1 ? 'edit' : 'edits'})`))
56 const status = lines(h.gitStatus, GIT_LINES)
57 if (status.length) out.push('', '## git status --short', '```', ...status, '```')
58 const diff = lines(h.gitDiffStat, GIT_LINES)
59 if (diff.length) out.push('', '## git diff --stat', '```', ...diff, '```')
60 if (h.lastAnswer?.trim()) out.push('', '## Last thing Claude said', ...clip(h.lastAnswer, ANSWER_CHARS).split('\n').map(l => `> ${l}`))
61 out.push(
62 '',
63 '## Next',
64 'Continue the latest request. Read the files above as you need them rather than all at once, check the todo list, and ask the user if anything here is unclear.',
65 )
66 const doc = out.join('\n')
67 // The cap leaves the latest prompt out of the count: it comes before the
68 // rest, so the cut never reaches it.
69 const max = MAX_CHARS + (h.prompts.at(-1)?.trim().length ?? 0)
70 return doc.length > max ? `${doc.slice(0, max - 40)}\n\n… (handoff cut at ${max.toLocaleString('en-US')} characters)` : doc
71}
72
73/** The one message that starts the fresh session: the handoff itself, so nothing needs reading first. */
74export const continueMessage = (doc: string, path: string) =>
75 `Continue the work from my previous session. The handoff is below (also saved at ${path}).\n\n${doc}`
76