Compacts when the work moves on from one sub-task to the next instead of when the context window fills, including in the middle of a long turn. Once the…

A Claude Code plugin that compacts the conversation when the work moves on from one sub-task to the next, instead of when the context window fills.
without taskcut with taskcut
─────────────────────────────────── ───────────────────────────────────
you: "do A, B and C" you: "do A, B and C"
... A's work ... ... A's work ...
... B's work ... A done, on to B
... C's work, half-way ── compacted ──
── window full ── ... B's work ...
── compacted ── B done, on to C
C's live detail ── compacted ──
summarised away ... C's work ...
C done: kept, it is yours now
taskcut changes when Claude Code compacts, not how. Once the context is past a floor, a model judges each step Claude takes for whether the work is moving on from a finished piece to another, reading what you asked for and where Claude has got to, never any tool output. If it is, taskcut runs Claude Code's own compaction, the one /compact runs. A piece being finished is not enough: when the last thing you asked for is done, nothing has moved on yet, and what you say next may well be about it. Hand over twenty tasks in one message and walk away, and it compacts between them, not after the last. Once a turn has ended the work is back with you, and so is /compact. Below the floor it does nothing at all and costs nothing, and nothing in your prompts has to mention it.
Needs Claude Code 2.1.278 or newer; developed and measured on 2.1.280. In a project you work on — installed for you alone, nothing committed:
claude plugin marketplace add wasd96040501/taskcut --scope local
claude plugin install taskcut@taskcut --scope local
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude
The install notes that options are "not yet set". That is fine: unset, each takes its default.
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 is needed on every launch while function hooks are in early access. Without it the plugin is listed as installed but never runs, and nothing says so. To set it once, put "CLAUDE_CODE_ENABLE_FUNCTION_HOOKS": "1" under env in ~/.claude/settings.json. To check that taskcut loads:
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude -p ok --debug-file /tmp/taskcut.log >/dev/null
grep 'taskcut@taskcut loaded' /tmp/taskcut.log
Then work as usual. Nothing happens until the context passes 35% — the ctx figure in the status line. From then on, when Claude finishes one of the things you asked for and moves on to the next, a dim line says so:
taskcut: context at 43%, compacting before the next piece
taskcut ends the turn before Claude's next request, compacts exactly as /compact would, and sends Continue. in your place — the transcript shows it as a message from the taskcut plugin — and the work carries on. A step that asks you something, works on a piece not yet finished, or finishes the last thing you asked for is left alone. Nothing you type mentions taskcut.
To remove it:
claude plugin uninstall taskcut@taskcut --scope local
claude plugin marketplace remove taskcut --scope local
Claude Code compacts when the window fills. On a job that runs for hours or days, that moment almost never lines up with the shape of the work: it fires in the middle of a sub-task, summarising away detail that is still live while keeping detail from work that finished an hour ago.
A long job is a sequence of shorter ones, and the moment when what still matters has a clean answer is the end of a sub-task. taskcut compacts there instead, and leaves what a compaction keeps to Claude Code. The job that needs it most is the one nobody watches: a list of tasks handed over in one message, worked through in one turn that runs for hours.
# this repository, for everyone who clones it
claude plugin marketplace add wasd96040501/taskcut --scope project
claude plugin install taskcut@taskcut --scope project
# every session on this machine
claude plugin marketplace add wasd96040501/taskcut
claude plugin install taskcut@taskcut --scope user
Start sessions with CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude, or put "CLAUDE_CODE_ENABLE_FUNCTION_HOOKS": "1" under env in ~/.claude/settings.json and start them as usual. Off for one session: TASKCUT=0 claude. To remove it, uninstall with the same --scope.
| Below the floor | Nothing: no model call, no tool, nothing written. |
| Each step Claude explains past it | One sonnet call and a sentence out. What goes in is your messages, Claude's latest messages, what its latest commands touched and the step it is taking, never any output: about two thousand tokens however long the session has run, uncached ($.model.complete marks no cache point) — about half a cent. It runs while the step's tools run, so it rarely adds a wait. |
| Each compaction | Whatever /compact costs, because it is /compact. |
Past the floor is a short stretch: a compaction takes the context back under it, and judging stops until it fills again.
/taskcut says what it has spent so far in the session. /cost does not count the judge's calls -- Claude Code keeps no ledger of a plugin's model calls -- so they are counted there, from the token counts the API returns.
Thirty-six real sqlglot changes, handed to Sonnet 5 in one message and left to run: taskcut compacted once inside the turn, after issue 20, from 431,567 tokens to 10,511, and carried on. The peak context fell from 71% of the window to 43%, and the cost by about a fifth; both runs solved all thirty-six. Whether compacting at boundaries makes long sessions work better is what the benchmark is for: docs/measurement.md.
Both have a working default. Set one at install time with --config KEY=VALUE, or change it later from a session with /plugin. Settings are yours, not a scope's: Claude Code keeps plugin settings in your user settings, so they apply wherever taskcut is installed on this machine, whatever --scope it was installed with.
| Setting | In /plugin | Default | What it controls |
|---|---|---|---|
floorPercent | Context % before taskcut acts | 35 | A number from 0 to 100. Below this context fill, taskcut does nothing: no model is asked and nothing is compacted. Lower compacts sooner and more often; each compaction clears the prompt cache. 0 checks every step. |
model | Model that spots task switches | sonnet | The model that decides whether Claude has moved on to the next task — the model auto mode's permission classifier uses by default. haiku costs about a third but misses about half the task switches. An alias or a full model id, resolved the way a --model value is. |
Left empty in /plugin, a setting keeps its default.
After each step of the main conversation, taskcut reads the context fill the status line shows. Below floorPercent it stops there. Past it, it asks model, through the hooks API's $.model.complete, whether the work moves on from a finished piece to another at that step. The judge reads what says so and little else: every message you sent (the first and the latest few whole, the rest cut to a line), Claude's latest messages, what its latest calls touched — a file, or what a command says it does, never the call in full — its task list if it keeps one, and the step it is judging: what Claude just said and the calls it is making. Never any tool output, and not CLAUDE.md. A step that asks you something is never a boundary, and the end of a turn is never judged.
On a yes, taskcut waits for that step's tools to finish, ends the turn with $.turn.abort before the next request goes out, calls $.session.compact() — the call /compact makes — and submits Continue. with $.prompt.submit: what you would do yourself with Esc, /compact and "continue". On anything else it leaves the conversation alone.
The source is two files: hooks/register.ts (the hooks) and hooks/judge.ts (what the judge reads and asks). docs/design.md has the reasoning, and the designs this one replaced.
claude -p and the SDK transport cannot compact, and taskcut never ends a turn there.Continue., which the transcript shows as a message from the plugin. What the compaction keeps is the same as /compact keeps./compact before it is yours to run; Claude Code's own threshold still applies./compact typed by hand included — answers that message instead of summarising the conversation, so taskcut compacts only after a step that followed tool results.make validate reports anything the engine would refuse.See CONTRIBUTING.md. make check runs everything CI runs.
MIT. See LICENSE.
hooks/register.ts 234 lines1/**
2 * taskcut: compaction at sub-task boundaries.
3 *
4 * Claude Code compacts when the context window fills, which is rarely the moment
5 * a piece of work ends: the compaction lands in the middle of one. taskcut moves
6 * it to the point where the work moves on from one piece to the next. Once the
7 * context is past the floor, a model judges each step the working model makes
8 * for whether the work moves on from a finished piece to another. When it
9 * does, taskcut calls `$.session.compact()`, the same call `/compact` makes,
10 * and Claude Code compacts as it always does.
11 *
12 * The end of a turn is left alone: the work is back with the person, and what
13 * they say next may well be about the piece just finished. When they move on,
14 * `/compact` is theirs.
15 *
16 * A compaction can only run between turns. So taskcut ends the turn before its
17 * next model request, compacts, and submits one line that picks the work back
18 * up: what a person watching would do with Esc, `/compact` and "continue".
19 *
20 * taskcut decides when; the engine decides what is kept. Below the floor it
21 * does nothing at all: no model is asked, nothing is written, and the working
22 * model is never told taskcut exists.
23 *
24 * Everything that touches `$` is declared here, at the top level of this file:
25 * the loader follows `$` into a function declared in the same module and refuses
26 * one imported from another.
27 */
28
29import type { EngineInterface, Register } from 'claude-code'
30
31import { ENV_VAR, INERT, decideActivation, type Activation } from './activation'
32import { judgingFrom, readConfig, type Config } from './config'
33import { ANSWER_TOKENS, JUDGE_SYSTEM, acts, judgeable, judgePrompt, readReply, saysNext, type Step } from './judge'
34import { NOTHING, addCompaction, addJudgement, judgementLine, readUsage, spendReport, type Spend } from './spend'
35
36/**
37 * Resolved once, at `session.start`. It starts inert so that a session in which
38 * that hook never runs does nothing at all, rather than everything.
39 */
40let activation: Activation = INERT
41
42/**
43 * Whether a person is at the prompt. Only then can a compaction run, and only
44 * then is it safe to end a turn early: a `-p` run would end with it.
45 */
46let interactive = false
47
48/**
49 * The context fill at which a step was judged to move the work on to another
50 * piece. The turn is ended before its next request.
51 */
52let movedOnAt: number | undefined
53
54/** Set when taskcut itself ended the running turn in order to compact. */
55let cutting = false
56
57/**
58 * Whether the last request the main loop sent ended on tool results -- every
59 * step's but the first of a turn, whose request ends on the prompt that opened
60 * it. Claude Code builds its summary request on the last request it sent, so
61 * when that ended on the person's own words, the summary instruction reads as
62 * part of them, and the model answers them instead of summarising (2.1.280,
63 * `/compact` typed by hand included). taskcut compacts only when it did not.
64 */
65let lastSentOnResults = false
66
67/**
68 * Whether the working model has changed anything since the last compaction.
69 * Until it has, no piece can have been finished since, and no step is judged.
70 */
71let actedSince = true
72
73/**
74 * The context fill the last compaction left, until there has been one. See
75 * judgingFrom.
76 */
77let leftAt: number | undefined
78
79/** What taskcut has spent in this session: its judgements, and the compactions it started. */
80let spend: Spend = NOTHING
81
82/** The command that says so. */
83const COMMAND = 'taskcut'
84
85/** What taskcut submits to pick the work back up after it ended a turn. */
86const CONTINUE = 'Continue.'
87
88/** Compile-time guard: the literal below must stay equal to the constant. */
89const ENV_VAR_LITERAL: typeof ENV_VAR = 'TASKCUT'
90void ENV_VAR_LITERAL
91
92/** The context fill as the status line shows it. Free: read off the last response. */
93async function contextPercent($: EngineInterface): Promise<number> {
94 const { context } = await $.session.usage()
95 return context.percent ?? 0
96}
97
98/** The context fill when a judgement is due, or undefined when it is not. */
99async function pastFloor($: EngineInterface, config: Config): Promise<number | undefined> {
100 const percent = await contextPercent($)
101 return percent >= judgingFrom(config.floorPercent, leftAt) ? percent : undefined
102}
103
104/**
105 * Whether the work moves on from a finished piece to another at this moment,
106 * as the judge sees it. Anything but a clear yes keeps the context: a
107 * compaction in the middle of a piece costs re-reading, and a missed boundary
108 * only waits for the next one.
109 *
110 * Every call is counted, answered or not, and logged to the debug log with
111 * what it cost.
112 */
113async function judge($: EngineInterface, step: Step, config: Config, percent: number): Promise<boolean> {
114 const started = Date.now()
115 let verdict: string
116 let movesOn = false
117 let answer: unknown
118 try {
119 const prompt = judgePrompt(await $.session.messages(), step)
120 // `unknown`: what the call resolves to has changed between releases, and
121 // readReply takes every shape it has had.
122 answer = await $.model.complete({ model: config.model, system: JUDGE_SYSTEM, prompt, maxTokens: ANSWER_TOKENS })
123 const reply = readReply(answer)
124 spend = addJudgement(spend, readUsage(answer), 'text' in reply)
125 if ('reason' in reply) verdict = `could not judge the step (${reply.reason})`
126 else {
127 movesOn = saysNext(reply.text)
128 verdict = `step judged ${movesOn ? 'a new piece' : 'the same work'}`
129 }
130 } catch (error) {
131 // Refused before it was sent: nothing was spent.
132 verdict = `could not judge the step (${String(error)})`
133 }
134 $.ui.log(judgementLine(percent, verdict, config.model, readUsage(answer), Date.now() - started), { to: 'debug' })
135 return movesOn
136}
137
138/**
139 * Compacts, and never throws. It is unavailable in a headless session and
140 * rejects while a turn runs; either way the conversation stays as it is.
141 */
142async function compact($: EngineInterface): Promise<void> {
143 try {
144 const result = await $.session.compact()
145 if (result.skip === undefined) {
146 spend = addCompaction(spend, result.tokensBefore, result.tokensAfter)
147 actedSince = false
148 } else {
149 $.ui.log(`compaction skipped (${result.skip})`)
150 }
151 } catch (error) {
152 $.ui.log(`compaction skipped (${String(error)})`)
153 }
154 leftAt = await contextPercent($)
155}
156
157export const register: Register = (on, options) => {
158 const config = readConfig(options)
159
160 on('session.start', async ($, e, next) => {
161 const result = await next(e)
162 activation = decideActivation({
163 // Spelled out: $.env.get takes a literal name so that the loader can list
164 // every variable a module reads. ENV_VAR_LITERAL fails the build if this
165 // string and the constant ever disagree.
166 env: await $.env.get('TASKCUT'),
167 })
168 interactive = e.isInteractive
169 try {
170 // Immediate: a long unattended turn is when what taskcut has spent is
171 // worth asking, and the answer needs nothing from the turn.
172 await $.command.register({ name: COMMAND, description: 'What taskcut has judged and compacted in this session, and what the judging cost', immediate: true })
173 } catch (error) {
174 $.ui.log(`/${COMMAND} not registered (${String(error)})`, { to: 'debug' })
175 }
176 return result
177 })
178
179 on('command.run', async ($, e, next) => {
180 if (e.command !== COMMAND) return next(e)
181 const percent = activation.active && interactive ? await contextPercent($) : undefined
182 return { text: spendReport(spend, { active: activation.active, interactive, floorPercent: config.floorPercent, model: config.model, percent }) }
183 })
184
185 // Inside a turn: each step the main loop makes is judged once its response
186 // is in, while the engine runs the step's tools, and a step that moves on to
187 // another piece ends the turn before the next request goes out -- with every
188 // tool result of that step already in. The judgement is a `$` call made
189 // inside the step's own hook, so it runs beside the tools and costs the
190 // hook's budget nothing; the engine sends the next request once both are done.
191 on('turn.step', async function* ($, e, next) {
192 if (e.agentId !== undefined || !activation.active || !interactive) return yield* next(e)
193
194 if (movedOnAt !== undefined) {
195 $.ui.log(`context at ${movedOnAt}%, compacting before the next piece`)
196 movedOnAt = undefined
197 try {
198 await $.turn.abort({ turnId: e.turnId })
199 cutting = true
200 return { turnId: e.turnId, index: e.index, answer: '', toolUses: [], stopReason: null, usage: null }
201 } catch (error) {
202 $.ui.log(`compaction skipped (${String(error)})`)
203 }
204 }
205
206 lastSentOnResults = e.index > 0
207 const result = yield* next(e)
208 // A step that ends the turn hands the work back to the person.
209 const step: Step = { text: result.answer, calls: result.toolUses }
210 const acted = actedSince
211 if (acts(step.calls)) actedSince = true
212 if (result.stopReason !== 'tool_use' || !acted || !lastSentOnResults || !judgeable(step)) return result
213 const percent = await pastFloor($, config)
214 if (percent === undefined) return result
215 if (await judge($, step, config, percent)) movedOnAt = percent
216 return result
217 })
218
219 on('turn.complete', async ($, e, next) => {
220 // `next` first, so the engine has finished settling the turn before a
221 // compaction is raised against it.
222 const result = await next(e)
223 if (e.agentId !== undefined) return result
224 movedOnAt = undefined
225 if (!cutting) return result
226
227 cutting = false
228 await compact($)
229 // Whether or not it compacted, the work taskcut interrupted goes on.
230 void $.prompt.submit({ text: CONTINUE })
231 return result
232 })
233}
234hooks/activation.ts 73 lines1/**
2 * Whether taskcut does anything at all in this session.
3 *
4 * There is one rule, and Claude Code already owns most of it. Installing a
5 * plugin is where you say where it runs: `--scope user` puts it in every
6 * session on the machine, `--scope project` in one repository, `--scope local`
7 * in one repository for you alone. taskcut does not second-guess that. A
8 * session that loaded it is a session you asked for.
9 *
10 * What the platform cannot express is "not this one", so the plugin owns
11 * exactly that: `TASKCUT=0` switches it off for a single session, and
12 * `TASKCUT=1` switches it back on if some outer setting had turned it off.
13 * Nothing else. An earlier version also had a marker file and a mode setting,
14 * which between them answered the question the install scope had already
15 * answered, and the three of them had to be read together to know what would
16 * happen.
17 *
18 * Nothing here depends on `CLAUDE_CODE_ENABLE_FUNCTION_HOOKS`. That flag gates
19 * every hooks module today, but it is an early access flag: when function hooks
20 * graduate it will default to on or disappear. A plugin whose consent rested on
21 * it would become active on an unrelated Claude Code release. Consent here is
22 * the install.
23 */
24
25/**
26 * The environment variable that switches a single session off, or back on.
27 *
28 * `as const` so that the call site, which has to spell the name as a literal for
29 * the loader to list what this module reads, can assert against it and fail the
30 * build rather than drift.
31 */
32export const ENV_VAR = 'TASKCUT' as const
33
34export type ActivationInputs = {
35 /** The value of ENV_VAR, if it is set. */
36 env: string | undefined
37}
38
39export type Activation = {
40 active: boolean
41 /** Why, in a few words, for the log line and for tests. */
42 reason: string
43}
44
45/** Values of ENV_VAR that mean "off". */
46const OFF_VALUES = new Set(['0', 'off', 'false', 'no'])
47
48/** Values of ENV_VAR that mean "on". */
49const ON_VALUES = new Set(['1', 'on', 'true', 'yes'])
50
51/**
52 * Resolves activation from plain inputs, so the rule can be read and tested
53 * without a session.
54 *
55 * An unrecognised value is not an off switch. Someone who writes `TASKCUT=yes`
56 * meant yes, and someone who writes `TASKCUT=maybe` has said nothing useful --
57 * treating either as "off" would make a typo silently disable the plugin, which
58 * is the failure that is hardest to notice.
59 */
60export function decideActivation(inputs: ActivationInputs): Activation {
61 const env = inputs.env?.trim().toLowerCase()
62 if (env !== undefined && OFF_VALUES.has(env)) {
63 return { active: false, reason: `${ENV_VAR} is off` }
64 }
65 if (env !== undefined && ON_VALUES.has(env)) {
66 return { active: true, reason: `${ENV_VAR} is on` }
67 }
68 return { active: true, reason: 'the plugin is installed for this session' }
69}
70
71/** The state before `session.start` has run: inert, so a missed hook fails closed. */
72export const INERT: Activation = { active: false, reason: 'session.start has not run yet' }
73hooks/config.ts 55 lines1/** The plugin's settings, as `/plugin` collects them and `register` receives them. */
2
3import type { PluginOptions } from 'claude-code'
4
5export type Config = {
6 /** Context fill, as a percentage, below which taskcut does nothing at all. */
7 floorPercent: number
8 /** The model that judges whether a step finished a piece of the work, as a `--model` value. */
9 model: string
10}
11
12export const DEFAULTS: Config = {
13 floorPercent: 35,
14 model: 'sonnet',
15}
16
17function numberOr(value: unknown, fallback: number, min: number): number {
18 // Only a number, or a string that actually spells one. `Number(null)`,
19 // `Number('')` and `Number([])` are all 0, so a blank left in a settings file
20 // would otherwise read as a deliberate zero.
21 const parsed =
22 typeof value === 'number'
23 ? value
24 : typeof value === 'string' && value.trim() !== ''
25 ? Number(value)
26 : Number.NaN
27 if (!Number.isFinite(parsed) || parsed < min) return fallback
28 return parsed
29}
30
31/** Reads the settings, falling back to the defaults for anything unset or out of range. */
32export function readConfig(options: PluginOptions): Config {
33 return {
34 floorPercent: numberOr(options['floorPercent'], DEFAULTS.floorPercent, 0),
35 model: typeof options['model'] === 'string' && options['model'] ? options['model'] : DEFAULTS.model,
36 }
37}
38
39/**
40 * How far past the context a compaction left it the next judgement waits, when
41 * that compaction could not bring it back under the floor: so that the same
42 * finished work is not compacted again and again.
43 */
44export const REGROWTH = 5
45
46/**
47 * The context fill from which steps are judged: the floor, or -- after a
48 * compaction that left the context at or over it -- REGROWTH points past what
49 * it left. Before any compaction there is nothing it left, and the floor
50 * alone decides: a floor of 0 judges from the first step.
51 */
52export function judgingFrom(floorPercent: number, leftAt: number | undefined): number {
53 return leftAt !== undefined && leftAt >= floorPercent ? leftAt + REGROWTH : floorPercent
54}
55hooks/judge.ts 343 lines1/**
2 * What the judge reads, and how its answer is read back, as pure functions of
3 * plain data. A hooks module may only pass `$` to a function declared in the
4 * same file, never across an import, which is why the code that fetches these
5 * inputs lives in register.ts.
6 *
7 * The judge is asked one thing -- does the work move on here from a finished
8 * piece to another? -- and most of a session says nothing about it. What does,
9 * in order of how much:
10 *
11 * 1. the step itself, which says in its own words that a piece is complete and
12 * names the next;
13 * 2. what the person asked for, which says whether there is a next;
14 * 3. a trail of what the assistant has been doing -- what its recent calls
15 * touched, what it said, its task list -- which says where it has got to.
16 *
17 * So the judge reads those and nothing else: every message the person sent,
18 * the first and the latest whole and the rest cut to a line; the assistant's
19 * latest messages; what its latest calls touched, a file or what a command
20 * says it does, never the call in full; its task list as it last wrote it; and
21 * the step. Never any tool output, and not CLAUDE.md, which says how to work
22 * and not how far the work has got.
23 *
24 * Commands in full were nine tenths of what an earlier judge read, and the
25 * reason its prompt grew with the session. Without them the prompt is a few
26 * thousand tokens however long the session runs, and the judge is as right as
27 * it was: see docs/measurement.md.
28 */
29
30import type { SessionMessage } from 'claude-code'
31
32/**
33 * The two answers the judge gives, as it writes them: the work moves on from a
34 * finished piece to another, or it stays where it is.
35 */
36export const NEXT = 'NEXT'
37export const SAME = 'SAME'
38
39/**
40 * Lookups that change nothing. What the assistant read says nothing about
41 * whether a piece of work is done, and a step made only of them never follows
42 * the finishing of one.
43 */
44export const READ_ONLY_TOOLS: ReadonlySet<string> = new Set([
45 'Read',
46 'Grep',
47 'Glob',
48 'LS',
49 'NotebookRead',
50 'WebSearch',
51 'ToolSearch',
52])
53
54/** How many of each the judge reads, and how much of each. */
55export const RECENT_REQUESTS = 3
56export const REQUEST_HEAD = 1_500
57export const REQUEST_TAIL = 500
58export const REQUEST_LINE = 150
59export const SUMMARY_HEAD = 1_500
60export const SUMMARY_TAIL = 1_000
61export const RECENT_SAYINGS = 10
62export const SAYING_LIMIT = 400
63export const RECENT_ACTIONS = 12
64export const TOUCH_LIMIT = 100
65export const TASK_LIMIT = 160
66export const STEP_LIMIT = 2_000
67export const CALL_LIMIT = 300
68/** Room for one sentence and the verdict. */
69export const ANSWER_TOKENS = 300
70
71/**
72 * What is judged: a step the assistant takes inside a turn -- what it just
73 * said, and the calls it is making now.
74 */
75export type Step = { text: string; calls: readonly { name: string; input: unknown }[] }
76
77/**
78 * The question. A compaction pays for itself only when the work moves on to a
79 * piece that does not need the detail of the one before, so that is what is
80 * asked -- not whether a piece is done. The last piece of a list is done, and
81 * nothing follows it yet: whatever the person says next may well be about it.
82 *
83 * One sentence of reasoning before the verdict is what makes a small answer
84 * reliable: asked for the word alone, the judge reads "task 2 is done; now task
85 * 3" as the same work and misses the boundary it was asked to find.
86 */
87export const JUDGE_SYSTEM = [
88 'You watch an assistant working through what a person asked for. You are shown what the person',
89 "said (older messages shortened), the assistant's latest messages, what its latest commands",
90 'touched, its task list if it keeps one, and the latest step the assistant is taking: what it just',
91 'said and the calls it is making now. The output of every command is left out.',
92 '',
93 'The conversation may open with a summary Claude Code wrote when it compacted what came before: it',
94 'is a record of past work, not a request.',
95 '',
96 'Decide whether, at this step, the work moves on from a finished piece to another piece. A piece',
97 'is one of the things the person asked for -- one task in a list, one fix, one feature, one',
98 'question answered -- or one clearly separate phase of a single task.',
99 '',
100 'Apply these rules in order, and stop at the first that fits:',
101 `1. The step asks the person something, waits for their decision, or proposes work it has not`,
102 ` done: ${SAME}.`,
103 `2. The step itself says, in its own words, that a piece is complete (checked, where it could be`,
104 ` checked), and another piece the person asked for is still to do -- the step names or starts`,
105 ` it, or the person's request plainly lists more: ${NEXT}. A piece finished before this step,`,
106 ` whatever came after it, does not count.`,
107 `3. Anything else: ${SAME}. That includes the last piece asked for being complete, wrapping up`,
108 ' once every piece is done (a final check, a summary, a commit), and work on a piece not yet',
109 ' said to be complete.',
110 '',
111 `First say in one sentence what the step does. Then, on its own last line, write ${NEXT} or ${SAME}.`,
112].join('\n')
113
114function head(text: string, limit: number): string {
115 return text.length <= limit ? text : `${text.slice(0, limit)} [...]`
116}
117
118function tail(text: string, limit: number): string {
119 return text.length <= limit ? text : `[...] ${text.slice(-limit)}`
120}
121
122function ends(text: string, first: number, last: number): string {
123 return text.length <= first + last ? text : `${text.slice(0, first)} [...] ${text.slice(-last)}`
124}
125
126function oneLine(text: string): string {
127 return text.replace(/\s+/g, ' ').trim()
128}
129
130/** How Claude Code opens the message a compaction leaves in place of what it summarised. */
131const SUMMARY_OPENING = 'This session is being continued from a previous conversation'
132
133/**
134 * User messages nobody typed: a background task reporting in, a prompt a
135 * plugin submitted (taskcut's own `Continue.` among them), the record of a
136 * local command, an interruption. None of them asks for anything.
137 */
138const NOT_ASKED = [
139 /^<task-notification>/,
140 /^The \S+ plugin sent a message/,
141 /^<local-command-/,
142 /^Caveat: The messages below were generated by the user while running local commands/,
143 /^\[Request interrupted by user/,
144]
145
146/** Whether a user message is the person asking for something. */
147export function isRequest(message: SessionMessage): boolean {
148 if (message.role !== 'user' || (message.toolResults?.length ?? 0) > 0) return false
149 const text = message.text.trim()
150 return text !== '' && !NOT_ASKED.some((pattern) => pattern.test(text))
151}
152
153/**
154 * What a call touched: the file it read or wrote, or what a command says it
155 * does. Enough to tell one piece of work from the next; the call in full is
156 * what made an earlier judge's prompt grow with the session.
157 */
158export function touched(tool: string, input: unknown): string {
159 const fields = (typeof input === 'object' && input !== null ? input : {}) as Record<string, unknown>
160 if (typeof fields.file_path === 'string') return `${tool} ${fields.file_path}`
161 if (tool === 'Bash') return `Bash: ${head(oneLine(String(fields.description || fields.command || '')), TOUCH_LIMIT)}`
162 return `${tool}: ${head(oneLine(JSON.stringify(input) ?? ''), TOUCH_LIMIT)}`
163}
164
165/**
166 * The assistant's task list as it last wrote it, through `TodoWrite` or the
167 * `TaskCreate` and `TaskUpdate` tools, the step's own calls included. Empty
168 * when it keeps none. `TaskCreate` numbers its tasks from 1, in order.
169 */
170export function taskList(messages: readonly SessionMessage[], step: Step): string[] {
171 const calls = [
172 ...messages.flatMap((message) => (message.role === 'assistant' ? message.toolUses.map((use) => ({ name: use.tool, input: use.input })) : [])),
173 ...step.calls,
174 ]
175 let todos: { content: string; status: string }[] | undefined
176 const tasks = new Map<string, { subject: string; status: string }>()
177 for (const { name, input } of calls) {
178 const fields = (typeof input === 'object' && input !== null ? input : {}) as Record<string, unknown>
179 if (name === 'TodoWrite' && Array.isArray(fields.todos)) {
180 todos = (fields.todos as Record<string, unknown>[]).map((todo) => ({ content: String(todo.content ?? ''), status: String(todo.status ?? '') }))
181 } else if (name === 'TaskCreate' && typeof fields.subject === 'string') {
182 tasks.set(String(tasks.size + 1), { subject: fields.subject, status: 'pending' })
183 } else if (name === 'TaskUpdate' && typeof fields.status === 'string') {
184 const task = tasks.get(String(fields.taskId))
185 if (task) task.status = fields.status
186 }
187 }
188 const mark = (status: string) => (status === 'completed' ? '[done]' : status === 'in_progress' ? '[doing]' : '[todo]')
189 const items = todos?.map((todo) => ({ subject: todo.content, status: todo.status })) ?? [...tasks.values()].filter((task) => task.status !== 'deleted')
190 return items.map((item) => `${mark(item.status)} ${head(oneLine(item.subject), TASK_LIMIT)}`)
191}
192
193/**
194 * The messages before the step, without the step.
195 *
196 * When the step is judged, `$.session.messages()` (2.1.280) already holds it,
197 * and as more than one message: a response is recorded a block at a time, so
198 * its words are one message and each call another -- and a call that has
199 * already run is followed by its result. Left in, the step would be read
200 * twice, once as the latest thing the assistant said, and whether it was
201 * would depend on how fast its tools ran.
202 *
203 * So the step is the shortest run of messages at the end whose words, joined,
204 * are the step's and whose calls are the step's, with nothing between them
205 * but the results of those calls. When there is none -- the engine does not
206 * hold the step yet -- every message is before it.
207 */
208export function beforeStep(messages: readonly SessionMessage[], step: Step): readonly SessionMessage[] {
209 const words = oneLine(step.text)
210 const calls = step.calls.map((call) => `${call.name} ${JSON.stringify(call.input)}`)
211 const said: string[] = []
212 const made: string[] = []
213 const madeIds = new Set<string>()
214 const resultIds: string[] = []
215 for (let i = messages.length - 1; i >= 0; i--) {
216 const message = messages[i]!
217 if (message.role === 'user') {
218 if (!message.toolResults?.length) return messages
219 resultIds.push(...message.toolResults.map((result) => result.tool_use_id))
220 continue
221 }
222 said.unshift(message.text)
223 made.unshift(...message.toolUses.map((use) => `${use.tool} ${JSON.stringify(use.input)}`))
224 for (const use of message.toolUses) madeIds.add(use.tool_use_id)
225 if (made.length > calls.length) return messages
226 const isStep = oneLine(said.join(' ')) === words && made.length === calls.length && made.every((call, k) => call === calls[k])
227 // Nothing may come between the step's messages but its own calls' results.
228 if (isStep) return resultIds.every((id) => madeIds.has(id)) ? messages.slice(0, i) : messages
229 }
230 return messages
231}
232
233/**
234 * The conversation as the judge reads it, oldest first: every request, the
235 * first and the latest few whole; the assistant's latest messages; what its
236 * latest calls other than lookups touched. Nothing else, and no output.
237 */
238export function conversationLines(messages: readonly SessionMessage[]): string[] {
239 const indexed = messages.map((message, i) => ({ message, i }))
240 const requests = indexed.filter(({ message }) => isRequest(message)).map(({ i }) => i)
241 const asked = new Set(requests)
242 const whole = new Set([...requests.slice(0, 1), ...requests.slice(-RECENT_REQUESTS)])
243 const sayings = new Set(indexed.filter(({ message }) => message.role === 'assistant' && message.text.trim()).map(({ i }) => i).slice(-RECENT_SAYINGS))
244 const actions = new Set(
245 indexed.filter(({ message }) => message.role === 'assistant' && message.toolUses.some((use) => !READ_ONLY_TOOLS.has(use.tool))).map(({ i }) => i).slice(-RECENT_ACTIONS),
246 )
247
248 const lines: string[] = []
249 let unsaid = 0
250 const skipped = () => {
251 if (unsaid > 0) lines.push(`(${unsaid} earlier assistant message${unsaid > 1 ? 's' : ''} left out)`)
252 unsaid = 0
253 }
254 for (const [i, message] of messages.entries()) {
255 if (message.role === 'user') {
256 if (!asked.has(i)) continue
257 skipped()
258 const text = message.text.trim()
259 if (text.startsWith(SUMMARY_OPENING)) lines.push(`Summary of the conversation before it was compacted: ${ends(text, SUMMARY_HEAD, SUMMARY_TAIL)}`)
260 else if (whole.has(i)) lines.push(`Person: ${ends(text, REQUEST_HEAD, REQUEST_TAIL)}`)
261 else lines.push(`Person (shortened): ${head(oneLine(text), REQUEST_LINE)}`)
262 continue
263 }
264 if (sayings.has(i)) {
265 skipped()
266 lines.push(`Assistant: ${head(message.text.trim(), SAYING_LIMIT)}`)
267 } else if (message.text.trim()) {
268 unsaid++
269 }
270 if (actions.has(i)) {
271 skipped()
272 for (const use of message.toolUses) if (!READ_ONLY_TOOLS.has(use.tool)) lines.push(`Assistant ran ${touched(use.tool, use.input)}`)
273 }
274 }
275 skipped()
276 return lines
277}
278
279/**
280 * Whether the step is worth asking about. A step says a piece is complete or
281 * it is not a boundary, so a step that says nothing is not one -- and after a
282 * compaction, whose summary reads as a message saying which pieces are done, a
283 * silent first step would otherwise look like the move to the next.
284 */
285export function judgeable(step: Step): boolean {
286 return step.text.trim() !== ''
287}
288
289/**
290 * Whether calls change anything. Until a step after a compaction has, nothing
291 * can have been finished since it -- and the summary, which says which pieces
292 * are done, makes the first step of the next one look like the move to it.
293 */
294export function acts(calls: readonly { name: string }[]): boolean {
295 return calls.some((call) => !READ_ONLY_TOOLS.has(call.name))
296}
297
298/** Everything the judge is shown: the conversation, the task list, and the step last. */
299export function judgePrompt(messages: readonly SessionMessage[], step: Step): string {
300 const before = beforeStep(messages, step)
301 const tasks = taskList(before, step)
302 return [
303 '--- the conversation, oldest first (recent commands shortened, output left out) ---',
304 ...conversationLines(before),
305 ...(tasks.length > 0 ? ['', "--- the assistant's task list, as it last wrote it ---", ...tasks] : []),
306 '',
307 '--- latest step: what the assistant is doing now ---',
308 tail(step.text.trim(), STEP_LIMIT) || '(no text)',
309 ...step.calls.map((call) => `Calls ${call.name}: ${head(JSON.stringify(call.input) ?? '', CALL_LIMIT)}`),
310 ].join('\n')
311}
312
313/**
314 * The judge's reply, whichever shape `$.model.complete` resolved: the reply's
315 * text up to Claude Code 2.1.278, `{ isAnswered, text }` or
316 * `{ isAnswered: false, reason }` from 2.1.280. Either is read here, where the
317 * engine's answer enters, and nothing past this point knows there were two.
318 */
319export type Reply = { text: string } | { reason: string }
320
321export function readReply(answer: unknown): Reply {
322 if (typeof answer === 'string') return { text: answer }
323 if (typeof answer === 'object' && answer !== null && 'isAnswered' in answer) {
324 const { isAnswered, text, reason, status, error } = answer as { isAnswered: unknown; text?: unknown; reason?: unknown; status?: unknown; error?: unknown }
325 if (isAnswered === true && typeof text === 'string') return { text }
326 // An API error says which: a spent rate limit and a refused request
327 // otherwise read the same.
328 if (isAnswered === false) return { reason: [reason, status, error].filter((part) => part !== undefined && part !== null).join(' ') }
329 }
330 return { reason: 'a reply of no shape taskcut knows' }
331}
332
333/**
334 * Whether the judge's answer says the work moves on: its last line, stripped
335 * of the emphasis a model sometimes puts round it. Anything else, including no
336 * answer, is not.
337 */
338export function saysNext(answer: string): boolean {
339 const lines = answer.trim().split('\n')
340 const last = lines[lines.length - 1] ?? ''
341 return last.replace(/[^A-Za-z]/g, '').toUpperCase() === NEXT
342}
343hooks/spend.ts 160 lines1/**
2 * What taskcut has spent in a session, as pure functions of plain data.
3 *
4 * The judge's calls are taskcut's own cost, and nothing else counts them:
5 * `$.model.complete` is not on the session's cost ledger, so `/cost` and the
6 * status line leave them out, and they are not in the transcript. Each call
7 * resolves the API's own token counts, so the tally is exact and free to keep:
8 * a sum of numbers already in hand, held in memory, never written anywhere.
9 *
10 * It is shown two ways. `/taskcut` gives the session's totals. The debug log
11 * (`claude --debug`, or `--debug-file`) gets one line per judgement, in a
12 * format kept stable so that a harness can add the lines up: see
13 * `judgementLine`.
14 *
15 * Tokens, not dollars: a price is the host's to know -- it differs by model,
16 * by plan and by contract -- and a price list in a plugin would be wrong the
17 * day it changed.
18 */
19
20/** The four token counts a call reports, as the API does. */
21export type Usage = {
22 input_tokens: number
23 output_tokens: number
24 cache_read_input_tokens: number
25 cache_creation_input_tokens: number
26}
27
28export type Spend = {
29 /** Calls made to the judge. */
30 judgements: number
31 /** Of those, the calls that came back without an answer. */
32 unanswered: number
33 /** Of those, the calls whose cost the engine did not report (before 2.1.280). */
34 unmetered: number
35 input: number
36 cacheRead: number
37 cacheWrite: number
38 output: number
39 /** Compactions taskcut started, and the context before and after them, summed. */
40 compactions: number
41 compactedFrom: number
42 compactedTo: number
43}
44
45export const NOTHING: Spend = {
46 judgements: 0,
47 unanswered: 0,
48 unmetered: 0,
49 input: 0,
50 cacheRead: 0,
51 cacheWrite: 0,
52 output: 0,
53 compactions: 0,
54 compactedFrom: 0,
55 compactedTo: 0,
56}
57
58function count(value: unknown): number | undefined {
59 return typeof value === 'number' && Number.isFinite(value) && value >= 0 ? value : undefined
60}
61
62/**
63 * The token counts a model call resolved with, whichever shape it had; none
64 * when it reported none, as a reply that was only text did up to 2.1.278.
65 */
66export function readUsage(answer: unknown): Usage | undefined {
67 if (typeof answer !== 'object' || answer === null || !('usage' in answer)) return undefined
68 const usage = (answer as { usage: unknown }).usage
69 if (typeof usage !== 'object' || usage === null) return undefined
70 const fields = usage as Record<string, unknown>
71 const input = count(fields.input_tokens)
72 const output = count(fields.output_tokens)
73 if (input === undefined || output === undefined) return undefined
74 return {
75 input_tokens: input,
76 output_tokens: output,
77 cache_read_input_tokens: count(fields.cache_read_input_tokens) ?? 0,
78 cache_creation_input_tokens: count(fields.cache_creation_input_tokens) ?? 0,
79 }
80}
81
82/** The tally with one more judgement in it. */
83export function addJudgement(spend: Spend, usage: Usage | undefined, answered: boolean): Spend {
84 return {
85 ...spend,
86 judgements: spend.judgements + 1,
87 unanswered: spend.unanswered + (answered ? 0 : 1),
88 unmetered: spend.unmetered + (usage ? 0 : 1),
89 input: spend.input + (usage?.input_tokens ?? 0),
90 cacheRead: spend.cacheRead + (usage?.cache_read_input_tokens ?? 0),
91 cacheWrite: spend.cacheWrite + (usage?.cache_creation_input_tokens ?? 0),
92 output: spend.output + (usage?.output_tokens ?? 0),
93 }
94}
95
96/** The tally with one more compaction in it, and the context either side of it when the engine said. */
97export function addCompaction(spend: Spend, before: number | undefined, after: number | undefined): Spend {
98 const known = count(before) !== undefined && count(after) !== undefined
99 return {
100 ...spend,
101 compactions: spend.compactions + 1,
102 compactedFrom: spend.compactedFrom + (known ? before! : 0),
103 compactedTo: spend.compactedTo + (known ? after! : 0),
104 }
105}
106
107/**
108 * The debug line for one judgement:
109 *
110 * context at 41%, step judged the same work [judge sonnet: in=1834 cache_read=0 cache_write=0 out=52 ms=2140]
111 *
112 * The bracket is the record: the model asked, the four token counts and how
113 * long the call took. It is left out when the engine reported no counts. Its
114 * format is part of what taskcut promises, because the benchmark reads it.
115 */
116export function judgementLine(percent: number, verdict: string, model: string, usage: Usage | undefined, ms: number): string {
117 const record = usage
118 ? ` [judge ${model}: in=${usage.input_tokens} cache_read=${usage.cache_read_input_tokens} cache_write=${usage.cache_creation_input_tokens} out=${usage.output_tokens} ms=${Math.round(ms)}]`
119 : ''
120 return `context at ${percent}%, ${verdict}${record}`
121}
122
123function thousands(n: number): string {
124 return Math.round(n).toLocaleString('en-US')
125}
126
127function plural(n: number, one: string, many = `${one}s`): string {
128 return `${thousands(n)} ${n === 1 ? one : many}`
129}
130
131/** What `/taskcut` says: whether taskcut is at work in this session, and what it has spent. */
132export function spendReport(
133 spend: Spend,
134 state: { active: boolean; interactive: boolean; floorPercent: number; model: string; percent: number | undefined },
135): string {
136 const lines: string[] = []
137 if (!state.active) lines.push('taskcut is off in this session (TASKCUT=0).')
138 else if (!state.interactive) lines.push('taskcut does nothing in a session nobody is at: it cannot compact there.')
139 else {
140 const now = state.percent === undefined ? '' : ` The context is at ${state.percent}%.`
141 lines.push(`taskcut is on: past ${state.floorPercent}% of the context, ${state.model} judges each step.${now}`)
142 }
143
144 if (spend.judgements === 0) {
145 lines.push('Nothing judged yet, so nothing spent.')
146 return lines.join('\n')
147 }
148 const read = spend.input + spend.cacheRead + spend.cacheWrite
149 const cached = spend.cacheRead > 0 ? ` (${thousands(spend.cacheRead)} from the cache)` : ''
150 const failed = spend.unanswered > 0 ? `, ${thousands(spend.unanswered)} of them unanswered` : ''
151 lines.push(`Judged ${plural(spend.judgements, 'step')}${failed}: ${thousands(read)} tokens read${cached}, ${thousands(spend.output)} written.`)
152 if (spend.unmetered > 0) lines.push(`${plural(spend.unmetered, 'call')} reported no token counts, and ${spend.unmetered === 1 ? 'is' : 'are'} not in these.`)
153 if (spend.compactions > 0) {
154 const sizes = spend.compactedFrom > 0 ? `, ${thousands(spend.compactedFrom)} tokens of context down to ${thousands(spend.compactedTo)}` : ''
155 lines.push(`Compacted ${plural(spend.compactions, 'time')}${sizes}.`)
156 }
157 lines.push("/cost does not count the judge's calls: they are all here. It does count the compactions, which are Claude Code's own.")
158 return lines.join('\n')
159}
160