Has a small model rate how hard each prompt is and runs that turn and its subagents at the matching effort, within the levels you allow, on the models whose…

One effort setting never fits a whole session: max wastes minutes and tokens on "thanks", low rushes an architecture question. This mod has a small model rate how hard each of your prompts is, and runs that turn, with the subagents it starts, at the matching effort. You can limit the levels a rated turn runs at. The next turn starts again from the session's own effort, so nothing has to be set back.
low, medium, high, xhigh or max. It reads the first 4,000 characters of the prompt. Measured on 2.1.284: an answer in about 0.6 seconds.named and the level, and the mod keeps that answer only when the prompt holds the level's word, because haiku also answered named for prompts that only called the work hard. Measured on 2.1.284, three runs per prompt: eight prompts that name a level got it 24 of 24 times; eight that call the work hard or easy, or use a level word in another sense ("a medium sized image", "the max value of the counter"), were rated 24 of 24 times. Haiku answered named for "yüksek effort" and "maximum effort" 6 of 6 times, and the word check read those answers as ratings./effort-auto levels low medium max limits the levels a rated turn runs at. A rating outside them moves to the nearest allowed level, the higher of two equally near ones: with low medium max, a high rating runs at medium and an xhigh rating at max. A level the prompt names runs even when it is not allowed. Haiku still rates on the whole scale and is not told the allowed levels: told only low and medium, it answered named max for "the max value of the counter overflows at high load; find why" 3 of 3 times, while with the whole scale it rated the same prompt 6 of 6 times. /effort-auto levels all lifts the limit, which is the default.low, green for medium, yellow for high and red for xhigh and max, the same colours session-watch gives the effort. While a turn runs, the first line shows that turn's effort beside the session's own; between turns, the effort of the last turn that ended. The second line shows the allowed levels:this turn high · session low allowed low · medium · high · xhigh · max
last turn max (named) · session low allowed low · medium
A level the prompt named has a faint (named) after it, and a turn that is not rated shows the session's effort with a faint (session) after it. /effort-auto off takes the section down. With the sidebar closed, a rated turn writes one transcript line at its start:
effort-auto: this turn max (named) · session low
In the live check on Sonnet 5.5 with levels low medium and the session at low, a prompt that called the work hard and asked for an Explore subagent ran at medium, and so did the subagent's requests. "bunu max ile çöz" with a general-purpose subagent ran at max, and so did the subagent's requests.
An effort change can rewrite the whole prompt cache, and rewriting a long conversation costs more than the turn gains. So the mod changes the effort only on the models that keep the cache across an effort change: Opus 5.5, Sonnet 5.5 and Fable 5.1. On every other model it changes nothing, and once the session's first request names such a model, no prompt is rated at all. A subagent's request is checked by its own model the same way. In a live check on Sonnet 5, every request went out at the session's effort, a subagent's too.
Measured on Claude Code 2.1.283, with the same conversation and an effort change from one turn to the next:
Opus 5.5 high → low, low → max cache read 58,408 each time, as at the same effort Sonnet 5 high → low, low → max cache read 0, the whole conversation of about 73,700 tokens written again
Measured on Claude Code 2.1.284 the same way, a medium turn first as the control:
Sonnet 5.5 medium → high, high → low cache read 58,417 each time and about 5,500 written, as at the same effort Sonnet 5 medium → high cache read 0, about 73,700 tokens written again
In the live check a greeting was rated low and a design question max. The max turn read the cache the low turn had written (74,079 tokens), and the next turn, back at low, read 91,755.
/effort-auto on or off, and the allowed levels /effort-auto on | off on by default /effort-auto levels the allowed levels /effort-auto levels <level ...> allow only these, in any order: levels low medium max /effort-auto levels all allow every level, the default
claude plugin marketplace add KilimcininKorOglu/claude-code-mods claude plugin install effort-auto@kilimcininkoroglu-mods
Function hooks are early access. Claude Code 2.1.288 and later load them by default, so there is nothing to switch on.
Validated with claude plugin validate on Claude Code 2.1.284:
❯ ./register.ts hooks: session.start, command.run{command=effort-auto}, prompt.submit, turn.step, agent.spawn, turn.complete ❯ ./register.ts calls: $.command.register, $.model.complete (via rate), $.sidebar.clear (via dropLine), $.sidebar.set (via toPerson), $.store.get (via readSettings), $.store.set (via setLevels, switchTo), $.ui.log (via rate, toPerson)
Reach L3: one model request per prompt.
low even in the middle of hard work.max turn thinks much longer and costs more. In the live check a five-point design summary took 7 minutes and 42,683 output tokens.CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS, the Claude Code docs say an effort change still rewrites the cache on every model.make install # eslint, typescript-eslint, typescript make lint # complexity limit 10, the build fails above it make typecheck # needs .claude/types/ from /plugin-types make validate make test # claude plugin test
hooks/register.ts 245 lines1import type { EngineInterface, Register, TurnStepInput } from 'claude-code'
2import { allowedLine, effortLine, keepsCacheAcrossEffort, levelFor, levelsOf, parseLevels, RATE_TIMEOUT_MS, RATER_MODEL, RATER_SYSTEM, raterPrompt, ratingOf, type Level, type Line, type Rating, type TurnEffort } from './effort.ts'
3
4const ENABLED_KEY = 'enabled'
5const LEVELS_KEY = 'levels'
6
7const USAGE = 'expects nothing (the status), on, off, or levels all | <level ...>'
8
9/** A subagent started in a rated turn: that turn's level, and whether its first request was checked. */
10type Pin = { level: Level; checked: boolean }
11
12/**
13 * The on/off setting and the allowed levels as the store held them at the last read, the rating of the
14 * person's last prompt, the effort of the running and of the last ended main-loop turn, the session's
15 * effort as the last main-loop request carried it, whether the sidebar holds the section now, the model
16 * of the last main-loop request, unknown until the first one, and the level of each subagent a rated
17 * turn started.
18 */
19type State = {
20 enabled: boolean
21 allowed: Level[]
22 rating?: Rating
23 current?: TurnEffort
24 last?: TurnEffort
25 session?: string | number
26 inSidebar: boolean
27 model?: string
28 pins: Map<string, Pin>
29}
30
31const SECTION = { consumer: 'effort-auto', key: 'effort' }
32
33/**
34 * The person's section: standing in the sidebar while the pane is open, else one transcript line when
35 * `log` is given, which only a rated turn's start gives, so a closed pane gets no line per turn.
36 */
37async function toPerson($: EngineInterface, state: State, lines: Line[], log: string | undefined): Promise<void> {
38 try {
39 const taken = await $.sidebar.set({ ...SECTION, title: 'effort', lines, until: 'session', order: 6 })
40 state.inSidebar = taken
41 if (taken) return
42 } catch {
43 // The sidebar mod is not installed.
44 }
45 if (log !== undefined) $.ui.log(log)
46}
47
48/** The effort line once a request told the session's effort, and the allowed levels under it. */
49function sectionLines(state: State): Line[] {
50 const known = state.current !== undefined || state.last !== undefined || state.session !== undefined
51 const allowed = allowedLine(state.allowed)
52 return known ? [effortLine(state.current, state.last, state.session), allowed] : [allowed]
53}
54
55/** Draws the section as the state holds it now; `logs` writes the effort line when the pane is closed. */
56async function drawLine($: EngineInterface, state: State, logs: boolean): Promise<void> {
57 const lines = sectionLines(state)
58 await toPerson($, state, lines, logs ? lines[0].text : undefined)
59}
60
61/** Drops the section, when the mod is turned off. */
62async function dropLine($: EngineInterface, state: State): Promise<void> {
63 if (!state.inSidebar) return
64 state.inSidebar = false
65 try {
66 await $.sidebar.clear(SECTION)
67 } catch {
68 // The sidebar mod went away since the section was drawn; there is nothing to drop.
69 }
70}
71
72/** Forgets the rating, the running turn and every subagent's level, and drops the section. */
73async function turnOff($: EngineInterface, state: State): Promise<void> {
74 state.rating = undefined
75 state.current = undefined
76 state.pins.clear()
77 await dropLine($, state)
78}
79
80/**
81 * Reads the on/off setting and the allowed levels from the store, which every window shares, so a change
82 * made in another window applies here at the next hook that acts on it. A mod turned off there drops its
83 * rating and its section, as `off` does; new allowed levels redraw the section.
84 */
85async function readSettings($: EngineInterface, state: State): Promise<void> {
86 const was = state.enabled
87 const allowed = levelsOf(await $.store.get(LEVELS_KEY))
88 const changed = allowed.join() !== state.allowed.join()
89 state.allowed = allowed
90 state.enabled = (await $.store.get(ENABLED_KEY)) !== false
91 if (!state.enabled && was) return turnOff($, state)
92 if (changed && state.enabled && state.inSidebar) await drawLine($, state, false)
93}
94
95/**
96 * The level a prompt's turn runs at and whether the prompt named it, or undefined when the rater did not
97 * answer with a level. The rater rates on the whole scale; a rated level the person does not allow is
98 * moved to the nearest allowed one, and so is a level the rater calls named that the prompt does not hold.
99 */
100async function rate($: EngineInterface, text: string, allowed: Level[]): Promise<Rating | undefined> {
101 const r = await $.model.complete({ model: RATER_MODEL, system: RATER_SYSTEM, prompt: raterPrompt(text), effort: 'low', maxTokens: 16, timeoutMs: RATE_TIMEOUT_MS })
102 if (!r.isAnswered) {
103 $.ui.log(`the rater did not answer (${r.reason}), so this turn keeps the session's effort`)
104 return undefined
105 }
106 const rating = ratingOf(r.text, text)
107 if (rating === undefined) {
108 $.ui.log(`the rater answered "${r.text.slice(0, 40)}", so this turn keeps the session's effort`)
109 return undefined
110 }
111 return { level: levelFor(rating, allowed), named: rating.named }
112}
113
114async function switchTo($: EngineInterface, state: State, on: boolean): Promise<string> {
115 await $.store.set(ENABLED_KEY, on)
116 state.enabled = on
117 if (!on) {
118 await turnOff($, state)
119 return "off: every turn runs at the session's effort"
120 }
121 state.rating = undefined
122 state.current = undefined
123 return 'on: each prompt is rated and its turn runs at that effort'
124}
125
126/** `levels` shows the allowed levels; `levels all` or `levels <level ...>` stores them. */
127async function setLevels($: EngineInterface, state: State, args: string): Promise<string> {
128 if (args === '') {
129 await readSettings($, state)
130 return `allowed ${state.allowed.join(', ')}`
131 }
132 const levels = parseLevels(args)
133 if (typeof levels === 'string') return levels
134 await $.store.set(LEVELS_KEY, levels)
135 state.allowed = levels
136 if (state.enabled) await drawLine($, state, false)
137 return `allowed ${levels.join(', ')}: a rated turn runs at the nearest of these, and a level a prompt names still applies`
138}
139
140async function runCommand($: EngineInterface, state: State, args: string): Promise<string> {
141 const text = args.trim()
142 if (text === 'on' || text === 'off') return switchTo($, state, text === 'on')
143 const [word, ...rest] = text.split(/\s+/)
144 if (word === 'levels') return setLevels($, state, rest.join(' '))
145 if (text !== '') return USAGE
146 await readSettings($, state)
147 return state.enabled ? `on · allowed ${state.allowed.join(', ')}` : 'off'
148}
149
150/** A rated main-loop turn's level, which the subagents it starts run at. */
151const ratedLevel = (t: TurnEffort | undefined): Level | undefined => (t?.rated === true ? t.level : undefined)
152
153/** Pins a started subagent to the level its parent runs at: the rated main-loop turn's, or its parent subagent's. */
154function pinAgent(state: State, agentId: string, parent: string | undefined): void {
155 const level = parent === undefined ? ratedLevel(state.current) : state.pins.get(parent)?.level
156 if (level !== undefined) state.pins.set(agentId, { level, checked: false })
157}
158
159/**
160 * A subagent's request goes out at its pinned level, on a model whose cache survives the change. Measured
161 * on 2.1.284: a subagent's requests carry the session's effort. So one whose first request carries another
162 * has its own effort in its definition, keeps it, and loses the pin.
163 */
164function subagentStep(state: State, e: TurnStepInput, agentId: string): TurnStepInput {
165 const pin = state.pins.get(agentId)
166 if (pin === undefined || !keepsCacheAcrossEffort(e.model)) return e
167 if (!pin.checked) {
168 pin.checked = true
169 if (e.effort !== state.session) {
170 state.pins.delete(agentId)
171 return e
172 }
173 }
174 return { ...e, effort: pin.level }
175}
176
177export const register: Register = on => {
178 const state: State = { enabled: true, allowed: levelsOf(undefined), inSidebar: false, pins: new Map() }
179
180 on('session.start', async ($, e, next) => {
181 const r = await next(e)
182 await $.command.register({ name: 'effort-auto', description: "Run each turn at the effort its prompt needs: status, on, off, levels (effort-auto)", argumentHint: '[on | off | levels all | levels <level ...>]' })
183 await readSettings($, state)
184 return r
185 })
186
187 // The engine prints the plugin name in front of command text, so the texts do not repeat it.
188 on('command.run', { command: 'effort-auto' }, async ($, e) => ({ text: await runCommand($, state, String(e.args ?? '')) }))
189
190 // Only the person's own prompt, typed while the session is idle, is rated: a prompt typed over a running
191 // turn, a task notification or a plugin's prompt leaves the next turn at the session's effort. So is a
192 // prompt of a session whose model would lose its cache, as the last main-loop request named it. Every
193 // prompt reads the settings, so the turn it starts runs as the store says; turn.step, which runs at each
194 // model request, does not read the store.
195 on('prompt.submit', async ($, e, next) => {
196 await readSettings($, state)
197 const typed = e.origin.kind === 'composer' || e.origin.kind === 'bridge'
198 const cacheSafe = state.model === undefined || keepsCacheAcrossEffort(state.model)
199 if (state.enabled && typed && e.turnId === undefined && cacheSafe) {
200 state.rating = await rate($, e.text, state.allowed)
201 }
202 return next(e)
203 })
204
205 // A main-loop request of a rated turn runs at the rated effort, on a model whose cache survives the change.
206 // Nothing is stored, so the next turn starts from the session's own effort. A turn's first request draws
207 // the section, a turn that was not rated with the session's effort. A subagent's request runs at the
208 // level of the rated turn that started it.
209 on('turn.step', async function* ($, e, next) {
210 if (!state.enabled) return yield* next(e)
211 if (e.agentId !== undefined) return yield* next(subagentStep(state, e, e.agentId))
212 state.model = e.model
213 state.session = e.effort
214 const rating = keepsCacheAcrossEffort(e.model) ? state.rating : undefined
215 if (state.current === undefined) {
216 state.current = rating === undefined ? { rated: false, value: e.effort } : { rated: true, level: rating.level, named: rating.named }
217 await drawLine($, state, rating !== undefined)
218 }
219 return yield* next(rating === undefined ? e : { ...e, effort: rating.level })
220 })
221
222 // A subagent started while a rated turn runs, directly or by one of its subagents, takes that level.
223 on('agent.spawn', async (_, e, next) => {
224 const r = await next(e)
225 if (state.enabled && r.agentId !== undefined) pinAgent(state, r.agentId, e.parentAgentId)
226 return r
227 })
228
229 // The rated effort ends with the turn: the section keeps it as the last turn's, beside the session's
230 // effort. A subagent's level ends with its own run.
231 on('turn.complete', async ($, e, next) => {
232 if (e.agentId !== undefined) {
233 state.pins.delete(e.agentId)
234 return next(e)
235 }
236 state.rating = undefined
237 if (state.current !== undefined) {
238 state.last = state.current
239 state.current = undefined
240 await drawLine($, state, false)
241 }
242 return next(e)
243 })
244}
245hooks/effort.ts 147 lines1/**
2 * How hard a prompt is, as a small model rates it, the levels the person allows a turn to run at, and the
3 * models whose prompt cache survives an effort change.
4 */
5
6export const LEVELS = ['low', 'medium', 'high', 'xhigh', 'max'] as const
7export type Level = (typeof LEVELS)[number]
8
9/** The model that rates each prompt, asked for one word at its lowest effort. */
10export const RATER_MODEL = 'haiku'
11
12/** How long the rating may take before the turn runs at the session's own effort. */
13export const RATE_TIMEOUT_MS = 8000
14
15/** How much of the prompt the rater reads. */
16const MAX_PROMPT_CHARS = 4000
17
18/**
19 * The models whose prompt cache survives an effort change. Measured on 2.1.283: on Opus 5.5 an effort
20 * change kept the cache read whole; on Sonnet 5 it rewrote the whole conversation. Measured on 2.1.284:
21 * Sonnet 5.5 keeps it as Opus 5.5 does. Fable 5.1 keeps it since 2.1.260, as the Claude Code docs say.
22 */
23const CACHE_SAFE = /opus-5-5|fable-5-1|sonnet-5-5/
24
25export function keepsCacheAcrossEffort(model: string): boolean {
26 return CACHE_SAFE.test(model)
27}
28
29/** What each level is for, as the rater reads it. */
30const MEANING: Record<Level, string> = {
31 low: 'a greeting, a thank-you, a yes or no, a lookup, a one-line change.',
32 medium: 'a small edit, a clear question about code, running a known command.',
33 high: 'a change across several files, or a bug whose place is known.',
34 xhigh: 'a bug whose cause is unknown, a feature with design choices, a review.',
35 max: 'an architecture decision, a subtle concurrency or security bug, work that must be right the first time.',
36}
37
38/** `low, medium, high, xhigh or max`. */
39const SCALE = `${LEVELS.slice(0, -1).join(', ')} or ${LEVELS[LEVELS.length - 1]}`
40
41/**
42 * The rater's instructions, over the whole scale: the allowed levels stay out of them, because a rater
43 * told only low and medium answered `named max` for a request that used "max" in another sense (measured
44 * on haiku), and the mod moves a rating into the allowed levels itself. A level the request asks for by
45 * its word comes back as `named <level>`; how hard the request calls the work names no level.
46 */
47export const RATER_SYSTEM = [
48 'You rate how much reasoning a coding agent needs for the request it just received.',
49 'The levels, from least to most reasoning:',
50 ...LEVELS.map(l => `${l}: ${MEANING[l]}`),
51 `A request names a level only when it holds one of the words ${SCALE} and asks for that level as the effort, reasoning or thinking to spend, as in "use max effort" or "bunu max ile çöz"; the rest of the request may be in any language. Calling the work hard, easy, simple or important, or asking to think carefully, to take time or to be thorough, names no level: rate that request. A level word used for something else, as in "a medium sized image" or "high traffic", names no level either.`,
52 'When the request names a level, answer "named" and that word, for example "named high".',
53 `Otherwise answer one word: ${SCALE}.`,
54 'Answer with those words alone, whatever language the request is in: no sentence, no translation, no explanation.',
55].join('\n')
56
57/** The rater's prompt: the person's request, cut to what the rater needs to judge it. */
58export function raterPrompt(text: string): string {
59 const cut = text.length > MAX_PROMPT_CHARS ? `${text.slice(0, MAX_PROMPT_CHARS)}\n[cut]` : text
60 return `The request:\n<request>\n${cut}\n</request>\nWhen the request names a level, answer "named" and that word.\nOtherwise answer ${SCALE}.\nThe level:`
61}
62
63/** What the rater answered: a level, and whether the request named it itself. */
64export type Rating = { level: Level; named: boolean }
65
66/**
67 * The rating a rater's reply holds, or undefined when it names no level. A `named` answer stands only
68 * when the request holds that level's word; else the level is read as a rating, because the rater also
69 * answers `named` for a request that only calls the work hard.
70 */
71export function ratingOf(reply: string, request: string): Rating | undefined {
72 const words = reply.toLowerCase().match(/[a-z]+/g) ?? []
73 const named = words[0] === 'named'
74 const level = LEVELS.find(l => l === words[named ? 1 : 0])
75 if (level === undefined) return undefined
76 return { level, named: named && new RegExp(`\\b${level}\\b`, 'i').test(request) }
77}
78
79/**
80 * The level a turn runs at: a level the request named as it is, a rated one moved to the nearest allowed
81 * level, the higher of two equally near ones.
82 */
83export function levelFor(rating: Rating, allowed: readonly Level[]): Level {
84 if (rating.named || allowed.includes(rating.level)) return rating.level
85 const at = LEVELS.indexOf(rating.level)
86 const distance = (l: Level) => Math.abs(LEVELS.indexOf(l) - at)
87 return allowed.reduce((best, l) => (distance(l) < distance(best) || (distance(l) === distance(best) && LEVELS.indexOf(l) > LEVELS.indexOf(best)) ? l : best))
88}
89
90/** The allowed levels the store holds, in scale order; every level when it holds none. */
91export function levelsOf(value: unknown): Level[] {
92 const kept = Array.isArray(value) ? LEVELS.filter(l => value.includes(l)) : []
93 return kept.length === 0 ? [...LEVELS] : kept
94}
95
96/** The allowed levels a `levels` command names, in scale order, or why they are refused. */
97export function parseLevels(args: string): Level[] | string {
98 const words = args.toLowerCase().split(/[\s,]+/).filter(w => w !== '')
99 if (words.length === 1 && words[0] === 'all') return [...LEVELS]
100 const unknown = words.find(w => !(LEVELS as readonly string[]).includes(w))
101 if (words.length === 0 || unknown !== undefined) return `levels takes all, or one or more of ${LEVELS.join(', ')}${unknown === undefined ? '' : `; "${unknown}" is none of them`}`
102 return LEVELS.filter(l => words.includes(l))
103}
104
105/** How the sidebar colours a line or a part of one. */
106type Tone = 'ok' | 'warn' | 'error' | 'dim'
107export type Part = { text: string; kind?: Tone }
108export type Line = { text: string; kind?: Tone; parts?: Part[] }
109
110/** The colours session-watch gives the same levels (`effortTone`), so both lines read alike. */
111const TONE: Record<Level, Tone> = { low: 'dim', medium: 'ok', high: 'warn', xhigh: 'error', max: 'error' }
112
113/** A level as a part in its own colour; a budget or no setting stays plain. */
114function levelPart(value: string | number | undefined): Part {
115 const text = String(value ?? 'default')
116 const kind = typeof value === 'string' ? TONE[value as Level] : undefined
117 return kind === undefined ? { text } : { text, kind }
118}
119
120/**
121 * The effort a main-loop turn ran at: the rated level and whether the prompt named it itself, or the
122 * session's own when the turn was not rated.
123 */
124export type TurnEffort = { rated: true; level: Level; named: boolean } | { rated: false; value: string | number | undefined }
125
126/** A turn's effort as parts: the level coloured, and a faint ` (session)` or ` (named)` after it. */
127function turnParts(t: TurnEffort): Part[] {
128 if (!t.rated) return [levelPart(t.value), { text: ' (session)', kind: 'dim' }]
129 return t.named ? [levelPart(t.level), { text: ' (named)', kind: 'dim' }] : [levelPart(t.level)]
130}
131
132const lineOf = (parts: Part[]): Line => ({ text: parts.map(p => p.text).join(''), parts })
133
134/**
135 * The person's line: `this turn` while a turn runs, else `last turn` once one ended, and the session's
136 * effort; only the levels coloured, each by how high it is, the words plain.
137 */
138export function effortLine(current: TurnEffort | undefined, last: TurnEffort | undefined, session: string | number | undefined): Line {
139 const shown = current === undefined ? (last === undefined ? [] : [{ text: 'last turn ' }, ...turnParts(last), { text: ' · ' }]) : [{ text: 'this turn ' }, ...turnParts(current), { text: ' · ' }]
140 return lineOf([...shown, { text: 'session ' }, levelPart(session)])
141}
142
143/** The levels the person allows a rated turn to run at, each in its own colour. */
144export function allowedLine(allowed: readonly Level[]): Line {
145 return lineOf([{ text: 'allowed ' }, ...allowed.flatMap((l, i) => (i === 0 ? [levelPart(l)] : [{ text: ' · ' }, levelPart(l)]))])
146}
147