Picks low, medium or high effort for each prompt, keeping the prompt cache warm

A Claude Code mod that picks the effort level for each prompt you send. On a subscription, lower effort on easy turns means less thinking, so your usage limits last longer.
When you send a prompt, Haiku sorts it into one of three levels:
| Level | For |
|---|---|
low | Questions, explanations, chat, running a command, git chores. No file edits. |
medium | Ordinary code edits and multi-step work. |
high | Hard debugging, design, large refactors, or "still broken" after a failed attempt. |
The level holds for every step of that turn. A turn you didn't start, such as a background task finishing, keeps the previous level so effort doesn't flip back and forth. With context-meter installed, its band shows the current level; the mod draws nothing itself.
Explore subagents run at low unless Claude asks for a specific effort. A subagent starts with an empty context, so this costs no cache.
The mod never goes above high. If you set xhigh or max with /effort, or write "ultrathink", it leaves that turn alone.
Once any of your usage windows passes 80%, either the 5-hour or the 7-day one, turns the mod would run at high run at medium instead. context-meter then shows (budget) after the level. A level you pinned yourself isn't capped.
Anthropic's API can change effort in two ways. A per-message change keeps the cached conversation. A top-level change throws it away, and the next request pays to cache the whole conversation again. I haven't confirmed which way Claude Code sends the change this mod makes. So the mod checks: if a switch writes more to the cache than it reads, you get a toast saying so. If that toast keeps appearing, the switches cost more than they save, and you should run /auto-effort off.
/auto-effort show the current state
/auto-effort off stop changing effort
/auto-effort on back to automatic
/auto-effort high pin one level (low, medium or high)
/auto-effort budget 70 start budget mode at 70% instead of 80%
/auto-effort budget off turn budget mode off (budget on brings back 80%)
Each prompt costs one small Haiku call to classify it.
/plugin install auto-effort --marketplace noash-xrc/claude-tools
Answer y to add the marketplace, then pick a scope.
claude plugin marketplace update noash-tools
claude plugin update auto-effort@noash-tools
Restart Claude Code to load the new version.
MIT
hooks/register.ts 92 lines1import { atom, update } from 'claude-code'
2import type { Register } from 'claude-code'
3
4type Level = 'low' | 'medium' | 'high'
5
6// Haiku answers with one label verbatim; the word before the colon is the level.
7// ponytail: a low turn that edits files anyway stays low, add a mid-turn bump if that shows up
8const LABELS = [
9 'low: a question, explanation, chat, running a command or a git chore, no file edits',
10 'medium: ordinary code edits or multi-step work',
11 'high: hard debugging, design, a large refactor, or a retry after a failed attempt',
12]
13const LEVELS: readonly string[] = ['low', 'medium', 'high']
14// context-meter reads this to show the effort the mod picked
15const effortAtom = atom({ plugin: 'auto-effort', key: 'effort' } as const, null)
16
17let mode: 'auto' | 'off' | Level = 'auto'
18// the newest prompt's level; a turn with no prompt of its own keeps it, so effort doesn't flap
19let level: Level | undefined
20let last: { model: string; effort: Level } | undefined
21let warned = false
22// budget mode: past this share of any usage window, high turns run at medium
23let budget: number | 'off' = 80
24let used = 0
25const budgetOn = () => budget !== 'off' && used >= budget
26
27export const register: Register = on => {
28 on('session.start', async ($, e, next) => {
29 await $.command.register({
30 name: 'auto-effort',
31 description: 'Show or set automatic effort: on, off, pin low|medium|high, or budget on|off|<percent>',
32 argumentHint: '[on | off | low | medium | high | budget on|off|<percent>]',
33 })
34 return next(e)
35 })
36
37 on('command.run', { command: 'auto-effort' }, async ($, e) => {
38 const arg = e.args.trim().toLowerCase().replace(/^pin\s+/, '')
39 const b = /^budget(?:\s+(on|off|\d{1,3})%?)?$/.exec(arg)
40 if (b) budget = b[1] === 'off' ? 'off' : b[1] === 'on' || !b[1] ? 80 : Math.min(100, Number(b[1]))
41 else if (arg === 'on') mode = 'auto'
42 else if (arg === 'off' || LEVELS.includes(arg)) mode = arg as typeof mode
43 else if (arg) return { text: `Unknown argument "${arg}". Use on, off, low, medium or high.` }
44 if (mode === 'off') await update($, effortAtom, () => null)
45 const now = mode === 'auto' ? `on${level ? `, last pick ${level}` : ''}` : mode === 'off' ? 'off' : `pinned to ${mode}`
46 const limit = budget === 'off' ? 'off' : `at ${budget}%, ${Math.round(used)}% of the fullest usage window used`
47 return { text: `auto-effort is ${now}. Budget mode is ${limit}.` }
48 })
49
50 // Typed prompts only: a message delivered into a running turn shouldn't re-pick its effort.
51 on('prompt.submit', async ($, e, next) => {
52 if (mode === 'auto' && e.turnId === undefined && !e.text.startsWith('/')) {
53 const label = await $.model.classify(e.text.slice(0, 4000), LABELS, { model: 'haiku' }).catch(() => undefined)
54 const picked = label?.split(':')[0]
55 if (picked && LEVELS.includes(picked)) level = picked as Level
56 }
57 return next(e)
58 }).catch(($, e, next) => next(e))
59
60 // rateLimits is the whole list each time, so the fullest window is recomputed, not accumulated
61 on('session.measure', ($, e, next) => {
62 used = Math.max(0, ...e.rateLimits.map(r => r.percentUsed))
63 return next(e)
64 })
65
66 // Search subagents start with no context, so a lower effort there loses no cache.
67 on('tool.call', { tool: 'Agent' }, ($, e, next) =>
68 mode !== 'off' && e.subagent_type === 'Explore' && e.effort === undefined ? next({ ...e, effort: 'low' }) : next(e),
69 ).catch(($, e, next) => next(e))
70
71 on('turn.step', async function* ($, e, next) {
72 // xhigh, max or a number came from /effort or ultrathink: more than the mod would pick, so it stands
73 if (e.agentId !== undefined || mode === 'off' || !LEVELS.includes(String(e.effort))) return yield* next(e)
74 // a pin is the person's choice, so budget mode only caps automatic picks
75 const effort = mode !== 'auto' ? mode : level === 'high' && budgetOn() ? 'medium' : level
76 if (!effort) return yield* next(e)
77
78 await update($, effortAtom, () => `${effort}${mode !== 'auto' ? ' (pinned)' : budgetOn() ? ' (budget)' : ''}`)
79 const result = yield* next({ ...e, effort })
80
81 // Writing more than it read right after a switch means the change went out top-level and rebuilt the cache.
82 const u = result.usage
83 const switched = last !== undefined && last.model === e.model && last.effort !== effort
84 if (switched && u && !warned && u.cache_creation_input_tokens > u.cache_read_input_tokens) {
85 warned = true
86 $.ui.toast(`auto-effort: switching to ${effort} rebuilt the prompt cache. If this keeps happening, run /auto-effort off.`)
87 }
88 last = { model: e.model, effort }
89 return result
90 })
91}
92types/index.d.ts 9 lines1// the effort the last main-thread step ran at, with " (pinned)" or " (budget)"; null while the mod stands aside
2export type AutoEffortEffort = string | null
3
4declare module 'claude-code' {
5 interface PluginState {
6 'auto-effort': { effort: AutoEffortEffort }
7 }
8}
9