SLOPSHOPPER

auto-effort

Picks /effort per prompt: Jev (TypeSafe AI) grades how hard the task is, and the turn runs at that level.

newbandcommandpromptprocessnetwork
★ 1v0.1.0MITupdated 2026-10-08peterlimg/auto-effort
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · auto-effort
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /auto-effort ⎿ auto-effort: no turns logged yet. » auto-effort kept session effort · no TYPESAFE_API_KEY [ Turn off ] ⟨Claude Code's own drawing⟩ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Band
» auto-effort kept session effort · no TYPESAFE_API_KEY [ Turn off ] ⟨Claude Code's own drawing⟩
README

auto-effort

A Claude Code mod that picks /effort for each prompt. Jev (TypeSafe AI) grades how hard the request is, and the turn's model requests run at that level: low, medium, high, xhigh or max. By default it only lowers effort below your own setting; raising is opt-in (see Settings).

» auto-effort ▰▰▱▱▱ effort MEDIUM ✓ sent  (was high) · Jev 84% sure · 300ms [Turn off]

Install

/plugin install auto-effort --marketplace peterlimg/auto-effort

Answer y to add the marketplace, then pick a scope. It needs a TypeSafe API key in the environment Claude Code starts with:

export TYPESAFE_API_KEY=...        # fish: set -Ux TYPESAFE_API_KEY ...

Without a key, or when Jev fails or takes over 3s, your own /effort applies.

Use

  • The band above the prompt shows the level in use, whether the request was actually sent at it (✓ sent), Jev's confidence, and a Turn off / Turn on button.
  • /auto-effort prints the last turns: what Jev picked, what each turn's requests ran at, and how many turns were lowered or raised from your session's effort. /auto-effort off and /auto-effort on switch it, like the band's button.
  • ~/.claude/auto-effort/decisions.jsonl keeps one JSON line per judged prompt and per turn.

Settings

  • raise (default off): let Jev raise effort above your session setting for hard tasks. Off, a higher pick is capped at your setting and the band shows (Jev: xhigh, capped). Set it in /config, or in settings under pluginConfigs.auto-effort.raise.

In the benchmark below, on the tasks Jev lowered from high, thinking fell by about half (1,164 to 502 tokens a run) and every run still passed. Those were small tasks, so the bill barely moved. Raising from medium cost 4-15% more and changed no outcomes. So it's off by default.

How it decides

  • Each prompt you write is sent to Jev as one Score question, with your previous prompt and the end of the last reply as context, so "yes, fix" is judged against what it approves.
  • Only the main conversation is changed. Subagents, slash-command turns and models without an effort setting keep your own effort.
  • A message sent while a task runs can only raise its effort (up to your setting, unless raise is on). A prompt queued with ctrl+x enter is judged for its own turn.
  • Messages you didn't write (subagent reports, notifications) aren't judged.

Benchmark

bench/run.py runs the same tasks with the mod on and off and compares cost, tokens and pass rate. Each run copies a task's fixture repo to a temp dir, runs claude -p on its prompt, and grades the result with a hidden check.py. Both arms share the model, session effort, prompt and settings; only --plugin-dir differs.

bench/run.py                          # 12 tasks x 2 arms x 3 runs at --effort high
bench/run.py --tasks 02,10 --runs 1   # a quick smoke run
bench/run.py --effort medium          # let the mod raise as well as lower

Costs come from Claude Code's own total_cost_usd; the effort each request ran at comes from the session transcript. --max-cost (default $60) stops new runs once the total is spent.

Develop

claude --plugin-dir .        # run it from this folder
claude plugin validate .
claude plugin test .
Source 2 files
hooks/register.tsx 341 lines
1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, Register } from 'claude-code'
3
4import type { Effort, Judgement } from '../types'
5
6const JEV_URL = 'https://api.typesafe.ai/v1/systemone'
7const TIMEOUT_MS = 3000
8const LOG_LINES = 1000
9const LOG_TRIM_BYTES = 2_000_000 // under fs.read's 4 MiB cap
10const REPORT_TURNS = 12
11// A fresh TLS connection to Jev costs ~550ms on top of its ~250ms answer, and an idle one is dropped within a few
12// minutes. So while the session is in use, an unauthenticated HEAD (a free 405) keeps the host's connection open.
13const WARM_EVERY_MS = 60_000
14const WARM_IDLE_MS = 15 * 60_000
15// Where a prompt a person wrote comes from: typed, Remote Control, `claude -p`, a /loop or routine, a chat channel,
16// or a channel the engine can't attest. Every other origin is model- or agent-authored.
17const PERSON_ORIGINS = new Set<string | undefined>(['composer', 'bridge', 'sdk', 'scheduled-trigger', 'channel', 'unclassified'])
18const LEVELS: readonly Effort[] = ['low', 'medium', 'high', 'xhigh', 'max']
19const COLORS: Record<Effort, string> = { low: 'green', medium: 'cyan', high: 'blue', xhigh: 'magenta', max: 'red' }
20
21const judgement = atom({ plugin: 'auto-effort', key: 'judgement' } as const, null)
22const isOff = atom({ plugin: 'auto-effort', key: 'isOff' } as const, false)
23const previous = atom({ plugin: 'auto-effort', key: 'previous' } as const, null)
24const sessionEffort = atom({ plugin: 'auto-effort', key: 'sessionEffort' } as const, null)
25const queued = atom({ plugin: 'auto-effort', key: 'queued' } as const, null)
26const run = atom({ plugin: 'auto-effort', key: 'run' } as const, null)
27const lastReply = atom({ plugin: 'auto-effort', key: 'lastReply' } as const, null)
28
29// One Score question; its criteria are ordered like LEVELS, so round(score) indexes it.
30const QUESTION = {
31  type: 'score',
32  instructions:
33    'How much reasoning effort does a coding agent need for `request`? A short follow-up ("yes", "do it", "1", "continue") approves or continues the work proposed in `last_reply` or asked for in `previous_request`.',
34  criteria: [
35    'Trivial: a quick question, a lookup, or a one-line edit',
36    'Small: a focused change in one or two files',
37    'Moderate: a feature or bug fix across several files, a non-obvious bug, or running a multi-step workflow such as a code review',
38    'Hard: a cross-cutting refactor, tricky concurrency/security/performance work, or debugging with little to go on',
39    'Very hard: architecture or algorithm design where a wrong call is expensive',
40  ],
41}
42
43// Resolves `work`, or null once `ms` pass. Only a real timeout wins: the abort after the race settles nothing.
44// ponytail: $.http.fetch takes no signal, so a timed-out request runs on and is ignored
45function withTimeout<T>($: EngineInterface, ms: number, work: Promise<T>): Promise<T | null> {
46  const stop = new AbortController()
47  const timeout = $.clock.sleep(ms, { signal: stop.signal }).then(() => null, () => new Promise<never>(() => {}))
48  return Promise.race([work, timeout]).finally(() => stop.abort())
49}
50
51// ponytail: module-level, reset by a reload; at worst one prompt pays the handshake again
52let lastActive = 0
53let warming: unknown = null
54
55async function warm($: EngineInterface) {
56  if (!(await $.env.get('TYPESAFE_API_KEY')) || (await read($, isOff))) return
57  await $.http.fetch(JEV_URL, { method: 'HEAD' }).catch(() => {})
58}
59
60// The band's button and `/auto-effort on|off`: a stale pick must not apply once switched.
61async function setOff($: EngineInterface, off: boolean) {
62  await update($, isOff, () => off)
63  await update($, judgement, () => null)
64}
65
66// The session is in use: keep Jev's connection open until it has been idle for WARM_IDLE_MS.
67// Started by the first sign of use, so a mod loaded into a running session warms too.
68async function active($: EngineInterface) {
69  lastActive = await $.clock.now()
70  warming ??= $.clock.every(WARM_EVERY_MS, async () => {
71    if ((await $.clock.now()) - lastActive < WARM_IDLE_MS) await warm($)
72  })
73}
74
75async function judge($: EngineInterface, text: string, before: string | null, reply: string | null): Promise<Judgement> {
76  const key = await $.env.get('TYPESAFE_API_KEY')
77  if (!key) return { phase: 'kept', reason: 'no TYPESAFE_API_KEY' }
78
79  const startedAt = await $.clock.now()
80  const res = await withTimeout($, TIMEOUT_MS, $.http.fetch(JEV_URL, {
81    method: 'POST',
82    headers: { Authorization: `Bearer ${key}`, 'Content-Type': 'application/json' },
83    // ponytail: chars not tokens; 4k + 1.5k + 16k chars stays well under Jev's 32k-token state budget
84    body: JSON.stringify({
85      model: 'jev-latest',
86      state: { previous_request: before?.slice(0, 4000) ?? null, last_reply: reply, request: text.slice(0, 16000) },
87      questions: { effort: QUESTION },
88    }),
89  }))
90  if (!res) return { phase: 'kept', reason: `Jev timed out (${TIMEOUT_MS / 1000}s)` }
91  if (!res.ok) return { phase: 'kept', reason: `Jev HTTP ${res.status}` }
92
93  const { score, confidence } = JSON.parse(res.text).answers?.effort ?? {}
94  // ponytail: round to the nearest level; a split answer (2.51) flips between neighbours
95  const pick = typeof score === 'number' && typeof confidence === 'number' ? LEVELS[Math.round(score)] : undefined
96  if (!pick) return { phase: 'kept', reason: 'Jev gave no score' }
97  return { phase: 'picked', pick, score, confidence, ms: (await $.clock.now()) - startedAt }
98}
99
100// One JSON line per judged prompt, for monitoring; outside the mod folder so writing it never reloads the mod.
101async function logPath($: EngineInterface) {
102  return `${await $.env.get('HOME')}/.claude/auto-effort/decisions.jsonl`
103}
104
105// Appended (fs has no append), so sessions logging at once don't drop each other's lines.
106async function log($: EngineInterface, entry: object) {
107  const path = await logPath($)
108  if (!(await $.fs.exists(path))) await $.fs.write(path, '') // makes the folder
109  await $.process.run(['sh', '-c', 'cat >> "$1"', 'sh', path], { stdin: JSON.stringify(entry) + '\n' })
110  // ponytail: a trim racing another session's append can drop that one line; once per LOG_TRIM_BYTES of log
111  if ((await $.fs.stat(path)).size < LOG_TRIM_BYTES) return
112  const lines = (await $.fs.read(path)).split('\n').filter(Boolean).slice(-LOG_LINES)
113  await $.fs.write(path, lines.join('\n') + '\n')
114}
115
116// The /auto-effort report: the last turns (all sessions), newest last, with what Jev picked and what ran.
117export function report(logText: string): string {
118  const turns = logText.split('\n').filter(Boolean).map(l => JSON.parse(l)).filter(x => x.kind === 'turn').slice(-REPORT_TURNS)
119  if (!turns.length) return 'no turns logged yet.'
120  const level = (x: unknown) => LEVELS.indexOf(x as Effort)
121  const moved = { lowered: 0, raised: 0, same: 0 }
122  const rows = turns.map(t => {
123    const ran = (t.efforts as unknown[]).map(String).join('/') || '·' // not '-': the reply is markdown, and a leading '- ' makes a list
124    const top = Math.max(...(t.efforts as unknown[]).map(level))
125    if (t.session !== undefined && t.efforts.length) {
126      const d = top - level(t.session)
127      moved[d < 0 ? 'lowered' : d > 0 ? 'raised' : 'same']++
128    }
129    const jev = t.jev === 'off' ? 'off'
130      : t.jev?.phase === 'picked' ? `${t.jev.pick} ${t.jev.score.toFixed(2)} (${Math.round(t.jev.confidence * 100)}%)`
131      : t.jev?.phase === 'kept' ? t.jev.reason : '·'
132    return `${ran.padEnd(14)}${String(t.session ?? '·').padEnd(9)}${String(t.requests).padStart(4)}${String(t.outputTokens).padStart(8)}  ${jev.padEnd(28)}${String(t.prompt).replace(/\s+/g, ' ').slice(0, 50)}`
133  })
134  return [
135    `last ${turns.length} turns: ${moved.lowered} lowered, ${moved.raised} raised, ${moved.same} unchanged vs the session's effort`,
136    '',
137    `${'ran at'.padEnd(14)}${'session'.padEnd(9)}${'reqs'.padStart(4)}${'out tok'.padStart(8)}  ${'jev'.padEnd(28)}prompt`,
138    ...rows,
139  ].join('\n')
140}
141
142// The effort a pick runs at: capped at the session's own effort unless raising is allowed.
143function applied(pick: Effort, session: unknown, canRaise: boolean): Effort {
144  const cap = LEVELS.indexOf(session as Effort)
145  return !canRaise && cap >= 0 && LEVELS.indexOf(pick) > cap ? LEVELS[cap]! : pick
146}
147
148export const register: Register = (on, options) => {
149  // Benchmarks (bench/) found raising cost 4-15% more with no change in pass rate, so it's opt-in.
150  const canRaise = options.raise === true
151
152  // Grade the prompt before it enters, so the turn's first request already uses the pick.
153  on('prompt.submit', async ($, e, next) => {
154    // Only requests a person wrote are judged. A subagent's report, a task notification or a peer's message
155    // continues the work already running: it keeps the current pick and isn't the "previous request" either.
156    if (!PERSON_ORIGINS.has(e.origin?.kind)) return next(e)
157    await active($)
158    // Typed during a running turn (e.turnId): with ctrl+x enter (e.wait) it waits to run as its own turn, so its
159    // pick waits too and turn.start applies it; otherwise it's an aside delivered into the running task.
160    const waits = Boolean(e.turnId && e.wait)
161    const isAside = Boolean(e.turnId) && !waits
162    const isCommand = e.text.trim().startsWith('/')
163    const before = await read($, previous)
164    // The task a later "yes"/"continue" refers to: kept while off too; not a slash command, and not an aside
165    // sent mid-task, so "continue" after an interrupt is judged against the task, not the aside.
166    if (!isCommand && !e.turnId) await update($, previous, () => e.text)
167
168    if (isCommand || (await read($, isOff))) {
169      // Not judged: its turn runs at the session's effort, not the last task's pick.
170      if (waits) await update($, queued, () => ({ text: e.text, judgement: null }))
171      else if (!isAside) await update($, judgement, () => null)
172      return next(e)
173    }
174
175    const current = await read($, judgement)
176    // A message typed during a running turn must not unset the running task's pick while Jev answers.
177    if (!e.turnId) {
178      await update($, judgement, () => ({ phase: 'asking' }))
179    }
180    const judged = await judge($, e.text, before, await read($, lastReply)).catch((): Judgement => ({ phase: 'kept', reason: 'Jev unreachable' }))
181    // An aside can only raise a running pick. With none running the turn is at the session's effort,
182    // which we can't compare against, so the aside leaves it alone.
183    let j: Judgement | null = judged
184    if (isAside && (current?.phase !== 'picked' || j.phase !== 'picked' || LEVELS.indexOf(j.pick) < LEVELS.indexOf(current.pick))) {
185      j = current
186    }
187    if (waits) await update($, queued, () => ({ text: e.text, judgement: judged }))
188    else await update($, judgement, () => j)
189    await log($, {
190      kind: 'decision',
191      at: new Date(await $.clock.now()).toISOString(),
192      midTurn: isAside,
193      queued: waits,
194      prompt: e.text.slice(0, 120),
195      jev: judged,
196      applies: waits ? 'at its own turn' : j?.phase === 'picked' ? applied(j.pick, await read($, sessionEffort), canRaise) : 'session',
197    }).catch(() => {})
198
199    return next(e)
200  }).catch(async ($, e, next) => {
201    // Anything else broke: the prompt goes through untouched; between turns, no stale pick either.
202    if (!e.turnId) await update($, judgement, () => null)
203    return next(e)
204  })
205
206  // A queued prompt's turn begins: its held pick (or none, for a slash command) becomes the running one.
207  on('turn.start', async ($, e, next) => {
208    // ponytail: one turn record at a time; assumes turn.start is the main thread's (its input has no agentId)
209    await update($, run, () => ({ turnId: e.turnId, prompt: e.text, requests: 0, efforts: [], outputTokens: 0 }))
210    const q = await read($, queued)
211    if (q && e.text.includes(q.text)) {
212      await update($, queued, () => null)
213      await update($, judgement, () => q.judgement)
214      if (!q.text.trim().startsWith('/')) await update($, previous, () => q.text)
215    }
216    return next(e)
217  })
218
219  // Main thread only; subagents keep their own effort. Models without effort (e.effort absent) untouched.
220  on('turn.step', async function* ($, e, next) {
221    if (e.agentId || e.effort === undefined) return yield* next(e)
222    if ((await read($, sessionEffort)) !== e.effort) await update($, sessionEffort, () => e.effort ?? null)
223    const j = (await read($, isOff)) ? null : await read($, judgement)
224    const pick = j?.phase === 'picked' ? j.pick : undefined
225
226    const result = yield* next(pick ? { ...e, effort: applied(pick, e.effort, canRaise) } : e)
227    // Proof, not intent: the effort the bottom of the chain (the engine) actually received.
228    const sent = next.trace.at(-1)?.received.effort
229    // Written onto the latest pick only if it is still the one this step sent: a mid-turn raise
230    // that landed during the step must not be overwritten; the next step proves it instead.
231    if (pick) await update($, judgement, cur => (cur?.phase === 'picked' && cur.pick === pick ? { ...cur, from: e.effort, sent } : cur))
232    await update($, run, r => (r?.turnId !== e.turnId ? r : {
233      ...r,
234      session: e.effort,
235      requests: r.requests + 1,
236      efforts: sent === undefined || r.efforts.includes(sent) ? r.efforts : [...r.efforts, sent],
237      outputTokens: r.outputTokens + (result.usage?.output_tokens ?? 0),
238    }))
239    return result
240  })
241
242  // One log line per main-thread turn: what Jev said, and what the requests actually ran at.
243  on('turn.complete', async ($, e, next) => {
244    // What the agent last proposed or asked: a "yes, fix" approves that, so Jev needs it. Its end carries the ask.
245    if (!e.agentId && e.answer) await update($, lastReply, () => e.answer.slice(-1500))
246    await active($)
247    const r = await read($, run)
248    if (!e.agentId && r?.turnId === e.turnId) {
249      await update($, run, () => null)
250      const { turnId, ...turn } = r
251      await log($, {
252        kind: 'turn',
253        at: new Date(await $.clock.now()).toISOString(),
254        ...turn,
255        prompt: r.prompt.slice(0, 120),
256        jev: (await read($, isOff)) ? 'off' : await read($, judgement),
257      }).catch(() => {})
258    }
259    return next(e)
260  })
261
262  on('session.start', async ($, e, next) => {
263    const started = await next(e)
264    await $.command.register({ name: 'auto-effort', description: 'Show what auto-effort did on the last turns; `on`/`off` to switch it' })
265    // Open the connection now for the first prompt, then keep it open while the session is in use.
266    await active($)
267    void warm($)
268    return started
269  })
270
271  on('command.run', { command: 'auto-effort' }, async ($, e) => {
272    const arg = e.args.trim()
273    if (arg === 'on' || arg === 'off') {
274      await setOff($, arg === 'off')
275      return { text: arg === 'on' ? 'on: Jev picks the effort' : 'off: your session effort applies' }
276    }
277    return { text: report(await $.fs.read(await logPath($)).catch(() => '')) }
278  })
279
280  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
281    if (e.props.hasSurvey) return next(e)
282    // The band is shared: other mods draw here too, so stack our line on whatever the rest of the chain draws.
283    const below = await next(e)
284    const j = await read($, judgement)
285    const off = await read($, isOff)
286    const session = await read($, sessionEffort)
287
288    const { Box, Button, Text } = $.ui.resolve(e)
289    const toggle = (
290      <Button
291        key="toggle"
292        label={off ? 'Turn on' : 'Turn off'}
293        onPress={() => setOff($, !off)}
294      />
295    )
296    // One line: `<>...</>` is a column Box here, so every group of parts is a row Box.
297    const row = (body: JSX.Element) => (
298      <Box flexDirection="column">
299        <Box flexDirection="row" paddingX={1}><Text bold>» auto-effort </Text>{body}{toggle}</Box>
300        {below}
301      </Box>
302    )
303    // A 5-cell gauge, filled up to the level the turn runs at: ▰▰▰▱▱ = high.
304    const level = (use: Effort, tail: JSX.Element) => {
305      const n = LEVELS.indexOf(use) + 1
306      return row(
307        <Box flexDirection="row">
308          <Text color={COLORS[use]}>{'▰'.repeat(n) + '▱'.repeat(LEVELS.length - n)}</Text>
309          <Text color={COLORS[use]} bold>{` effort ${use.toUpperCase()}`}</Text>
310          {tail}
311        </Box>,
312      )
313    }
314    // No pick applies: the session's own effort, drawn like a pick once a request has shown it.
315    const unpicked = (note: JSX.Element) => LEVELS.includes(session as Effort)
316      ? level(session as Effort, <Box flexDirection="row"><Text dimColor>{' · '}</Text>{note}</Box>)
317      : row(note)
318
319    if (off) return unpicked(<Text dimColor>{'off, session effort applies '}</Text>)
320    if (!j) return unpicked(<Text dimColor>{'session effort until the next prompt '}</Text>)
321    if (j.phase === 'asking') return row(<Text dimColor>◌ asking Jev... </Text>)
322    if (j.phase === 'kept') {
323      return unpicked(<Box flexDirection="row"><Text color="yellow">kept session effort</Text><Text dimColor>{` · ${j.reason} `}</Text></Box>)
324    }
325
326    const use = applied(j.pick, session, canRaise)
327    const note = (j.from === undefined || j.from === use ? '' : `  (was ${j.from})`) + (use === j.pick ? '' : `  (Jev: ${j.pick}, capped)`)
328    return level(
329      use,
330      <Box flexDirection="row">
331        {j.sent === undefined
332          ? <Text dimColor>{' · not sent yet'}</Text>
333          : j.sent === use
334            ? <Text color="green">{' ✓ sent'}</Text>
335            : <Text color="yellow">{` ⚠ model got ${j.sent}`}</Text>}
336        <Text dimColor>{`${note} · Jev ${Math.round(j.confidence * 100)}% sure · ${Math.round(j.ms)}ms `}</Text>
337      </Box>,
338    )
339  })
340}
341
types/index.d.ts 26 lines
1export type Effort = 'low' | 'medium' | 'high' | 'xhigh' | 'max'
2
3/** The latest prompt's judgement: asking, a pick, or why there is none. */
4export type Judgement =
5  | { phase: 'asking' }
6  | { phase: 'picked'; pick: Effort; score: number; confidence: number; ms: number; from?: Effort | number; sent?: Effort | number }
7  | { phase: 'kept'; reason: string }
8
9declare module 'claude-code' {
10  interface PluginState {
11    'auto-effort': {
12      judgement: Judgement | null
13      isOff: boolean
14      previous: string | null
15      /** The session's own effort, as the last main-thread request carried it before any rewrite. */
16      sessionEffort: Effort | number | null
17      /** A prompt queued (ctrl+x enter) to run as its own turn, and its pick, held until that turn starts. */
18      queued: { text: string; judgement: Judgement | null } | null
19      /** The running main-thread turn, as its requests went out; logged at turn.complete. */
20      /** The end of the main thread's last reply: what a short "yes, fix" approves. */
21      lastReply: string | null
22      run: { turnId: string; prompt: string; session?: Effort | number; requests: number; efforts: (Effort | number)[]; outputTokens: number } | null
23    }
24  }
25}
26