Picks /effort per prompt: Jev (TypeSafe AI) grades how hard the task is, and the turn runs at that level.

A Claude Code mod that picks /effort for each prompt. Jev (TypeSafe AI) grades how hard the request is, and the turn's model requests run at that level: low, medium, high, xhigh or max. By default it only lowers effort below your own setting; raising is opt-in (see Settings).
» auto-effort ▰▰▱▱▱ effort MEDIUM ✓ sent (was high) · Jev 84% sure · 300ms [Turn off]
/plugin install auto-effort --marketplace peterlimg/auto-effort
Answer y to add the marketplace, then pick a scope. It needs a TypeSafe API key in the environment Claude Code starts with:
export TYPESAFE_API_KEY=... # fish: set -Ux TYPESAFE_API_KEY ...
Without a key, or when Jev fails or takes over 3s, your own /effort applies.
✓ sent), Jev's confidence, and a Turn off / Turn on button./auto-effort prints the last turns: what Jev picked, what each turn's requests ran at, and how many turns were lowered or raised from your session's effort. /auto-effort off and /auto-effort on switch it, like the band's button.~/.claude/auto-effort/decisions.jsonl keeps one JSON line per judged prompt and per turn.raise (default off): let Jev raise effort above your session setting for hard tasks. Off, a higher pick is capped at your setting and the band shows (Jev: xhigh, capped). Set it in /config, or in settings under pluginConfigs.auto-effort.raise.In the benchmark below, on the tasks Jev lowered from high, thinking fell by about half (1,164 to 502 tokens a run) and every run still passed. Those were small tasks, so the bill barely moved. Raising from medium cost 4-15% more and changed no outcomes. So it's off by default.
raise is on). A prompt queued with ctrl+x enter is judged for its own turn.bench/run.py runs the same tasks with the mod on and off and compares cost, tokens and pass rate. Each run copies a task's fixture repo to a temp dir, runs claude -p on its prompt, and grades the result with a hidden check.py. Both arms share the model, session effort, prompt and settings; only --plugin-dir differs.
bench/run.py # 12 tasks x 2 arms x 3 runs at --effort high
bench/run.py --tasks 02,10 --runs 1 # a quick smoke run
bench/run.py --effort medium # let the mod raise as well as lower
Costs come from Claude Code's own total_cost_usd; the effort each request ran at comes from the session transcript. --max-cost (default $60) stops new runs once the total is spent.
claude --plugin-dir . # run it from this folder
claude plugin validate .
claude plugin test .hooks/register.tsx 341 lines1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, Register } from 'claude-code'
3
4import type { Effort, Judgement } from '../types'
5
6const JEV_URL = 'https://api.typesafe.ai/v1/systemone'
7const TIMEOUT_MS = 3000
8const LOG_LINES = 1000
9const LOG_TRIM_BYTES = 2_000_000 // under fs.read's 4 MiB cap
10const REPORT_TURNS = 12
11// A fresh TLS connection to Jev costs ~550ms on top of its ~250ms answer, and an idle one is dropped within a few
12// minutes. So while the session is in use, an unauthenticated HEAD (a free 405) keeps the host's connection open.
13const WARM_EVERY_MS = 60_000
14const WARM_IDLE_MS = 15 * 60_000
15// Where a prompt a person wrote comes from: typed, Remote Control, `claude -p`, a /loop or routine, a chat channel,
16// or a channel the engine can't attest. Every other origin is model- or agent-authored.
17const PERSON_ORIGINS = new Set<string | undefined>(['composer', 'bridge', 'sdk', 'scheduled-trigger', 'channel', 'unclassified'])
18const LEVELS: readonly Effort[] = ['low', 'medium', 'high', 'xhigh', 'max']
19const COLORS: Record<Effort, string> = { low: 'green', medium: 'cyan', high: 'blue', xhigh: 'magenta', max: 'red' }
20
21const judgement = atom({ plugin: 'auto-effort', key: 'judgement' } as const, null)
22const isOff = atom({ plugin: 'auto-effort', key: 'isOff' } as const, false)
23const previous = atom({ plugin: 'auto-effort', key: 'previous' } as const, null)
24const sessionEffort = atom({ plugin: 'auto-effort', key: 'sessionEffort' } as const, null)
25const queued = atom({ plugin: 'auto-effort', key: 'queued' } as const, null)
26const run = atom({ plugin: 'auto-effort', key: 'run' } as const, null)
27const lastReply = atom({ plugin: 'auto-effort', key: 'lastReply' } as const, null)
28
29// One Score question; its criteria are ordered like LEVELS, so round(score) indexes it.
30const QUESTION = {
31 type: 'score',
32 instructions:
33 'How much reasoning effort does a coding agent need for `request`? A short follow-up ("yes", "do it", "1", "continue") approves or continues the work proposed in `last_reply` or asked for in `previous_request`.',
34 criteria: [
35 'Trivial: a quick question, a lookup, or a one-line edit',
36 'Small: a focused change in one or two files',
37 'Moderate: a feature or bug fix across several files, a non-obvious bug, or running a multi-step workflow such as a code review',
38 'Hard: a cross-cutting refactor, tricky concurrency/security/performance work, or debugging with little to go on',
39 'Very hard: architecture or algorithm design where a wrong call is expensive',
40 ],
41}
42
43// Resolves `work`, or null once `ms` pass. Only a real timeout wins: the abort after the race settles nothing.
44// ponytail: $.http.fetch takes no signal, so a timed-out request runs on and is ignored
45function withTimeout<T>($: EngineInterface, ms: number, work: Promise<T>): Promise<T | null> {
46 const stop = new AbortController()
47 const timeout = $.clock.sleep(ms, { signal: stop.signal }).then(() => null, () => new Promise<never>(() => {}))
48 return Promise.race([work, timeout]).finally(() => stop.abort())
49}
50
51// ponytail: module-level, reset by a reload; at worst one prompt pays the handshake again
52let lastActive = 0
53let warming: unknown = null
54
55async function warm($: EngineInterface) {
56 if (!(await $.env.get('TYPESAFE_API_KEY')) || (await read($, isOff))) return
57 await $.http.fetch(JEV_URL, { method: 'HEAD' }).catch(() => {})
58}
59
60// The band's button and `/auto-effort on|off`: a stale pick must not apply once switched.
61async function setOff($: EngineInterface, off: boolean) {
62 await update($, isOff, () => off)
63 await update($, judgement, () => null)
64}
65
66// The session is in use: keep Jev's connection open until it has been idle for WARM_IDLE_MS.
67// Started by the first sign of use, so a mod loaded into a running session warms too.
68async function active($: EngineInterface) {
69 lastActive = await $.clock.now()
70 warming ??= $.clock.every(WARM_EVERY_MS, async () => {
71 if ((await $.clock.now()) - lastActive < WARM_IDLE_MS) await warm($)
72 })
73}
74
75async function judge($: EngineInterface, text: string, before: string | null, reply: string | null): Promise<Judgement> {
76 const key = await $.env.get('TYPESAFE_API_KEY')
77 if (!key) return { phase: 'kept', reason: 'no TYPESAFE_API_KEY' }
78
79 const startedAt = await $.clock.now()
80 const res = await withTimeout($, TIMEOUT_MS, $.http.fetch(JEV_URL, {
81 method: 'POST',
82 headers: { Authorization: `Bearer ${key}`, 'Content-Type': 'application/json' },
83 // ponytail: chars not tokens; 4k + 1.5k + 16k chars stays well under Jev's 32k-token state budget
84 body: JSON.stringify({
85 model: 'jev-latest',
86 state: { previous_request: before?.slice(0, 4000) ?? null, last_reply: reply, request: text.slice(0, 16000) },
87 questions: { effort: QUESTION },
88 }),
89 }))
90 if (!res) return { phase: 'kept', reason: `Jev timed out (${TIMEOUT_MS / 1000}s)` }
91 if (!res.ok) return { phase: 'kept', reason: `Jev HTTP ${res.status}` }
92
93 const { score, confidence } = JSON.parse(res.text).answers?.effort ?? {}
94 // ponytail: round to the nearest level; a split answer (2.51) flips between neighbours
95 const pick = typeof score === 'number' && typeof confidence === 'number' ? LEVELS[Math.round(score)] : undefined
96 if (!pick) return { phase: 'kept', reason: 'Jev gave no score' }
97 return { phase: 'picked', pick, score, confidence, ms: (await $.clock.now()) - startedAt }
98}
99
100// One JSON line per judged prompt, for monitoring; outside the mod folder so writing it never reloads the mod.
101async function logPath($: EngineInterface) {
102 return `${await $.env.get('HOME')}/.claude/auto-effort/decisions.jsonl`
103}
104
105// Appended (fs has no append), so sessions logging at once don't drop each other's lines.
106async function log($: EngineInterface, entry: object) {
107 const path = await logPath($)
108 if (!(await $.fs.exists(path))) await $.fs.write(path, '') // makes the folder
109 await $.process.run(['sh', '-c', 'cat >> "$1"', 'sh', path], { stdin: JSON.stringify(entry) + '\n' })
110 // ponytail: a trim racing another session's append can drop that one line; once per LOG_TRIM_BYTES of log
111 if ((await $.fs.stat(path)).size < LOG_TRIM_BYTES) return
112 const lines = (await $.fs.read(path)).split('\n').filter(Boolean).slice(-LOG_LINES)
113 await $.fs.write(path, lines.join('\n') + '\n')
114}
115
116// The /auto-effort report: the last turns (all sessions), newest last, with what Jev picked and what ran.
117export function report(logText: string): string {
118 const turns = logText.split('\n').filter(Boolean).map(l => JSON.parse(l)).filter(x => x.kind === 'turn').slice(-REPORT_TURNS)
119 if (!turns.length) return 'no turns logged yet.'
120 const level = (x: unknown) => LEVELS.indexOf(x as Effort)
121 const moved = { lowered: 0, raised: 0, same: 0 }
122 const rows = turns.map(t => {
123 const ran = (t.efforts as unknown[]).map(String).join('/') || '·' // not '-': the reply is markdown, and a leading '- ' makes a list
124 const top = Math.max(...(t.efforts as unknown[]).map(level))
125 if (t.session !== undefined && t.efforts.length) {
126 const d = top - level(t.session)
127 moved[d < 0 ? 'lowered' : d > 0 ? 'raised' : 'same']++
128 }
129 const jev = t.jev === 'off' ? 'off'
130 : t.jev?.phase === 'picked' ? `${t.jev.pick} ${t.jev.score.toFixed(2)} (${Math.round(t.jev.confidence * 100)}%)`
131 : t.jev?.phase === 'kept' ? t.jev.reason : '·'
132 return `${ran.padEnd(14)}${String(t.session ?? '·').padEnd(9)}${String(t.requests).padStart(4)}${String(t.outputTokens).padStart(8)} ${jev.padEnd(28)}${String(t.prompt).replace(/\s+/g, ' ').slice(0, 50)}`
133 })
134 return [
135 `last ${turns.length} turns: ${moved.lowered} lowered, ${moved.raised} raised, ${moved.same} unchanged vs the session's effort`,
136 '',
137 `${'ran at'.padEnd(14)}${'session'.padEnd(9)}${'reqs'.padStart(4)}${'out tok'.padStart(8)} ${'jev'.padEnd(28)}prompt`,
138 ...rows,
139 ].join('\n')
140}
141
142// The effort a pick runs at: capped at the session's own effort unless raising is allowed.
143function applied(pick: Effort, session: unknown, canRaise: boolean): Effort {
144 const cap = LEVELS.indexOf(session as Effort)
145 return !canRaise && cap >= 0 && LEVELS.indexOf(pick) > cap ? LEVELS[cap]! : pick
146}
147
148export const register: Register = (on, options) => {
149 // Benchmarks (bench/) found raising cost 4-15% more with no change in pass rate, so it's opt-in.
150 const canRaise = options.raise === true
151
152 // Grade the prompt before it enters, so the turn's first request already uses the pick.
153 on('prompt.submit', async ($, e, next) => {
154 // Only requests a person wrote are judged. A subagent's report, a task notification or a peer's message
155 // continues the work already running: it keeps the current pick and isn't the "previous request" either.
156 if (!PERSON_ORIGINS.has(e.origin?.kind)) return next(e)
157 await active($)
158 // Typed during a running turn (e.turnId): with ctrl+x enter (e.wait) it waits to run as its own turn, so its
159 // pick waits too and turn.start applies it; otherwise it's an aside delivered into the running task.
160 const waits = Boolean(e.turnId && e.wait)
161 const isAside = Boolean(e.turnId) && !waits
162 const isCommand = e.text.trim().startsWith('/')
163 const before = await read($, previous)
164 // The task a later "yes"/"continue" refers to: kept while off too; not a slash command, and not an aside
165 // sent mid-task, so "continue" after an interrupt is judged against the task, not the aside.
166 if (!isCommand && !e.turnId) await update($, previous, () => e.text)
167
168 if (isCommand || (await read($, isOff))) {
169 // Not judged: its turn runs at the session's effort, not the last task's pick.
170 if (waits) await update($, queued, () => ({ text: e.text, judgement: null }))
171 else if (!isAside) await update($, judgement, () => null)
172 return next(e)
173 }
174
175 const current = await read($, judgement)
176 // A message typed during a running turn must not unset the running task's pick while Jev answers.
177 if (!e.turnId) {
178 await update($, judgement, () => ({ phase: 'asking' }))
179 }
180 const judged = await judge($, e.text, before, await read($, lastReply)).catch((): Judgement => ({ phase: 'kept', reason: 'Jev unreachable' }))
181 // An aside can only raise a running pick. With none running the turn is at the session's effort,
182 // which we can't compare against, so the aside leaves it alone.
183 let j: Judgement | null = judged
184 if (isAside && (current?.phase !== 'picked' || j.phase !== 'picked' || LEVELS.indexOf(j.pick) < LEVELS.indexOf(current.pick))) {
185 j = current
186 }
187 if (waits) await update($, queued, () => ({ text: e.text, judgement: judged }))
188 else await update($, judgement, () => j)
189 await log($, {
190 kind: 'decision',
191 at: new Date(await $.clock.now()).toISOString(),
192 midTurn: isAside,
193 queued: waits,
194 prompt: e.text.slice(0, 120),
195 jev: judged,
196 applies: waits ? 'at its own turn' : j?.phase === 'picked' ? applied(j.pick, await read($, sessionEffort), canRaise) : 'session',
197 }).catch(() => {})
198
199 return next(e)
200 }).catch(async ($, e, next) => {
201 // Anything else broke: the prompt goes through untouched; between turns, no stale pick either.
202 if (!e.turnId) await update($, judgement, () => null)
203 return next(e)
204 })
205
206 // A queued prompt's turn begins: its held pick (or none, for a slash command) becomes the running one.
207 on('turn.start', async ($, e, next) => {
208 // ponytail: one turn record at a time; assumes turn.start is the main thread's (its input has no agentId)
209 await update($, run, () => ({ turnId: e.turnId, prompt: e.text, requests: 0, efforts: [], outputTokens: 0 }))
210 const q = await read($, queued)
211 if (q && e.text.includes(q.text)) {
212 await update($, queued, () => null)
213 await update($, judgement, () => q.judgement)
214 if (!q.text.trim().startsWith('/')) await update($, previous, () => q.text)
215 }
216 return next(e)
217 })
218
219 // Main thread only; subagents keep their own effort. Models without effort (e.effort absent) untouched.
220 on('turn.step', async function* ($, e, next) {
221 if (e.agentId || e.effort === undefined) return yield* next(e)
222 if ((await read($, sessionEffort)) !== e.effort) await update($, sessionEffort, () => e.effort ?? null)
223 const j = (await read($, isOff)) ? null : await read($, judgement)
224 const pick = j?.phase === 'picked' ? j.pick : undefined
225
226 const result = yield* next(pick ? { ...e, effort: applied(pick, e.effort, canRaise) } : e)
227 // Proof, not intent: the effort the bottom of the chain (the engine) actually received.
228 const sent = next.trace.at(-1)?.received.effort
229 // Written onto the latest pick only if it is still the one this step sent: a mid-turn raise
230 // that landed during the step must not be overwritten; the next step proves it instead.
231 if (pick) await update($, judgement, cur => (cur?.phase === 'picked' && cur.pick === pick ? { ...cur, from: e.effort, sent } : cur))
232 await update($, run, r => (r?.turnId !== e.turnId ? r : {
233 ...r,
234 session: e.effort,
235 requests: r.requests + 1,
236 efforts: sent === undefined || r.efforts.includes(sent) ? r.efforts : [...r.efforts, sent],
237 outputTokens: r.outputTokens + (result.usage?.output_tokens ?? 0),
238 }))
239 return result
240 })
241
242 // One log line per main-thread turn: what Jev said, and what the requests actually ran at.
243 on('turn.complete', async ($, e, next) => {
244 // What the agent last proposed or asked: a "yes, fix" approves that, so Jev needs it. Its end carries the ask.
245 if (!e.agentId && e.answer) await update($, lastReply, () => e.answer.slice(-1500))
246 await active($)
247 const r = await read($, run)
248 if (!e.agentId && r?.turnId === e.turnId) {
249 await update($, run, () => null)
250 const { turnId, ...turn } = r
251 await log($, {
252 kind: 'turn',
253 at: new Date(await $.clock.now()).toISOString(),
254 ...turn,
255 prompt: r.prompt.slice(0, 120),
256 jev: (await read($, isOff)) ? 'off' : await read($, judgement),
257 }).catch(() => {})
258 }
259 return next(e)
260 })
261
262 on('session.start', async ($, e, next) => {
263 const started = await next(e)
264 await $.command.register({ name: 'auto-effort', description: 'Show what auto-effort did on the last turns; `on`/`off` to switch it' })
265 // Open the connection now for the first prompt, then keep it open while the session is in use.
266 await active($)
267 void warm($)
268 return started
269 })
270
271 on('command.run', { command: 'auto-effort' }, async ($, e) => {
272 const arg = e.args.trim()
273 if (arg === 'on' || arg === 'off') {
274 await setOff($, arg === 'off')
275 return { text: arg === 'on' ? 'on: Jev picks the effort' : 'off: your session effort applies' }
276 }
277 return { text: report(await $.fs.read(await logPath($)).catch(() => '')) }
278 })
279
280 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
281 if (e.props.hasSurvey) return next(e)
282 // The band is shared: other mods draw here too, so stack our line on whatever the rest of the chain draws.
283 const below = await next(e)
284 const j = await read($, judgement)
285 const off = await read($, isOff)
286 const session = await read($, sessionEffort)
287
288 const { Box, Button, Text } = $.ui.resolve(e)
289 const toggle = (
290 <Button
291 key="toggle"
292 label={off ? 'Turn on' : 'Turn off'}
293 onPress={() => setOff($, !off)}
294 />
295 )
296 // One line: `<>...</>` is a column Box here, so every group of parts is a row Box.
297 const row = (body: JSX.Element) => (
298 <Box flexDirection="column">
299 <Box flexDirection="row" paddingX={1}><Text bold>» auto-effort </Text>{body}{toggle}</Box>
300 {below}
301 </Box>
302 )
303 // A 5-cell gauge, filled up to the level the turn runs at: ▰▰▰▱▱ = high.
304 const level = (use: Effort, tail: JSX.Element) => {
305 const n = LEVELS.indexOf(use) + 1
306 return row(
307 <Box flexDirection="row">
308 <Text color={COLORS[use]}>{'▰'.repeat(n) + '▱'.repeat(LEVELS.length - n)}</Text>
309 <Text color={COLORS[use]} bold>{` effort ${use.toUpperCase()}`}</Text>
310 {tail}
311 </Box>,
312 )
313 }
314 // No pick applies: the session's own effort, drawn like a pick once a request has shown it.
315 const unpicked = (note: JSX.Element) => LEVELS.includes(session as Effort)
316 ? level(session as Effort, <Box flexDirection="row"><Text dimColor>{' · '}</Text>{note}</Box>)
317 : row(note)
318
319 if (off) return unpicked(<Text dimColor>{'off, session effort applies '}</Text>)
320 if (!j) return unpicked(<Text dimColor>{'session effort until the next prompt '}</Text>)
321 if (j.phase === 'asking') return row(<Text dimColor>◌ asking Jev... </Text>)
322 if (j.phase === 'kept') {
323 return unpicked(<Box flexDirection="row"><Text color="yellow">kept session effort</Text><Text dimColor>{` · ${j.reason} `}</Text></Box>)
324 }
325
326 const use = applied(j.pick, session, canRaise)
327 const note = (j.from === undefined || j.from === use ? '' : ` (was ${j.from})`) + (use === j.pick ? '' : ` (Jev: ${j.pick}, capped)`)
328 return level(
329 use,
330 <Box flexDirection="row">
331 {j.sent === undefined
332 ? <Text dimColor>{' · not sent yet'}</Text>
333 : j.sent === use
334 ? <Text color="green">{' ✓ sent'}</Text>
335 : <Text color="yellow">{` ⚠ model got ${j.sent}`}</Text>}
336 <Text dimColor>{`${note} · Jev ${Math.round(j.confidence * 100)}% sure · ${Math.round(j.ms)}ms `}</Text>
337 </Box>,
338 )
339 })
340}
341types/index.d.ts 26 lines1export type Effort = 'low' | 'medium' | 'high' | 'xhigh' | 'max'
2
3/** The latest prompt's judgement: asking, a pick, or why there is none. */
4export type Judgement =
5 | { phase: 'asking' }
6 | { phase: 'picked'; pick: Effort; score: number; confidence: number; ms: number; from?: Effort | number; sent?: Effort | number }
7 | { phase: 'kept'; reason: string }
8
9declare module 'claude-code' {
10 interface PluginState {
11 'auto-effort': {
12 judgement: Judgement | null
13 isOff: boolean
14 previous: string | null
15 /** The session's own effort, as the last main-thread request carried it before any rewrite. */
16 sessionEffort: Effort | number | null
17 /** A prompt queued (ctrl+x enter) to run as its own turn, and its pick, held until that turn starts. */
18 queued: { text: string; judgement: Judgement | null } | null
19 /** The running main-thread turn, as its requests went out; logged at turn.complete. */
20 /** The end of the main thread's last reply: what a short "yes, fix" approves. */
21 lastReply: string | null
22 run: { turnId: string; prompt: string; session?: Effort | number; requests: number; efforts: (Effort | number)[]; outputTokens: number } | null
23 }
24 }
25}
26