Picks Claude's effort level for each prompt with a System One decision model

A Claude Code mod that picks Claude's effort level for each prompt, using a System One decision model such as Jev. It raises effort for work that needs it, lowers it for quick edits, and otherwise leaves the model's default alone.
Tested with Claude Code v2.1.287.
The repository is its own plugin marketplace:
claude plugin marketplace add arthur-fontaine/cc-mod-auto-effort
claude plugin install auto-effort@auto-effort-dev
Then start a new session (or run /reload-plugins) and choose a provider with /auto-effort setup.
Inside a session, /plugin does the same: add the marketplace arthur-fontaine/cc-mod-auto-effort, then install auto-effort.
To try it without installing, load a checkout for one session; it reloads when you save a file:
git clone https://github.com/arthur-fontaine/cc-mod-auto-effort
claude --plugin-dir ./cc-mod-auto-effort
Add this to the repository's .claude/settings.local.json, and make sure that file is gitignored. Use .claude/settings.json to share it with everyone in the repository, but keep the key out of that file.
{
"extraKnownMarketplaces": {
"auto-effort-dev": { "source": { "source": "github", "repo": "arthur-fontaine/cc-mod-auto-effort" } }
},
"enabledPlugins": { "auto-effort@auto-effort-dev": true }
}
To develop against a checkout, use { "source": "directory", "path": "/path/to/cc-mod-auto-effort" } as the source: edits apply after /reload-plugins. Claude Code honors these entries only once you trust the folder.
The mod asks a decision model how much effort each prompt needs. Run /auto-effort setup to pick one; there's no default, so the mod calls nothing until you do. Your choice is saved across sessions.
| Model | Runs | Right level | Latency, median | You need | | :- | :- | -: | -: | :- | | Nisev 1.7B | On your machine | 65.4% | 129 ms | llama.cpp, and a one-time 1.9 GB download | | Jev 1.13 | Cloud | 61.8% | 588 ms | An API key from a provider that serves it | | Clef-Flash 9B, Clef 27B | Cloud | 56.6%, 51.5% | 309, 495 ms | An API key from a provider that serves it |
"Right level" is how often the model picked the effort that two Claude teachers agreed on, over 136 real turns; always keeping the default gets 51.5%. Latency is from an M4 Pro; for cloud models it depends on the provider. Details in the benchmark.
llama-server or the unified llama CLI. Get it from its releases or a package manager.127.0.0.1:8765 for the session, using about 2 GB of memory, and stops it with the session.Any provider that serves the System One API works. Setup has OpenCode Zen and TypeSafe built in; for any other, choose Other and enter its endpoint and model. Then set the provider's key as AUTO_EFFORT_API_KEY (where). The mod never stores it.
Some providers that work:
| Provider | Endpoint | Model | | :- | :- | :- | | OpenCode Zen | https://opencode.ai/zen/v1/systemone | jev-1.13 | | TypeSafe | https://api.typesafe.ai/v1/systemone | jev-latest | | OpenRouter | https://openrouter.ai/api/v1/systemone | typesafe/jev-1.13 | | Cloudflare Workers AI | see below | clef or clef-flash |
cc-mod-auto-effort.AUTO_EFFORT_API_KEY.AUTO_EFFORT_API_KEY.dash.cloudflare.com/<account id>/…./auto-effort setup, choose Other, enter this URL, then clef or clef-flash as the model, matching the end of the URL: https://api.cloudflare.com/client/v4/accounts/<account id>/ai/run/@cf/cloudflare/clef
A decision model you run yourself works the same way. For example, start llama-server -hf ggml-org/Kev-4B-GGUF:Q8_0 --port 8080 and choose Other with http://127.0.0.1:8080/v1/systemone.
Privacy: each prompt goes to the endpoint, truncated to 6,000 characters, with up to 1,500 characters of Claude's previous reply and the session's model name. Set AUTO_EFFORT_INCLUDE_CONTEXT=false to send the prompt alone. Your provider's data policy applies; OpenCode says Jev inputs aren't used for training.
| Command | What it does | | :- | :- | | /auto-effort | Shows the provider, what's missing, Nisev's state, and the last decision. | | /auto-effort setup | Chooses the provider. | | /auto-effort off / on | Stops or resumes the overrides. Saved across sessions. |
Variables take precedence over what setup saved, so you can also configure the mod with them alone. Set them in your shell before starting claude, or in the env block of ~/.claude/settings.json (every session) or a repository's .claude/settings.local.json. Never put the key in a settings file that's committed.
{
"env": {
"AUTO_EFFORT_ENDPOINT": "https://opencode.ai/zen/v1/systemone",
"AUTO_EFFORT_API_KEY": "oc_…",
"AUTO_EFFORT_MODEL": "jev-1.13"
}
}
The mod reads them on each prompt; after editing a settings file, start a new session or run /reload-plugins.
Provider
| Variable | With a cloud endpoint | With Nisev | | :- | :- | :- | | AUTO_EFFORT_PROVIDER | jev, for any System One endpoint (Jev, Clef…). Implied by AUTO_EFFORT_ENDPOINT. | nisev | | AUTO_EFFORT_ENDPOINT | The System One URL. Use https://: the key is sent as a bearer token. | The local URL to serve it at. Defaults to http://127.0.0.1:8765/v1/systemone; if Nisev already answers there, the mod uses that server. | | AUTO_EFFORT_MODEL | The model name the endpoint expects. | The model llama.cpp serves: a Hugging Face repo:quant or a .gguf path. Defaults to arthur-fontaine/nisev-1.7b-GGUF:Q8_0. | | AUTO_EFFORT_API_KEY | The endpoint's key. Only ever read from the environment. | Not used. | | AUTO_EFFORT_LLAMA_SERVER | Not used. | The llama.cpp binary, llama-server or llama. Defaults to the first on your PATH. |
What setup saved applies only to the provider it was saved for. The mod changes nothing, and the status line and /auto-effort say why, while:
AUTO_EFFORT_ENDPOINT or AUTO_EFFORT_MODEL holds a cloud endpoint's value, such as one left in a repository's settings. Unset them to use Nisev's defaults.Behavior
| Variable | Default | What it is | | :- | :- | :- | | AUTO_EFFORT_MIN_CONFIDENCE | 0.5 | Below this confidence, from 0 to 1, the session's effort stands. | | AUTO_EFFORT_TIMEOUT_MS | 4000 | How long to wait for the provider. | | AUTO_EFFORT_MIN_EFFORT / _MAX_EFFORT | low / xhigh | The levels the mod may pick, from low, medium, high, xhigh, max. | | AUTO_EFFORT_INCLUDE_CONTEXT | true | false sends the prompt without Claude's previous reply. |
Anthropic's guidance is that the model's default effort is right for most tasks, so effort should change only for a clear reason. Two models are involved: the decision model you chose in setup (Nisev, Jev or Clef), which picks a level, and the Claude model of your session, which does the work at that level. For each prompt, the mod:
medium, most others to high.low, medium, default, high, xhigh and max, each described with the blog post's rubric.trivial to exhaustive, and a table in hooks/policy.js turns it into a level for your Claude model. For example, multi-step work becomes high, and an ordinary request becomes the model's default.AUTO_EFFORT_MIN_EFFORT–AUTO_EFFORT_MAX_EFFORT (low–xhigh by default). The next prompt is decided afresh.It keeps your session's effort instead, whether that's the Claude model's default or what you set with /effort, when:
AUTO_EFFORT_MIN_CONFIDENCE (50% by default);The prompt footer shows each decision, beside Claude Code's own mode labels, with the decision model's confidence (in the terminal and the desktop app):
| Footer | Meaning | | :- | :- | | effort high · 82% | It picked high, 82% sure, so this turn runs at high. | | effort default · 70% | It picked the Claude model's default, so your session's effort stands. | | effort default (unsure: high · 41%) | It leaned to high, but only 41% sure, so your session's effort stands. | | effort xhigh (picked max) · 75% | It picked max, above the allowed range, so the turn runs at xhigh. |
The test: 136 held-out turns from 35 real coding sessions. The answer key is the effort level two Claude teachers, Sonnet 5.5 and Opus 5.5, agreed on after seeing what happened in the turn. Always keeping the model's default gets 51.5% right.
The setup: each model got the request the mod sends it, one at a time. The local models ran in llama.cpp b11406 on an Apple M4 Pro with 24 GB. Cloud latency includes the round trip from that Mac. Clef and Clef-Flash ran on Workers AI, since neither is practical on a 24 GB Mac: Clef-Flash in llama.cpp took 3.4 s per prompt at the median, at 4-bit.
The green zone is where a picker should be. Both edges are set by a rule, not read off the results:
plot.py.| Model | Runs | Right level | Right level on Opus 5.5 | Right level applied | Latency p50 / p95 | | :- | :- | -: | -: | -: | -: | | Nisev 1.7B | Local · llama.cpp · Q8_0, 1.8 GB | 65.4% | 73.5% | 66.9% | 129 / 364 ms | | Kev-4B | Local · llama.cpp · Q8_0, 4.5 GB | 55.1% | 60.3% | 52.2% | 1750 / 3236 ms | | Clef-Flash 9B | Cloud · Cloudflare Workers AI | 56.6% | 64.7% | 51.5% | 309 / 909 ms | | Clef 27B | Cloud · Cloudflare Workers AI | 51.5% | 65.4% | 51.5% | 495 / 880 ms | | Jev 1.13 | Cloud · OpenCode Zen | 61.8% | 70.6% | 62.5% | 588 / 770 ms |
medium.What stands out:
medium (40% of turns), right on Opus 5.5 and a step too low on models that default to high. It also answers max on 12 turns.To rerun it, run uv run python pipeline/benchmark.py in training/, which also explains how Nisev was built.
pnpm test # claude plugin test: hooks with stubbed providers and processes, no network
pnpm validate # claude plugin validate --strict
pnpm eval # sample prompts against the live provider, reading .env
pnpm eval needs a cloud endpoint's three variables, in the environment or a gitignored .env (see .env.example), or AUTO_EFFORT_PROVIDER=nisev with Nisev already served. It reports the session model as claude-opus-5-5; set AUTO_EFFORT_EVAL_CLAUDE_MODEL to try another.jev-1.13 and Nisev. The Jev run took 0.4–1 s per call and about 7k input tokens ($0.0003).claude --plugin-dir . --debug-file /tmp/cc.log logs each override as [auto-effort] … effort low → high.hooks/register.js 402 lines1import { atom, read, update } from 'claude-code'
2
3import { buildRequest, decide, describe, JEV_PRESETS, LEVELS, NISEV, resolveConfig } from './policy.js'
4import { parseBuild, parseServerOutput, serverArgs, serverCommand, servesOurModel } from './nisev.js'
5
6const USER_ORIGINS = ['composer', 'bridge', 'sdk']
7
8// Decision for the prompt whose turn hasn't started yet.
9let pending = null
10// turnId -> decision, for turns in flight.
11const decisions = new Map()
12let enabled = true
13let lastDecision = null
14let warnedMissing = false
15
16// What the mod reports without asking anything of you (the last decision, Nisev starting) goes
17// to the footer's mode labels; the status line, which the engine draws as a notice, is kept for
18// what needs action.
19const label = atom({ plugin: 'auto-effort', key: 'label' }, null)
20
21function show($, text) {
22 return update($, label, () => text).catch(() => {})
23}
24// The llama-server this module started, if any: { state: starting | downloading | ready | stopped }.
25let server = null
26// Set once a request reached Nisev, so a prompt doesn't probe the port every time.
27let nisevConfirmed = false
28
29function nisevState() {
30 return server && server.state !== 'stopped' ? server.state : nisevConfirmed ? 'ready' : 'stopped'
31}
32
33async function llamaServerBuild($, binary) {
34 try {
35 const { stdout, stderr } = await $.process.run([binary, '--version'], { timeoutMs: 10_000 })
36 return parseBuild(stdout + stderr)
37 } catch {
38 return null // Not installed.
39 }
40}
41
42// The configured binary, else the first of llama-server and llama that runs: { binary, build }.
43async function findServer($, config) {
44 for (const binary of config.llamaServer ? [config.llamaServer] : NISEV.servers) {
45 const build = await llamaServerBuild($, binary)
46 if (build !== null) return { binary, build }
47 }
48 return { binary: config.llamaServer ?? NISEV.servers[0], build: null }
49}
50
51async function isServing($, config) {
52 try {
53 const response = await $.http.fetch(`http://127.0.0.1:${config.port}/v1/models`)
54 return response.ok && servesOurModel(response.text)
55 } catch {
56 return false
57 }
58}
59
60// Starts llama-server for the rest of the session unless our model is already served on the
61// port; never waits. The first start downloads the model (-hf), reported on the status line.
62async function startNisev($, config) {
63 if (server && server.state !== 'stopped') return server.state
64 if (await isServing($, config)) {
65 nisevConfirmed = true
66 return 'ready'
67 }
68 const current = { state: 'starting' }
69 server = current
70 const { binary } = await findServer($, config)
71 // llama.cpp may print nothing while it downloads, so say up front that the first start can take minutes.
72 void show($, 'Nisev starting · the first start downloads ' + NISEV.sizeLabel)
73 void (async () => {
74 try {
75 // The loop is the child's life: it ends with the child or with this module.
76 for await (const { text } of $.process.spawn({ argv: serverArgs(config, binary) })) {
77 const seen = parseServerOutput(text)
78 if (seen.ready && current.state !== 'ready') {
79 current.state = 'ready'
80 nisevConfirmed = true
81 void show($, 'Nisev ready')
82 } else if (seen.percent !== undefined && current.state !== 'ready') {
83 current.state = 'downloading'
84 void show($, 'Nisev downloading · ' + seen.percent + '%')
85 }
86 $.ui.log(text.trimEnd(), { to: 'debug' })
87 }
88 } catch (error) {
89 $.ui.log('llama-server failed to start: ' + error.message, { to: 'debug' })
90 void show($, null)
91 $.ui.status('auto-effort: llama-server failed to start, see the debug log')
92 }
93 current.state = 'stopped'
94 nisevConfirmed = false
95 })()
96 return current.state
97}
98
99async function storedConfig($) {
100 try {
101 const saved = await $.store.get('config')
102 return saved && typeof saved === 'object' ? saved : {}
103 } catch {
104 return {}
105 }
106}
107
108async function loadConfig($) {
109 const env = {
110 provider: await $.env.get('AUTO_EFFORT_PROVIDER'),
111 apiKey: await $.env.get('AUTO_EFFORT_API_KEY'),
112 endpoint: await $.env.get('AUTO_EFFORT_ENDPOINT'),
113 model: await $.env.get('AUTO_EFFORT_MODEL'),
114 minConfidence: await $.env.get('AUTO_EFFORT_MIN_CONFIDENCE'),
115 timeoutMs: await $.env.get('AUTO_EFFORT_TIMEOUT_MS'),
116 minEffort: await $.env.get('AUTO_EFFORT_MIN_EFFORT'),
117 maxEffort: await $.env.get('AUTO_EFFORT_MAX_EFFORT'),
118 includeContext: await $.env.get('AUTO_EFFORT_INCLUDE_CONTEXT'),
119 llamaServer: await $.env.get('AUTO_EFFORT_LLAMA_SERVER'),
120 }
121 return resolveConfig(env, await storedConfig($))
122}
123
124async function previousReply($) {
125 try {
126 const messages = await $.session.messages()
127 for (let i = messages.length - 1; i >= 0; i--) {
128 if (messages[i].role === 'assistant' && messages[i].text) return messages[i].text
129 }
130 } catch {
131 // No transcript to read; classify the prompt alone.
132 }
133 return undefined
134}
135
136async function sessionModel($) {
137 try {
138 return await $.session.model()
139 } catch {
140 return undefined
141 }
142}
143
144// The same System One call for every provider: a hosted Jev, or llama-server on this machine.
145async function askSystemOne($, config, body) {
146 const request = $.http.fetch(config.endpoint, {
147 method: 'POST',
148 headers: { Authorization: 'Bearer ' + config.apiKey, 'Content-Type': 'application/json' },
149 body: JSON.stringify(body),
150 })
151 // $.http.fetch takes no timeout, and a slow classifier must never hold up the prompt.
152 const timeout = $.clock.sleep(config.timeoutMs).then(() => null)
153 const response = await Promise.race([request, timeout])
154 if (response === null) throw new Error('timed out after ' + config.timeoutMs + 'ms')
155 if (response.status === 404 && config.provider === 'nisev') {
156 throw new Error('llama-server has no /v1/systemone; it needs build ' + NISEV.minBuild + ' or later')
157 }
158 if (!response.ok) throw new Error('HTTP ' + response.status + ' ' + response.text.slice(0, 200))
159 return JSON.parse(response.text)
160}
161
162async function classify($, prompt) {
163 const config = await loadConfig($)
164 if (!config.provider) {
165 if (!warnedMissing) {
166 warnedMissing = true
167 $.ui.status('auto-effort: run /auto-effort setup, or set ' + config.missing.join(', '))
168 }
169 return null
170 }
171 if (config.missing.length || config.problems.length) {
172 if (!warnedMissing) {
173 warnedMissing = true
174 $.ui.status(config.problems.length
175 ? 'auto-effort: Nisev is selected, but ' + config.problems.join('; ') + '. Unset them to use its defaults.'
176 : 'auto-effort: set ' + config.missing.join(', '))
177 }
178 return null
179 }
180 if (config.provider === 'nisev' && nisevState() !== 'ready') {
181 // Start it for the next prompts; this one keeps the session's effort.
182 const state = await startNisev($, config)
183 if (state !== 'ready') return { effort: null, reason: 'Nisev ' + state }
184 }
185 const model = await sessionModel($)
186 const body = buildRequest({ prompt, previousReply: await previousReply($), model }, config)
187 try {
188 const decision = decide(await askSystemOne($, config, body), config, model)
189 if (config.provider === 'nisev') nisevConfirmed = true
190 return decision
191 } catch (error) {
192 $.ui.log('effort request failed: ' + error.message, { to: 'debug' })
193 if (config.provider === 'nisev') nisevConfirmed = false
194 return { effort: null, reason: config.provider === 'nisev' ? 'Nisev unavailable' : 'endpoint unavailable' }
195 }
196}
197
198async function saveConfig($, config) {
199 try {
200 await $.store.set('config', config)
201 return true
202 } catch {
203 return false
204 }
205}
206
207async function ask($, question, options) {
208 try {
209 return await $.ui.ask(question, options)
210 } catch {
211 return null // Dismissed, or no one to ask (a -p run).
212 }
213}
214
215async function setupJev($, endpoint, model) {
216 const saved = await saveConfig($, { provider: 'jev', endpoint, model })
217 const key = await $.env.get('AUTO_EFFORT_API_KEY')
218 return {
219 text: [
220 saved ? 'Saved: ' + endpoint + ', model ' + model + '.' : 'Could not save the setting for later sessions.',
221 key
222 ? 'AUTO_EFFORT_API_KEY is set, so the next prompt uses it.'
223 : 'Now set AUTO_EFFORT_API_KEY to the key for that endpoint, in your shell or in the env block of a ' +
224 'settings file that is not committed. The key is never stored by the mod.',
225 'AUTO_EFFORT_* environment variables, when set, take precedence over this setup.',
226 ].join('\n'),
227 }
228}
229
230async function setupNisev($) {
231 const env = { provider: 'nisev', llamaServer: await $.env.get('AUTO_EFFORT_LLAMA_SERVER') }
232 const config = resolveConfig(env, await storedConfig($))
233 const { binary, build } = await findServer($, config)
234 if (build === null) {
235 return {
236 text:
237 'llama.cpp was not found. Install build ' + NISEV.minBuild + ' or later, from ' +
238 'https://github.com/ggml-org/llama.cpp/releases, so that llama-server or llama is on your PATH, or ' +
239 'set AUTO_EFFORT_LLAMA_SERVER to its path. Then run /auto-effort setup again.',
240 }
241 }
242 if (build < NISEV.minBuild) {
243 return {
244 text:
245 binary + ' is build ' + build + '; decision models need build ' + NISEV.minBuild + ' or later. ' +
246 'Update it from https://github.com/ggml-org/llama.cpp/releases and run /auto-effort setup again.',
247 }
248 }
249 config.llamaServer = binary
250 const serving = await isServing($, config)
251 if (!serving) {
252 const go = await ask($, 'Download ' + config.model + ' (' + NISEV.sizeLabel + ') and run it with llama.cpp?', {
253 header: 'Download',
254 options: ['Download and start', 'Cancel'],
255 })
256 if (go !== 'Download and start') return { text: 'Setup cancelled; nothing changed.' }
257 }
258 await saveConfig($, { provider: 'nisev', model: config.model, endpoint: config.endpoint, llamaServer: config.llamaServer })
259 await startNisev($, config)
260 return {
261 text: [
262 'Saved: Nisev, served by ' + serverCommand(binary).join(' ') + ' (build ' + build + ') at ' + config.endpoint + '.',
263 serving
264 ? 'It is already running.'
265 : 'The first start downloads the model; the status line shows the progress. Until it is ready, ' +
266 'prompts keep the session effort.',
267 'llama-server runs while this session does; the next session starts it again from the cache.',
268 'AUTO_EFFORT_* environment variables, when set, take precedence over this setup.',
269 ].join('\n'),
270 }
271}
272
273async function setup($) {
274 const choice = await ask($, 'Which classifier should pick the effort of each prompt?', {
275 header: 'Classifier',
276 options: ['Nisev (local)', ...Object.keys(JEV_PRESETS)],
277 })
278 if (choice === null) return { text: 'Setup cancelled; nothing changed.' }
279 if (choice === 'Nisev (local)') return setupNisev($)
280 if (JEV_PRESETS[choice]) return setupJev($, JEV_PRESETS[choice].endpoint, JEV_PRESETS[choice].model)
281 // "Other": any System One endpoint.
282 if (!/^https?:\/\/\S+$/.test(choice.trim())) {
283 return { text: 'Under Other, enter the URL of a System One endpoint, such as https://example.com/v1/systemone.' }
284 }
285 const model = await ask($, 'Which model name does that endpoint expect?', { header: 'Model', options: ['jev-latest', 'jev-1.13'] })
286 if (model === null) return { text: 'Setup cancelled; nothing changed.' }
287 return setupJev($, choice.trim(), model.trim())
288}
289
290async function status($) {
291 const config = await loadConfig($)
292 const lines = ['Effort picker: ' + (enabled ? 'on' : 'off')]
293 if (config.provider === 'nisev') {
294 lines.push('Provider: Nisev ' + config.model + ' · llama.cpp at ' + config.endpoint + ' (' + nisevState() + ')')
295 for (const problem of config.problems) lines.push('Problem: ' + problem)
296 } else {
297 lines.push('Provider: ' + (config.provider ? 'System One endpoint' : 'not set, run /auto-effort setup'))
298 lines.push('Endpoint: ' + (config.endpoint ?? 'unset') + ' · model ' + (config.model ?? 'unset'))
299 lines.push('API key: ' + (config.apiKey ? 'set' : 'unset'))
300 if (config.missing.length) lines.push('Missing: ' + config.missing.join(', '))
301 }
302 lines.push(
303 'Range: ' + config.minEffort + '–' + config.maxEffort + ' of ' + LEVELS.join(', ') +
304 ' · min confidence ' + config.minConfidence,
305 'Last decision: ' + (lastDecision ? describe(lastDecision) : 'none yet'),
306 )
307 return { text: lines.join('\n') }
308}
309
310export function register(on) {
311 on('session.start', async ($, e, next) => {
312 try {
313 const saved = await $.store.get('enabled')
314 if (typeof saved === 'boolean') enabled = saved
315 } catch {
316 // Keep the default.
317 }
318 try {
319 await $.command.register({
320 name: 'auto-effort',
321 description: 'Set up, show or toggle the effort picker',
322 argumentHint: '[setup|on|off|status]',
323 immediate: true,
324 })
325 } catch (error) {
326 $.ui.log('could not register /auto-effort: ' + error.message, { to: 'debug' })
327 }
328 const result = await next(e)
329 // Warm Nisev up before the first prompt needs it.
330 const config = await loadConfig($)
331 if (enabled && config.provider === 'nisev' && !config.problems.length) await startNisev($, config)
332 return result
333 })
334
335 on('prompt.submit', async ($, e, next) => {
336 // A delivery into a running turn (a peer's message, a notification) isn't a
337 // new request from the user, so the running turn keeps its decision. Prompts
338 // the user types mid-turn are queued and get a turn of their own.
339 const intoRunningTurn = e.turnId && !USER_ORIGINS.includes(e.origin?.kind)
340 if (!enabled || intoRunningTurn || !e.text.trim()) return next(e)
341 const decision = await classify($, e.text)
342 if (decision) {
343 lastDecision = decision
344 await show($, describe(decision))
345 $.ui.status(undefined)
346 pending = decision
347 }
348 const result = await next(e)
349 // A later hook dropped the prompt: don't let its decision reach another turn.
350 if (result.drop) pending = null
351 return result
352 })
353
354 on('turn.start', async ($, e, next) => {
355 if (pending) {
356 decisions.set(e.turnId, pending)
357 pending = null
358 }
359 return next(e)
360 })
361
362 on('ui.render', { component: 'SessionMode' }, async ($, e, next) => {
363 const text = await read($, label)
364 return text ? next({ ...e, props: { ...e.props, modes: [...e.props.modes, text] } }) : next(e)
365 })
366
367 on('turn.step', async function* ($, e, next) {
368 const decision = decisions.get(e.turnId)
369 // Subagents keep their own effort; a model without effort, or a numeric
370 // budget set by hand, is left alone.
371 if (!enabled || e.agentId || !decision?.effort || typeof e.effort !== 'string') {
372 return yield* next(e)
373 }
374 if (e.index === 0) $.ui.log('effort ' + e.effort + ' → ' + decision.effort, { to: 'debug' })
375 return yield* next({ ...e, effort: decision.effort })
376 })
377
378 on('turn.complete', async ($, e, next) => {
379 decisions.delete(e.turnId)
380 return next(e)
381 })
382
383 on('command.run', { command: 'auto-effort' }, async ($, e) => {
384 const arg = e.args.trim().toLowerCase()
385 if (arg === 'setup') return setup($)
386 if (arg === 'on' || arg === 'off') {
387 enabled = arg === 'on'
388 try {
389 await $.store.set('enabled', enabled)
390 } catch {
391 // Applies to this session only.
392 }
393 $.ui.status(undefined)
394 await show($, enabled ? null : 'effort picker off')
395 return { text: 'Effort picker turned ' + arg + '.' }
396 }
397 if (arg && arg !== 'status') return { text: 'Usage: /auto-effort [setup|on|off|status]' }
398 return status($)
399 })
400}
401
402hooks/policy.js 295 lines1// Pure decision logic: no `$`, so it can be unit tested on its own.
2
3export const LEVELS = ['low', 'medium', 'high', 'xhigh', 'max']
4
5export const DEFAULTS = {
6 minConfidence: 0.5,
7 timeoutMs: 4000,
8 minEffort: 'low',
9 maxEffort: 'xhigh',
10 includeContext: true,
11}
12
13// The fine-tuned classifier, served by llama.cpp's llama-server as a decision model.
14export const NISEV = {
15 // A Hugging Face repo for `llama-server -hf`, or a path to a .gguf file.
16 model: 'arthur-fontaine/nisev-1.7b-GGUF:Q8_0',
17 sizeLabel: 'about 1.9 GB',
18 endpoint: 'http://127.0.0.1:8765/v1/systemone',
19 // Tried in order when none is configured: the server binary, or the unified CLI (`llama serve`).
20 servers: ['llama-server', 'llama'],
21 alias: 'nisev',
22 // The first llama.cpp build with /v1/systemone (ggml-org/llama.cpp#29818).
23 minBuild: 11361,
24}
25
26export const JEV_PRESETS = {
27 'OpenCode Zen': { endpoint: 'https://opencode.ai/zen/v1/systemone', model: 'jev-1.13' },
28 TypeSafe: { endpoint: 'https://api.typesafe.ai/v1/systemone', model: 'jev-latest' },
29}
30
31// Worded after https://claude.com/blog/claude-model-and-effort-level-in-claude-code:
32// the model's default is right for most tasks, effort controls how thorough
33// Claude is (files read, verification, how far it pushes before checking in),
34// and deviating should need a clear reason.
35export const EFFORT_QUESTION = {
36 type: 'choice',
37 instructions: {
38 task:
39 'A developer sent `latest_user_message` to Claude, an AI coding agent working in their repository. ' +
40 'Pick how much effort Claude should spend on it. Effort controls how many files Claude reads, ' +
41 'how much it verifies (running tests, double-checking), and how far it pushes through a multi-step ' +
42 'task before checking back in. It is not about how capable Claude is. Most requests should get ' +
43 '`default`. Pick another level only when the request clearly calls for less or more thoroughness. ' +
44 '`previous_assistant_reply`, when present, is what Claude last said; use it to understand short ' +
45 'follow-ups such as "yes, do it". `model`, when present, is the Claude model; `default` keeps that ' +
46 "model's own default effort.",
47 },
48 criteria: {
49 low:
50 'Routine work that needs no investigation: a precisely described edit, a rename, a typo, a one-line ' +
51 'change, a question about code already in context, a quick lookup or shell command, small talk.',
52 medium:
53 'Light work: a small, well-scoped change or a focused answer that needs a little reading, but no ' +
54 'multi-file investigation.',
55 default:
56 'A typical coding request: an ordinary feature, a normal bug fix, a focused explanation, or anything ' +
57 'unclear. The model default already scales the work to the task.',
58 high:
59 'Multi-step work that must be verified: changes across several files, a bug that needs several ' +
60 'hypotheses checked, a refactor that has to be finished completely, work where skipping a file or ' +
61 'not running the tests would make the result wrong.',
62 xhigh:
63 'Long, hard, high-stakes work: a large migration or refactor, a subtle bug across systems, an ' +
64 'architecture decision, a security or correctness audit, or a request that explicitly asks Claude to ' +
65 'be thorough and verify everything.',
66 max:
67 'Exhaustive work where cost does not matter: the request explicitly demands maximum effort or ' +
68 'leaving nothing unchecked, on a critical, very large task.',
69 },
70}
71
72// The question Nisev was trained on, word for word (training/pipeline/task.py): it
73// sizes the request, and the model's table below turns the size into a level.
74export const CATEGORIES = ['trivial', 'light', 'ordinary', 'multi_step', 'hard', 'exhaustive']
75export const CATEGORY_QUESTION = {
76 type: 'choice',
77 instructions:
78 'A developer sent latest_user_message to an AI coding agent working in their repository. ' +
79 'How much thoroughness does it need: how many files to read, how much to verify, and how far ' +
80 'to push before checking back in? previous_assistant_reply, when present, is the agent\'s last ' +
81 'reply; use it to size short follow-ups such as "yes, do it".',
82 criteria: {
83 trivial: 'Routine: a precise small edit, a lookup, a question about code in context, an acknowledgement.',
84 light: 'A small, well-scoped change or focused answer that needs a little reading.',
85 ordinary: 'A typical feature, bug fix or explanation, or anything unclear.',
86 multi_step: 'Several files, several hypotheses, or work that must be verified by running tests.',
87 hard: 'A large migration, a subtle cross-system bug, an audit, or an explicit ask to be thorough.',
88 exhaustive: 'An explicit demand for maximum effort on a critical, very large task.',
89 },
90}
91
92// The level each kind of request gets on each model (training/pipeline/models.py). Levels are
93// calibrated per model: Opus 5.5 defaults to medium, so verified multi-step work is where it
94// pays to raise effort there, while models that default to high already run it at high.
95const HIGH = { trivial: 'low', light: 'medium', ordinary: 'high', multi_step: 'high', hard: 'xhigh', exhaustive: 'max' }
96const MEDIUM = { ...HIGH, ordinary: 'medium' }
97const NO_XHIGH = { ...HIGH, hard: 'high' }
98const NO_MAX = { ...HIGH, hard: 'high', exhaustive: 'high' }
99export const PROFILES = {
100 'claude-opus-5-5': { default: 'medium', table: MEDIUM },
101 'claude-sonnet-5-5': { default: 'high', table: HIGH },
102 'claude-fable-5-1': { default: 'high', table: HIGH },
103 'claude-opus-5': { default: 'high', table: HIGH },
104 'claude-sonnet-5': { default: 'high', table: HIGH },
105 'claude-fable-5': { default: 'high', table: HIGH },
106 'claude-opus-4-8': { default: 'high', table: HIGH },
107 'claude-opus-4-7': { default: 'high', table: HIGH },
108 'claude-opus-4-6': { default: 'high', table: NO_XHIGH },
109 'claude-sonnet-4-6': { default: 'high', table: NO_XHIGH },
110 'claude-opus-4-5': { default: 'high', table: NO_MAX },
111}
112const FALLBACK = { default: 'high', table: HIGH }
113const ALIASES = { opus: 'claude-opus-5-5', sonnet: 'claude-sonnet-5-5', fable: 'claude-fable-5-1' }
114
115// `claude-opus-5-5[1m]`, `us.anthropic.claude-opus-5-5-v1`, `opus` -> `claude-opus-5-5`
116export function normalizeModel(model) {
117 if (!model) return undefined
118 const m = String(model).toLowerCase().split('/').pop()
119 .replace(/^((us|eu|apac|global)\.)?anthropic\./, '')
120 .replace(/\[.*?\]$|:.*$|-v\d+$|-\d{8}$|-fast$/g, '')
121 .replaceAll('.', '-')
122 return ALIASES[m] ?? m
123}
124
125export function profileOf(model) {
126 return PROFILES[normalizeModel(model)] ?? FALLBACK
127}
128
129// Haiku 4.5 takes no effort; an unknown or missing model gets the common profile.
130export function takesEffort(model) {
131 return !(normalizeModel(model) ?? '').includes('haiku')
132}
133
134// Keeps the request well under the classifier's context.
135const CLIP = {
136 jev: { prompt: 6000, context: 1500 },
137 // What Nisev was trained with (training/pipeline/task.py).
138 nisev: { prompt: 3000, context: 800 },
139}
140
141// Counts code points, as the Python that built the training data does, so a long prompt is
142// cut at the same place.
143export function clip(text, max) {
144 const chars = Array.from(text)
145 if (chars.length <= max) return text
146 // Keep both ends: the ask is usually at the start, pasted output at the end.
147 const half = Math.floor((max - 20) / 2)
148 return chars.slice(0, half).join('') + '\n[… truncated …]\n' + chars.slice(-half).join('')
149}
150
151function pick(...values) {
152 return values.find((v) => v !== undefined && v !== null && v !== '')
153}
154
155function toNumber(value, fallback) {
156 const n = typeof value === 'number' ? value : Number(value)
157 return Number.isFinite(n) ? n : fallback
158}
159
160function toLevel(value, fallback) {
161 return LEVELS.includes(value) ? value : fallback
162}
163
164function toBool(value, fallback) {
165 if (typeof value === 'boolean') return value
166 if (value === 'true' || value === '1') return true
167 if (value === 'false' || value === '0') return false
168 return fallback
169}
170
171// The endpoint, key, and model of a Jev provider have no default, so the mod never calls
172// a provider the user didn't choose.
173export const REQUIRED = {
174 endpoint: 'AUTO_EFFORT_ENDPOINT',
175 apiKey: 'AUTO_EFFORT_API_KEY',
176 model: 'AUTO_EFFORT_MODEL',
177}
178
179// `env` holds the AUTO_EFFORT_* values, keyed as in the returned config. `stored` holds what
180// `/auto-effort setup` saved: never a key, and the environment wins over it.
181// The port Nisev serves on, from its endpoint; null unless the endpoint is on this machine.
182export function localPort(endpoint) {
183 const match = /^http:\/\/(?:127\.0\.0\.1|localhost):(\d+)\/v1\/systemone$/.exec(endpoint ?? '')
184 return match ? Number(match[1]) : null
185}
186
187// What llama.cpp can serve: a .gguf path, or a Hugging Face repo with an optional :quant.
188export function isGgufSource(model) {
189 return /\.gguf$/i.test(model ?? '') || /^[\w.-]+\/[\w.-]+(:[\w.-]+)?$/.test(model ?? '')
190}
191
192export function resolveConfig(env = {}, stored = {}) {
193 // The environment wins: an endpoint there means a cloud endpoint unless AUTO_EFFORT_PROVIDER says otherwise.
194 const provider = [env.provider, env.endpoint && 'jev', stored.provider].find((p) => ['jev', 'nisev'].includes(p))
195 // Saved values belong to the provider they were saved for: a cloud model name is no GGUF to serve.
196 const saved = stored.provider === provider ? stored : {}
197 const nisev = provider === 'nisev'
198 const config = {
199 provider,
200 minConfidence: toNumber(pick(env.minConfidence), DEFAULTS.minConfidence),
201 timeoutMs: toNumber(pick(env.timeoutMs), DEFAULTS.timeoutMs),
202 minEffort: toLevel(pick(env.minEffort), DEFAULTS.minEffort),
203 maxEffort: toLevel(pick(env.maxEffort), DEFAULTS.maxEffort),
204 includeContext: toBool(pick(env.includeContext), DEFAULTS.includeContext),
205 // For Nisev, the model is what llama.cpp serves: a Hugging Face repo:quant or a .gguf path.
206 endpoint: pick(env.endpoint, saved.endpoint, nisev ? NISEV.endpoint : undefined),
207 model: pick(env.model, saved.model, nisev ? NISEV.model : undefined),
208 apiKey: nisev ? 'none' : pick(env.apiKey),
209 llamaServer: nisev ? pick(env.llamaServer, saved.llamaServer) : undefined,
210 }
211 if (nisev) {
212 config.port = localPort(config.endpoint)
213 // Usually a cloud endpoint's settings left in the environment, so say what they are.
214 config.problems = [
215 ...(config.port ? [] : ['AUTO_EFFORT_ENDPOINT is ' + config.endpoint + ', not a local URL such as ' + NISEV.endpoint]),
216 ...(isGgufSource(config.model) ? [] : ['AUTO_EFFORT_MODEL is ' + config.model + ', not a Hugging Face repo:quant or a .gguf path']),
217 ]
218 config.missing = []
219 } else {
220 config.problems = []
221 config.missing = Object.keys(REQUIRED).filter((key) => !config[key]).map((key) => REQUIRED[key])
222 }
223 return config
224}
225
226export function buildRequest({ prompt, previousReply, model }, config) {
227 const limits = CLIP[config.provider === 'nisev' ? 'nisev' : 'jev']
228 const state = { latest_user_message: clip(prompt, limits.prompt) }
229 if (config.includeContext && previousReply) {
230 state.previous_assistant_reply = clip(previousReply, limits.context)
231 }
232 // Effort levels are calibrated per model, so the classifier needs to know which one runs.
233 if (model) state.model = model
234 const question = config.provider === 'nisev' ? CATEGORY_QUESTION : EFFORT_QUESTION
235 // llama.cpp answers to the alias it serves Nisev under, whatever file it loaded.
236 return { model: config.provider === 'nisev' ? NISEV.alias : config.model, state, questions: { effort: question } }
237}
238
239export function clamp(level, min, max) {
240 const lo = LEVELS.indexOf(min)
241 const hi = Math.max(lo, LEVELS.indexOf(max))
242 const i = LEVELS.indexOf(level)
243 return LEVELS[Math.min(Math.max(i, lo), hi)]
244}
245
246// Nisev answers with category probabilities: sum them into levels through the
247// model's table, so the confidence is the probability of the level it picks.
248export function levelProbabilities(categoryProbabilities, model) {
249 const table = profileOf(model).table
250 const levels = Object.fromEntries(LEVELS.map((level) => [level, 0]))
251 for (const category of CATEGORIES) levels[table[category]] += categoryProbabilities?.[category] ?? 0
252 return levels
253}
254
255function decideNisev(answer, config, model) {
256 if (!takesEffort(model)) return { effort: null, reason: 'model takes no effort' }
257 const levels = levelProbabilities(answer.probabilities, model)
258 const level = LEVELS.reduce((best, l) => (levels[l] > levels[best] ? l : best), LEVELS[0])
259 const confidence = levels[level]
260 const category = answer.choice
261 // The model's default is what `default` keeps: the session's own effort stands.
262 if (level === profileOf(model).default) return { effort: null, choice: 'default', confidence, category, reason: 'default' }
263 if (confidence < config.minConfidence) return { effort: null, choice: level, confidence, category, reason: 'low confidence' }
264 return { effort: clamp(level, config.minEffort, config.maxEffort), choice: level, confidence, category, reason: 'jev' }
265}
266
267// Returns `{ effort, choice, confidence, reason }`. `effort` is null when the
268// session's own effort should stand.
269export function decide(response, config, model) {
270 // Cloudflare Workers AI wraps the System One response in `result`.
271 const answer = (response?.answers ?? response?.result?.answers)?.effort
272 if (!answer || answer.type !== 'choice' || typeof answer.choice !== 'string') {
273 return { effort: null, reason: 'no effort answer' }
274 }
275 if (config.provider === 'nisev') return decideNisev(answer, config, model)
276 const { choice, confidence } = answer
277 if (choice === 'default') return { effort: null, choice, confidence, reason: 'default' }
278 if (!LEVELS.includes(choice)) return { effort: null, choice, confidence, reason: 'unknown choice ' + choice }
279 if (typeof confidence !== 'number' || confidence < config.minConfidence) {
280 return { effort: null, choice, confidence, reason: 'low confidence' }
281 }
282 return { effort: clamp(choice, config.minEffort, config.maxEffort), choice, confidence, reason: 'jev' }
283}
284
285export function describe(decision) {
286 const pct = typeof decision.confidence === 'number' ? ' · ' + Math.round(decision.confidence * 100) + '%' : ''
287 if (decision.effort) {
288 const capped = decision.choice && decision.choice !== decision.effort ? ' (picked ' + decision.choice + ')' : ''
289 return 'effort ' + decision.effort + capped + pct
290 }
291 if (decision.reason === 'default') return 'effort default' + pct
292 if (decision.reason === 'low confidence') return 'effort default (unsure: ' + decision.choice + pct + ')'
293 return 'effort default (' + decision.reason + ')'
294}
295hooks/nisev.js 43 lines1// The Nisev provider: the fine-tuned classifier, served by llama.cpp.
2// Pure helpers; register.js does the calls ($ is only passed within a file).
3
4import { NISEV } from './policy.js'
5
6// `llama` is llama.cpp's unified CLI, whose server is `llama serve`; it takes the same flags.
7export function serverCommand(binary) {
8 return /(^|\/)llama$/.test(binary) ? [binary, 'serve'] : [binary]
9}
10
11export function serverArgs(config, binary) {
12 const source = /\.gguf$/i.test(config.model) ? ['-m', config.model] : ['-hf', config.model]
13 return [
14 ...serverCommand(binary), ...source,
15 '--host', '127.0.0.1', '--port', String(config.port),
16 '--alias', NISEV.alias,
17 // Prompts are at most about 1,300 tokens; two slots let two prompts run at once.
18 '-c', '8192', '-np', '2',
19 ]
20}
21
22// `llama-server --version` prints e.g. "version: 0.5.0-dev (build 11408, commit 9f12cd4a4)".
23export function parseBuild(output) {
24 const match = output.match(/build (\d+)/)
25 return match ? Number(match[1]) : null
26}
27
28// What a piece of llama-server's output says about its start: { ready }, { percent } or {}.
29export function parseServerOutput(text) {
30 if (/listening on/i.test(text)) return { ready: true }
31 const percent = /download/i.test(text) && text.match(/(\d{1,3}(?:\.\d+)?)%/)
32 return percent ? { percent: Math.round(Number(percent[1])) } : {}
33}
34
35// True when /v1/models lists our alias, whoever started the server.
36export function servesOurModel(text) {
37 try {
38 return (JSON.parse(text).data ?? []).some((m) => m.id === NISEV.alias)
39 } catch {
40 return false
41 }
42}
43types/index.d.ts 9 lines1// The values the mod keeps in $.state: the footer label it draws.
2export type AutoEffortLabel = string | null
3
4declare module 'claude-code' {
5 interface PluginState {
6 'auto-effort': { label: AutoEffortLabel }
7 }
8}
9