SLOPSHOPPER

Auto Effort

Picks Claude's effort level for each prompt with a System One decision model

newspinnercommandstatuspromptprocess
★ 1v0.1.1MITupdated 2026-10-05arthur-fontaine/cc-mod-auto-effort
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · auto-effort
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /auto-effort ⎿ auto-effort: Effort picker: on ⎿ auto-effort: Provider: not set, run /auto-effort setup ⎿ auto-effort: Endpoint: unset · model unset ⎿ auto-effort: API key: unset ⎿ auto-effort: Missing: AUTO_EFFORT_ENDPOINT, AUTO_EFFORT_API_KEY, AUTO_EFFORT_MODEL ⎿ auto-effort: Range: low–xhigh of low, medium, high, xhigh, max · min confidence 0.5 ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts ⚠ auto-effort: auto-effort: run /auto-effort setup, or set AUTO_EFFORT_ENDPOINT, AUTO_EFFORT_API_KEY, AUTO_EFFORT_
README

auto-effort

A Claude Code mod that picks Claude's effort level for each prompt, using a System One decision model such as Jev. It raises effort for work that needs it, lowers it for quick edits, and otherwise leaves the model's default alone.

Tested with Claude Code v2.1.287.

Install

The repository is its own plugin marketplace:

claude plugin marketplace add arthur-fontaine/cc-mod-auto-effort
claude plugin install auto-effort@auto-effort-dev

Then start a new session (or run /reload-plugins) and choose a provider with /auto-effort setup.

Inside a session, /plugin does the same: add the marketplace arthur-fontaine/cc-mod-auto-effort, then install auto-effort.

To try it without installing, load a checkout for one session; it reloads when you save a file:

git clone https://github.com/arthur-fontaine/cc-mod-auto-effort
claude --plugin-dir ./cc-mod-auto-effort

Add this to the repository's .claude/settings.local.json, and make sure that file is gitignored. Use .claude/settings.json to share it with everyone in the repository, but keep the key out of that file.

{
  "extraKnownMarketplaces": {
    "auto-effort-dev": { "source": { "source": "github", "repo": "arthur-fontaine/cc-mod-auto-effort" } }
  },
  "enabledPlugins": { "auto-effort@auto-effort-dev": true }
}

To develop against a checkout, use { "source": "directory", "path": "/path/to/cc-mod-auto-effort" } as the source: edits apply after /reload-plugins. Claude Code honors these entries only once you trust the folder.

Choose a provider

The mod asks a decision model how much effort each prompt needs. Run /auto-effort setup to pick one; there's no default, so the mod calls nothing until you do. Your choice is saved across sessions.

  • Nisev is our own model, made for this one question: a Qwen3-1.7B fine-tuned on about 2,400 prompts from real coding-agent sessions, labeled by Claude. It runs on your machine through llama.cpp, so prompts never leave it. See how it was built.
  • Jev (by TypeSafe) and Clef (by Cloudflare) are general decision models, served in the cloud. They answer the mod's question without training for it.

| Model | Runs | Right level | Latency, median | You need | | :- | :- | -: | -: | :- | | Nisev 1.7B | On your machine | 65.4% | 129 ms | llama.cpp, and a one-time 1.9 GB download | | Jev 1.13 | Cloud | 61.8% | 588 ms | An API key from a provider that serves it | | Clef-Flash 9B, Clef 27B | Cloud | 56.6%, 51.5% | 309, 495 ms | An API key from a provider that serves it |

"Right level" is how often the model picked the effort that two Claude teachers agreed on, over 136 real turns; always keeping the default gets 51.5%. Latency is from an M4 Pro; for cloud models it depends on the provider. Details in the benchmark.

Nisev (local)

  • Needs llama.cpp build b11361 or later, as llama-server or the unified llama CLI. Get it from its releases or a package manager.
  • First start: setup asks before downloading. llama.cpp then fetches the model and caches it, which takes about 3 minutes on a fast connection. Prompts keep the session's effort until the prompt footer says Nisev is ready.
  • While it runs: the mod serves it on 127.0.0.1:8765 for the session, using about 2 GB of memory, and stops it with the session.
  • Privacy: nothing leaves your machine, apart from that one-time download.

A cloud model (Jev, Clef)

Any provider that serves the System One API works. Setup has OpenCode Zen and TypeSafe built in; for any other, choose Other and enter its endpoint and model. Then set the provider's key as AUTO_EFFORT_API_KEY (where). The mod never stores it.

Some providers that work:

| Provider | Endpoint | Model | | :- | :- | :- | | OpenCode Zen | https://opencode.ai/zen/v1/systemone | jev-1.13 | | TypeSafe | https://api.typesafe.ai/v1/systemone | jev-latest | | OpenRouter | https://openrouter.ai/api/v1/systemone | typesafe/jev-1.13 | | Cloudflare Workers AI | see below | clef or clef-flash |

  1. Sign in at <https://opencode.ai/auth> and open your workspace.
  2. Open Keys, which lists service accounts, and click Add Service Account. Name it, for example cc-mod-auto-effort.
  3. On the service account, click Add API Key and set Permissions to Inference only. The expiry date is optional.
  4. Copy the key, since it is shown once, and set it as AUTO_EFFORT_API_KEY.
  1. In the Cloudflare dashboard, open My Profile > API Tokens and click Create Token.
  2. Use the Workers AI template. Set a TTL if you want the token to expire.
  3. Copy the token, since it is shown once, and set it as AUTO_EFFORT_API_KEY.
  4. Find your account ID in the dashboard's URL: dash.cloudflare.com/<account id>/….
  5. In /auto-effort setup, choose Other, enter this URL, then clef or clef-flash as the model, matching the end of the URL:
   https://api.cloudflare.com/client/v4/accounts/<account id>/ai/run/@cf/cloudflare/clef

A decision model you run yourself works the same way. For example, start llama-server -hf ggml-org/Kev-4B-GGUF:Q8_0 --port 8080 and choose Other with http://127.0.0.1:8080/v1/systemone.

Privacy: each prompt goes to the endpoint, truncated to 6,000 characters, with up to 1,500 characters of Claude's previous reply and the session's model name. Set AUTO_EFFORT_INCLUDE_CONTEXT=false to send the prompt alone. Your provider's data policy applies; OpenCode says Jev inputs aren't used for training.

Commands

| Command | What it does | | :- | :- | | /auto-effort | Shows the provider, what's missing, Nisev's state, and the last decision. | | /auto-effort setup | Chooses the provider. | | /auto-effort off / on | Stops or resumes the overrides. Saved across sessions. |

Environment variables

Variables take precedence over what setup saved, so you can also configure the mod with them alone. Set them in your shell before starting claude, or in the env block of ~/.claude/settings.json (every session) or a repository's .claude/settings.local.json. Never put the key in a settings file that's committed.

{
  "env": {
    "AUTO_EFFORT_ENDPOINT": "https://opencode.ai/zen/v1/systemone",
    "AUTO_EFFORT_API_KEY": "oc_…",
    "AUTO_EFFORT_MODEL": "jev-1.13"
  }
}

The mod reads them on each prompt; after editing a settings file, start a new session or run /reload-plugins.

Provider

| Variable | With a cloud endpoint | With Nisev | | :- | :- | :- | | AUTO_EFFORT_PROVIDER | jev, for any System One endpoint (Jev, Clef…). Implied by AUTO_EFFORT_ENDPOINT. | nisev | | AUTO_EFFORT_ENDPOINT | The System One URL. Use https://: the key is sent as a bearer token. | The local URL to serve it at. Defaults to http://127.0.0.1:8765/v1/systemone; if Nisev already answers there, the mod uses that server. | | AUTO_EFFORT_MODEL | The model name the endpoint expects. | The model llama.cpp serves: a Hugging Face repo:quant or a .gguf path. Defaults to arthur-fontaine/nisev-1.7b-GGUF:Q8_0. | | AUTO_EFFORT_API_KEY | The endpoint's key. Only ever read from the environment. | Not used. | | AUTO_EFFORT_LLAMA_SERVER | Not used. | The llama.cpp binary, llama-server or llama. Defaults to the first on your PATH. |

What setup saved applies only to the provider it was saved for. The mod changes nothing, and the status line and /auto-effort say why, while:

  • a cloud endpoint lacks its endpoint, key or model;
  • Nisev is selected but AUTO_EFFORT_ENDPOINT or AUTO_EFFORT_MODEL holds a cloud endpoint's value, such as one left in a repository's settings. Unset them to use Nisev's defaults.

Behavior

| Variable | Default | What it is | | :- | :- | :- | | AUTO_EFFORT_MIN_CONFIDENCE | 0.5 | Below this confidence, from 0 to 1, the session's effort stands. | | AUTO_EFFORT_TIMEOUT_MS | 4000 | How long to wait for the provider. | | AUTO_EFFORT_MIN_EFFORT / _MAX_EFFORT | low / xhigh | The levels the mod may pick, from low, medium, high, xhigh, max. | | AUTO_EFFORT_INCLUDE_CONTEXT | true | false sends the prompt without Claude's previous reply. |

How it decides

Anthropic's guidance is that the model's default effort is right for most tasks, so effort should change only for a clear reason. Two models are involved: the decision model you chose in setup (Nisev, Jev or Clef), which picks a level, and the Claude model of your session, which does the work at that level. For each prompt, the mod:

  1. Asks the decision model how much effort the prompt needs. It sends:
  2. your prompt;
  3. Claude's previous reply, because a short follow-up means nothing alone. After Claude asks "Want me to migrate all 40 API handlers and update their tests?", the prompt "yes, do it" is a big task, and only the previous reply says so.
  4. the name of the Claude model, because effort is calibrated per model: Opus 5.5 defaults to medium, most others to high.
  5. Gets a pick and a confidence back: a level, and how sure the decision model is of it, from 0 to 100%.
  6. A cloud model picks one of low, medium, default, high, xhigh and max, each described with the blog post's rubric.
  7. Nisev picks a kind of task, from trivial to exhaustive, and a table in hooks/policy.js turns it into a level for your Claude model. For example, multi-step work becomes high, and an ordinary request becomes the model's default.
  8. Applies the pick to that turn only, within AUTO_EFFORT_MIN_EFFORT–AUTO_EFFORT_MAX_EFFORT (low–xhigh by default). The next prompt is decided afresh.

It keeps your session's effort instead, whether that's the Claude model's default or what you set with /effort, when:

  • the decision model picks the Claude model's default, which means nothing calls for a change;
  • its confidence is below AUTO_EFFORT_MIN_CONFIDENCE (50% by default);
  • it times out or errors, or Nisev is still starting;
  • no provider is set up, the Claude model takes no effort, or the request comes from a subagent.

The prompt footer shows each decision, beside Claude Code's own mode labels, with the decision model's confidence (in the terminal and the desktop app):

| Footer | Meaning | | :- | :- | | effort high · 82% | It picked high, 82% sure, so this turn runs at high. | | effort default · 70% | It picked the Claude model's default, so your session's effort stands. | | effort default (unsure: high · 41%) | It leaned to high, but only 41% sure, so your session's effort stands. | | effort xhigh (picked max) · 75% | It picked max, above the allowed range, so the turn runs at xhigh. |

Benchmark

Right effort level against median latency, one point per model. Nisev is the only one in the fast and accurate zone.

The test: 136 held-out turns from 35 real coding sessions. The answer key is the effort level two Claude teachers, Sonnet 5.5 and Opus 5.5, agreed on after seeing what happened in the turn. Always keeping the model's default gets 51.5% right.

The setup: each model got the request the mod sends it, one at a time. The local models ran in llama.cpp b11406 on an Apple M4 Pro with 24 GB. Cloud latency includes the round trip from that Mac. Clef and Clef-Flash ran on Workers AI, since neither is practical on a 24 GB Mac: Clef-Flash in llama.cpp took 3.4 s per prompt at the median, at 4-bit.

The green zone is where a picker should be. Both edges are set by a rule, not read off the results:

  • Fast: under 400 ms at the median, the Doherty threshold, below which people stay engaged instead of waiting.
  • Accurate: at least 59.6%, the lowest score that beats always keeping the default by more than chance on these turns (one-sided exact binomial test, p < 0.05), computed by plot.py.

| Model | Runs | Right level | Right level on Opus 5.5 | Right level applied | Latency p50 / p95 | | :- | :- | -: | -: | -: | -: | | Nisev 1.7B | Local · llama.cpp · Q8_0, 1.8 GB | 65.4% | 73.5% | 66.9% | 129 / 364 ms | | Kev-4B | Local · llama.cpp · Q8_0, 4.5 GB | 55.1% | 60.3% | 52.2% | 1750 / 3236 ms | | Clef-Flash 9B | Cloud · Cloudflare Workers AI | 56.6% | 64.7% | 51.5% | 309 / 909 ms | | Clef 27B | Cloud · Cloudflare Workers AI | 51.5% | 65.4% | 51.5% | 495 / 880 ms | | Jev 1.13 | Cloud · OpenCode Zen | 61.8% | 70.6% | 62.5% | 588 / 770 ms |

  • Right level: picked the teachers' level, for the Claude model that ran the turn.
  • Right level on Opus 5.5: the same turns, scored as if they ran on Opus 5.5, whose default is medium.
  • Right level applied: what the mod ends up doing. It applies a pick only at a confidence of 0.5 or more, and otherwise keeps the session's effort.

What stands out:

  • Only Nisev and Jev beat the default by more than chance, compared turn by turn (one-sided exact McNemar test): Nisev p = 0.009, Jev p = 0.04, Kev 0.18, Clef-Flash 0.21, Clef 0.55.
  • Kev and Clef are rarely confident: 0.5 or more on at most 2 of 136 turns (median 0.12 to 0.18), so the mod almost never applies their picks. Jev clears it on half the turns.
  • Clef leans to medium (40% of turns), right on Opus 5.5 and a step too low on models that default to high. It also answers max on 12 turns.
  • Read gaps of a few points as noise at 136 turns. Kev, Clef and Jev answer zero-shot; Nisev was trained on this question, from the same kind of sessions.

To rerun it, run uv run python pipeline/benchmark.py in training/, which also explains how Nisev was built.

Develop

pnpm test        # claude plugin test: hooks with stubbed providers and processes, no network
pnpm validate    # claude plugin validate --strict
pnpm eval        # sample prompts against the live provider, reading .env
  • pnpm eval needs a cloud endpoint's three variables, in the environment or a gitignored .env (see .env.example), or AUTO_EFFORT_PROVIDER=nisev with Nisev already served. It reports the session model as claude-opus-5-5; set AUTO_EFFORT_EVAL_CLAUDE_MODEL to try another.
  • On 2026-10-03 it matched the expected level on 6 of 9 sample prompts, with both jev-1.13 and Nisev. The Jev run took 0.4–1 s per call and about 7k input tokens ($0.0003).
  • claude --plugin-dir . --debug-file /tmp/cc.log logs each override as [auto-effort] … effort low → high.
Source 4 files
hooks/register.js 402 lines
1import { atom, read, update } from 'claude-code'
2
3import { buildRequest, decide, describe, JEV_PRESETS, LEVELS, NISEV, resolveConfig } from './policy.js'
4import { parseBuild, parseServerOutput, serverArgs, serverCommand, servesOurModel } from './nisev.js'
5
6const USER_ORIGINS = ['composer', 'bridge', 'sdk']
7
8// Decision for the prompt whose turn hasn't started yet.
9let pending = null
10// turnId -> decision, for turns in flight.
11const decisions = new Map()
12let enabled = true
13let lastDecision = null
14let warnedMissing = false
15
16// What the mod reports without asking anything of you (the last decision, Nisev starting) goes
17// to the footer's mode labels; the status line, which the engine draws as a notice, is kept for
18// what needs action.
19const label = atom({ plugin: 'auto-effort', key: 'label' }, null)
20
21function show($, text) {
22  return update($, label, () => text).catch(() => {})
23}
24// The llama-server this module started, if any: { state: starting | downloading | ready | stopped }.
25let server = null
26// Set once a request reached Nisev, so a prompt doesn't probe the port every time.
27let nisevConfirmed = false
28
29function nisevState() {
30  return server && server.state !== 'stopped' ? server.state : nisevConfirmed ? 'ready' : 'stopped'
31}
32
33async function llamaServerBuild($, binary) {
34  try {
35    const { stdout, stderr } = await $.process.run([binary, '--version'], { timeoutMs: 10_000 })
36    return parseBuild(stdout + stderr)
37  } catch {
38    return null // Not installed.
39  }
40}
41
42// The configured binary, else the first of llama-server and llama that runs: { binary, build }.
43async function findServer($, config) {
44  for (const binary of config.llamaServer ? [config.llamaServer] : NISEV.servers) {
45    const build = await llamaServerBuild($, binary)
46    if (build !== null) return { binary, build }
47  }
48  return { binary: config.llamaServer ?? NISEV.servers[0], build: null }
49}
50
51async function isServing($, config) {
52  try {
53    const response = await $.http.fetch(`http://127.0.0.1:${config.port}/v1/models`)
54    return response.ok && servesOurModel(response.text)
55  } catch {
56    return false
57  }
58}
59
60// Starts llama-server for the rest of the session unless our model is already served on the
61// port; never waits. The first start downloads the model (-hf), reported on the status line.
62async function startNisev($, config) {
63  if (server && server.state !== 'stopped') return server.state
64  if (await isServing($, config)) {
65    nisevConfirmed = true
66    return 'ready'
67  }
68  const current = { state: 'starting' }
69  server = current
70  const { binary } = await findServer($, config)
71  // llama.cpp may print nothing while it downloads, so say up front that the first start can take minutes.
72  void show($, 'Nisev starting · the first start downloads ' + NISEV.sizeLabel)
73  void (async () => {
74    try {
75      // The loop is the child's life: it ends with the child or with this module.
76      for await (const { text } of $.process.spawn({ argv: serverArgs(config, binary) })) {
77        const seen = parseServerOutput(text)
78        if (seen.ready && current.state !== 'ready') {
79          current.state = 'ready'
80          nisevConfirmed = true
81          void show($, 'Nisev ready')
82        } else if (seen.percent !== undefined && current.state !== 'ready') {
83          current.state = 'downloading'
84          void show($, 'Nisev downloading · ' + seen.percent + '%')
85        }
86        $.ui.log(text.trimEnd(), { to: 'debug' })
87      }
88    } catch (error) {
89      $.ui.log('llama-server failed to start: ' + error.message, { to: 'debug' })
90      void show($, null)
91      $.ui.status('auto-effort: llama-server failed to start, see the debug log')
92    }
93    current.state = 'stopped'
94    nisevConfirmed = false
95  })()
96  return current.state
97}
98
99async function storedConfig($) {
100  try {
101    const saved = await $.store.get('config')
102    return saved && typeof saved === 'object' ? saved : {}
103  } catch {
104    return {}
105  }
106}
107
108async function loadConfig($) {
109  const env = {
110    provider: await $.env.get('AUTO_EFFORT_PROVIDER'),
111    apiKey: await $.env.get('AUTO_EFFORT_API_KEY'),
112    endpoint: await $.env.get('AUTO_EFFORT_ENDPOINT'),
113    model: await $.env.get('AUTO_EFFORT_MODEL'),
114    minConfidence: await $.env.get('AUTO_EFFORT_MIN_CONFIDENCE'),
115    timeoutMs: await $.env.get('AUTO_EFFORT_TIMEOUT_MS'),
116    minEffort: await $.env.get('AUTO_EFFORT_MIN_EFFORT'),
117    maxEffort: await $.env.get('AUTO_EFFORT_MAX_EFFORT'),
118    includeContext: await $.env.get('AUTO_EFFORT_INCLUDE_CONTEXT'),
119    llamaServer: await $.env.get('AUTO_EFFORT_LLAMA_SERVER'),
120  }
121  return resolveConfig(env, await storedConfig($))
122}
123
124async function previousReply($) {
125  try {
126    const messages = await $.session.messages()
127    for (let i = messages.length - 1; i >= 0; i--) {
128      if (messages[i].role === 'assistant' && messages[i].text) return messages[i].text
129    }
130  } catch {
131    // No transcript to read; classify the prompt alone.
132  }
133  return undefined
134}
135
136async function sessionModel($) {
137  try {
138    return await $.session.model()
139  } catch {
140    return undefined
141  }
142}
143
144// The same System One call for every provider: a hosted Jev, or llama-server on this machine.
145async function askSystemOne($, config, body) {
146  const request = $.http.fetch(config.endpoint, {
147    method: 'POST',
148    headers: { Authorization: 'Bearer ' + config.apiKey, 'Content-Type': 'application/json' },
149    body: JSON.stringify(body),
150  })
151  // $.http.fetch takes no timeout, and a slow classifier must never hold up the prompt.
152  const timeout = $.clock.sleep(config.timeoutMs).then(() => null)
153  const response = await Promise.race([request, timeout])
154  if (response === null) throw new Error('timed out after ' + config.timeoutMs + 'ms')
155  if (response.status === 404 && config.provider === 'nisev') {
156    throw new Error('llama-server has no /v1/systemone; it needs build ' + NISEV.minBuild + ' or later')
157  }
158  if (!response.ok) throw new Error('HTTP ' + response.status + ' ' + response.text.slice(0, 200))
159  return JSON.parse(response.text)
160}
161
162async function classify($, prompt) {
163  const config = await loadConfig($)
164  if (!config.provider) {
165    if (!warnedMissing) {
166      warnedMissing = true
167      $.ui.status('auto-effort: run /auto-effort setup, or set ' + config.missing.join(', '))
168    }
169    return null
170  }
171  if (config.missing.length || config.problems.length) {
172    if (!warnedMissing) {
173      warnedMissing = true
174      $.ui.status(config.problems.length
175        ? 'auto-effort: Nisev is selected, but ' + config.problems.join('; ') + '. Unset them to use its defaults.'
176        : 'auto-effort: set ' + config.missing.join(', '))
177    }
178    return null
179  }
180  if (config.provider === 'nisev' && nisevState() !== 'ready') {
181    // Start it for the next prompts; this one keeps the session's effort.
182    const state = await startNisev($, config)
183    if (state !== 'ready') return { effort: null, reason: 'Nisev ' + state }
184  }
185  const model = await sessionModel($)
186  const body = buildRequest({ prompt, previousReply: await previousReply($), model }, config)
187  try {
188    const decision = decide(await askSystemOne($, config, body), config, model)
189    if (config.provider === 'nisev') nisevConfirmed = true
190    return decision
191  } catch (error) {
192    $.ui.log('effort request failed: ' + error.message, { to: 'debug' })
193    if (config.provider === 'nisev') nisevConfirmed = false
194    return { effort: null, reason: config.provider === 'nisev' ? 'Nisev unavailable' : 'endpoint unavailable' }
195  }
196}
197
198async function saveConfig($, config) {
199  try {
200    await $.store.set('config', config)
201    return true
202  } catch {
203    return false
204  }
205}
206
207async function ask($, question, options) {
208  try {
209    return await $.ui.ask(question, options)
210  } catch {
211    return null // Dismissed, or no one to ask (a -p run).
212  }
213}
214
215async function setupJev($, endpoint, model) {
216  const saved = await saveConfig($, { provider: 'jev', endpoint, model })
217  const key = await $.env.get('AUTO_EFFORT_API_KEY')
218  return {
219    text: [
220      saved ? 'Saved: ' + endpoint + ', model ' + model + '.' : 'Could not save the setting for later sessions.',
221      key
222        ? 'AUTO_EFFORT_API_KEY is set, so the next prompt uses it.'
223        : 'Now set AUTO_EFFORT_API_KEY to the key for that endpoint, in your shell or in the env block of a ' +
224          'settings file that is not committed. The key is never stored by the mod.',
225      'AUTO_EFFORT_* environment variables, when set, take precedence over this setup.',
226    ].join('\n'),
227  }
228}
229
230async function setupNisev($) {
231  const env = { provider: 'nisev', llamaServer: await $.env.get('AUTO_EFFORT_LLAMA_SERVER') }
232  const config = resolveConfig(env, await storedConfig($))
233  const { binary, build } = await findServer($, config)
234  if (build === null) {
235    return {
236      text:
237        'llama.cpp was not found. Install build ' + NISEV.minBuild + ' or later, from ' +
238        'https://github.com/ggml-org/llama.cpp/releases, so that llama-server or llama is on your PATH, or ' +
239        'set AUTO_EFFORT_LLAMA_SERVER to its path. Then run /auto-effort setup again.',
240    }
241  }
242  if (build < NISEV.minBuild) {
243    return {
244      text:
245        binary + ' is build ' + build + '; decision models need build ' + NISEV.minBuild + ' or later. ' +
246        'Update it from https://github.com/ggml-org/llama.cpp/releases and run /auto-effort setup again.',
247    }
248  }
249  config.llamaServer = binary
250  const serving = await isServing($, config)
251  if (!serving) {
252    const go = await ask($, 'Download ' + config.model + ' (' + NISEV.sizeLabel + ') and run it with llama.cpp?', {
253      header: 'Download',
254      options: ['Download and start', 'Cancel'],
255    })
256    if (go !== 'Download and start') return { text: 'Setup cancelled; nothing changed.' }
257  }
258  await saveConfig($, { provider: 'nisev', model: config.model, endpoint: config.endpoint, llamaServer: config.llamaServer })
259  await startNisev($, config)
260  return {
261    text: [
262      'Saved: Nisev, served by ' + serverCommand(binary).join(' ') + ' (build ' + build + ') at ' + config.endpoint + '.',
263      serving
264        ? 'It is already running.'
265        : 'The first start downloads the model; the status line shows the progress. Until it is ready, ' +
266          'prompts keep the session effort.',
267      'llama-server runs while this session does; the next session starts it again from the cache.',
268      'AUTO_EFFORT_* environment variables, when set, take precedence over this setup.',
269    ].join('\n'),
270  }
271}
272
273async function setup($) {
274  const choice = await ask($, 'Which classifier should pick the effort of each prompt?', {
275    header: 'Classifier',
276    options: ['Nisev (local)', ...Object.keys(JEV_PRESETS)],
277  })
278  if (choice === null) return { text: 'Setup cancelled; nothing changed.' }
279  if (choice === 'Nisev (local)') return setupNisev($)
280  if (JEV_PRESETS[choice]) return setupJev($, JEV_PRESETS[choice].endpoint, JEV_PRESETS[choice].model)
281  // "Other": any System One endpoint.
282  if (!/^https?:\/\/\S+$/.test(choice.trim())) {
283    return { text: 'Under Other, enter the URL of a System One endpoint, such as https://example.com/v1/systemone.' }
284  }
285  const model = await ask($, 'Which model name does that endpoint expect?', { header: 'Model', options: ['jev-latest', 'jev-1.13'] })
286  if (model === null) return { text: 'Setup cancelled; nothing changed.' }
287  return setupJev($, choice.trim(), model.trim())
288}
289
290async function status($) {
291  const config = await loadConfig($)
292  const lines = ['Effort picker: ' + (enabled ? 'on' : 'off')]
293  if (config.provider === 'nisev') {
294    lines.push('Provider: Nisev ' + config.model + ' · llama.cpp at ' + config.endpoint + ' (' + nisevState() + ')')
295    for (const problem of config.problems) lines.push('Problem: ' + problem)
296  } else {
297    lines.push('Provider: ' + (config.provider ? 'System One endpoint' : 'not set, run /auto-effort setup'))
298    lines.push('Endpoint: ' + (config.endpoint ?? 'unset') + ' · model ' + (config.model ?? 'unset'))
299    lines.push('API key: ' + (config.apiKey ? 'set' : 'unset'))
300    if (config.missing.length) lines.push('Missing: ' + config.missing.join(', '))
301  }
302  lines.push(
303    'Range: ' + config.minEffort + '–' + config.maxEffort + ' of ' + LEVELS.join(', ') +
304      ' · min confidence ' + config.minConfidence,
305    'Last decision: ' + (lastDecision ? describe(lastDecision) : 'none yet'),
306  )
307  return { text: lines.join('\n') }
308}
309
310export function register(on) {
311  on('session.start', async ($, e, next) => {
312    try {
313      const saved = await $.store.get('enabled')
314      if (typeof saved === 'boolean') enabled = saved
315    } catch {
316      // Keep the default.
317    }
318    try {
319      await $.command.register({
320        name: 'auto-effort',
321        description: 'Set up, show or toggle the effort picker',
322        argumentHint: '[setup|on|off|status]',
323        immediate: true,
324      })
325    } catch (error) {
326      $.ui.log('could not register /auto-effort: ' + error.message, { to: 'debug' })
327    }
328    const result = await next(e)
329    // Warm Nisev up before the first prompt needs it.
330    const config = await loadConfig($)
331    if (enabled && config.provider === 'nisev' && !config.problems.length) await startNisev($, config)
332    return result
333  })
334
335  on('prompt.submit', async ($, e, next) => {
336    // A delivery into a running turn (a peer's message, a notification) isn't a
337    // new request from the user, so the running turn keeps its decision. Prompts
338    // the user types mid-turn are queued and get a turn of their own.
339    const intoRunningTurn = e.turnId && !USER_ORIGINS.includes(e.origin?.kind)
340    if (!enabled || intoRunningTurn || !e.text.trim()) return next(e)
341    const decision = await classify($, e.text)
342    if (decision) {
343      lastDecision = decision
344      await show($, describe(decision))
345      $.ui.status(undefined)
346      pending = decision
347    }
348    const result = await next(e)
349    // A later hook dropped the prompt: don't let its decision reach another turn.
350    if (result.drop) pending = null
351    return result
352  })
353
354  on('turn.start', async ($, e, next) => {
355    if (pending) {
356      decisions.set(e.turnId, pending)
357      pending = null
358    }
359    return next(e)
360  })
361
362  on('ui.render', { component: 'SessionMode' }, async ($, e, next) => {
363    const text = await read($, label)
364    return text ? next({ ...e, props: { ...e.props, modes: [...e.props.modes, text] } }) : next(e)
365  })
366
367  on('turn.step', async function* ($, e, next) {
368    const decision = decisions.get(e.turnId)
369    // Subagents keep their own effort; a model without effort, or a numeric
370    // budget set by hand, is left alone.
371    if (!enabled || e.agentId || !decision?.effort || typeof e.effort !== 'string') {
372      return yield* next(e)
373    }
374    if (e.index === 0) $.ui.log('effort ' + e.effort + ' → ' + decision.effort, { to: 'debug' })
375    return yield* next({ ...e, effort: decision.effort })
376  })
377
378  on('turn.complete', async ($, e, next) => {
379    decisions.delete(e.turnId)
380    return next(e)
381  })
382
383  on('command.run', { command: 'auto-effort' }, async ($, e) => {
384    const arg = e.args.trim().toLowerCase()
385    if (arg === 'setup') return setup($)
386    if (arg === 'on' || arg === 'off') {
387      enabled = arg === 'on'
388      try {
389        await $.store.set('enabled', enabled)
390      } catch {
391        // Applies to this session only.
392      }
393      $.ui.status(undefined)
394      await show($, enabled ? null : 'effort picker off')
395      return { text: 'Effort picker turned ' + arg + '.' }
396    }
397    if (arg && arg !== 'status') return { text: 'Usage: /auto-effort [setup|on|off|status]' }
398    return status($)
399  })
400}
401
402
hooks/policy.js 295 lines
1// Pure decision logic: no `$`, so it can be unit tested on its own.
2
3export const LEVELS = ['low', 'medium', 'high', 'xhigh', 'max']
4
5export const DEFAULTS = {
6  minConfidence: 0.5,
7  timeoutMs: 4000,
8  minEffort: 'low',
9  maxEffort: 'xhigh',
10  includeContext: true,
11}
12
13// The fine-tuned classifier, served by llama.cpp's llama-server as a decision model.
14export const NISEV = {
15  // A Hugging Face repo for `llama-server -hf`, or a path to a .gguf file.
16  model: 'arthur-fontaine/nisev-1.7b-GGUF:Q8_0',
17  sizeLabel: 'about 1.9 GB',
18  endpoint: 'http://127.0.0.1:8765/v1/systemone',
19  // Tried in order when none is configured: the server binary, or the unified CLI (`llama serve`).
20  servers: ['llama-server', 'llama'],
21  alias: 'nisev',
22  // The first llama.cpp build with /v1/systemone (ggml-org/llama.cpp#29818).
23  minBuild: 11361,
24}
25
26export const JEV_PRESETS = {
27  'OpenCode Zen': { endpoint: 'https://opencode.ai/zen/v1/systemone', model: 'jev-1.13' },
28  TypeSafe: { endpoint: 'https://api.typesafe.ai/v1/systemone', model: 'jev-latest' },
29}
30
31// Worded after https://claude.com/blog/claude-model-and-effort-level-in-claude-code:
32// the model's default is right for most tasks, effort controls how thorough
33// Claude is (files read, verification, how far it pushes before checking in),
34// and deviating should need a clear reason.
35export const EFFORT_QUESTION = {
36  type: 'choice',
37  instructions: {
38    task:
39      'A developer sent `latest_user_message` to Claude, an AI coding agent working in their repository. ' +
40      'Pick how much effort Claude should spend on it. Effort controls how many files Claude reads, ' +
41      'how much it verifies (running tests, double-checking), and how far it pushes through a multi-step ' +
42      'task before checking back in. It is not about how capable Claude is. Most requests should get ' +
43      '`default`. Pick another level only when the request clearly calls for less or more thoroughness. ' +
44      '`previous_assistant_reply`, when present, is what Claude last said; use it to understand short ' +
45      'follow-ups such as "yes, do it". `model`, when present, is the Claude model; `default` keeps that ' +
46      "model's own default effort.",
47  },
48  criteria: {
49    low:
50      'Routine work that needs no investigation: a precisely described edit, a rename, a typo, a one-line ' +
51      'change, a question about code already in context, a quick lookup or shell command, small talk.',
52    medium:
53      'Light work: a small, well-scoped change or a focused answer that needs a little reading, but no ' +
54      'multi-file investigation.',
55    default:
56      'A typical coding request: an ordinary feature, a normal bug fix, a focused explanation, or anything ' +
57      'unclear. The model default already scales the work to the task.',
58    high:
59      'Multi-step work that must be verified: changes across several files, a bug that needs several ' +
60      'hypotheses checked, a refactor that has to be finished completely, work where skipping a file or ' +
61      'not running the tests would make the result wrong.',
62    xhigh:
63      'Long, hard, high-stakes work: a large migration or refactor, a subtle bug across systems, an ' +
64      'architecture decision, a security or correctness audit, or a request that explicitly asks Claude to ' +
65      'be thorough and verify everything.',
66    max:
67      'Exhaustive work where cost does not matter: the request explicitly demands maximum effort or ' +
68      'leaving nothing unchecked, on a critical, very large task.',
69  },
70}
71
72// The question Nisev was trained on, word for word (training/pipeline/task.py): it
73// sizes the request, and the model's table below turns the size into a level.
74export const CATEGORIES = ['trivial', 'light', 'ordinary', 'multi_step', 'hard', 'exhaustive']
75export const CATEGORY_QUESTION = {
76  type: 'choice',
77  instructions:
78    'A developer sent latest_user_message to an AI coding agent working in their repository. ' +
79    'How much thoroughness does it need: how many files to read, how much to verify, and how far ' +
80    'to push before checking back in? previous_assistant_reply, when present, is the agent\'s last ' +
81    'reply; use it to size short follow-ups such as "yes, do it".',
82  criteria: {
83    trivial: 'Routine: a precise small edit, a lookup, a question about code in context, an acknowledgement.',
84    light: 'A small, well-scoped change or focused answer that needs a little reading.',
85    ordinary: 'A typical feature, bug fix or explanation, or anything unclear.',
86    multi_step: 'Several files, several hypotheses, or work that must be verified by running tests.',
87    hard: 'A large migration, a subtle cross-system bug, an audit, or an explicit ask to be thorough.',
88    exhaustive: 'An explicit demand for maximum effort on a critical, very large task.',
89  },
90}
91
92// The level each kind of request gets on each model (training/pipeline/models.py). Levels are
93// calibrated per model: Opus 5.5 defaults to medium, so verified multi-step work is where it
94// pays to raise effort there, while models that default to high already run it at high.
95const HIGH = { trivial: 'low', light: 'medium', ordinary: 'high', multi_step: 'high', hard: 'xhigh', exhaustive: 'max' }
96const MEDIUM = { ...HIGH, ordinary: 'medium' }
97const NO_XHIGH = { ...HIGH, hard: 'high' }
98const NO_MAX = { ...HIGH, hard: 'high', exhaustive: 'high' }
99export const PROFILES = {
100  'claude-opus-5-5': { default: 'medium', table: MEDIUM },
101  'claude-sonnet-5-5': { default: 'high', table: HIGH },
102  'claude-fable-5-1': { default: 'high', table: HIGH },
103  'claude-opus-5': { default: 'high', table: HIGH },
104  'claude-sonnet-5': { default: 'high', table: HIGH },
105  'claude-fable-5': { default: 'high', table: HIGH },
106  'claude-opus-4-8': { default: 'high', table: HIGH },
107  'claude-opus-4-7': { default: 'high', table: HIGH },
108  'claude-opus-4-6': { default: 'high', table: NO_XHIGH },
109  'claude-sonnet-4-6': { default: 'high', table: NO_XHIGH },
110  'claude-opus-4-5': { default: 'high', table: NO_MAX },
111}
112const FALLBACK = { default: 'high', table: HIGH }
113const ALIASES = { opus: 'claude-opus-5-5', sonnet: 'claude-sonnet-5-5', fable: 'claude-fable-5-1' }
114
115// `claude-opus-5-5[1m]`, `us.anthropic.claude-opus-5-5-v1`, `opus` -> `claude-opus-5-5`
116export function normalizeModel(model) {
117  if (!model) return undefined
118  const m = String(model).toLowerCase().split('/').pop()
119    .replace(/^((us|eu|apac|global)\.)?anthropic\./, '')
120    .replace(/\[.*?\]$|:.*$|-v\d+$|-\d{8}$|-fast$/g, '')
121    .replaceAll('.', '-')
122  return ALIASES[m] ?? m
123}
124
125export function profileOf(model) {
126  return PROFILES[normalizeModel(model)] ?? FALLBACK
127}
128
129// Haiku 4.5 takes no effort; an unknown or missing model gets the common profile.
130export function takesEffort(model) {
131  return !(normalizeModel(model) ?? '').includes('haiku')
132}
133
134// Keeps the request well under the classifier's context.
135const CLIP = {
136  jev: { prompt: 6000, context: 1500 },
137  // What Nisev was trained with (training/pipeline/task.py).
138  nisev: { prompt: 3000, context: 800 },
139}
140
141// Counts code points, as the Python that built the training data does, so a long prompt is
142// cut at the same place.
143export function clip(text, max) {
144  const chars = Array.from(text)
145  if (chars.length <= max) return text
146  // Keep both ends: the ask is usually at the start, pasted output at the end.
147  const half = Math.floor((max - 20) / 2)
148  return chars.slice(0, half).join('') + '\n[… truncated …]\n' + chars.slice(-half).join('')
149}
150
151function pick(...values) {
152  return values.find((v) => v !== undefined && v !== null && v !== '')
153}
154
155function toNumber(value, fallback) {
156  const n = typeof value === 'number' ? value : Number(value)
157  return Number.isFinite(n) ? n : fallback
158}
159
160function toLevel(value, fallback) {
161  return LEVELS.includes(value) ? value : fallback
162}
163
164function toBool(value, fallback) {
165  if (typeof value === 'boolean') return value
166  if (value === 'true' || value === '1') return true
167  if (value === 'false' || value === '0') return false
168  return fallback
169}
170
171// The endpoint, key, and model of a Jev provider have no default, so the mod never calls
172// a provider the user didn't choose.
173export const REQUIRED = {
174  endpoint: 'AUTO_EFFORT_ENDPOINT',
175  apiKey: 'AUTO_EFFORT_API_KEY',
176  model: 'AUTO_EFFORT_MODEL',
177}
178
179// `env` holds the AUTO_EFFORT_* values, keyed as in the returned config. `stored` holds what
180// `/auto-effort setup` saved: never a key, and the environment wins over it.
181// The port Nisev serves on, from its endpoint; null unless the endpoint is on this machine.
182export function localPort(endpoint) {
183  const match = /^http:\/\/(?:127\.0\.0\.1|localhost):(\d+)\/v1\/systemone$/.exec(endpoint ?? '')
184  return match ? Number(match[1]) : null
185}
186
187// What llama.cpp can serve: a .gguf path, or a Hugging Face repo with an optional :quant.
188export function isGgufSource(model) {
189  return /\.gguf$/i.test(model ?? '') || /^[\w.-]+\/[\w.-]+(:[\w.-]+)?$/.test(model ?? '')
190}
191
192export function resolveConfig(env = {}, stored = {}) {
193  // The environment wins: an endpoint there means a cloud endpoint unless AUTO_EFFORT_PROVIDER says otherwise.
194  const provider = [env.provider, env.endpoint && 'jev', stored.provider].find((p) => ['jev', 'nisev'].includes(p))
195  // Saved values belong to the provider they were saved for: a cloud model name is no GGUF to serve.
196  const saved = stored.provider === provider ? stored : {}
197  const nisev = provider === 'nisev'
198  const config = {
199    provider,
200    minConfidence: toNumber(pick(env.minConfidence), DEFAULTS.minConfidence),
201    timeoutMs: toNumber(pick(env.timeoutMs), DEFAULTS.timeoutMs),
202    minEffort: toLevel(pick(env.minEffort), DEFAULTS.minEffort),
203    maxEffort: toLevel(pick(env.maxEffort), DEFAULTS.maxEffort),
204    includeContext: toBool(pick(env.includeContext), DEFAULTS.includeContext),
205    // For Nisev, the model is what llama.cpp serves: a Hugging Face repo:quant or a .gguf path.
206    endpoint: pick(env.endpoint, saved.endpoint, nisev ? NISEV.endpoint : undefined),
207    model: pick(env.model, saved.model, nisev ? NISEV.model : undefined),
208    apiKey: nisev ? 'none' : pick(env.apiKey),
209    llamaServer: nisev ? pick(env.llamaServer, saved.llamaServer) : undefined,
210  }
211  if (nisev) {
212    config.port = localPort(config.endpoint)
213    // Usually a cloud endpoint's settings left in the environment, so say what they are.
214    config.problems = [
215      ...(config.port ? [] : ['AUTO_EFFORT_ENDPOINT is ' + config.endpoint + ', not a local URL such as ' + NISEV.endpoint]),
216      ...(isGgufSource(config.model) ? [] : ['AUTO_EFFORT_MODEL is ' + config.model + ', not a Hugging Face repo:quant or a .gguf path']),
217    ]
218    config.missing = []
219  } else {
220    config.problems = []
221    config.missing = Object.keys(REQUIRED).filter((key) => !config[key]).map((key) => REQUIRED[key])
222  }
223  return config
224}
225
226export function buildRequest({ prompt, previousReply, model }, config) {
227  const limits = CLIP[config.provider === 'nisev' ? 'nisev' : 'jev']
228  const state = { latest_user_message: clip(prompt, limits.prompt) }
229  if (config.includeContext && previousReply) {
230    state.previous_assistant_reply = clip(previousReply, limits.context)
231  }
232  // Effort levels are calibrated per model, so the classifier needs to know which one runs.
233  if (model) state.model = model
234  const question = config.provider === 'nisev' ? CATEGORY_QUESTION : EFFORT_QUESTION
235  // llama.cpp answers to the alias it serves Nisev under, whatever file it loaded.
236  return { model: config.provider === 'nisev' ? NISEV.alias : config.model, state, questions: { effort: question } }
237}
238
239export function clamp(level, min, max) {
240  const lo = LEVELS.indexOf(min)
241  const hi = Math.max(lo, LEVELS.indexOf(max))
242  const i = LEVELS.indexOf(level)
243  return LEVELS[Math.min(Math.max(i, lo), hi)]
244}
245
246// Nisev answers with category probabilities: sum them into levels through the
247// model's table, so the confidence is the probability of the level it picks.
248export function levelProbabilities(categoryProbabilities, model) {
249  const table = profileOf(model).table
250  const levels = Object.fromEntries(LEVELS.map((level) => [level, 0]))
251  for (const category of CATEGORIES) levels[table[category]] += categoryProbabilities?.[category] ?? 0
252  return levels
253}
254
255function decideNisev(answer, config, model) {
256  if (!takesEffort(model)) return { effort: null, reason: 'model takes no effort' }
257  const levels = levelProbabilities(answer.probabilities, model)
258  const level = LEVELS.reduce((best, l) => (levels[l] > levels[best] ? l : best), LEVELS[0])
259  const confidence = levels[level]
260  const category = answer.choice
261  // The model's default is what `default` keeps: the session's own effort stands.
262  if (level === profileOf(model).default) return { effort: null, choice: 'default', confidence, category, reason: 'default' }
263  if (confidence < config.minConfidence) return { effort: null, choice: level, confidence, category, reason: 'low confidence' }
264  return { effort: clamp(level, config.minEffort, config.maxEffort), choice: level, confidence, category, reason: 'jev' }
265}
266
267// Returns `{ effort, choice, confidence, reason }`. `effort` is null when the
268// session's own effort should stand.
269export function decide(response, config, model) {
270  // Cloudflare Workers AI wraps the System One response in `result`.
271  const answer = (response?.answers ?? response?.result?.answers)?.effort
272  if (!answer || answer.type !== 'choice' || typeof answer.choice !== 'string') {
273    return { effort: null, reason: 'no effort answer' }
274  }
275  if (config.provider === 'nisev') return decideNisev(answer, config, model)
276  const { choice, confidence } = answer
277  if (choice === 'default') return { effort: null, choice, confidence, reason: 'default' }
278  if (!LEVELS.includes(choice)) return { effort: null, choice, confidence, reason: 'unknown choice ' + choice }
279  if (typeof confidence !== 'number' || confidence < config.minConfidence) {
280    return { effort: null, choice, confidence, reason: 'low confidence' }
281  }
282  return { effort: clamp(choice, config.minEffort, config.maxEffort), choice, confidence, reason: 'jev' }
283}
284
285export function describe(decision) {
286  const pct = typeof decision.confidence === 'number' ? ' · ' + Math.round(decision.confidence * 100) + '%' : ''
287  if (decision.effort) {
288    const capped = decision.choice && decision.choice !== decision.effort ? ' (picked ' + decision.choice + ')' : ''
289    return 'effort ' + decision.effort + capped + pct
290  }
291  if (decision.reason === 'default') return 'effort default' + pct
292  if (decision.reason === 'low confidence') return 'effort default (unsure: ' + decision.choice + pct + ')'
293  return 'effort default (' + decision.reason + ')'
294}
295
hooks/nisev.js 43 lines
1// The Nisev provider: the fine-tuned classifier, served by llama.cpp.
2// Pure helpers; register.js does the calls ($ is only passed within a file).
3
4import { NISEV } from './policy.js'
5
6// `llama` is llama.cpp's unified CLI, whose server is `llama serve`; it takes the same flags.
7export function serverCommand(binary) {
8  return /(^|\/)llama$/.test(binary) ? [binary, 'serve'] : [binary]
9}
10
11export function serverArgs(config, binary) {
12  const source = /\.gguf$/i.test(config.model) ? ['-m', config.model] : ['-hf', config.model]
13  return [
14    ...serverCommand(binary), ...source,
15    '--host', '127.0.0.1', '--port', String(config.port),
16    '--alias', NISEV.alias,
17    // Prompts are at most about 1,300 tokens; two slots let two prompts run at once.
18    '-c', '8192', '-np', '2',
19  ]
20}
21
22// `llama-server --version` prints e.g. "version: 0.5.0-dev (build 11408, commit 9f12cd4a4)".
23export function parseBuild(output) {
24  const match = output.match(/build (\d+)/)
25  return match ? Number(match[1]) : null
26}
27
28// What a piece of llama-server's output says about its start: { ready }, { percent } or {}.
29export function parseServerOutput(text) {
30  if (/listening on/i.test(text)) return { ready: true }
31  const percent = /download/i.test(text) && text.match(/(\d{1,3}(?:\.\d+)?)%/)
32  return percent ? { percent: Math.round(Number(percent[1])) } : {}
33}
34
35// True when /v1/models lists our alias, whoever started the server.
36export function servesOurModel(text) {
37  try {
38    return (JSON.parse(text).data ?? []).some((m) => m.id === NISEV.alias)
39  } catch {
40    return false
41  }
42}
43
types/index.d.ts 9 lines
1// The values the mod keeps in $.state: the footer label it draws.
2export type AutoEffortLabel = string | null
3
4declare module 'claude-code' {
5  interface PluginState {
6    'auto-effort': { label: AutoEffortLabel }
7  }
8}
9