Routes each turn's effort from your prompt: git housekeeping replies ("merged", "commit and push", "push it") run at low effort, deep asks (review, audit…

Claude Code mods (function-hook plugins) I find helpful. Each folder is one plugin.
Requires a Claude Code build with function-hook plugins (2.1.289 or newer). machine-guard reads macOS tools (sysctl, memory_pressure, ioreg).
git clone https://github.com/joeldg/claude-mods ~/Projects/claude-mods
A pane of your project's running dev servers, so you don't have to ask Claude to restart them.
/servers opens the Servers pane:package.json dev/start/serve/preview scripts (run with your lockfile's package manager), .claude/launch.json and Procfile~/.claude/dev-servers/<project>/.Port 4000 is held by node (pid 123, up 2h, in /Users/me/other).servers: :4000 :5173.lsof and ps every 15s.Puts files you just downloaded into your prompt with one click.
~/Downloads (top level). When a new file arrives (PDF, Markdown, images, 3MF/STL/OBJ, zip, video…), a band appears above the prompt: New in Downloads: paper.pdf, model-b.3mf · 2m ago [Attach] [Dismiss].@"/Users/you/Downloads/paper.pdf" mentions in your prompt. Dismiss hides those files./downloads lists the 10 newest files, numbered. /downloads attach 1 3 (or 2-4) adds those, and /downloads clear dismisses everything new.Settings: folder (~/Downloads), extensions, pollSeconds (5), maxAgeMinutes (120).
Sets effort per message, so you don't have to switch it by hand.
avoidModel (a regex such as fable) to send those requests to fallbackModel instead, subagents included./route shows the last decision and the session's counts. /route off and /route on toggle it; /route deep and /route routine force the next turn.effort: low (routine).Settings: routineEffort (low), deepEffort (max), routinePattern, deepPattern, avoidModel, fallbackModel (opus), stickyTurns (2), freeSwitchTokens (30000), cacheTtlMinutes (60).
A Jobs pane for long-running work: training runs, downloads, extractions.
nohup … > log & launches by itself./watch <log> [label] adds any other log file./ and /Volumes/* (the NAS).jobs: 2 running · 1 stalled./jobs opens the pane, /unwatch <label|done|all> removes jobs.tail, checks processes with ps, and runs df.Settings (in /config): stall minutes (10), refresh seconds (10), how long finished jobs stay (120 min), auto-open (on), which disks to show.
Memory, swap and GPU on the status line. It refuses heavy local jobs when the Mac can't take them.
RAM tight 12% free · swap 7.9/8G · top python 31G · GPU 87%./busy.When memory is only tight, the job runs and Claude gets a note to start one heavy job at a time.
/busy 3h training a vision model reserves the Mac in every Claude session. /busy off lifts it. The reservation lives in ~/.claude/machine-guard.json, so a training script can write it too: ``bash echo '{"reason":"overnight training","until":'$(( ($(date +%s) + 8*3600) * 1000 ))'}' > ~/.claude/machine-guard.json ``/guard shows what it sees. /guard pause 15m lets heavy jobs through in this session; /guard on resumes the guard.modal run, ssh), tests (pytest) and installs are never treated as heavy.overnight_|nightly_run\.sh.Watches how the other mods behave in real use, without changing them. It is listed first in CLAUDE_CODE_PLUGIN_DIRS, so the other mods' hooks run beneath it.
next.trace), with the mod's name, the event and how long it ran. Slow hooks (over 1.5 s) are recorded too. The first failure of each mod in a session raises a toast./second-opinion, /recall ask) and file writes (folders only, never contents)./mods: a pane with one row per mod: ✓ active, ⚠ failing, ✗ not seen this session, · seen but idle. Each row shows today's counts and last activity, with Details for its recent events. It also says which mods it can't see, if any of them run above it./mods report [24h|7d|30d]: a per-mod report across all sessions, also written to ~/.claude/mods/monitor/report-latest.md for a scheduled review or Claude to read./mods failures [7d]: failures and failed subprocesses only.~/.claude/mods/monitor/<date>/<session>.jsonl, flushed every minute and at session end, with secrets masked and old days removed after 30 days.$.ui.log with wording like "failed" or "could not"): shown in Details and in /mods failures. Three in an hour mark the mod ⚠ and raise one toast. That is how effort-router's per-request hook, which runs inside the response stream where no monitor should sit, reports a failure. It also always sends the request on unchanged./secrets records there) are watched for failures and slow runs, but not counted per run.Settings: alerts (on), slowMs (1500), watchRender (on), watchCommands (on; off stops "mod-monitor" appearing beside other mods' command output), watchAppend (on), retentionDays (30), flushSeconds (60).
Keeps an eye on Modal so idle GPU containers don't burn credits.
Modal: 1 running (2 containers). Deployed apps with no containers cost nothing, so they stay off it.alertMinutes (30), repeated at most every 30 minutes./modal opens a pane of apps with state, containers and uptime. Stop asks for Confirm, then runs modal app stop. Nothing is stopped any other way.budgetToday where the Modal CLI supports billing report (1.3.3+, Team/Enterprise workspaces). Otherwise /modal says why spend isn't shown.modal or python3 -m modal. It checks PATH first rather than running a command that can only fail, and stays silent when Modal isn't set up.Does the "merged #219, clean up branches and start #214" round trip for you, and surfaces CI failures with their logs.
gh pr create Claude runs. It polls gh pr view every 60s.PRs: #219 ✓ · #220 CI… · #221 ✗. Toasts when CI fails (with the failing check names) or passes.git fetch --prune, switch to the default branch (only from the PR's own branch) and git pull --ff-only.--force, never other branches.gh pr checks and the tail of the failed log attached, so you don't paste it./prs lists watched PRs. /prs watch <n|url> and /prs forget <n|all> add and remove them.gh and git, at about one GitHub API call per open PR per minute.Settings:
pollSeconds (60)attachCiLogs (on)logLines (120)deleteRemoteBranch (off): deletes the branch on GitHub too, only while it still points at the merged commit. GitHub's own "Automatically delete head branches" setting does the same job.It never closes issues; put "Closes #N" in PR bodies for that.
Search everything you've done with coding agents, from Claude or from /recall. It replaces the broken agent-memory plugin.
/remember notes. Routine (scheduled) runs are left out unless you add routines:include to a query.grep and cat are kept but ranked low.search, expand, recap and list, which run without permission prompts. It checks them when you say "like last time" or "what did we decide", and before asking you something you already settled./recall <query> opens a pane of hits grouped by session. Open shows the conversation around a hit, Attach sends it with your next message, and Copy resume command copies claude --resume <id>./recall last [n] recaps your last session in this repo: last asks, last answer, PRs, commits, open tasks and decisions. Send to Claude attaches it./recall timeline [7d|30d|90d] [all]/recall decisions|commands|files|prs|commits|issues|urls|tasks|notes [query]/recall ask <question> answers from your history with Haiku 4.5, citing sessions. It costs a little usage and sends the matching excerpts to the model./recall stats, /recall reindex, /recall forget session <id>|project <name>|before <date> (asks you to confirm), /recall help./remember <fact>, /remember list, /remember forget <ref>.Last session here (2d ago): "…" · PR #99 · 3 open tasks [Recap].#214, ABC-12, a file name or a quoted phrase seen in past sessions, a band offers what happened then. Nothing is sent unless you click.OR gives alternatives, "quotes" an exact phrase, and -word excludes. Filters: project:name, kind:decision, since:7d, until:2026-09-30, source:codex, routines:include.~/.claude/recall/index.db, readable only by you, and never goes in a repo.~/.zshrc, ~/.zprofile, ~/.bashrc and ~/.bash_profile, and any literal strings you list in ~/.claude/recall/redact.txt (one per line). Editing that list re-masks the existing index on the next update./recall ask. The first index takes about 2 minutes in the background, with progress on the status line. After that it updates incrementally (about 1s) at session start and every 10 minutes./usr/bin/python3 (Command Line Tools), whose SQLite has FTS5. Nothing else to install.Settings: dbPath, python, sources, includeSubagents (on), includeRoutines (off), updateMinutes (10), relatedBand (on), lastSessionBand (on), maxResults (8), askModel (claude-haiku-4-5-20251001).
Catches Claude up on the repo when a session starts, so you don't have to ask "check the recent commits/PRs and issues".
owner, todo, P0 or blockedmain ↑1 · 3 changed · PRs #123 ✗ #124 ✓ · 2 owner issues · 2 stale branches · last commit 2h ago. Hide dismisses it./brief re-gathers now and prints the full summary.git and gh. The band refreshes after a turn at most every 2 minutes.Settings: focus labels, refresh minutes, and whether to brief Claude.
Keeps scheduled routines (daily digests, newsletters) from silently stalling while you're away.
AskUserQuestion, you get a Mac notification and a toast, and the status line shows routine: daily-report · waiting on you 3m.Routine daily-report finished after 23m · waited on you 2 times.notifyCommand runs a command on the same events, e.g. curl -s -d {message} ntfy.sh/your-topic. {title} and {message} are filled in as single arguments, never through a shell.allowWebReads (off by default): lets routines use WebFetch and WebSearch without asking. It only replaces a prompt; your deny rules still apply, and nothing else is ever auto-allowed./routine shows the routine's name, how long it has run, its waits, and the settings.A Fable review in the background, without switching your session's model. Each run is one Fable call against your usage.
/second-opinion: reviews recent work. On a feature branch that's the branch against the default branch; otherwise the last 12 commits, plus the diff and git status, capped at 60k characters./second-opinion commits 5/second-opinion diff (uncommitted changes)/second-opinion file docs/ADR-007.md/second-opinion <question>: adds a question for Fable to answer first.second opinion: reviewing…. When the review is ready you get a toast, and a pane opens with it, ranked: wrong assumptions, bugs and risks, what's missing, what to do next.~/.claude/second-opinions/<project>/. /second-opinion list lists them, and /second-opinion show [n] reopens one.Settings: model (claude-fable-5-1), effort (high), maxContextChars (60000).
Keeps your "always / never / don't / from now on" instructions alive across compaction.
Keep as a standing order? [Project] [This session] [No]. Nothing is saved without a click.~/.claude/standing-orders/<repo>.json and apply to every session in that repo. Session orders and your active /goal last for the session./clear, so the prompt cache isn't disturbed. A newly saved order also rides along once with your next message./orders lists them. /orders add [project|session] <text>, /orders forget <n>, /orders clear session|project, and /orders export (a Markdown block for CLAUDE.md).Stops keys and passwords from going into a prompt, and so into your transcripts, and turns them into env vars instead.
$NAME references, placeholders, plain URLs, paths, git SHAs and ordinary prose about passwords.…vxrm) with a suggested name such as OPENDATALAB_SECRET_ACCESS_KEY, which you can edit:export NAME='…' to ~/.zshrc (reusing an existing identical export) and replaces the secret in your prompt with $NAME./secrets test <text> output is masked too./secrets test <text> shows what would be caught. /secrets off and /secrets on toggle it for the session.Settings: enabled (on), extraPatterns (a regex), zshrcPath (~/.zshrc).
Makes Claude's open commands hand 3D files to the right slicer.
full.?spectrum|snapmaker-only|-fs\.3mf$|-u1[-.], orAn open -a BambuStudio … for one becomes open -b com.snapmaker.snapmaker-orca …, with the rest of the command untouched. You get a toast, and Claude gets a note so it doesn't try again.
/slice <file> [bambu|snapmaker|orca] opens a file yourself, with the same rules.open -a <app>, open -a /Applications/X.app and open -b <bundle id>, including variables set earlier in the command (S=… && open -a BambuStudio "$S/x.3mf") and files copied in the same command.Settings:
closePrevious (on): turn it off if you keep your own slicer window open, since the quit request reaches your windows too.fullSpectrumPattern (the regex above)checkContents (on)The quit request goes out when Claude issues the command, before any permission prompt for it.
--plugin-dir once per mod, e.g. claude --plugin-dir ~/Projects/claude-mods/job-watch --plugin-dir ~/Projects/claude-mods/pr-autopilot~/.claude/settings.json. Put mod-monitor first so it sees the others; CLAUDE_CODE_PLUGIN_DIR_WATCH makes desktop sessions pick up edits and show mod failures: ``json { "env": { "CLAUDE_CODE_PLUGIN_DIR_WATCH": "1", "CLAUDE_CODE_PLUGIN_DIRS": "~/Projects/claude-mods/mod-monitor:~/Projects/claude-mods/job-watch:~/Projects/claude-mods/machine-guard:~/Projects/claude-mods/repo-brief:~/Projects/claude-mods/slicer-handoff:~/Projects/claude-mods/pr-autopilot:~/Projects/claude-mods/routine-watch:~/Projects/claude-mods/modal-meter:~/Projects/claude-mods/second-opinion:~/Projects/claude-mods/downloads-drop:~/Projects/claude-mods/dev-servers:~/Projects/claude-mods/standing-orders:~/Projects/claude-mods/effort-router:~/Projects/claude-mods/secret-guard:~/Projects/claude-mods/recall" } } ``Run with Claude Code 2.1.289 or newer; older CLIs ignore per-test settings, so a few tests fall back to defaults.
claude plugin validate job-watch && claude plugin test job-watch
claude plugin validate machine-guard && claude plugin test machine-guard
claude plugin validate repo-brief && claude plugin test repo-brief
claude plugin validate slicer-handoff && claude plugin test slicer-handoff
claude plugin validate pr-autopilot && claude plugin test pr-autopilot
claude plugin validate routine-watch && claude plugin test routine-watch
claude plugin validate modal-meter && claude plugin test modal-meter
claude plugin validate second-opinion && claude plugin test second-opinion
claude plugin validate downloads-drop && claude plugin test downloads-drop
claude plugin validate dev-servers && claude plugin test dev-servers
claude plugin validate standing-orders && claude plugin test standing-orders
claude plugin validate effort-router && claude plugin test effort-router
claude plugin validate secret-guard && claude plugin test secret-guard
claude plugin validate recall && claude plugin test recall
claude plugin validate mod-monitor && claude plugin test mod-monitor
(cd recall/engine && /usr/bin/python3 -m unittest)hooks/register.ts 295 lines1/**
2 * effort-router: sends git housekeeping replies ("merged", "#219 merged", "commit and push", "push it")
3 * at low effort and deep asks ("review", "audit", "why does ...") at max, by rewriting the effort of
4 * each model request (`turn.step`). Approvals and replies that start new work ("yes", "continue",
5 * "merged, go ahead with #214") keep the session's own effort. The prompt itself passes through
6 * unchanged; no model call is made.
7 *
8 * The prompt cache. Claude Code sends effort as the request's top-level `output_config.effort`
9 * (this build has no per-message effort beta), and the API renders effort into the prompt: changing
10 * it between requests invalidates the cached messages (on some models the system prompt and tools
11 * too), so the next request writes the whole conversation to the cache again at the write price
12 * instead of reading it at a tenth of the input price. On a large conversation one switch can cost
13 * more than low effort saves on a short reply, and switching down and back up costs it twice. So:
14 *
15 * - The effort is chosen once per turn, at its first request, and every later request of the turn
16 * (each tool round) carries the same effort: never a switch mid-turn.
17 * - A switch is made at once when it is cheap or asked for: the context is small (`freeSwitchTokens`),
18 * the cache has lapsed (idle past `cacheTtlMinutes`; Claude Code caches the main thread for an hour
19 * on a subscription, five minutes with an API key), the model or the session's own /effort changed,
20 * compaction or /clear rewrote the history, `/route deep|routine` forced it, or the turn wants MORE
21 * effort (quality never waits for the cache).
22 * - Less effort over a warm, large cache waits until `stickyTurns` turns in a row have wanted it
23 * (hysteresis): one "push it" between two real tasks is not worth two cache rewrites.
24 *
25 * The model guard (`avoidModel` → `fallbackModel`) rewrites every request the same way, main loop and
26 * subagents alike, so the cache is rebuilt once for the new model and then reused.
27 */
28import { atom, read, update } from 'claude-code'
29import type { EngineInterface, PromptOrigin, Register, TurnStepInput, TurnStepResult } from 'claude-code'
30
31import {
32 EMPTY_STATE,
33 cleared,
34 compilePatterns,
35 describeRoute,
36 guardModel,
37 isLevel,
38 planStep,
39 resolveModel,
40 responded,
41 started,
42 submitted,
43} from './route'
44import type { Level, Patterns, Planned, Settings } from './route'
45
46type Engine = EngineInterface
47
48const router = atom({ plugin: 'effort-router', key: 'router' } as const, EMPTY_STATE)
49
50/** Prompts the person wrote: typed, through Remote Control, a host's own turn, a scheduled routine, a Slack ping. */
51const PERSONAL = new Set<string>(['composer', 'bridge', 'sdk', 'scheduled-trigger', 'slack-ping'])
52
53const isPersonal = (origin: PromptOrigin): boolean =>
54 PERSONAL.has(origin.kind) || (origin.kind === 'plugin' && origin.asUser === true)
55
56const USAGE = 'Usage: /route (status) | /route on | /route off | /route deep | /route routine (forces the next prompt)'
57
58type Config = {
59 enabled: boolean
60 settings: Settings
61 patterns: Patterns
62 errors: string[]
63 avoid: RegExp | null
64 fallbackModel: string
65}
66
67const DEFAULT_SETTINGS: Settings = {
68 routineEffort: 'low',
69 deepEffort: 'max',
70 guard: { stickyTurns: 2, freeSwitchTokens: 30_000, cacheTtlMs: 60 * 60_000 },
71}
72
73let config: Config = {
74 enabled: true,
75 settings: DEFAULT_SETTINGS,
76 patterns: compilePatterns('', '').patterns,
77 errors: [],
78 avoid: null,
79 fallbackModel: 'opus',
80}
81
82const numberIn = (value: unknown, fallback: number, min: number, max: number): number => {
83 const n = typeof value === 'number' ? value : Number(value)
84 return Number.isFinite(n) ? Math.min(max, Math.max(min, n)) : fallback
85}
86
87const textOf = (value: unknown): string => (typeof value === 'string' ? value : '')
88
89const modelGuardLine = (): string | null =>
90 config.avoid === null ? null : `models matching /${config.avoid.source}/i run on ${resolveModel(config.fallbackModel, '')}`
91
92/** The model a request goes to; the first rewrite of each kind is announced once. */
93async function guardedModel($: Engine, model: string): Promise<string> {
94 if (config.avoid === null || !config.avoid.test(model)) {
95 return model
96 }
97 const target = guardModel(model, config.avoid, config.fallbackModel)
98 const key = `${model}→${target}`
99 let isFirst = false
100 await update($, router, state => {
101 isFirst = !state.rewrites.includes(key)
102 return isFirst ? { ...state, rewrites: [...state.rewrites, key] } : state
103 })
104 if (isFirst) {
105 $.ui.toast(
106 target === model
107 ? `effort-router: ${model} matches avoidModel, but so does the fallback "${config.fallbackModel}"; requests stay on ${model}.`
108 : `effort-router: requests for ${model} go to ${target} (avoidModel).`,
109 )
110 }
111 return target
112}
113
114/** The context's size when the mod first sees a request (it may have loaded mid-session); 0 if unknown. */
115async function contextNow($: Engine): Promise<number> {
116 try {
117 const usage = await $.session.usage()
118 return usage.context.tokens ?? 0
119 } catch {
120 return 0
121 }
122}
123
124/** The effort this main-loop request is sent at: decided at the turn's first request, then kept. */
125async function chooseEffort($: Engine, e: TurnStepInput, model: string, base: Level): Promise<Level> {
126 const before = await read($, router)
127 if (before.turn !== null && before.turn.id === e.turnId && before.turn.effort !== null) {
128 return before.turn.effort
129 }
130 const now = await $.clock.now()
131 const contextTokens = before.cache === null ? await contextNow($) : 0
132 const plan: { planned: Planned | null } = { planned: null }
133 await update($, router, state => {
134 plan.planned = planStep(state, { turnId: e.turnId, model, base, messageCount: e.messageCount, now, contextTokens }, config.settings)
135 return plan.planned.state
136 })
137 if (plan.planned === null) {
138 return base
139 }
140 if (plan.planned.isNew) {
141 $.ui.status(plan.planned.status)
142 }
143 return plan.planned.effort
144}
145
146/** Remembers how large the cached conversation is now, and when it was last written. */
147async function noteResponse($: Engine, messageCount: number, model: string, result: TurnStepResult) {
148 try {
149 const now = await $.clock.now()
150 await update($, router, state => responded(state, { model, messageCount, now, usage: result.usage }))
151 } catch {
152 // The next turn decides from what was known before; nothing to undo.
153 }
154}
155
156/** A failure in effort-router's own code, logged (to the debug log) for mod-monitor to count. */
157function reportFailure($: Engine, what: string, error: unknown) {
158 const message = error instanceof Error ? error.message : String(error)
159 $.ui.log(`effort-router: ${what} failed, so the request went out unchanged: ${message.slice(0, 200)}`, { to: 'debug' })
160}
161
162export const register: Register = (on, options) => {
163 const compiled = compilePatterns(textOf(options.routinePattern), textOf(options.deepPattern))
164 const errors = [...compiled.errors]
165 let avoid: RegExp | null = null
166 const avoidSource = textOf(options.avoidModel).trim()
167 if (avoidSource !== '') {
168 try {
169 avoid = new RegExp(avoidSource, 'i')
170 } catch (error) {
171 errors.push(`avoidModel is not a valid regular expression (${error instanceof Error ? error.message : String(error)}); the model guard is off.`)
172 }
173 }
174 config = {
175 enabled: options.enabled !== false,
176 settings: {
177 routineEffort: isLevel(options.routineEffort) ? options.routineEffort : 'low',
178 deepEffort: isLevel(options.deepEffort) ? options.deepEffort : 'max',
179 guard: {
180 stickyTurns: Math.round(numberIn(options.stickyTurns, 2, 1, 20)),
181 freeSwitchTokens: numberIn(options.freeSwitchTokens, 30_000, 0, 10_000_000),
182 cacheTtlMs: numberIn(options.cacheTtlMinutes, 60, 0, 24 * 60) * 60_000,
183 },
184 },
185 patterns: compiled.patterns,
186 errors,
187 avoid,
188 fallbackModel: textOf(options.fallbackModel).trim() || 'opus',
189 }
190
191 on('session.start', async ($, e, next) => {
192 await $.command.register({
193 name: 'route',
194 description: 'Effort routing: status, on/off for this session, or force the next prompt deep or routine',
195 argumentHint: '[on | off | deep | routine]',
196 immediate: true,
197 })
198 return next(e)
199 })
200
201 on('command.run', { command: 'route' }, async ($, e) => {
202 if (!config.enabled) {
203 return { text: 'effort-router is turned off in its settings (enabled: false); requests are left as they are.' }
204 }
205 const verb = e.args.trim().toLowerCase()
206 if (verb === 'off') {
207 await update($, router, state => ({ ...state, isOn: false }))
208 return { text: "Effort routing is off for this session: turns run at the session's own effort. /route on turns it back on." }
209 }
210 if (verb === 'on') {
211 await update($, router, state => ({ ...state, isOn: true }))
212 return { text: 'Effort routing is on.' }
213 }
214 if (verb === 'deep' || verb === 'routine') {
215 const forced: 'deep' | 'routine' = verb
216 await update($, router, state => ({ ...state, forced }))
217 const effort = forced === 'deep' ? config.settings.deepEffort : config.settings.routineEffort
218 return { text: `Your next prompt runs at ${effort} (${forced}), whatever it says.` }
219 }
220 if (verb !== '' && verb !== 'status') {
221 return { text: USAGE }
222 }
223 const state = await read($, router)
224 return { text: describeRoute(state, { settings: config.settings, modelGuard: modelGuardLine(), errors: config.errors }) }
225 })
226
227 if (!config.enabled) {
228 return
229 }
230
231 // Classifies the prompt for the turn it starts; the prompt itself goes on unchanged.
232 on('prompt.submit', async ($, e, next) => {
233 try {
234 const isMine = isPersonal(e.origin)
235 await update($, router, state => submitted(state, e.text, isMine, config.patterns))
236 } catch {
237 // Unclassified: the turn runs at the session's own effort.
238 }
239 return next(e)
240 })
241
242 on('turn.start', async ($, e, next) => {
243 await update($, router, state => started(state, e.turnId, e.text, config.patterns))
244 return next(e)
245 })
246
247 on('turn.step', async function* ($, e, next) {
248 // Routing must never cost a request: if choosing fails, the request goes out exactly as it came, and
249 // the failure is logged where mod-monitor reads it (this stream is not a hook it can watch).
250 let routed = e
251 let model = e.model
252 let isRouted = false
253 try {
254 model = await guardedModel($, e.model)
255 if (e.agentId !== undefined || !isLevel(e.effort)) {
256 // A subagent keeps its own effort; a model without effort levels has none to route.
257 routed = model === e.model ? e : { ...e, model }
258 } else {
259 const effort = await chooseEffort($, e, model, e.effort)
260 routed = effort === e.effort && model === e.model ? e : { ...e, model, effort }
261 isRouted = true
262 }
263 } catch (error) {
264 routed = e
265 model = e.model
266 isRouted = false
267 reportFailure($, 'choosing the effort', error)
268 }
269 const result = yield* next(routed)
270 if (isRouted) {
271 try {
272 await noteResponse($, e.messageCount, model, result)
273 } catch (error) {
274 reportFailure($, 'noting the response', error)
275 }
276 }
277 return result
278 })
279
280 on('turn.complete', async ($, e, next) => {
281 const done = await next(e)
282 if (e.agentId === undefined) {
283 $.ui.status(undefined)
284 }
285 return done
286 })
287
288 on('session.end', async ($, e, next) => {
289 if (e.reason === 'clear') {
290 await update($, router, cleared)
291 }
292 return next(e)
293 })
294}
295hooks/route.ts 435 lines1import type { ModelEffort, ModelUsage } from 'claude-code'
2
3import type { RouteCache, RouteClass, RouteCounts, RoutePending, RouteState, RouteTurn } from '../types'
4
5export type Level = ModelEffort
6
7export const LEVELS: readonly Level[] = ['low', 'medium', 'high', 'xhigh', 'max']
8
9export const isLevel = (value: unknown): value is Level =>
10 typeof value === 'string' && (LEVELS as readonly string[]).includes(value)
11
12export const rank = (level: Level): number => LEVELS.indexOf(level)
13
14/** A prompt longer than this is never routine, whatever it starts with. */
15export const ROUTINE_MAX_CHARS = 120
16/** How many words may follow the routine phrases ("both are merged now" leaves one), none of them new work. */
17export const ROUTINE_MAX_EXTRA_WORDS = 6
18
19/**
20 * Routine replies are git housekeeping, matched at the start of the prompt: "merged", "#123 merged",
21 * "both are merged", "commit and push", "commit it", "push it", "open a PR", "close it". Approvals and
22 * resumptions ("yes", "go ahead", "continue", "keep going") are not: they start real work.
23 */
24export const DEFAULT_ROUTINE = String.raw`merged|(?:pr\s*)?#\d+\s+(?:is\s+|was\s+|got\s+)?merged|(?:all\s+|both\s+|everything\s+)?(?:is\s+|are\s+)?merged|commit(?:\s+and)?\s+push|commit(?:\s+it)?|push(?:\s+it)?|open\s+(?:a\s+)?pr|close\s+(?:it|the\s+issue)`
25
26/** Words that ask for deep thought, anywhere in the prompt (word starts: "audit" covers "auditing"). */
27export const DEFAULT_DEEP = String.raw`\b(?:audit|review|plan(?:s|ned|ning)?\b|architect|design|deep[\s-]*dive|assess|investigat|root[\s-]*cause|research|why\s+(?:is|does|did)\b|figure\s+out|what(?:['’]s|\s+is)\s+wrong)`
28
29export type Patterns = {
30 /** Anchored at the start; one routine phrase. */
31 routine: RegExp
32 /** Unanchored; any deep wording. */
33 deep: RegExp
34}
35
36export type PatternSet = { patterns: Patterns; errors: string[] }
37
38const leading = (source: string): RegExp => new RegExp(String.raw`^(?:${source})(?![A-Za-z0-9_])`, 'i')
39const anywhere = (source: string): RegExp => new RegExp(source, 'i')
40
41const compileOne = (
42 source: string,
43 fallback: string,
44 make: (source: string) => RegExp,
45 name: string,
46 errors: string[],
47): RegExp => {
48 const trimmed = source.trim()
49 if (trimmed !== '') {
50 try {
51 return make(trimmed)
52 } catch (error) {
53 errors.push(`${name} is not a valid regular expression (${error instanceof Error ? error.message : String(error)}); the built-in one is used.`)
54 }
55 }
56 return make(fallback)
57}
58
59/** The routine and deep patterns from settings; an empty or broken one falls back to the built-in. */
60export function compilePatterns(routineSource: string, deepSource: string): PatternSet {
61 const errors: string[] = []
62 return {
63 patterns: {
64 routine: compileOne(routineSource, DEFAULT_ROUTINE, leading, 'routinePattern', errors),
65 deep: compileOne(deepSource, DEFAULT_DEEP, anywhere, 'deepPattern', errors),
66 },
67 errors,
68 }
69}
70
71/** Punctuation and quoting between and around routine phrases. */
72const FILLER = /^[\s,.;:!?…—–\-"'`*_()[\]>]+/
73
74const wordCount = (text: string): number => text.split(/\s+/).filter(word => /[A-Za-z0-9]/.test(word)).length
75
76/** The routine phrases a prompt opens with ("merged, commit and push") and what follows them. */
77export function leadingRoutine(text: string, routine: RegExp): { phrases: string[]; rest: string } | null {
78 let rest = text.replace(FILLER, '')
79 const phrases: string[] = []
80 for (let i = 0; i < 8; i++) {
81 const found = routine.exec(rest)
82 if (!found || found[0] === '') {
83 break
84 }
85 phrases.push(found[0])
86 rest = rest.slice(found[0].length).replace(FILLER, '')
87 }
88 return phrases.length > 0 ? { phrases, rest } : null
89}
90
91/**
92 * Words after the routine phrases that start new work ("merged, go ahead with #214", "push it and fix
93 * the lint"): such a prompt is not routine, whatever it opens with.
94 */
95export const NEW_WORK =
96 /\b(?:start|begin|work\s+on|take|tackle|fix|implement|build|add|write|create|make|proceed|move\s+on|run|deploy|refactor|update|change|rename|remove|delete|investigate|look|check|debug|test|handle|address|file|install|download|train|generate|render|print|go\s+(?:ahead\s+)?(?:with|on)|continue|keep\s+going|do\b|try)\b|\b(?:with|on|to)\s+#?[A-Za-z]*-?\d/i
97
98export type Classified = { cls: RouteClass; why: string }
99
100/**
101 * Routine: at most 120 characters, opening with routine phrases and little else. Deep: deep wording
102 * anywhere. Both at once ("commit and push, then review the diff") is no opinion: the session's own effort.
103 */
104export function classify(text: string, patterns: Patterns): Classified {
105 const trimmed = text.trim()
106 const lead = leadingRoutine(trimmed, patterns.routine)
107 const isRoutine =
108 lead !== null &&
109 trimmed.length <= ROUTINE_MAX_CHARS &&
110 wordCount(lead.rest) <= ROUTINE_MAX_EXTRA_WORDS &&
111 !NEW_WORK.test(lead.rest)
112 const deep = patterns.deep.exec(trimmed)
113 if (isRoutine && deep) {
114 return { cls: 'neutral', why: `both routine ("${lead.phrases.join(', ')}") and deep ("${deep[0]}")` }
115 }
116 if (isRoutine) {
117 return { cls: 'routine', why: `starts with "${lead.phrases.join(', ')}"` }
118 }
119 if (deep) {
120 return { cls: 'deep', why: `mentions "${deep[0]}"` }
121 }
122 if (lead !== null) {
123 return { cls: 'neutral', why: trimmed.length > ROUTINE_MAX_CHARS ? 'too long to be routine' : 'more than a routine reply' }
124 }
125 return { cls: 'neutral', why: 'no routine or deep wording' }
126}
127
128export type Guard = {
129 /** Turns in a row a lower effort must be wanted before it is sent over a warm, large cache. */
130 stickyTurns: number
131 /** Below this many context tokens a switch costs little: it is made at once. */
132 freeSwitchTokens: number
133 /** Idle longer than this and the cache has lapsed: a switch costs nothing extra. */
134 cacheTtlMs: number
135}
136
137export type Why = 'same' | 'forced' | 'off' | 'model' | 'base' | 'history' | 'cold' | 'small' | 'up' | 'streak' | 'hold'
138
139export type Ask = {
140 want: Level
141 /** The model this request names (after the model guard). */
142 model: string
143 /** The session's own effort, as the engine offers it on this request. */
144 base: Level
145 messageCount: number
146 now: number
147 isForced: boolean
148}
149
150export type Decision = { effort: Level; why: Why; lowerStreak: number }
151
152/**
153 * The effort a new turn is sent at, given what the prompt cache was written with.
154 *
155 * Claude Code sends effort as the request's top-level `output_config.effort`, and the API renders it
156 * into the prompt: a change invalidates the cached messages (on some models the system prompt and
157 * tools too), so the next request re-writes the whole conversation at the cache-write price. So a
158 * switch is made at once only when it is cheap or asked for: the person forced it, the model or the
159 * session's own effort changed (the cache is cold for it anyway), the history shrank (compaction,
160 * /clear), the cache has lapsed (idle past its TTL), the context is small, or the turn wants MORE
161 * effort (quality never waits). A turn wanting LESS effort over a warm, large cache is held at the
162 * current effort until `stickyTurns` turns in a row have wanted less.
163 */
164export function decide(cache: RouteCache, ask: Ask, guard: Guard): Decision {
165 const switchTo = (why: Why): Decision => ({ effort: ask.want, why, lowerStreak: 0 })
166 if (ask.want === cache.applied) {
167 return switchTo('same')
168 }
169 if (ask.isForced) {
170 return switchTo('forced')
171 }
172 if (ask.model !== cache.model) {
173 return switchTo('model')
174 }
175 if (ask.base !== cache.base) {
176 return switchTo('base')
177 }
178 if (ask.messageCount < cache.messageCount) {
179 return switchTo('history')
180 }
181 if (cache.lastAt !== null && ask.now - cache.lastAt > guard.cacheTtlMs) {
182 return switchTo('cold')
183 }
184 if (cache.contextTokens < guard.freeSwitchTokens) {
185 return switchTo('small')
186 }
187 if (rank(ask.want) > rank(cache.applied)) {
188 return switchTo('up')
189 }
190 const lowerStreak = cache.lowerStreak + 1
191 if (lowerStreak >= Math.max(1, guard.stickyTurns)) {
192 return switchTo('streak')
193 }
194 return { effort: cache.applied, why: 'hold', lowerStreak }
195}
196
197export const EMPTY_COUNTS: RouteCounts = { routine: 0, deep: 0, neutral: 0, switches: 0, holds: 0 }
198
199export const EMPTY_STATE: RouteState = {
200 isOn: true,
201 forced: null,
202 pending: null,
203 turn: null,
204 cache: null,
205 counts: EMPTY_COUNTS,
206 rewrites: [],
207}
208
209/** A prompt was submitted: classify it (or take the class /route forced) for the turn it starts. */
210export function submitted(state: RouteState, text: string, isPersonal: boolean, patterns: Patterns): RouteState {
211 if (!isPersonal) {
212 return { ...state, pending: { cls: 'neutral', why: 'not typed by you', isForced: false, text } }
213 }
214 if (state.forced !== null) {
215 return { ...state, forced: null, pending: { cls: state.forced, why: 'forced with /route', isForced: true, text } }
216 }
217 return { ...state, pending: { ...classify(text, patterns), isForced: false, text } }
218}
219
220/** A main-loop turn started: it takes the pending prompt's class. */
221export function started(state: RouteState, turnId: string, text: string, patterns: Patterns): RouteState {
222 const pending: RoutePending =
223 state.pending ??
224 (text.trim() === ''
225 ? { cls: 'neutral', why: 'no prompt', isForced: false, text: '' }
226 : { ...classify(text, patterns), isForced: false, text })
227 const turn: RouteTurn = { ...pending, id: turnId, effort: null, want: null, decision: null, streak: 0 }
228 return { ...state, pending: null, turn }
229}
230
231export type Settings = { routineEffort: Level; deepEffort: Level; guard: Guard }
232
233/** What a main-loop request carries, as the engine is about to send it. */
234export type StepFacts = {
235 turnId: string
236 /** The model it names, after the model guard. */
237 model: string
238 /** The session's own effort. */
239 base: Level
240 messageCount: number
241 now: number
242 /** The context's size, used only when nothing was seen before (the mod loaded mid-session). */
243 contextTokens: number
244}
245
246export type Planned = { state: RouteState; effort: Level; isNew: boolean; status: string | undefined }
247
248/** What the status line says while a turn runs; undefined when the turn runs as the session would. */
249export function statusText(cls: RouteClass, decision: Decision, isForced: boolean, stickyTurns: number): string | undefined {
250 if (decision.why === 'hold') {
251 return `effort: ${decision.effort} (${cls === 'neutral' ? '' : `${cls}, `}held for the prompt cache ${decision.lowerStreak}/${stickyTurns})`
252 }
253 if (cls === 'neutral' || decision.why === 'off') {
254 return undefined
255 }
256 return `effort: ${decision.effort} (${cls}${isForced ? ', forced' : ''})`
257}
258
259/**
260 * The effort a main-loop request is sent at. The turn's first request decides; every later
261 * request of the same turn carries the same effort, so a turn never switches midway.
262 */
263export function planStep(state: RouteState, facts: StepFacts, settings: Settings): Planned {
264 const current = state.turn
265 if (current !== null && current.id === facts.turnId && current.effort !== null) {
266 return { state, effort: current.effort, isNew: false, status: undefined }
267 }
268 const turn: RouteTurn =
269 current !== null && current.id === facts.turnId
270 ? current
271 : { id: facts.turnId, cls: 'neutral', why: 'its prompt was not seen', isForced: false, text: '', effort: null, want: null, decision: null, streak: 0 }
272 const cache: RouteCache = state.cache ?? {
273 applied: facts.base,
274 model: facts.model,
275 base: facts.base,
276 messageCount: facts.messageCount,
277 contextTokens: facts.contextTokens,
278 lastAt: null,
279 lowerStreak: 0,
280 }
281 const isRouted = state.isOn || turn.isForced
282 const want = !isRouted || turn.cls === 'neutral' ? facts.base : turn.cls === 'routine' ? settings.routineEffort : settings.deepEffort
283 const decision: Decision = isRouted
284 ? decide(cache, { want, model: facts.model, base: facts.base, messageCount: facts.messageCount, now: facts.now, isForced: turn.isForced }, settings.guard)
285 : { effort: want, why: 'off', lowerStreak: 0 }
286 const counts: RouteCounts = {
287 ...state.counts,
288 [turn.cls]: state.counts[turn.cls] + 1,
289 switches: state.counts.switches + (decision.effort === cache.applied ? 0 : 1),
290 holds: state.counts.holds + (decision.why === 'hold' ? 1 : 0),
291 }
292 return {
293 state: {
294 ...state,
295 turn: { ...turn, effort: decision.effort, want, decision: decision.why, streak: decision.lowerStreak },
296 cache: {
297 ...cache,
298 applied: decision.effort,
299 model: facts.model,
300 base: facts.base,
301 messageCount: facts.messageCount,
302 lowerStreak: decision.lowerStreak,
303 },
304 counts,
305 },
306 effort: decision.effort,
307 isNew: true,
308 status: statusText(turn.cls, decision, turn.isForced, Math.max(1, settings.guard.stickyTurns)),
309 }
310}
311
312/** Prompt tokens a response was answered over plus its output: what the next request re-sends. */
313export const contextOf = (usage: ModelUsage): number =>
314 usage.input_tokens + usage.cache_read_input_tokens + usage.cache_creation_input_tokens + usage.output_tokens
315
316/** A main-loop response arrived: remember how big the cached prefix now is, and when it was written. */
317export function responded(
318 state: RouteState,
319 facts: { model: string; messageCount: number; now: number; usage: ModelUsage | null },
320): RouteState {
321 if (state.cache === null) {
322 return state
323 }
324 const cache: RouteCache =
325 facts.usage === null
326 ? { ...state.cache, model: facts.model, messageCount: facts.messageCount }
327 : { ...state.cache, model: facts.model, messageCount: facts.messageCount, contextTokens: contextOf(facts.usage), lastAt: facts.now }
328 return { ...state, cache }
329}
330
331/** /clear: a new conversation whose first request writes a fresh cache, so the next switch is free. */
332export function cleared(state: RouteState): RouteState {
333 return {
334 ...state,
335 pending: null,
336 turn: null,
337 counts: EMPTY_COUNTS,
338 cache: state.cache === null ? null : { ...state.cache, messageCount: 0, contextTokens: 0, lastAt: null, lowerStreak: 0 },
339 }
340}
341
342/** Current model ids for the family aliases, so a rewrite never depends on the engine resolving an alias. */
343export const MODEL_IDS: Readonly<Record<string, string>> = {
344 opus: 'claude-opus-5-5',
345 sonnet: 'claude-sonnet-5-5',
346 haiku: 'claude-haiku-5-5',
347 fable: 'claude-fable-5-1',
348}
349
350/**
351 * The model id a fallback names: a family alias becomes its current id, spelled with the provider
352 * prefix the engine's own id carries (`us.anthropic.`); anything else is taken as written.
353 */
354export function resolveModel(fallback: string, current: string): string {
355 const name = fallback.trim()
356 const id = MODEL_IDS[name.toLowerCase()]
357 if (id === undefined) {
358 return name
359 }
360 const at = current.indexOf('claude-')
361 return (at > 0 ? current.slice(0, at) : '') + id
362}
363
364/** The model a request goes to: the fallback when the engine's matches `avoid` (and the fallback does not). */
365export function guardModel(model: string, avoid: RegExp | null, fallback: string): string {
366 if (avoid === null || !avoid.test(model)) {
367 return model
368 }
369 const target = resolveModel(fallback, model)
370 return target === '' || avoid.test(target) ? model : target
371}
372
373const clip = (text: string, max: number): string => {
374 const line = text.replace(/\s+/g, ' ').trim()
375 return line.length > max ? `${line.slice(0, max - 1)}…` : line
376}
377
378const WHY_WORDS: Readonly<Record<string, string>> = {
379 same: 'no change',
380 forced: 'forced',
381 off: 'routing off',
382 model: 'switched with the model',
383 base: 'your own /effort changed',
384 history: 'after compaction',
385 cold: 'cache had lapsed',
386 small: 'small context',
387 up: 'more effort never waits',
388 streak: 'lower effort wanted turns in a row',
389 hold: 'held for the prompt cache',
390}
391
392export type Describe = {
393 settings: Settings
394 /** e.g. `models matching /fable/i run on claude-opus-5-5`, or null when off. */
395 modelGuard: string | null
396 errors: readonly string[]
397}
398
399const formatTokens = (n: number): string => (n >= 1000 ? `${Math.round(n / 1000)}k` : String(n))
400
401/** What /route prints. */
402export function describeRoute(state: RouteState, info: Describe): string {
403 const { settings } = info
404 const effortFor = (cls: 'routine' | 'deep'): Level => (cls === 'routine' ? settings.routineEffort : settings.deepEffort)
405 const lines = [state.isOn ? 'effort-router: on' : 'effort-router: off for this session (/route on resumes; /route deep or routine still apply)']
406 if (state.forced !== null) {
407 lines.push(`Next prompt: forced ${state.forced} (${effortFor(state.forced)})`)
408 }
409 const turn = state.turn
410 if (turn === null) {
411 lines.push('No turn routed yet.')
412 } else {
413 const quoted = turn.text === '' ? '' : ` "${clip(turn.text, 60)}"`
414 const effort =
415 turn.effort === null
416 ? 'effort not chosen yet'
417 : turn.decision === 'hold'
418 ? `${turn.effort}, held for the prompt cache (wanted ${turn.want ?? '?'})`
419 : `${turn.effort} (${WHY_WORDS[turn.decision ?? 'same'] ?? turn.decision})`
420 lines.push(`Last turn: ${turn.cls}${quoted}, ${turn.why} → ${effort}`)
421 }
422 const c = state.counts
423 lines.push(
424 `This session: ${c.routine} routine, ${c.deep} deep, ${c.neutral} neutral turns · ${c.switches} effort ${c.switches === 1 ? 'switch' : 'switches'} · ${c.holds} held for the prompt cache`,
425 `Routine → ${settings.routineEffort}, deep → ${settings.deepEffort}, neutral → the session's own effort.`,
426 `Cache guard: a switch is made at once under ${formatTokens(settings.guard.freeSwitchTokens)} context tokens, after ${Math.round(settings.guard.cacheTtlMs / 60_000)} min idle, or toward more effort; less effort waits for ${Math.max(1, settings.guard.stickyTurns)} turns in a row.`,
427 )
428 if (state.cache !== null) {
429 lines.push(`Prompt cache: written at ${state.cache.applied} over about ${formatTokens(state.cache.contextTokens)} tokens.`)
430 }
431 lines.push(info.modelGuard === null ? 'Model guard: off' : `Model guard: ${info.modelGuard}`)
432 lines.push(...info.errors)
433 return lines.join('\n')
434}
435types/index.d.ts 76 lines1/** The effort levels a `turn.step` request can name, lowest first. */
2export type RouteLevel = 'low' | 'medium' | 'high' | 'xhigh' | 'max'
3
4/** What a prompt asks for: little thought, deep thought, or no opinion. */
5export type RouteClass = 'routine' | 'deep' | 'neutral'
6
7/** A prompt's class, waiting for the turn it starts. */
8export type RoutePending = {
9 cls: RouteClass
10 /** Why it got that class, for /route ("starts with \"merged\""). */
11 why: string
12 /** True when /route deep or /route routine chose the class. */
13 isForced: boolean
14 text: string
15}
16
17/** The main loop's current (or last) turn and the effort chosen for it. */
18export type RouteTurn = RoutePending & {
19 id: string
20 /** The effort every request of the turn carries; null until its first request. */
21 effort: RouteLevel | null
22 /** The effort the class asked for; differs from `effort` while the cache guard holds. */
23 want: RouteLevel | null
24 /** Why `effort` was chosen (`same`, `small`, `up`, `hold`, ...). */
25 decision: string | null
26 /** Turns in a row that wanted less effort than was being sent, this one included. */
27 streak: number
28}
29
30/** What the prompt cache was last written with on the main loop, so a switch is only made when it pays. */
31export type RouteCache = {
32 /** The effort the main loop's last request carried. */
33 applied: RouteLevel
34 /** The model the main loop's last request named. */
35 model: string
36 /** The session's own effort, as the engine last offered it. */
37 base: RouteLevel
38 /** How many messages the main loop's last request carried. */
39 messageCount: number
40 /** Prompt tokens the last response was answered over, plus its output: what the next request re-sends. */
41 contextTokens: number
42 /** When the last main-loop response arrived; null before one has. */
43 lastAt: number | null
44 /** Turns in a row that wanted less effort than `applied`. */
45 lowerStreak: number
46}
47
48export type RouteCounts = {
49 routine: number
50 deep: number
51 neutral: number
52 /** Turns whose effort differed from the turn before (each one rewrites the prompt cache). */
53 switches: number
54 /** Turns kept at the previous effort to spare the prompt cache. */
55 holds: number
56}
57
58export type RouteState = {
59 /** /route off turns automatic routing off for the session; /route deep and /route routine still apply. */
60 isOn: boolean
61 /** A class /route forced on the next prompt. */
62 forced: 'routine' | 'deep' | null
63 pending: RoutePending | null
64 turn: RouteTurn | null
65 cache: RouteCache | null
66 counts: RouteCounts
67 /** Model rewrites already announced with a toast (`from→to`). */
68 rewrites: string[]
69}
70
71declare module 'claude-code' {
72 interface PluginState {
73 'effort-router': { router: RouteState }
74 }
75}
76