SLOPSHOPPER

effort-router

Routes each turn's effort from your prompt: git housekeeping replies ("merged", "commit and push", "push it") run at low effort, deep asks (review, audit…

newcommandtoaststatusprompt
v0.1.0no licenseupdated 2026-10-08joeldg/claude-mods/effort-router
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · effort-router
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /route ⎿ effort-router: effort-router: on ⎿ effort-router: Last turn: deep "fix the failing auth test and add an audit log call", mentions "audit" → effort not chosen ye ⎿ effort-router: This session: 0 routine, 0 deep, 0 neutral turns · 0 effort switches · 0 held for the prompt cache ⎿ effort-router: Routine → low, deep → max, neutral → the session's own effort. ⎿ effort-router: Cache guard: a switch is made at once under 30k context tokens, after 60 min idle, or toward more effort; less ⎿ effort-router: Model guard: off ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts
README

claude-mods

Claude Code mods (function-hook plugins) I find helpful. Each folder is one plugin.

Requires a Claude Code build with function-hook plugins (2.1.289 or newer). machine-guard reads macOS tools (sysctl, memory_pressure, ioreg).

git clone https://github.com/joeldg/claude-mods ~/Projects/claude-mods

dev-servers

A pane of your project's running dev servers, so you don't have to ask Claude to restart them.

  • /servers opens the Servers pane:
  • running servers whose working folder is in this repo: name, port, pid and uptime, with Restart, Stop and Log
  • known start commands that aren't running, with Start: package.json dev/start/serve/preview scripts (run with your lockfile's package manager), .claude/launch.json and Procfile
  • a count of other listeners on the Mac
  • Only button presses start or stop anything:
  • Start runs the command detached, logging to ~/.claude/dev-servers/<project>/.
  • Stop sends SIGTERM. If the server ignores it, pressing again within 10s force-stops it.
  • Restart stops the server, waits for the port to free up, then starts it.
  • Before any signal, it checks the pid still runs the same command.
  • When a command fails with "address already in use", Claude gets a note (and you a toast) naming the holder, e.g. Port 4000 is held by node (pid 123, up 2h, in /Users/me/other).
  • Status line: servers: :4000 :5173.
  • Makes no model calls: lsof and ps every 15s.

downloads-drop

Puts files you just downloaded into your prompt with one click.

  • Watches ~/Downloads (top level). When a new file arrives (PDF, Markdown, images, 3MF/STL/OBJ, zip, video…), a band appears above the prompt: New in Downloads: paper.pdf, model-b.3mf · 2m ago [Attach] [Dismiss].
  • Attach puts @"/Users/you/Downloads/paper.pdf" mentions in your prompt. Dismiss hides those files.
  • Waits until a file has finished downloading (skips partial downloads and files still growing), and ignores hidden and zero-byte files.
  • /downloads lists the 10 newest files, numbered. /downloads attach 1 3 (or 2-4) adds those, and /downloads clear dismisses everything new.
  • Stacks with other mods' bands (repo-brief, standing-orders, secret-guard) instead of hiding them.
  • Makes no model calls.

Settings: folder (~/Downloads), extensions, pollSeconds (5), maxAgeMinutes (120).

effort-router

Sets effort per message, so you don't have to switch it by hand.

  • Git chores ("merged", "#219 merged", "commit and push", "push it", "open a PR", "close the issue") run at low effort and come back faster.
  • Deep asks (audit, review, plan, design, investigate, root cause, "why does…", "figure out") run at max.
  • Everything else, including approvals like "yes", "go ahead" and "continue" and anything that starts new work ("merged 219, go ahead with #214"), keeps your session's own effort.
  • Prompt cache: changing effort makes the whole conversation get cached again. So it never switches mid-turn, raises effort at once, and over a large, warm cache lowers it only after 2 routine turns in a row. In small contexts, or once the cache has lapsed, it switches right away.
  • Model guard: set avoidModel (a regex such as fable) to send those requests to fallbackModel instead, subagents included.
  • /route shows the last decision and the session's counts. /route off and /route on toggle it; /route deep and /route routine force the next turn.
  • Status line while a turn is routed: effort: low (routine).
  • Makes no model calls.

Settings: routineEffort (low), deepEffort (max), routinePattern, deepPattern, avoidModel, fallbackModel (opus), stickyTurns (2), freeSwitchTokens (30000), cacheTtlMinutes (60).

job-watch

A Jobs pane for long-running work: training runs, downloads, extractions.

  • Picks up background Bash tasks and detached nohup … > log & launches by itself.
  • /watch <log> [label] adds any other log file.
  • Shows progress, ETA and the last log line, and flags a job as stalled when its log goes quiet.
  • Shows free space on / and /Volumes/* (the NAS).
  • Toasts when a job finishes or stalls. The status line shows jobs: 2 running · 1 stalled.
  • /jobs opens the pane, /unwatch <label|done|all> removes jobs.
  • Makes no model calls: it reads logs with tail, checks processes with ps, and runs df.

Settings (in /config): stall minutes (10), refresh seconds (10), how long finished jobs stay (120 min), auto-open (on), which disks to show.

machine-guard

Memory, swap and GPU on the status line. It refuses heavy local jobs when the Mac can't take them.

  • Status line: RAM tight 12% free · swap 7.9/8G · top python 31G · GPU 87%.
  • It refuses heavy jobs (training, inference, rendering, extraction, Blender, Docker, ffmpeg) when:
  • macOS reports critical memory pressure, or
  • the Mac is reserved with /busy.

When memory is only tight, the job runs and Claude gets a note to start one heavy job at a time.

  • /busy 3h training a vision model reserves the Mac in every Claude session. /busy off lifts it. The reservation lives in ~/.claude/machine-guard.json, so a training script can write it too: ``bash echo '{"reason":"overnight training","until":'$(( ($(date +%s) + 8*3600) * 1000 ))'}' > ~/.claude/machine-guard.json ``
  • /guard shows what it sees. /guard pause 15m lets heavy jobs through in this session; /guard on resumes the guard.
  • Remote runs (modal run, ssh), tests (pytest) and installs are never treated as heavy.
  • Add your own heavy commands with the "Also heavy" setting (a regex), e.g. overnight_|nightly_run\.sh.

mod-monitor

Watches how the other mods behave in real use, without changing them. It is listed first in CLAUDE_CODE_PLUGIN_DIRS, so the other mods' hooks run beneath it.

  • Failures: any mod hook that throws, times out or rejects, read from the hook chain's results (next.trace), with the mod's name, the event and how long it ran. Slow hooks (over 1.5 s) are recorded too. The first failure of each mod in a session raises a toast.
  • What each mod did: its toasts ("#219 merged → …", "Blocked: …"), status-line changes, the mod commands you used (never their arguments), and failed subprocesses (a burst of 5 in 10 minutes raises a toast). Git checks run outside a repository are logged as expected, not as failures. It also records model calls (the only usage the mods cost: /second-opinion, /recall ask) and file writes (folders only, never contents).
  • /mods: a pane with one row per mod: ✓ active, ⚠ failing, ✗ not seen this session, · seen but idle. Each row shows today's counts and last activity, with Details for its recent events. It also says which mods it can't see, if any of them run above it.
  • /mods report [24h|7d|30d]: a per-mod report across all sessions, also written to ~/.claude/mods/monitor/report-latest.md for a scheduled review or Claude to read.
  • /mods failures [7d]: failures and failed subprocesses only.
  • Logs: ~/.claude/mods/monitor/<date>/<session>.jsonl, flushed every minute and at session end, with secrets masked and old days removed after 30 days.
  • Makes no model calls and adds no measurable latency.
  • Error lines mods log themselves ($.ui.log with wording like "failed" or "could not"): shown in Details and in /mods failures. Three in an hour mark the mod ⚠ and raise one toast. That is how effort-router's per-request hook, which runs inside the response stream where no monitor should sit, reports a failure. It also always sends the request on unchanged.
  • Transcript-row hooks (secret-guard masks /secrets records there) are watched for failures and slow runs, but not counted per run.

Settings: alerts (on), slowMs (1500), watchRender (on), watchCommands (on; off stops "mod-monitor" appearing beside other mods' command output), watchAppend (on), retentionDays (30), flushSeconds (60).

modal-meter

Keeps an eye on Modal so idle GPU containers don't burn credits.

  • Status line while containers run: Modal: 1 running (2 containers). Deployed apps with no containers cost nothing, so they stay off it.
  • A toast when an app has had containers up longer than alertMinutes (30), repeated at most every 30 minutes.
  • /modal opens a pane of apps with state, containers and uptime. Stop asks for Confirm, then runs modal app stop. Nothing is stopped any other way.
  • Shows today's spend and alerts on a budgetToday where the Modal CLI supports billing report (1.3.3+, Team/Enterprise workspaces). Otherwise /modal says why spend isn't shown.
  • Finds the CLI as modal or python3 -m modal. It checks PATH first rather than running a command that can only fail, and stays silent when Modal isn't set up.
  • Makes no model calls: only the Modal CLI, every 60s.

pr-autopilot

Does the "merged #219, clean up branches and start #214" round trip for you, and surfaces CI failures with their logs.

  • Watches your open PRs in the session's repo: it adopts them at session start, and picks up every gh pr create Claude runs. It polls gh pr view every 60s.
  • Status line: PRs: #219 ✓ · #220 CI… · #221 ✗. Toasts when CI fails (with the failing check names) or passes.
  • When a PR merges, it cleans up with plain local git, then toasts the outcome and suggests carrying on (Tab to accept):
  • git fetch --prune, switch to the default branch (only from the PR's own branch) and git pull --ff-only.
  • Never with uncommitted changes, never --force, never other branches.
  • Deletes the local branch only if it points at exactly the commit GitHub merged, so nothing local is lost. It also leaves a branch checked out in another worktree alone.
  • A merge seen mid-turn is cleaned up when the turn ends, so git never races Claude.
  • When you mention failing CI ("#258 is failing", "CI failed, fix it"), your message goes to Claude with gh pr checks and the tail of the failed log attached, so you don't paste it.
  • /prs lists watched PRs. /prs watch <n|url> and /prs forget <n|all> add and remove them.
  • Makes no model calls: only gh and git, at about one GitHub API call per open PR per minute.

Settings:

  • pollSeconds (60)
  • attachCiLogs (on)
  • logLines (120)
  • deleteRemoteBranch (off): deletes the branch on GitHub too, only while it still points at the merged commit. GitHub's own "Automatically delete head branches" setting does the same job.

It never closes issues; put "Closes #N" in PR bodies for that.

recall

Search everything you've done with coding agents, from Claude or from /recall. It replaces the broken agent-memory plugin.

  • What it searches: Claude Code sessions, Codex sessions, subagent and workflow runs, Claude's memory files, standing orders, second-opinion reviews, and your /remember notes. Routine (scheduled) runs are left out unless you add routines:include to a query.
  • What it keeps: prompts, answers, compaction summaries, session titles, commands, files touched, commits, PRs, issues, URLs, tasks and decisions (what you approved or ruled out). Read-only look-ups like grep and cat are kept but ranked low.
  • History survives cleanup: extracts stay searchable after Claude Code deletes old transcripts.
  • Claude searches it itself with four read-only tools, search, expand, recap and list, which run without permission prompts. It checks them when you say "like last time" or "what did we decide", and before asking you something you already settled.
  • Commands:
  • /recall <query> opens a pane of hits grouped by session. Open shows the conversation around a hit, Attach sends it with your next message, and Copy resume command copies claude --resume <id>.
  • /recall last [n] recaps your last session in this repo: last asks, last answer, PRs, commits, open tasks and decisions. Send to Claude attaches it.
  • /recall timeline [7d|30d|90d] [all]
  • /recall decisions|commands|files|prs|commits|issues|urls|tasks|notes [query]
  • /recall ask <question> answers from your history with Haiku 4.5, citing sessions. It costs a little usage and sends the matching excerpts to the model.
  • /recall stats, /recall reindex, /recall forget session <id>|project <name>|before <date> (asks you to confirm), /recall help.
  • /remember <fact>, /remember list, /remember forget <ref>.
  • Bands:
  • Once per session: Last session here (2d ago): "…" · PR #99 · 3 open tasks [Recap].
  • When a prompt mentions #214, ABC-12, a file name or a quoted phrase seen in past sessions, a band offers what happened then. Nothing is sent unless you click.
  • Query syntax: words must all match. OR gives alternatives, "quotes" an exact phrase, and -word excludes. Filters: project:name, kind:decision, since:7d, until:2026-09-30, source:codex, routines:include.
  • Privacy:
  • The index lives at ~/.claude/recall/index.db, readable only by you, and never goes in a repo.
  • Secrets are masked before anything is stored: known token shapes, labelled values ("password: …"), the values of secret-named exports in ~/.zshrc, ~/.zprofile, ~/.bashrc and ~/.bash_profile, and any literal strings you list in ~/.claude/recall/redact.txt (one per line). Editing that list re-masks the existing index on the next update.
  • Cost: no model calls except /recall ask. The first index takes about 2 minutes in the background, with progress on the status line. After that it updates incrementally (about 1s) at session start and every 10 minutes.
  • Requires macOS's /usr/bin/python3 (Command Line Tools), whose SQLite has FTS5. Nothing else to install.

Settings: dbPath, python, sources, includeSubagents (on), includeRoutines (off), updateMinutes (10), relatedBand (on), lastSessionBand (on), maxResults (8), askModel (claude-haiku-4-5-20251001).

repo-brief

Catches Claude up on the repo when a session starts, so you don't have to ask "check the recent commits/PRs and issues".

  • Gathers in the background at session start:
  • branch, ahead/behind and uncommitted files
  • the last 8 commits
  • open PRs with CI ✓/✗/…
  • issues labelled owner, todo, P0 or blocked
  • stale branches (merged, or upstream gone)
  • A one-line band above the prompt, e.g. main ↑1 · 3 changed · PRs #123 ✗ #124 ✓ · 2 owner issues · 2 stale branches · last commit 2h ago. Hide dismisses it.
  • Claude gets the same summary once, in its first message, so the prompt cache stays warm. It refreshes after compaction.
  • /brief re-gathers now and prints the full summary.
  • Makes no model calls: only git and gh. The band refreshes after a turn at most every 2 minutes.

Settings: focus labels, refresh minutes, and whether to brief Claude.

routine-watch

Keeps scheduled routines (daily digests, newsletters) from silently stalling while you're away.

  • Knows a session is a routine from its scheduled-task prompt, and does nothing in your other sessions.
  • When a routine stops to wait for your OK on a permission prompt or an AskUserQuestion, you get a Mac notification and a toast, and the status line shows routine: daily-report · waiting on you 3m.
  • When a turn ends in an error, or the routine finishes, you get a notification: Routine daily-report finished after 23m · waited on you 2 times.
  • Phone push (optional): notifyCommand runs a command on the same events, e.g. curl -s -d {message} ntfy.sh/your-topic. {title} and {message} are filled in as single arguments, never through a shell.
  • allowWebReads (off by default): lets routines use WebFetch and WebSearch without asking. It only replaces a prompt; your deny rules still apply, and nothing else is ever auto-allowed.
  • /routine shows the routine's name, how long it has run, its waits, and the settings.
  • Makes no model calls.

second-opinion

A Fable review in the background, without switching your session's model. Each run is one Fable call against your usage.

  • /second-opinion: reviews recent work. On a feature branch that's the branch against the default branch; otherwise the last 12 commits, plus the diff and git status, capped at 60k characters.
  • Other forms:
  • /second-opinion commits 5
  • /second-opinion diff (uncommitted changes)
  • /second-opinion file docs/ADR-007.md
  • /second-opinion <question>: adds a question for Fable to answer first.
  • The command returns at once, and the status line shows second opinion: reviewing…. When the review is ready you get a toast, and a pane opens with it, ranked: wrong assumptions, bugs and risks, what's missing, what to do next.
  • Send to Claude attaches the review to your next prompt (once) and drafts "What do you agree with, and what would you act on?".
  • Reviews are saved in ~/.claude/second-opinions/<project>/. /second-opinion list lists them, and /second-opinion show [n] reopens one.

Settings: model (claude-fable-5-1), effort (high), maxContextChars (60000).

standing-orders

Keeps your "always / never / don't / from now on" instructions alive across compaction.

  • When you write an instruction like "never open bambu with full spectrum files", a band asks: Keep as a standing order? [Project] [This session] [No]. Nothing is saved without a click.
  • Project orders live in ~/.claude/standing-orders/<repo>.json and apply to every session in that repo. Session orders and your active /goal last for the session.
  • Claude gets them at the start of every conversation and again after each compaction or /clear, so the prompt cache isn't disturbed. A newly saved order also rides along once with your next message.
  • /orders lists them. /orders add [project|session] <text>, /orders forget <n>, /orders clear session|project, and /orders export (a Markdown block for CLAUDE.md).
  • Makes no model calls.

secret-guard

Stops keys and passwords from going into a prompt, and so into your transcripts, and turns them into env vars instead.

  • Catches known token shapes: AWS, GitHub, Anthropic, OpenAI, Slack, Google, Hugging Face, GitLab, npm, Stripe, private keys and bearer tokens.
  • Also catches labelled values ("password: …", "api key = …", "the wifi password is …") and the two-line "Access Key ID / Secret Access Key" paste.
  • Leaves alone $NAME references, placeholders, plain URLs, paths, git SHAs and ordinary prose about passwords.
  • On a hit, the prompt isn't sent and goes back in the box. A band shows the secret masked (…vxrm) with a suggested name such as OPENDATALAB_SECRET_ACCESS_KEY, which you can edit:
  • Save as env var appends export NAME='…' to ~/.zshrc (reusing an existing identical export) and replaces the secret in your prompt with $NAME.
  • Send anyway lets exactly that text through once.
  • Edit dismisses the band.
  • The value is never shown in toasts, status, state or the transcript, and /secrets test <text> output is masked too.
  • /secrets test <text> shows what would be caught. /secrets off and /secrets on toggle it for the session.
  • Makes no model calls.

Settings: enabled (on), extraPatterns (a regex), zshrcPath (~/.zshrc).

slicer-handoff

Makes Claude's open commands hand 3D files to the right slicer.

  • Full-spectrum files go to Snapmaker Orca. Bambu Studio and OrcaSlicer can't open them. A file counts as full-spectrum when:
  • its name or folder matches full.?spectrum|snapmaker-only|-fs\.3mf$|-u1[-.], or
  • its 3MF names a Full Spectrum filament profile.

An open -a BambuStudio … for one becomes open -b com.snapmaker.snapmaker-orca …, with the rest of the command untouched. You get a toast, and Claude gets a note so it doesn't try again.

  • Earlier windows close first. Before opening a file, it asks the running slicer to quit (a normal quit, never forced), so windows don't pile up. If one won't close, for example because it's waiting on a save prompt, it stops trying and tells Claude to leave it alone.
  • /slice <file> [bambu|snapmaker|orca] opens a file yourself, with the same rules.
  • Recognizes open -a <app>, open -a /Applications/X.app and open -b <bundle id>, including variables set earlier in the command (S=… && open -a BambuStudio "$S/x.3mf") and files copied in the same command.
  • Makes no model calls.

Settings:

  • closePrevious (on): turn it off if you keep your own slicer window open, since the quit request reaches your windows too.
  • fullSpectrumPattern (the regex above)
  • checkContents (on)

The quit request goes out when Claude issues the command, before any permission prompt for it.

Loading

  • One session from a terminal: pass --plugin-dir once per mod, e.g. claude --plugin-dir ~/Projects/claude-mods/job-watch --plugin-dir ~/Projects/claude-mods/pr-autopilot
  • Every session, including the desktop app: add to ~/.claude/settings.json. Put mod-monitor first so it sees the others; CLAUDE_CODE_PLUGIN_DIR_WATCH makes desktop sessions pick up edits and show mod failures: ``json { "env": { "CLAUDE_CODE_PLUGIN_DIR_WATCH": "1", "CLAUDE_CODE_PLUGIN_DIRS": "~/Projects/claude-mods/mod-monitor:~/Projects/claude-mods/job-watch:~/Projects/claude-mods/machine-guard:~/Projects/claude-mods/repo-brief:~/Projects/claude-mods/slicer-handoff:~/Projects/claude-mods/pr-autopilot:~/Projects/claude-mods/routine-watch:~/Projects/claude-mods/modal-meter:~/Projects/claude-mods/second-opinion:~/Projects/claude-mods/downloads-drop:~/Projects/claude-mods/dev-servers:~/Projects/claude-mods/standing-orders:~/Projects/claude-mods/effort-router:~/Projects/claude-mods/secret-guard:~/Projects/claude-mods/recall" } } ``

Checking

Run with Claude Code 2.1.289 or newer; older CLIs ignore per-test settings, so a few tests fall back to defaults.

claude plugin validate job-watch && claude plugin test job-watch
claude plugin validate machine-guard && claude plugin test machine-guard
claude plugin validate repo-brief && claude plugin test repo-brief
claude plugin validate slicer-handoff && claude plugin test slicer-handoff
claude plugin validate pr-autopilot && claude plugin test pr-autopilot
claude plugin validate routine-watch && claude plugin test routine-watch
claude plugin validate modal-meter && claude plugin test modal-meter
claude plugin validate second-opinion && claude plugin test second-opinion
claude plugin validate downloads-drop && claude plugin test downloads-drop
claude plugin validate dev-servers && claude plugin test dev-servers
claude plugin validate standing-orders && claude plugin test standing-orders
claude plugin validate effort-router && claude plugin test effort-router
claude plugin validate secret-guard && claude plugin test secret-guard
claude plugin validate recall && claude plugin test recall
claude plugin validate mod-monitor && claude plugin test mod-monitor
(cd recall/engine && /usr/bin/python3 -m unittest)
Source 3 files
hooks/register.ts 295 lines
1/**
2 * effort-router: sends git housekeeping replies ("merged", "#219 merged", "commit and push", "push it")
3 * at low effort and deep asks ("review", "audit", "why does ...") at max, by rewriting the effort of
4 * each model request (`turn.step`). Approvals and replies that start new work ("yes", "continue",
5 * "merged, go ahead with #214") keep the session's own effort. The prompt itself passes through
6 * unchanged; no model call is made.
7 *
8 * The prompt cache. Claude Code sends effort as the request's top-level `output_config.effort`
9 * (this build has no per-message effort beta), and the API renders effort into the prompt: changing
10 * it between requests invalidates the cached messages (on some models the system prompt and tools
11 * too), so the next request writes the whole conversation to the cache again at the write price
12 * instead of reading it at a tenth of the input price. On a large conversation one switch can cost
13 * more than low effort saves on a short reply, and switching down and back up costs it twice. So:
14 *
15 * - The effort is chosen once per turn, at its first request, and every later request of the turn
16 *   (each tool round) carries the same effort: never a switch mid-turn.
17 * - A switch is made at once when it is cheap or asked for: the context is small (`freeSwitchTokens`),
18 *   the cache has lapsed (idle past `cacheTtlMinutes`; Claude Code caches the main thread for an hour
19 *   on a subscription, five minutes with an API key), the model or the session's own /effort changed,
20 *   compaction or /clear rewrote the history, `/route deep|routine` forced it, or the turn wants MORE
21 *   effort (quality never waits for the cache).
22 * - Less effort over a warm, large cache waits until `stickyTurns` turns in a row have wanted it
23 *   (hysteresis): one "push it" between two real tasks is not worth two cache rewrites.
24 *
25 * The model guard (`avoidModel` → `fallbackModel`) rewrites every request the same way, main loop and
26 * subagents alike, so the cache is rebuilt once for the new model and then reused.
27 */
28import { atom, read, update } from 'claude-code'
29import type { EngineInterface, PromptOrigin, Register, TurnStepInput, TurnStepResult } from 'claude-code'
30
31import {
32  EMPTY_STATE,
33  cleared,
34  compilePatterns,
35  describeRoute,
36  guardModel,
37  isLevel,
38  planStep,
39  resolveModel,
40  responded,
41  started,
42  submitted,
43} from './route'
44import type { Level, Patterns, Planned, Settings } from './route'
45
46type Engine = EngineInterface
47
48const router = atom({ plugin: 'effort-router', key: 'router' } as const, EMPTY_STATE)
49
50/** Prompts the person wrote: typed, through Remote Control, a host's own turn, a scheduled routine, a Slack ping. */
51const PERSONAL = new Set<string>(['composer', 'bridge', 'sdk', 'scheduled-trigger', 'slack-ping'])
52
53const isPersonal = (origin: PromptOrigin): boolean =>
54  PERSONAL.has(origin.kind) || (origin.kind === 'plugin' && origin.asUser === true)
55
56const USAGE = 'Usage: /route (status) | /route on | /route off | /route deep | /route routine (forces the next prompt)'
57
58type Config = {
59  enabled: boolean
60  settings: Settings
61  patterns: Patterns
62  errors: string[]
63  avoid: RegExp | null
64  fallbackModel: string
65}
66
67const DEFAULT_SETTINGS: Settings = {
68  routineEffort: 'low',
69  deepEffort: 'max',
70  guard: { stickyTurns: 2, freeSwitchTokens: 30_000, cacheTtlMs: 60 * 60_000 },
71}
72
73let config: Config = {
74  enabled: true,
75  settings: DEFAULT_SETTINGS,
76  patterns: compilePatterns('', '').patterns,
77  errors: [],
78  avoid: null,
79  fallbackModel: 'opus',
80}
81
82const numberIn = (value: unknown, fallback: number, min: number, max: number): number => {
83  const n = typeof value === 'number' ? value : Number(value)
84  return Number.isFinite(n) ? Math.min(max, Math.max(min, n)) : fallback
85}
86
87const textOf = (value: unknown): string => (typeof value === 'string' ? value : '')
88
89const modelGuardLine = (): string | null =>
90  config.avoid === null ? null : `models matching /${config.avoid.source}/i run on ${resolveModel(config.fallbackModel, '')}`
91
92/** The model a request goes to; the first rewrite of each kind is announced once. */
93async function guardedModel($: Engine, model: string): Promise<string> {
94  if (config.avoid === null || !config.avoid.test(model)) {
95    return model
96  }
97  const target = guardModel(model, config.avoid, config.fallbackModel)
98  const key = `${model}→${target}`
99  let isFirst = false
100  await update($, router, state => {
101    isFirst = !state.rewrites.includes(key)
102    return isFirst ? { ...state, rewrites: [...state.rewrites, key] } : state
103  })
104  if (isFirst) {
105    $.ui.toast(
106      target === model
107        ? `effort-router: ${model} matches avoidModel, but so does the fallback "${config.fallbackModel}"; requests stay on ${model}.`
108        : `effort-router: requests for ${model} go to ${target} (avoidModel).`,
109    )
110  }
111  return target
112}
113
114/** The context's size when the mod first sees a request (it may have loaded mid-session); 0 if unknown. */
115async function contextNow($: Engine): Promise<number> {
116  try {
117    const usage = await $.session.usage()
118    return usage.context.tokens ?? 0
119  } catch {
120    return 0
121  }
122}
123
124/** The effort this main-loop request is sent at: decided at the turn's first request, then kept. */
125async function chooseEffort($: Engine, e: TurnStepInput, model: string, base: Level): Promise<Level> {
126  const before = await read($, router)
127  if (before.turn !== null && before.turn.id === e.turnId && before.turn.effort !== null) {
128    return before.turn.effort
129  }
130  const now = await $.clock.now()
131  const contextTokens = before.cache === null ? await contextNow($) : 0
132  const plan: { planned: Planned | null } = { planned: null }
133  await update($, router, state => {
134    plan.planned = planStep(state, { turnId: e.turnId, model, base, messageCount: e.messageCount, now, contextTokens }, config.settings)
135    return plan.planned.state
136  })
137  if (plan.planned === null) {
138    return base
139  }
140  if (plan.planned.isNew) {
141    $.ui.status(plan.planned.status)
142  }
143  return plan.planned.effort
144}
145
146/** Remembers how large the cached conversation is now, and when it was last written. */
147async function noteResponse($: Engine, messageCount: number, model: string, result: TurnStepResult) {
148  try {
149    const now = await $.clock.now()
150    await update($, router, state => responded(state, { model, messageCount, now, usage: result.usage }))
151  } catch {
152    // The next turn decides from what was known before; nothing to undo.
153  }
154}
155
156/** A failure in effort-router's own code, logged (to the debug log) for mod-monitor to count. */
157function reportFailure($: Engine, what: string, error: unknown) {
158  const message = error instanceof Error ? error.message : String(error)
159  $.ui.log(`effort-router: ${what} failed, so the request went out unchanged: ${message.slice(0, 200)}`, { to: 'debug' })
160}
161
162export const register: Register = (on, options) => {
163  const compiled = compilePatterns(textOf(options.routinePattern), textOf(options.deepPattern))
164  const errors = [...compiled.errors]
165  let avoid: RegExp | null = null
166  const avoidSource = textOf(options.avoidModel).trim()
167  if (avoidSource !== '') {
168    try {
169      avoid = new RegExp(avoidSource, 'i')
170    } catch (error) {
171      errors.push(`avoidModel is not a valid regular expression (${error instanceof Error ? error.message : String(error)}); the model guard is off.`)
172    }
173  }
174  config = {
175    enabled: options.enabled !== false,
176    settings: {
177      routineEffort: isLevel(options.routineEffort) ? options.routineEffort : 'low',
178      deepEffort: isLevel(options.deepEffort) ? options.deepEffort : 'max',
179      guard: {
180        stickyTurns: Math.round(numberIn(options.stickyTurns, 2, 1, 20)),
181        freeSwitchTokens: numberIn(options.freeSwitchTokens, 30_000, 0, 10_000_000),
182        cacheTtlMs: numberIn(options.cacheTtlMinutes, 60, 0, 24 * 60) * 60_000,
183      },
184    },
185    patterns: compiled.patterns,
186    errors,
187    avoid,
188    fallbackModel: textOf(options.fallbackModel).trim() || 'opus',
189  }
190
191  on('session.start', async ($, e, next) => {
192    await $.command.register({
193      name: 'route',
194      description: 'Effort routing: status, on/off for this session, or force the next prompt deep or routine',
195      argumentHint: '[on | off | deep | routine]',
196      immediate: true,
197    })
198    return next(e)
199  })
200
201  on('command.run', { command: 'route' }, async ($, e) => {
202    if (!config.enabled) {
203      return { text: 'effort-router is turned off in its settings (enabled: false); requests are left as they are.' }
204    }
205    const verb = e.args.trim().toLowerCase()
206    if (verb === 'off') {
207      await update($, router, state => ({ ...state, isOn: false }))
208      return { text: "Effort routing is off for this session: turns run at the session's own effort. /route on turns it back on." }
209    }
210    if (verb === 'on') {
211      await update($, router, state => ({ ...state, isOn: true }))
212      return { text: 'Effort routing is on.' }
213    }
214    if (verb === 'deep' || verb === 'routine') {
215      const forced: 'deep' | 'routine' = verb
216      await update($, router, state => ({ ...state, forced }))
217      const effort = forced === 'deep' ? config.settings.deepEffort : config.settings.routineEffort
218      return { text: `Your next prompt runs at ${effort} (${forced}), whatever it says.` }
219    }
220    if (verb !== '' && verb !== 'status') {
221      return { text: USAGE }
222    }
223    const state = await read($, router)
224    return { text: describeRoute(state, { settings: config.settings, modelGuard: modelGuardLine(), errors: config.errors }) }
225  })
226
227  if (!config.enabled) {
228    return
229  }
230
231  // Classifies the prompt for the turn it starts; the prompt itself goes on unchanged.
232  on('prompt.submit', async ($, e, next) => {
233    try {
234      const isMine = isPersonal(e.origin)
235      await update($, router, state => submitted(state, e.text, isMine, config.patterns))
236    } catch {
237      // Unclassified: the turn runs at the session's own effort.
238    }
239    return next(e)
240  })
241
242  on('turn.start', async ($, e, next) => {
243    await update($, router, state => started(state, e.turnId, e.text, config.patterns))
244    return next(e)
245  })
246
247  on('turn.step', async function* ($, e, next) {
248    // Routing must never cost a request: if choosing fails, the request goes out exactly as it came, and
249    // the failure is logged where mod-monitor reads it (this stream is not a hook it can watch).
250    let routed = e
251    let model = e.model
252    let isRouted = false
253    try {
254      model = await guardedModel($, e.model)
255      if (e.agentId !== undefined || !isLevel(e.effort)) {
256        // A subagent keeps its own effort; a model without effort levels has none to route.
257        routed = model === e.model ? e : { ...e, model }
258      } else {
259        const effort = await chooseEffort($, e, model, e.effort)
260        routed = effort === e.effort && model === e.model ? e : { ...e, model, effort }
261        isRouted = true
262      }
263    } catch (error) {
264      routed = e
265      model = e.model
266      isRouted = false
267      reportFailure($, 'choosing the effort', error)
268    }
269    const result = yield* next(routed)
270    if (isRouted) {
271      try {
272        await noteResponse($, e.messageCount, model, result)
273      } catch (error) {
274        reportFailure($, 'noting the response', error)
275      }
276    }
277    return result
278  })
279
280  on('turn.complete', async ($, e, next) => {
281    const done = await next(e)
282    if (e.agentId === undefined) {
283      $.ui.status(undefined)
284    }
285    return done
286  })
287
288  on('session.end', async ($, e, next) => {
289    if (e.reason === 'clear') {
290      await update($, router, cleared)
291    }
292    return next(e)
293  })
294}
295
hooks/route.ts 435 lines
1import type { ModelEffort, ModelUsage } from 'claude-code'
2
3import type { RouteCache, RouteClass, RouteCounts, RoutePending, RouteState, RouteTurn } from '../types'
4
5export type Level = ModelEffort
6
7export const LEVELS: readonly Level[] = ['low', 'medium', 'high', 'xhigh', 'max']
8
9export const isLevel = (value: unknown): value is Level =>
10  typeof value === 'string' && (LEVELS as readonly string[]).includes(value)
11
12export const rank = (level: Level): number => LEVELS.indexOf(level)
13
14/** A prompt longer than this is never routine, whatever it starts with. */
15export const ROUTINE_MAX_CHARS = 120
16/** How many words may follow the routine phrases ("both are merged now" leaves one), none of them new work. */
17export const ROUTINE_MAX_EXTRA_WORDS = 6
18
19/**
20 * Routine replies are git housekeeping, matched at the start of the prompt: "merged", "#123 merged",
21 * "both are merged", "commit and push", "commit it", "push it", "open a PR", "close it". Approvals and
22 * resumptions ("yes", "go ahead", "continue", "keep going") are not: they start real work.
23 */
24export const DEFAULT_ROUTINE = String.raw`merged|(?:pr\s*)?#\d+\s+(?:is\s+|was\s+|got\s+)?merged|(?:all\s+|both\s+|everything\s+)?(?:is\s+|are\s+)?merged|commit(?:\s+and)?\s+push|commit(?:\s+it)?|push(?:\s+it)?|open\s+(?:a\s+)?pr|close\s+(?:it|the\s+issue)`
25
26/** Words that ask for deep thought, anywhere in the prompt (word starts: "audit" covers "auditing"). */
27export const DEFAULT_DEEP = String.raw`\b(?:audit|review|plan(?:s|ned|ning)?\b|architect|design|deep[\s-]*dive|assess|investigat|root[\s-]*cause|research|why\s+(?:is|does|did)\b|figure\s+out|what(?:['’]s|\s+is)\s+wrong)`
28
29export type Patterns = {
30  /** Anchored at the start; one routine phrase. */
31  routine: RegExp
32  /** Unanchored; any deep wording. */
33  deep: RegExp
34}
35
36export type PatternSet = { patterns: Patterns; errors: string[] }
37
38const leading = (source: string): RegExp => new RegExp(String.raw`^(?:${source})(?![A-Za-z0-9_])`, 'i')
39const anywhere = (source: string): RegExp => new RegExp(source, 'i')
40
41const compileOne = (
42  source: string,
43  fallback: string,
44  make: (source: string) => RegExp,
45  name: string,
46  errors: string[],
47): RegExp => {
48  const trimmed = source.trim()
49  if (trimmed !== '') {
50    try {
51      return make(trimmed)
52    } catch (error) {
53      errors.push(`${name} is not a valid regular expression (${error instanceof Error ? error.message : String(error)}); the built-in one is used.`)
54    }
55  }
56  return make(fallback)
57}
58
59/** The routine and deep patterns from settings; an empty or broken one falls back to the built-in. */
60export function compilePatterns(routineSource: string, deepSource: string): PatternSet {
61  const errors: string[] = []
62  return {
63    patterns: {
64      routine: compileOne(routineSource, DEFAULT_ROUTINE, leading, 'routinePattern', errors),
65      deep: compileOne(deepSource, DEFAULT_DEEP, anywhere, 'deepPattern', errors),
66    },
67    errors,
68  }
69}
70
71/** Punctuation and quoting between and around routine phrases. */
72const FILLER = /^[\s,.;:!?…—–\-"'`*_()[\]>]+/
73
74const wordCount = (text: string): number => text.split(/\s+/).filter(word => /[A-Za-z0-9]/.test(word)).length
75
76/** The routine phrases a prompt opens with ("merged, commit and push") and what follows them. */
77export function leadingRoutine(text: string, routine: RegExp): { phrases: string[]; rest: string } | null {
78  let rest = text.replace(FILLER, '')
79  const phrases: string[] = []
80  for (let i = 0; i < 8; i++) {
81    const found = routine.exec(rest)
82    if (!found || found[0] === '') {
83      break
84    }
85    phrases.push(found[0])
86    rest = rest.slice(found[0].length).replace(FILLER, '')
87  }
88  return phrases.length > 0 ? { phrases, rest } : null
89}
90
91/**
92 * Words after the routine phrases that start new work ("merged, go ahead with #214", "push it and fix
93 * the lint"): such a prompt is not routine, whatever it opens with.
94 */
95export const NEW_WORK =
96  /\b(?:start|begin|work\s+on|take|tackle|fix|implement|build|add|write|create|make|proceed|move\s+on|run|deploy|refactor|update|change|rename|remove|delete|investigate|look|check|debug|test|handle|address|file|install|download|train|generate|render|print|go\s+(?:ahead\s+)?(?:with|on)|continue|keep\s+going|do\b|try)\b|\b(?:with|on|to)\s+#?[A-Za-z]*-?\d/i
97
98export type Classified = { cls: RouteClass; why: string }
99
100/**
101 * Routine: at most 120 characters, opening with routine phrases and little else. Deep: deep wording
102 * anywhere. Both at once ("commit and push, then review the diff") is no opinion: the session's own effort.
103 */
104export function classify(text: string, patterns: Patterns): Classified {
105  const trimmed = text.trim()
106  const lead = leadingRoutine(trimmed, patterns.routine)
107  const isRoutine =
108    lead !== null &&
109    trimmed.length <= ROUTINE_MAX_CHARS &&
110    wordCount(lead.rest) <= ROUTINE_MAX_EXTRA_WORDS &&
111    !NEW_WORK.test(lead.rest)
112  const deep = patterns.deep.exec(trimmed)
113  if (isRoutine && deep) {
114    return { cls: 'neutral', why: `both routine ("${lead.phrases.join(', ')}") and deep ("${deep[0]}")` }
115  }
116  if (isRoutine) {
117    return { cls: 'routine', why: `starts with "${lead.phrases.join(', ')}"` }
118  }
119  if (deep) {
120    return { cls: 'deep', why: `mentions "${deep[0]}"` }
121  }
122  if (lead !== null) {
123    return { cls: 'neutral', why: trimmed.length > ROUTINE_MAX_CHARS ? 'too long to be routine' : 'more than a routine reply' }
124  }
125  return { cls: 'neutral', why: 'no routine or deep wording' }
126}
127
128export type Guard = {
129  /** Turns in a row a lower effort must be wanted before it is sent over a warm, large cache. */
130  stickyTurns: number
131  /** Below this many context tokens a switch costs little: it is made at once. */
132  freeSwitchTokens: number
133  /** Idle longer than this and the cache has lapsed: a switch costs nothing extra. */
134  cacheTtlMs: number
135}
136
137export type Why = 'same' | 'forced' | 'off' | 'model' | 'base' | 'history' | 'cold' | 'small' | 'up' | 'streak' | 'hold'
138
139export type Ask = {
140  want: Level
141  /** The model this request names (after the model guard). */
142  model: string
143  /** The session's own effort, as the engine offers it on this request. */
144  base: Level
145  messageCount: number
146  now: number
147  isForced: boolean
148}
149
150export type Decision = { effort: Level; why: Why; lowerStreak: number }
151
152/**
153 * The effort a new turn is sent at, given what the prompt cache was written with.
154 *
155 * Claude Code sends effort as the request's top-level `output_config.effort`, and the API renders it
156 * into the prompt: a change invalidates the cached messages (on some models the system prompt and
157 * tools too), so the next request re-writes the whole conversation at the cache-write price. So a
158 * switch is made at once only when it is cheap or asked for: the person forced it, the model or the
159 * session's own effort changed (the cache is cold for it anyway), the history shrank (compaction,
160 * /clear), the cache has lapsed (idle past its TTL), the context is small, or the turn wants MORE
161 * effort (quality never waits). A turn wanting LESS effort over a warm, large cache is held at the
162 * current effort until `stickyTurns` turns in a row have wanted less.
163 */
164export function decide(cache: RouteCache, ask: Ask, guard: Guard): Decision {
165  const switchTo = (why: Why): Decision => ({ effort: ask.want, why, lowerStreak: 0 })
166  if (ask.want === cache.applied) {
167    return switchTo('same')
168  }
169  if (ask.isForced) {
170    return switchTo('forced')
171  }
172  if (ask.model !== cache.model) {
173    return switchTo('model')
174  }
175  if (ask.base !== cache.base) {
176    return switchTo('base')
177  }
178  if (ask.messageCount < cache.messageCount) {
179    return switchTo('history')
180  }
181  if (cache.lastAt !== null && ask.now - cache.lastAt > guard.cacheTtlMs) {
182    return switchTo('cold')
183  }
184  if (cache.contextTokens < guard.freeSwitchTokens) {
185    return switchTo('small')
186  }
187  if (rank(ask.want) > rank(cache.applied)) {
188    return switchTo('up')
189  }
190  const lowerStreak = cache.lowerStreak + 1
191  if (lowerStreak >= Math.max(1, guard.stickyTurns)) {
192    return switchTo('streak')
193  }
194  return { effort: cache.applied, why: 'hold', lowerStreak }
195}
196
197export const EMPTY_COUNTS: RouteCounts = { routine: 0, deep: 0, neutral: 0, switches: 0, holds: 0 }
198
199export const EMPTY_STATE: RouteState = {
200  isOn: true,
201  forced: null,
202  pending: null,
203  turn: null,
204  cache: null,
205  counts: EMPTY_COUNTS,
206  rewrites: [],
207}
208
209/** A prompt was submitted: classify it (or take the class /route forced) for the turn it starts. */
210export function submitted(state: RouteState, text: string, isPersonal: boolean, patterns: Patterns): RouteState {
211  if (!isPersonal) {
212    return { ...state, pending: { cls: 'neutral', why: 'not typed by you', isForced: false, text } }
213  }
214  if (state.forced !== null) {
215    return { ...state, forced: null, pending: { cls: state.forced, why: 'forced with /route', isForced: true, text } }
216  }
217  return { ...state, pending: { ...classify(text, patterns), isForced: false, text } }
218}
219
220/** A main-loop turn started: it takes the pending prompt's class. */
221export function started(state: RouteState, turnId: string, text: string, patterns: Patterns): RouteState {
222  const pending: RoutePending =
223    state.pending ??
224    (text.trim() === ''
225      ? { cls: 'neutral', why: 'no prompt', isForced: false, text: '' }
226      : { ...classify(text, patterns), isForced: false, text })
227  const turn: RouteTurn = { ...pending, id: turnId, effort: null, want: null, decision: null, streak: 0 }
228  return { ...state, pending: null, turn }
229}
230
231export type Settings = { routineEffort: Level; deepEffort: Level; guard: Guard }
232
233/** What a main-loop request carries, as the engine is about to send it. */
234export type StepFacts = {
235  turnId: string
236  /** The model it names, after the model guard. */
237  model: string
238  /** The session's own effort. */
239  base: Level
240  messageCount: number
241  now: number
242  /** The context's size, used only when nothing was seen before (the mod loaded mid-session). */
243  contextTokens: number
244}
245
246export type Planned = { state: RouteState; effort: Level; isNew: boolean; status: string | undefined }
247
248/** What the status line says while a turn runs; undefined when the turn runs as the session would. */
249export function statusText(cls: RouteClass, decision: Decision, isForced: boolean, stickyTurns: number): string | undefined {
250  if (decision.why === 'hold') {
251    return `effort: ${decision.effort} (${cls === 'neutral' ? '' : `${cls}, `}held for the prompt cache ${decision.lowerStreak}/${stickyTurns})`
252  }
253  if (cls === 'neutral' || decision.why === 'off') {
254    return undefined
255  }
256  return `effort: ${decision.effort} (${cls}${isForced ? ', forced' : ''})`
257}
258
259/**
260 * The effort a main-loop request is sent at. The turn's first request decides; every later
261 * request of the same turn carries the same effort, so a turn never switches midway.
262 */
263export function planStep(state: RouteState, facts: StepFacts, settings: Settings): Planned {
264  const current = state.turn
265  if (current !== null && current.id === facts.turnId && current.effort !== null) {
266    return { state, effort: current.effort, isNew: false, status: undefined }
267  }
268  const turn: RouteTurn =
269    current !== null && current.id === facts.turnId
270      ? current
271      : { id: facts.turnId, cls: 'neutral', why: 'its prompt was not seen', isForced: false, text: '', effort: null, want: null, decision: null, streak: 0 }
272  const cache: RouteCache = state.cache ?? {
273    applied: facts.base,
274    model: facts.model,
275    base: facts.base,
276    messageCount: facts.messageCount,
277    contextTokens: facts.contextTokens,
278    lastAt: null,
279    lowerStreak: 0,
280  }
281  const isRouted = state.isOn || turn.isForced
282  const want = !isRouted || turn.cls === 'neutral' ? facts.base : turn.cls === 'routine' ? settings.routineEffort : settings.deepEffort
283  const decision: Decision = isRouted
284    ? decide(cache, { want, model: facts.model, base: facts.base, messageCount: facts.messageCount, now: facts.now, isForced: turn.isForced }, settings.guard)
285    : { effort: want, why: 'off', lowerStreak: 0 }
286  const counts: RouteCounts = {
287    ...state.counts,
288    [turn.cls]: state.counts[turn.cls] + 1,
289    switches: state.counts.switches + (decision.effort === cache.applied ? 0 : 1),
290    holds: state.counts.holds + (decision.why === 'hold' ? 1 : 0),
291  }
292  return {
293    state: {
294      ...state,
295      turn: { ...turn, effort: decision.effort, want, decision: decision.why, streak: decision.lowerStreak },
296      cache: {
297        ...cache,
298        applied: decision.effort,
299        model: facts.model,
300        base: facts.base,
301        messageCount: facts.messageCount,
302        lowerStreak: decision.lowerStreak,
303      },
304      counts,
305    },
306    effort: decision.effort,
307    isNew: true,
308    status: statusText(turn.cls, decision, turn.isForced, Math.max(1, settings.guard.stickyTurns)),
309  }
310}
311
312/** Prompt tokens a response was answered over plus its output: what the next request re-sends. */
313export const contextOf = (usage: ModelUsage): number =>
314  usage.input_tokens + usage.cache_read_input_tokens + usage.cache_creation_input_tokens + usage.output_tokens
315
316/** A main-loop response arrived: remember how big the cached prefix now is, and when it was written. */
317export function responded(
318  state: RouteState,
319  facts: { model: string; messageCount: number; now: number; usage: ModelUsage | null },
320): RouteState {
321  if (state.cache === null) {
322    return state
323  }
324  const cache: RouteCache =
325    facts.usage === null
326      ? { ...state.cache, model: facts.model, messageCount: facts.messageCount }
327      : { ...state.cache, model: facts.model, messageCount: facts.messageCount, contextTokens: contextOf(facts.usage), lastAt: facts.now }
328  return { ...state, cache }
329}
330
331/** /clear: a new conversation whose first request writes a fresh cache, so the next switch is free. */
332export function cleared(state: RouteState): RouteState {
333  return {
334    ...state,
335    pending: null,
336    turn: null,
337    counts: EMPTY_COUNTS,
338    cache: state.cache === null ? null : { ...state.cache, messageCount: 0, contextTokens: 0, lastAt: null, lowerStreak: 0 },
339  }
340}
341
342/** Current model ids for the family aliases, so a rewrite never depends on the engine resolving an alias. */
343export const MODEL_IDS: Readonly<Record<string, string>> = {
344  opus: 'claude-opus-5-5',
345  sonnet: 'claude-sonnet-5-5',
346  haiku: 'claude-haiku-5-5',
347  fable: 'claude-fable-5-1',
348}
349
350/**
351 * The model id a fallback names: a family alias becomes its current id, spelled with the provider
352 * prefix the engine's own id carries (`us.anthropic.`); anything else is taken as written.
353 */
354export function resolveModel(fallback: string, current: string): string {
355  const name = fallback.trim()
356  const id = MODEL_IDS[name.toLowerCase()]
357  if (id === undefined) {
358    return name
359  }
360  const at = current.indexOf('claude-')
361  return (at > 0 ? current.slice(0, at) : '') + id
362}
363
364/** The model a request goes to: the fallback when the engine's matches `avoid` (and the fallback does not). */
365export function guardModel(model: string, avoid: RegExp | null, fallback: string): string {
366  if (avoid === null || !avoid.test(model)) {
367    return model
368  }
369  const target = resolveModel(fallback, model)
370  return target === '' || avoid.test(target) ? model : target
371}
372
373const clip = (text: string, max: number): string => {
374  const line = text.replace(/\s+/g, ' ').trim()
375  return line.length > max ? `${line.slice(0, max - 1)}…` : line
376}
377
378const WHY_WORDS: Readonly<Record<string, string>> = {
379  same: 'no change',
380  forced: 'forced',
381  off: 'routing off',
382  model: 'switched with the model',
383  base: 'your own /effort changed',
384  history: 'after compaction',
385  cold: 'cache had lapsed',
386  small: 'small context',
387  up: 'more effort never waits',
388  streak: 'lower effort wanted turns in a row',
389  hold: 'held for the prompt cache',
390}
391
392export type Describe = {
393  settings: Settings
394  /** e.g. `models matching /fable/i run on claude-opus-5-5`, or null when off. */
395  modelGuard: string | null
396  errors: readonly string[]
397}
398
399const formatTokens = (n: number): string => (n >= 1000 ? `${Math.round(n / 1000)}k` : String(n))
400
401/** What /route prints. */
402export function describeRoute(state: RouteState, info: Describe): string {
403  const { settings } = info
404  const effortFor = (cls: 'routine' | 'deep'): Level => (cls === 'routine' ? settings.routineEffort : settings.deepEffort)
405  const lines = [state.isOn ? 'effort-router: on' : 'effort-router: off for this session (/route on resumes; /route deep or routine still apply)']
406  if (state.forced !== null) {
407    lines.push(`Next prompt: forced ${state.forced} (${effortFor(state.forced)})`)
408  }
409  const turn = state.turn
410  if (turn === null) {
411    lines.push('No turn routed yet.')
412  } else {
413    const quoted = turn.text === '' ? '' : ` "${clip(turn.text, 60)}"`
414    const effort =
415      turn.effort === null
416        ? 'effort not chosen yet'
417        : turn.decision === 'hold'
418          ? `${turn.effort}, held for the prompt cache (wanted ${turn.want ?? '?'})`
419          : `${turn.effort} (${WHY_WORDS[turn.decision ?? 'same'] ?? turn.decision})`
420    lines.push(`Last turn: ${turn.cls}${quoted}, ${turn.why} → ${effort}`)
421  }
422  const c = state.counts
423  lines.push(
424    `This session: ${c.routine} routine, ${c.deep} deep, ${c.neutral} neutral turns · ${c.switches} effort ${c.switches === 1 ? 'switch' : 'switches'} · ${c.holds} held for the prompt cache`,
425    `Routine → ${settings.routineEffort}, deep → ${settings.deepEffort}, neutral → the session's own effort.`,
426    `Cache guard: a switch is made at once under ${formatTokens(settings.guard.freeSwitchTokens)} context tokens, after ${Math.round(settings.guard.cacheTtlMs / 60_000)} min idle, or toward more effort; less effort waits for ${Math.max(1, settings.guard.stickyTurns)} turns in a row.`,
427  )
428  if (state.cache !== null) {
429    lines.push(`Prompt cache: written at ${state.cache.applied} over about ${formatTokens(state.cache.contextTokens)} tokens.`)
430  }
431  lines.push(info.modelGuard === null ? 'Model guard: off' : `Model guard: ${info.modelGuard}`)
432  lines.push(...info.errors)
433  return lines.join('\n')
434}
435
types/index.d.ts 76 lines
1/** The effort levels a `turn.step` request can name, lowest first. */
2export type RouteLevel = 'low' | 'medium' | 'high' | 'xhigh' | 'max'
3
4/** What a prompt asks for: little thought, deep thought, or no opinion. */
5export type RouteClass = 'routine' | 'deep' | 'neutral'
6
7/** A prompt's class, waiting for the turn it starts. */
8export type RoutePending = {
9  cls: RouteClass
10  /** Why it got that class, for /route ("starts with \"merged\""). */
11  why: string
12  /** True when /route deep or /route routine chose the class. */
13  isForced: boolean
14  text: string
15}
16
17/** The main loop's current (or last) turn and the effort chosen for it. */
18export type RouteTurn = RoutePending & {
19  id: string
20  /** The effort every request of the turn carries; null until its first request. */
21  effort: RouteLevel | null
22  /** The effort the class asked for; differs from `effort` while the cache guard holds. */
23  want: RouteLevel | null
24  /** Why `effort` was chosen (`same`, `small`, `up`, `hold`, ...). */
25  decision: string | null
26  /** Turns in a row that wanted less effort than was being sent, this one included. */
27  streak: number
28}
29
30/** What the prompt cache was last written with on the main loop, so a switch is only made when it pays. */
31export type RouteCache = {
32  /** The effort the main loop's last request carried. */
33  applied: RouteLevel
34  /** The model the main loop's last request named. */
35  model: string
36  /** The session's own effort, as the engine last offered it. */
37  base: RouteLevel
38  /** How many messages the main loop's last request carried. */
39  messageCount: number
40  /** Prompt tokens the last response was answered over, plus its output: what the next request re-sends. */
41  contextTokens: number
42  /** When the last main-loop response arrived; null before one has. */
43  lastAt: number | null
44  /** Turns in a row that wanted less effort than `applied`. */
45  lowerStreak: number
46}
47
48export type RouteCounts = {
49  routine: number
50  deep: number
51  neutral: number
52  /** Turns whose effort differed from the turn before (each one rewrites the prompt cache). */
53  switches: number
54  /** Turns kept at the previous effort to spare the prompt cache. */
55  holds: number
56}
57
58export type RouteState = {
59  /** /route off turns automatic routing off for the session; /route deep and /route routine still apply. */
60  isOn: boolean
61  /** A class /route forced on the next prompt. */
62  forced: 'routine' | 'deep' | null
63  pending: RoutePending | null
64  turn: RouteTurn | null
65  cache: RouteCache | null
66  counts: RouteCounts
67  /** Model rewrites already announced with a toast (`from→to`). */
68  rewrites: string[]
69}
70
71declare module 'claude-code' {
72  interface PluginState {
73    'effort-router': { router: RouteState }
74  }
75}
76