Before a subagent starts on Opus, a Haiku judge reads its task and moves it to Sonnet unless it needs Opus's judgment.

Three mods for Claude Code. I asked Claude to audit my last 30 Claude Code sessions and tell me where my time went. Each mod answers one thing it found.
| Mod | What it does | What the audit found |
|---|---|---|
| Cache Band | Shows how warm the prompt cache is, keeps it warm while you're away, and compacts long conversations at 400K tokens | My cache went cold 91 times |
| Git Sync Band | Branch, ahead/behind and changed files above the prompt, with Pull, Files and Commit & push buttons | 1,660 git checks done by hand |
| Scout Router | Has a Haiku judge decide whether each subagent needs Opus, and runs the rest on Sonnet | 243 of 309 subagents ran on Opus |
They run on Claude Code's function hooks. Those are in early access, so the API can change between releases. I built and tested these on Claude Code 2.1.286 and 2.1.287.
~/.claude/skills/, for example ~/.claude/skills/cache-band./reload-plugins.To check a mod before you load it, or to run its tests:
claude plugin validate ~/.claude/skills/cache-band
claude plugin test ~/.claude/skills/cache-band
When Claude Code loads a mod, it writes the claude-code type declarations into the mod's .claude-plugin/types folder, which tsconfig.json points at. That folder is git-ignored here.
Every message re-reads the whole conversation. The prompt cache keeps that cheap: a cache read costs about a tenth of the normal input price. On my plan a cache entry lives for an hour after it was last read. Once it goes cold, the next message writes the whole conversation into the cache again, and a one-hour cache write costs twice the input price. On a 400K conversation, that adds up.
The band sits above the prompt: a flame that cools into a snowflake as the hour runs out, the minutes left, two toggles and a Compact button.
Auto cache. Ten minutes before the cache would go cold, it pings it. The ping goes through $.model.fork, which sends the conversation's last request again from a background copy, with a one-line keep-alive at the end. The copy is word for word the same conversation, so the API serves it from the same cache, and that read keeps the cache for another hour. The reply is dropped and nothing is added to your chat. A "hi" typed into the chat would keep the cache warm too, but it stays in the conversation for good, and Claude might act on it.
It stops after 20 pings in a row. A ping reads the cache (0.1x) and a cold restart writes it (2x), so 20 pings cost about what the restart they prevent would, around 16 hours in. Your next message starts it again. If a ping finds the cache already gone, Auto cache pauses instead of paying full price every hour.
Auto compact. It compacts once the conversation passes 400K tokens, right after Claude's reply, while the cache is still warm, so reading the conversation for the summary is cheap. Why 400K:
Compact. The button runs /compact as if you typed it. The desktop app refuses a mod's direct compaction call, so this is the route that works there.
28 tests cover the gauge, the pings and Auto compact.
A band above the prompt with the branch, how far ahead or behind it is, the changed files and when you last pushed. It reads git status every minute and fetches every five, with no credential prompts and without taking git's index lock, so it never gets in the way of Claude's own git.
Before your prompt, it pulls with --ff-only when that can't lose anything, and tells Claude what came in, or why nothing did. Pull, Files and Commit & push run git straight from the band, without a turn in the chat. Commit & push writes the message with one small model call and shows it to you before anything runs. It won't commit files that look like secrets (.env files, keys, credentials).
6 tests cover the band's layouts, the divider and the commit flow.
Before a subagent starts on Opus, Scout Router asks a judge: one small Haiku call that reads the subagent's task and the start of its instructions, and answers opus or sonnet. Opus is for deep judgment where a mistake is costly: planning, design decisions, audits and reviews, grading other work, and subtle debugging that spans many parts. Everything else runs on Sonnet: writing code, tests or data to a clear brief, searching and summarizing, running checks, mechanical edits. When the judge is unsure, it picks Sonnet. It decides even when the main model asked for Opus by name.
Forks keep their model, since they share the main conversation's cache. Explore agents go straight to Sonnet. Agents already headed for Sonnet or Haiku aren't judged. If the judge can't answer (an error, a time-out, a usage limit), a word rule decides: planning, audit and review tasks stay on Opus, and the rest go to Sonnet. Each decision is logged.
The judge costs a few hundred Haiku tokens per subagent and adds about a second before it starts. 5 tests cover the judge, its fallback and what it skips.
MIT
hooks/register.ts 85 lines1import type { AgentSpawnInput, EngineInterface, Register } from 'claude-code'
2
3type Verdict = 'opus' | 'sonnet'
4
5const words = (list: string) =>
6 new RegExp(`\\b(?:${list.trim().split(/\s+/).join('|')})\\b`, 'i')
7
8// When the judge can't answer: planning, audits and review stay on the main
9// model ("opus only plans and audits"); everything else goes to Sonnet.
10const KEEP = words(`
11 plan plans planning planner audit audits auditing review reviews reviewing
12 reviewer grade grades grading judge judging verify verifying verification
13 critique critic architect architecture design designing decide evaluate
14 evaluating evaluation score scoring rank ranking
15`)
16
17const JUDGE_SYSTEM = [
18 "You route a coding agent's subagent to a model. Reply with exactly one word: opus or sonnet.",
19 'opus: the task needs deep judgment where a mistake is costly: planning work or designing an ' +
20 'architecture, deciding between approaches, auditing or reviewing for correctness or security, ' +
21 'grading or evaluating other work, or debugging a subtle problem that spans many parts.',
22 'sonnet: everything else: writing code, tests, content or data to a clear brief; searching, ' +
23 'reading and summarizing; running commands and checks; mechanical edits, refactors and ' +
24 'migrations with clear instructions.',
25 'When unsure, sonnet.',
26].join('\n')
27
28// Asks Haiku whether the subagent needs Opus, from its task and the start of
29// its instructions. Undefined when it can't say (an error, a time-out, the
30// account's limit, a reply that is neither word).
31async function judge($: EngineInterface, e: AgentSpawnInput): Promise<Verdict | undefined> {
32 const reply = await $.model
33 .complete({
34 model: 'haiku',
35 system: JUDGE_SYSTEM,
36 prompt: `Subagent type: ${e.subagentType}\nTask: ${e.description}\n\nIts instructions, maybe cut:\n${e.prompt.slice(0, 4000)}`,
37 maxTokens: 5,
38 effort: 'low',
39 timeoutMs: 10_000,
40 })
41 .catch(() => undefined)
42 if (!reply?.isAnswered) return undefined
43 const word = reply.text.trim().toLowerCase()
44
45 return word.startsWith('opus') ? 'opus' : word.startsWith('sonnet') ? 'sonnet' : undefined
46}
47
48export const register: Register = on => {
49 // Clears the agent count an earlier version pinned to the status line.
50 on('session.start', ($, e, next) => {
51 $.ui.status(undefined)
52
53 return next(e)
54 })
55
56 on('agent.spawn', async ($, e, next) => {
57 const asked = e.model ?? ''
58 const runsOnOpus = /opus/i.test(asked) || (asked === '' && /opus/i.test(e.parentModel))
59 // A fork shares the main conversation's cache, so it keeps its model; a
60 // subagent not headed for Opus already costs less.
61 if (e.fork || !runsOnOpus) return next(e)
62
63 let verdict: Verdict
64 let why: string
65 if (e.subagentType === 'Explore') {
66 verdict = 'sonnet'
67 why = 'Explore only reads'
68 } else {
69 const ruled = await judge($, e)
70 verdict = ruled ?? (KEEP.test(e.description) ? 'opus' : 'sonnet')
71 why = ruled ? 'judge' : 'judge unavailable, by its words'
72 }
73
74 if (verdict === 'opus') {
75 $.ui.log(`scout-router: "${e.description}" stays on Opus (${why})`)
76
77 return next(e)
78 }
79 const ran = await next({ ...e, model: 'sonnet' })
80 if (ran.deny === undefined) $.ui.log(`scout-router: "${e.description}" runs on ${ran.model} (${why})`)
81
82 return ran
83 })
84}
85