SLOPSHOPPER

scout-router

Before a subagent starts on Opus, a Haiku judge reads its task and moves it to Sonnet unless it needs Opus's judgment.

newstatusmodelagents
A shopper browsing a rack in a slop shop
README

Claude Code mods

Three mods for Claude Code. I asked Claude to audit my last 30 Claude Code sessions and tell me where my time went. Each mod answers one thing it found.

ModWhat it doesWhat the audit found
Cache BandShows how warm the prompt cache is, keeps it warm while you're away, and compacts long conversations at 400K tokensMy cache went cold 91 times
Git Sync BandBranch, ahead/behind and changed files above the prompt, with Pull, Files and Commit & push buttons1,660 git checks done by hand
Scout RouterHas a Haiku judge decide whether each subagent needs Opus, and runs the rest on Sonnet243 of 309 subagents ran on Opus

They run on Claude Code's function hooks. Those are in early access, so the API can change between releases. I built and tested these on Claude Code 2.1.286 and 2.1.287.

Install

  1. Copy a mod's folder into ~/.claude/skills/, for example ~/.claude/skills/cache-band.
  2. New sessions load it. In a session that was already open, run /reload-plugins.

To check a mod before you load it, or to run its tests:

claude plugin validate ~/.claude/skills/cache-band
claude plugin test ~/.claude/skills/cache-band

When Claude Code loads a mod, it writes the claude-code type declarations into the mod's .claude-plugin/types folder, which tsconfig.json points at. That folder is git-ignored here.

Cache Band

Every message re-reads the whole conversation. The prompt cache keeps that cheap: a cache read costs about a tenth of the normal input price. On my plan a cache entry lives for an hour after it was last read. Once it goes cold, the next message writes the whole conversation into the cache again, and a one-hour cache write costs twice the input price. On a 400K conversation, that adds up.

The band sits above the prompt: a flame that cools into a snowflake as the hour runs out, the minutes left, two toggles and a Compact button.

Auto cache. Ten minutes before the cache would go cold, it pings it. The ping goes through $.model.fork, which sends the conversation's last request again from a background copy, with a one-line keep-alive at the end. The copy is word for word the same conversation, so the API serves it from the same cache, and that read keeps the cache for another hour. The reply is dropped and nothing is added to your chat. A "hi" typed into the chat would keep the cache warm too, but it stays in the conversation for good, and Claude might act on it.

It stops after 20 pings in a row. A ping reads the cache (0.1x) and a cold restart writes it (2x), so 20 pings cost about what the restart they prevent would, around 16 hours in. Your next message starts it again. If a ping finds the cache already gone, Auto cache pauses instead of paying full price every hour.

Auto compact. It compacts once the conversation passes 400K tokens, right after Claude's reply, while the cache is still warm, so reading the conversation for the summary is cheap. Why 400K:

  • Every message re-reads the whole conversation, so at 830K each message costs about twice what it does at 400K.
  • The longer the conversation, the more Claude misses (Chroma's context rot study).
  • Compacting too early summarizes a long task halfway through, and details get lost. Anthropic's API compaction defaults to 150K (docs). Claude Code's own auto-compact waits until about 83% of the window, around 830K on a 1M window.
  • 400K is twice the old 200K window, so a long task fits in one go, at under half the per-message cost of 830K.

Compact. The button runs /compact as if you typed it. The desktop app refuses a mod's direct compaction call, so this is the route that works there.

28 tests cover the gauge, the pings and Auto compact.

Git Sync Band

A band above the prompt with the branch, how far ahead or behind it is, the changed files and when you last pushed. It reads git status every minute and fetches every five, with no credential prompts and without taking git's index lock, so it never gets in the way of Claude's own git.

Before your prompt, it pulls with --ff-only when that can't lose anything, and tells Claude what came in, or why nothing did. Pull, Files and Commit & push run git straight from the band, without a turn in the chat. Commit & push writes the message with one small model call and shows it to you before anything runs. It won't commit files that look like secrets (.env files, keys, credentials).

6 tests cover the band's layouts, the divider and the commit flow.

Scout Router

Before a subagent starts on Opus, Scout Router asks a judge: one small Haiku call that reads the subagent's task and the start of its instructions, and answers opus or sonnet. Opus is for deep judgment where a mistake is costly: planning, design decisions, audits and reviews, grading other work, and subtle debugging that spans many parts. Everything else runs on Sonnet: writing code, tests or data to a clear brief, searching and summarizing, running checks, mechanical edits. When the judge is unsure, it picks Sonnet. It decides even when the main model asked for Opus by name.

Forks keep their model, since they share the main conversation's cache. Explore agents go straight to Sonnet. Agents already headed for Sonnet or Haiku aren't judged. If the judge can't answer (an error, a time-out, a usage limit), a word rule decides: planning, audit and review tasks stay on Opus, and the rest go to Sonnet. Each decision is logged.

The judge costs a few hundred Haiku tokens per subagent and adds about a second before it starts. 5 tests cover the judge, its fallback and what it skips.

License

MIT

Source 1 files
hooks/register.ts 85 lines
1import type { AgentSpawnInput, EngineInterface, Register } from 'claude-code'
2
3type Verdict = 'opus' | 'sonnet'
4
5const words = (list: string) =>
6  new RegExp(`\\b(?:${list.trim().split(/\s+/).join('|')})\\b`, 'i')
7
8// When the judge can't answer: planning, audits and review stay on the main
9// model ("opus only plans and audits"); everything else goes to Sonnet.
10const KEEP = words(`
11  plan plans planning planner audit audits auditing review reviews reviewing
12  reviewer grade grades grading judge judging verify verifying verification
13  critique critic architect architecture design designing decide evaluate
14  evaluating evaluation score scoring rank ranking
15`)
16
17const JUDGE_SYSTEM = [
18  "You route a coding agent's subagent to a model. Reply with exactly one word: opus or sonnet.",
19  'opus: the task needs deep judgment where a mistake is costly: planning work or designing an ' +
20    'architecture, deciding between approaches, auditing or reviewing for correctness or security, ' +
21    'grading or evaluating other work, or debugging a subtle problem that spans many parts.',
22  'sonnet: everything else: writing code, tests, content or data to a clear brief; searching, ' +
23    'reading and summarizing; running commands and checks; mechanical edits, refactors and ' +
24    'migrations with clear instructions.',
25  'When unsure, sonnet.',
26].join('\n')
27
28// Asks Haiku whether the subagent needs Opus, from its task and the start of
29// its instructions. Undefined when it can't say (an error, a time-out, the
30// account's limit, a reply that is neither word).
31async function judge($: EngineInterface, e: AgentSpawnInput): Promise<Verdict | undefined> {
32  const reply = await $.model
33    .complete({
34      model: 'haiku',
35      system: JUDGE_SYSTEM,
36      prompt: `Subagent type: ${e.subagentType}\nTask: ${e.description}\n\nIts instructions, maybe cut:\n${e.prompt.slice(0, 4000)}`,
37      maxTokens: 5,
38      effort: 'low',
39      timeoutMs: 10_000,
40    })
41    .catch(() => undefined)
42  if (!reply?.isAnswered) return undefined
43  const word = reply.text.trim().toLowerCase()
44
45  return word.startsWith('opus') ? 'opus' : word.startsWith('sonnet') ? 'sonnet' : undefined
46}
47
48export const register: Register = on => {
49  // Clears the agent count an earlier version pinned to the status line.
50  on('session.start', ($, e, next) => {
51    $.ui.status(undefined)
52
53    return next(e)
54  })
55
56  on('agent.spawn', async ($, e, next) => {
57    const asked = e.model ?? ''
58    const runsOnOpus = /opus/i.test(asked) || (asked === '' && /opus/i.test(e.parentModel))
59    // A fork shares the main conversation's cache, so it keeps its model; a
60    // subagent not headed for Opus already costs less.
61    if (e.fork || !runsOnOpus) return next(e)
62
63    let verdict: Verdict
64    let why: string
65    if (e.subagentType === 'Explore') {
66      verdict = 'sonnet'
67      why = 'Explore only reads'
68    } else {
69      const ruled = await judge($, e)
70      verdict = ruled ?? (KEEP.test(e.description) ? 'opus' : 'sonnet')
71      why = ruled ? 'judge' : 'judge unavailable, by its words'
72    }
73
74    if (verdict === 'opus') {
75      $.ui.log(`scout-router: "${e.description}" stays on Opus (${why})`)
76
77      return next(e)
78    }
79    const ran = await next({ ...e, model: 'sonnet' })
80    if (ran.deny === undefined) $.ui.log(`scout-router: "${e.description}" runs on ${ran.model} (${why})`)
81
82    return ran
83  })
84}
85