SLOPSHOPPER

smart-model-router

Routes each main-thread turn to a heavy or light model based on the prompt, and falls back to the light model near your rate limits.

newcommandstatusprompt
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · smart-model-router
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /route ⎿ smart-model-router: Router mode is "auto". Use /route auto | heavy | light | off. ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts ⚠ smart-model-router: router: claude-sonnet-5-5 (routine prompt)
README

cc-mods

Four Claude Code mods (function-hook plugins) that save tokens and keep you aware of your limits.

ModWhat it does
smart-model-routerPicks a heavy or light model for each turn
usage-barShows 5-hour and weekly rate-limit usage above the prompt
context-barShows how full the context window is, by category, above the prompt
research-offloaderKeeps research work out of the main context

Why: kaizen for your AI bill

Kaizen is continuous improvement through small changes, each removing one source of waste (muda). These mods apply it to a Claude Code workflow:

WasteMod that removes it
The most expensive model doing routine worksmart-model-router
Raw web pages filling the main contextresearch-offloader
Hitting a rate limit by surpriseusage-bar (and the router's usage guard)
Context filling up unnoticedcontext-bar

None of these is a big idea; each is one small fix. Run the cycle on your own usage:

  1. Plan: note your 5-hour and weekly usage over a normal week, using usage-bar.
  2. Do: install the mods and work as usual.
  3. Check: compare the next week's usage against the first.
  4. Adjust: tune usageGuard, the model options and the research trigger phrases, then repeat.

What you can see, you can improve.

Install

Try from a local folder (development)

claude --plugin-dir D:\ai\cc-mods\smart-model-router --plugin-dir D:\ai\cc-mods\usage-bar --plugin-dir D:\ai\cc-mods\context-bar --plugin-dir D:\ai\cc-mods\research-offloader

Saving a file reloads the mod in a running session.

Install from GitHub (timothylok/ClaudeCodeModsByTimLok)

/plugin install smart-model-router --marketplace timothylok/ClaudeCodeModsByTimLok
/plugin install usage-bar --marketplace timothylok/ClaudeCodeModsByTimLok
/plugin install context-bar --marketplace timothylok/ClaudeCodeModsByTimLok
/plugin install research-offloader --marketplace timothylok/ClaudeCodeModsByTimLok

Answer y to add the marketplace, then choose a scope. Each mod is active immediately.

Settings

Mods with options show them as rows in /config, or you can set them in settings.json under pluginConfigs.<mod>.options. Changing one reloads the mod.


smart-model-router

Chooses a model for every turn and sends the main thread's requests to it. Subagents keep their own model.

How a model is chosen, in order:

  1. If your 5-hour or weekly usage is at or above the usage guard (default 85%), use the light model.
  2. If the mode is heavy or light, use that model.
  3. In auto mode, a prompt that mentions architecture, designing a system, multi-file work, rewriting the whole thing, deep reasoning, trade-offs, migrations, root cause, or planning, or one longer than 2,500 characters, gets the heavy model. Anything else gets the light model.

The status line shows the choice and the reason, e.g. router: claude-opus-5-5 (complex prompt).

Usage

CommandEffect
/routeShow the current mode
/route autoPick per prompt (default)
/route heavyAlways use the heavy model
/route lightAlways use the light model
/route offRouter does nothing; the session's own model is used

The mode is remembered across sessions.

Options

OptionDefaultMeaning
heavyModelclaude-opus-5-5Model for complex turns
lightModelclaude-sonnet-5-5Model for everyday turns
usageGuard85Usage % above which the light model is always used

Examples

PromptModel
Design the architecture for a multi-tenant billing serviceheavy
Plan out a migration from REST to gRPC, with the trade-offsheavy
Find the root cause of this flaky testheavy
fix the typo in READMElight
write a unit test for parseDatelight

Switching models starts a new prompt cache, so flipping often costs extra input tokens. Use /route heavy or /route light to pin a model for a stretch of work.


usage-bar

A band above the prompt with a meter for each rate-limit window Claude Code reports:

Usage 5h █████░░░░░ 42% 2h30m   week █████████░ 91% 3d
  • The bar is green below 70%, yellow from 70%, and red from 90%.
  • The trailing time is how long until the window resets.
  • A toast appears once when a window crosses 80% and once more at 95%.
  • The band uses figures Claude Code already receives, so it needs no server. It only appears on a subscription, and only after the first response brings rate-limit figures. It is hidden while a survey is showing.

Usage

CommandEffect
/usage-barHide or show the band (toggle)

Options

None.


context-bar

A stacked bar above the prompt showing the context window, one colour per /context category, with a legend of the biggest categories:

████████▒▒▒▒░░░░░░░░░░░░░░░░░░░░ 38% · 76k/200k
■ Messages 41k  ■ System tools 18k  ■ Memory files 6k
  • █ is used space, ▒ is the autocompact buffer, ░ is free space.
  • Figures are estimated locally from the last response's usage, so refreshing the bar sends no token-count requests. It refreshes after each turn and after a compaction.
  • It sits in the same band as usage-bar and stacks with it. It is hidden while a survey is showing.

Usage

CommandEffect
/context-barHide or show the bar (toggle)

Options

None.


research-offloader

When you type a research-style prompt, the mod keeps the page fetching and reading out of the main thread, so only a summary enters your context.

What counts as research: prompts containing "research", "look up", "summarize this url/page/article/docs", "compare sources", "read this page/article/docs", "what is the latest", or "search the web". It only reacts to prompts you type yourself.

Two modes, depending on the endpoint option:

  • Endpoint set: the mod POSTs { "query": "<your prompt>" } to it, expects { "summary": "..." } back, and gives Claude that summary as context. If the endpoint fails or returns no summary, it falls back to the subagent mode below.
  • Endpoint empty (default): the mod tells Claude to do the research inside subagents and return only short summaries with source URLs. For that turn, subagents spawned without a model run on the subagentModel.

A toast tells you which mode was used.

Options

OptionDefaultMeaning
endpointemptyURL of your own research bridge (for example a NotebookLM wrapper). Leave empty to use subagents
subagentModelhaikuModel for research subagents: haiku, sonnet or opus

Examples

PromptResult
Research the best Rust ORMs in 2026offloaded
Can you look up how Vite handles HMR?offloaded
summarize this page for me: https://example.comoffloaded
compare sources on WebGPU supportoffloaded
fix the failing test in auth.tsuntouched

NotebookLM has no public API, so the endpoint is yours to provide. Any service that takes {query} and returns {summary} works.


Bug fixes

0.1.3: usage-bar hid other bands above the prompt

  • Symptom: with usage-bar installed, a separate context bar disappeared.
  • Cause: both mods draw into the same band above the prompt (ui.render on AbovePrompt). The outermost hook returned its own tree without calling next(e), so the hook beneath it never ran.
  • Fix: usage-bar now calls next(e) and stacks what comes back under its own row. context-bar does the same from 0.1.4, so the two show together in either order.
  • For mod authors: a ui.render hook on a shared site must await next(e) and include the result in its tree. In tests, add a stand-in ui.render hook beneath the mod so next has something to return.

Development

Each mod has hooks/register.ts(x), a manifest in .claude-plugin/plugin.json, and tests in tests/.

claude plugin validate <mod folder>
claude plugin test <mod folder>

License

MIT

Source 1 files
hooks/register.ts 93 lines
1import type { Register } from 'claude-code'
2
3type Mode = 'auto' | 'heavy' | 'light' | 'off'
4
5const HEAVY = [
6  /\barchitect(ure|ing)?\b/,
7  /\bdesign (a|the) (system|service|api|schema)\b/,
8  /\bmulti[- ]file\b/,
9  /\brefactor (everything|the (whole|entire))\b/,
10  /\b(re)?write the (whole|entire)\b/,
11  /\bdeep(ly)? (reason|think|dive)\b/,
12  /\btrade-?offs?\b/,
13  /\bmigrat(e|ion) (from|to|the)\b/,
14  /\broot cause\b/,
15  /\b(plan|blueprint) (for|out)\b/,
16]
17
18export const isHeavyPrompt = (text: string): boolean => {
19  const lower = text.toLowerCase()
20
21  return HEAVY.some(re => re.test(lower)) || lower.length > 2500
22}
23
24export const register: Register = (on, options) => {
25  const heavyModel = String(options.heavyModel ?? 'claude-opus-5-5')
26  const lightModel = String(options.lightModel ?? 'claude-sonnet-5-5')
27  const usageGuard = Number(options.usageGuard ?? 85)
28
29  let mode: Mode = 'auto'
30  // The model this turn's main-thread steps go to; null leaves the session's own.
31  let turnModel: string | null = null
32
33  on('session.start', async ($, e, next) => {
34    const stored = await $.store.get('mode')
35    if (stored === 'auto' || stored === 'heavy' || stored === 'light' || stored === 'off') mode = stored
36    await $.command.register({
37      name: 'route',
38      description: 'Model router: /route auto | heavy | light | off',
39    })
40
41    return next(e)
42  })
43
44  on('command.run', { command: 'route' }, async ($, e) => {
45    const want = e.args.trim().toLowerCase()
46    if (want !== 'auto' && want !== 'heavy' && want !== 'light' && want !== 'off') {
47      return { text: `Router mode is "${mode}". Use /route auto | heavy | light | off.` }
48    }
49    mode = want
50    await $.store.set('mode', mode)
51    if (mode === 'off') $.ui.status(undefined)
52
53    return { text: `Router mode set to "${mode}".` }
54  })
55
56  on('prompt.submit', async ($, e, next) => {
57    if (mode === 'off') {
58      turnModel = null
59
60      return next(e)
61    }
62
63    const { rateLimits } = await $.session.usage()
64    const peak = Math.max(0, ...rateLimits.map(r => r.percentUsed))
65    let why: string
66
67    if (peak >= usageGuard) {
68      turnModel = lightModel
69      why = `usage ${peak}%`
70    } else if (mode === 'heavy' || mode === 'light') {
71      turnModel = mode === 'heavy' ? heavyModel : lightModel
72      why = 'forced'
73    } else {
74      const heavy = isHeavyPrompt(e.text)
75      turnModel = heavy ? heavyModel : lightModel
76      why = heavy ? 'complex prompt' : 'routine prompt'
77    }
78    $.ui.status(`router: ${turnModel} (${why})`)
79
80    return next(e)
81  })
82
83  // Every model request of the main thread in this turn goes to the chosen model;
84  // subagents keep whatever model they were spawned with.
85  on('turn.step', async function* ($, e, next) {
86    if (turnModel === null || e.agentId !== undefined || e.model === turnModel) {
87      return yield* next(e)
88    }
89
90    return yield* next({ ...e, model: turnModel })
91  })
92}
93