Spots research prompts and keeps them out of the main context: sends them to your own research endpoint (e.g. a NotebookLM bridge) if configured, otherwise…

Four Claude Code mods (function-hook plugins) that save tokens and keep you aware of your limits.
| Mod | What it does |
|---|---|
smart-model-router | Picks a heavy or light model for each turn |
usage-bar | Shows 5-hour and weekly rate-limit usage above the prompt |
context-bar | Shows how full the context window is, by category, above the prompt |
research-offloader | Keeps research work out of the main context |
Kaizen is continuous improvement through small changes, each removing one source of waste (muda). These mods apply it to a Claude Code workflow:
| Waste | Mod that removes it |
|---|---|
| The most expensive model doing routine work | smart-model-router |
| Raw web pages filling the main context | research-offloader |
| Hitting a rate limit by surprise | usage-bar (and the router's usage guard) |
| Context filling up unnoticed | context-bar |
None of these is a big idea; each is one small fix. Run the cycle on your own usage:
usage-bar.usageGuard, the model options and the research trigger phrases, then repeat.What you can see, you can improve.
claude --plugin-dir D:\ai\cc-mods\smart-model-router --plugin-dir D:\ai\cc-mods\usage-bar --plugin-dir D:\ai\cc-mods\context-bar --plugin-dir D:\ai\cc-mods\research-offloader
Saving a file reloads the mod in a running session.
/plugin install smart-model-router --marketplace timothylok/ClaudeCodeModsByTimLok
/plugin install usage-bar --marketplace timothylok/ClaudeCodeModsByTimLok
/plugin install context-bar --marketplace timothylok/ClaudeCodeModsByTimLok
/plugin install research-offloader --marketplace timothylok/ClaudeCodeModsByTimLok
Answer y to add the marketplace, then choose a scope. Each mod is active immediately.
Mods with options show them as rows in /config, or you can set them in settings.json under pluginConfigs.<mod>.options. Changing one reloads the mod.
Chooses a model for every turn and sends the main thread's requests to it. Subagents keep their own model.
How a model is chosen, in order:
heavy or light, use that model.auto mode, a prompt that mentions architecture, designing a system, multi-file work, rewriting the whole thing, deep reasoning, trade-offs, migrations, root cause, or planning, or one longer than 2,500 characters, gets the heavy model. Anything else gets the light model.The status line shows the choice and the reason, e.g. router: claude-opus-5-5 (complex prompt).
| Command | Effect |
|---|---|
/route | Show the current mode |
/route auto | Pick per prompt (default) |
/route heavy | Always use the heavy model |
/route light | Always use the light model |
/route off | Router does nothing; the session's own model is used |
The mode is remembered across sessions.
| Option | Default | Meaning |
|---|---|---|
heavyModel | claude-opus-5-5 | Model for complex turns |
lightModel | claude-sonnet-5-5 | Model for everyday turns |
usageGuard | 85 | Usage % above which the light model is always used |
| Prompt | Model |
|---|---|
Design the architecture for a multi-tenant billing service | heavy |
Plan out a migration from REST to gRPC, with the trade-offs | heavy |
Find the root cause of this flaky test | heavy |
fix the typo in README | light |
write a unit test for parseDate | light |
Switching models starts a new prompt cache, so flipping often costs extra input tokens. Use /route heavy or /route light to pin a model for a stretch of work.
A band above the prompt with a meter for each rate-limit window Claude Code reports:
Usage 5h █████░░░░░ 42% 2h30m week █████████░ 91% 3d
| Command | Effect |
|---|---|
/usage-bar | Hide or show the band (toggle) |
None.
A stacked bar above the prompt showing the context window, one colour per /context category, with a legend of the biggest categories:
████████▒▒▒▒░░░░░░░░░░░░░░░░░░░░ 38% · 76k/200k
■ Messages 41k ■ System tools 18k ■ Memory files 6k
█ is used space, ▒ is the autocompact buffer, ░ is free space.usage-bar and stacks with it. It is hidden while a survey is showing.| Command | Effect |
|---|---|
/context-bar | Hide or show the bar (toggle) |
None.
When you type a research-style prompt, the mod keeps the page fetching and reading out of the main thread, so only a summary enters your context.
What counts as research: prompts containing "research", "look up", "summarize this url/page/article/docs", "compare sources", "read this page/article/docs", "what is the latest", or "search the web". It only reacts to prompts you type yourself.
Two modes, depending on the endpoint option:
{ "query": "<your prompt>" } to it, expects { "summary": "..." } back, and gives Claude that summary as context. If the endpoint fails or returns no summary, it falls back to the subagent mode below.subagentModel.A toast tells you which mode was used.
| Option | Default | Meaning |
|---|---|---|
endpoint | empty | URL of your own research bridge (for example a NotebookLM wrapper). Leave empty to use subagents |
subagentModel | haiku | Model for research subagents: haiku, sonnet or opus |
| Prompt | Result |
|---|---|
Research the best Rust ORMs in 2026 | offloaded |
Can you look up how Vite handles HMR? | offloaded |
summarize this page for me: https://example.com | offloaded |
compare sources on WebGPU support | offloaded |
fix the failing test in auth.ts | untouched |
NotebookLM has no public API, so the endpoint is yours to provide. Any service that takes {query} and returns {summary} works.
usage-bar hid other bands above the promptusage-bar installed, a separate context bar disappeared.ui.render on AbovePrompt). The outermost hook returned its own tree without calling next(e), so the hook beneath it never ran.usage-bar now calls next(e) and stacks what comes back under its own row. context-bar does the same from 0.1.4, so the two show together in either order.ui.render hook on a shared site must await next(e) and include the result in its tree. In tests, add a stand-in ui.render hook beneath the mod so next has something to return.Each mod has hooks/register.ts(x), a manifest in .claude-plugin/plugin.json, and tests in tests/.
claude plugin validate <mod folder>
claude plugin test <mod folder>
hooks/register.ts 78 lines1import type { Register } from 'claude-code'
2
3const TRIGGERS = [
4 /\bresearch\b/,
5 /\blook (it |this |that )?up\b/,
6 /\bsummari[sz]e (this|that|the) (url|page|article|site|link|docs?)\b/,
7 /\bcompare (the )?sources\b/,
8 /\bread (this|that|the) (page|article|docs?)\b/,
9 /\bwhat('s| is) the latest\b/,
10 /\bsearch the web\b/,
11]
12
13export const isResearchPrompt = (text: string): boolean => {
14 const lower = text.toLowerCase()
15
16 return TRIGGERS.some(re => re.test(lower))
17}
18
19const delegateNote = (model: string) =>
20 [
21 '[research-offloader] This prompt is a research task. To keep the main context small:',
22 `- Do the searching, fetching and reading inside one or more Agent subagents (model: "${model}"), not in this thread.`,
23 '- Ask each subagent to return only a concise summary with source URLs, not raw page content.',
24 '- Answer the user from those summaries.',
25 ].join('\n')
26
27export const register: Register = (on, options) => {
28 const endpoint = String(options.endpoint ?? '').trim()
29 const subagentModel = String(options.subagentModel ?? 'haiku') as 'haiku' | 'sonnet' | 'opus'
30 // True while the current turn came from a research prompt.
31 let isResearchTurn = false
32
33 on('prompt.submit', async ($, e, next) => {
34 isResearchTurn = e.origin.kind === 'composer' && isResearchPrompt(e.text)
35 if (!isResearchTurn) return next(e)
36
37 if (endpoint !== '') {
38 $.ui.toast('research-offloader: asking your research endpoint…')
39 const res = await $.http.fetch(endpoint, {
40 method: 'POST',
41 headers: { 'Content-Type': 'application/json' },
42 body: JSON.stringify({ query: e.text }),
43 })
44 let summary = ''
45 if (res.ok) {
46 try {
47 summary = String((JSON.parse(res.text) as { summary?: unknown }).summary ?? '')
48 } catch {
49 summary = ''
50 }
51 }
52 if (summary !== '') {
53 const note = `[research-offloader] Research summary from ${endpoint}. Use it as your source; only research further if it is clearly insufficient.\n\n${summary}`
54
55 return next({ ...e, context: [...(e.context ?? []), note] })
56 }
57 $.ui.toast(`research-offloader: endpoint gave no summary (HTTP ${res.status}); delegating to a subagent`)
58 } else {
59 $.ui.toast(`research-offloader: delegating research to a ${subagentModel} subagent`)
60 }
61
62 return next({ ...e, context: [...(e.context ?? []), delegateNote(subagentModel)] })
63 })
64
65 // During a research turn, subagents spawned without a model run on the cheap one.
66 on('tool.call', { tool: 'Agent' }, ($, e, next) =>
67 isResearchTurn && e.model === undefined && e.subagent_type !== 'fork'
68 ? next({ ...e, model: subagentModel })
69 : next(e),
70 )
71
72 on('turn.complete', ($, e, next) => {
73 isResearchTurn = false
74
75 return next(e)
76 })
77}
78