Local LLM monitor: shows whether Ollama and LM Studio are listening, counts this session's local requests, and lists local models.

A Claude Code mod that shows, right above your prompt:
It looks like this above the prompt (green, red and yellow follow your theme; every state also has a word):
✓ Ollama 0.40.2 · 3 models (1 loaded) ✗ LM Studio │ local reqs 12 (⟳1) · avg 2.3s · p95 5.1s · 41 tok/s
Below 60 columns it collapses to ✓O ✗L · 12 req · 2.3s.
Status: feature-complete for v0.1 (milestones M0–M7). See docs/PLAN.md.
claude --version; update with claude update)brew install gitleaksgit clone <repo-url> rmod && cd rmod
git config core.hooksPath .githooks # secret scanning before every commit
npm ci --ignore-scripts # dev-only: pinned TypeScript for the type-check
claude # dev session; CLAUDE.md + skills load automatically
In a second terminal, run the mod with hot reload (from outside the repo, with the clone's path):
claude --plugin-dir /path/to/rmod --debug
Useful slash commands in the dev session:
| Command | What it does | ||
|---|---|---|---|
/docs-check | refresh against the latest Claude Code, mods, Ollama and LM Studio docs | ||
/implement-milestone [id] | build the next PLAN item test-first | ||
/verify | validate, type-check, run tests, scan for secrets, check the call allowlist | ||
/probe-local-servers | inspect the real local servers and capture sanitized fixtures | ||
| `/security-review [diff\ | all]` | independent security review | |
| `/release [patch\ | minor\ | major]` | cut a release (you trigger it, never Claude) |
Terminal, for one session (also how you try a change):
claude --plugin-dir /path/to/rmod
Every session, including the Desktop app's Code tab: add the folder to env in ~/.claude/settings.json:
{ "env": { "CLAUDE_CODE_PLUGIN_DIRS": "/path/to/rmod" } }
A marketplace install line is added at release (.claude-plugin/marketplace.json).
| Command | What it does |
|---|---|
/rmod | opens the pane: status, request stats and model lists per backend. Keys: r Refresh, z Reset, h Hide/Show band, Esc closes. In claude -p it prints the text below instead |
/rmod-status | the same information as plain text (works in claude -p and VS Code) |
/rmod-refresh | re-reads the model lists now (it answers at once; run /rmod-status a moment later) |
/rmod-reset | resets this session's counters (requests still running are kept) |
What counts as a local request:
ANTHROPIC_BASE_URL is a local backend's port, or whose model id is in exactly one backend's model list.ollama run / lms chat or point a client (curl, wget, nc, inline python/node) at a backend port; WebFetch to a backend URL; and MCP tools whose server name contains "ollama" or "lm studio".A command that only mentions one in quotes, a comment or a heredoc doesn't count, and rmod's own health checks never count.
Settings (/plugin → rmod → configure, or /config): server URLs (loopback only), enable or disable each backend, poll interval (2–300 s), probe timeout (200–10000 ms), model refresh interval, the band's default visibility, the MCP server names that count as local (literal words, default ollama,lmstudio,lm-studio,lm_studio; narrow it if, say, an ollama-docs server isn't local), and an optional LM Studio API token (stored in the macOS Keychain). A bad URL or name list falls back to the default with one log line. An out-of-range number is refused by Claude Code itself when it loads the mod.
127.0.0.1, ::1 or localhost on the configured ports. Any other host in the config is rejected. There's no telemetry.tool.call and turn.step hooks pass every result through unchanged. claude plugin validate . lists the tool.call hook as gating hook without .catch: tool.call, because any hook there could refuse a call. rmod's never does, and if it ever failed, Claude Code would skip it and run the call.Authorization to your LM Studio URL, and only after LM Studio has asked for it with a 401/403. It's never logged or shown.$.http.fetch follows HTTP redirects and offers no way to turn that off. A program listening on Ollama's or LM Studio's port could redirect rmod's health check elsewhere, even to the internet. That request carries no token and no data of yours. The token is dropped on any cross-origin redirect.claude plugin validate .. Full model: docs/SECURITY.md.Operational advice:
127.0.0.1 by default. OLLAMA_HOST=0.0.0.0 exposes it to your network, and rmod notes that in /rmod-status.claude update).✗ not listening (no answer in time) although it's running again. A request to it never finished, and Claude Code can't cancel one. rmod keeps one request open per backend at most, so it waits. Run /reload-plugins, or restart the session./rmod-status shows ⚠ model list: unreadable response. The server answered with something rmod couldn't read. The last good list is kept, and /rmod-refresh tries again.Claude Code 2.1.296, Ollama 0.40.2 (live), LM Studio (response shapes from its docs; live check pending), macOS on Apple Silicon. Final versions are recorded in CHANGELOG.md at each release.
MIT © 2026 Radoslaw Szpila
Spec · Architecture · Plan · Testing · Security · Server APIs · References
hooks/register.tsx 694 lines1import type { EngineInterface, Register, Timer, ToolCallResult, TurnStepResult } from 'claude-code'
2import { atom, read, update } from 'claude-code'
3
4import { backendForModelStep, classifyToolCall, targetsOf } from './lib/attribution'
5import { bandLine } from './lib/band'
6import { applyInventory, inventoryDue, ollamaInventory, sameInventory } from './lib/inventory'
7import type { InventoryResult } from './lib/inventory'
8import { defaultConfig, parseConfig } from './lib/config'
9import type { Config } from './lib/config'
10import { applyProbe, initialStatus, lmstudioProbe, ollamaProbe } from './lib/probe'
11import type { LmStudioStep, ProbeOutcome, ProbeResult } from './lib/probe'
12import { isBackgroundBash, isFirstOutput, stepSucceeded, toolOutcome, usageOf } from './lib/requests'
13import { asStats, begin, cancel, emptyStats, end, reset, summarize } from './lib/stats'
14import { statusText } from './lib/status-text'
15import { freshMemory, transitionToast } from './lib/transitions'
16import type { ToastMemory } from './lib/transitions'
17import { parseJson } from './lib/models'
18import { parseTags } from './lib/parse-ollama'
19import { paneSection } from './lib/pane'
20import { buildUrl } from './lib/url-guard'
21import { bandTree } from './ui/band'
22import { paneTree } from './ui/pane'
23import type { BackendStatus, Health, LmStudioEndpoint, ModelInfo, Stats } from '../types'
24
25// Static analysis lets `$` reach only functions declared at the top of this file, and lets `read` /
26// `update` take only atoms defined in this file. So the work and the atoms live here, and the
27// module's mutable state is passed as `ctx`.
28
29const ollamaStatus = atom({ plugin: 'rmod', key: 'ollama' } as const, initialStatus())
30const lmstudioStatus = atom({ plugin: 'rmod', key: 'lmstudio' } as const, initialStatus())
31const statsOllama = atom({ plugin: 'rmod', key: 'statsOllama' } as const, emptyStats())
32const statsLmstudio = atom({ plugin: 'rmod', key: 'statsLmstudio' } as const, emptyStats())
33/** FR-3.6: drawn from ctx.bandHidden; this atom only makes the band and pane redraw when it changes. */
34const bandHidden = atom({ plugin: 'rmod', key: 'bandHidden' } as const, false)
35/** Which LM Studio endpoint answered this session (FR-2.2); null until known. */
36const lmstudioEndpoint = atom({ plugin: 'rmod', key: 'lmstudioEndpoint' } as const, null)
37
38type Backend = 'ollama' | 'lmstudio'
39
40type Ctx = {
41 config: Config
42 /** Kept out of Config so config can be logged or compared without leaking it. */
43 token: string | undefined
44 poll: Timer | undefined
45 /**
46 * No overlapping probes per backend (NFR-3). `pending`: a probe is running. `inflight`: fetches
47 * still running, including ones whose probe already timed out, since `$.http.fetch` can't be
48 * cancelled. A tick is skipped while either is set, so a hung server holds one request at most.
49 */
50 pending: { ollama: boolean; lmstudio: boolean }
51 inflight: { ollama: number; lmstudio: number }
52 /** FR-2.3: the token goes out only after a token-less probe to lmstudioUrl asked for auth. */
53 lmAskedForAuth: boolean
54 /** FR-3.4: what the person was last told about each backend. */
55 toasts: { ollama: ToastMemory; lmstudio: ToastMemory }
56 /** FR-5.1: when the Ollama inventory was last fetched, and whether /rmod-refresh asked for it now. */
57 inventoryAt: number | undefined
58 forceInventory: boolean
59 /** FR-4.1a: ANTHROPIC_BASE_URL, read at session.start. */
60 baseUrl: string | undefined
61 /** FR-4.1a: model ids per backend, cached as inventories are written so a model step reads no state. */
62 models: { ollama: readonly ModelInfo[]; lmstudio: readonly ModelInfo[] }
63 /**
64 * Stats writes in order, never awaited by a tool or model hook (NFR-2): begin and end land in the
65 * order they were made without delaying the request.
66 */
67 statsQueue: Promise<void>
68 /** Writes queued and not yet settled; past MAX_PENDING_WRITES new ones are dropped. */
69 pendingWrites: number
70 /**
71 * FR-3.6: the person's band choice, loaded from $.store at session.start (showBand is only the
72 * default). Kept here because /clear resets $.state: the band reads this, the tick re-seeds the atom.
73 */
74 bandHidden: boolean
75 /** Band-setting writes in order, so a double press leaves $.store with the last choice. */
76 bandQueue: Promise<void>
77 /** FR-3.7: `claude -p` or the SDK: panes are "placed" there but nothing draws them. */
78 headless: boolean
79 /** A refresh is scheduled or ran less than MIN_REFRESH_MS ago. */
80 refreshScheduled: boolean
81}
82
83const PANE_ID = 'rmod'
84/** The shortest gap between on-demand refreshes, the same floor as the poll interval (SECURITY.md §5). */
85const MIN_REFRESH_MS = 2000
86const STORE_BAND_HIDDEN = 'bandHidden'
87
88/** FR-5.1: re-read the inventories now, from a timer (the probes' sleeps must not run in a hook). */
89const requestRefresh = ($: EngineInterface, ctx: Ctx): void => {
90 ctx.forceInventory = true
91 // NFR-3: at most one on-demand refresh per MIN_REFRESH_MS, however fast the key repeats.
92 if (ctx.refreshScheduled) return
93 ctx.refreshScheduled = true
94 try {
95 $.clock.after(0, () => {
96 tick($, ctx)
97 $.clock.after(MIN_REFRESH_MS, () => {
98 ctx.refreshScheduled = false
99 })
100 })
101 } catch {
102 ctx.refreshScheduled = false
103 }
104}
105
106/** FR-4.5: clear the counters, keeping requests still running. */
107const resetCounters = ($: EngineInterface, ctx: Ctx): void => {
108 record($, ctx, 'ollama', async () => reset)
109 record($, ctx, 'lmstudio', async () => reset)
110}
111
112/** FR-3.6: flip the band, remember it across sessions, redraw the band and the pane. */
113const toggleBand = ($: EngineInterface, ctx: Ctx): void => {
114 ctx.bandHidden = !ctx.bandHidden
115 const hidden = ctx.bandHidden
116 ctx.bandQueue = ctx.bandQueue.then(async () => {
117 try {
118 await update($, bandHidden, () => hidden)
119 } catch {
120 // The band still reads ctx; the next tick re-seeds the atom.
121 }
122 // An equal-value write (right after /clear) may not redraw: ask for one.
123 try {
124 $.ui.invalidate('ui.render')
125 } catch {
126 // The next state write redraws.
127 }
128 try {
129 await $.store.set(STORE_BAND_HIDDEN, hidden)
130 } catch {
131 $.ui.log('rmod: could not save the band setting')
132 }
133 }).catch(() => {})
134}
135
136const MAX_PENDING_WRITES = 100
137const WRITE_TIMEOUT_MS = 2000
138
139/**
140 * Queues one stats change for a backend. `change` may await the clock reads the request hook
141 * started without awaiting, so no request waits on rmod. Each write is raced against a timer, so
142 * one that never settles can't stall the queue. Failures are dropped: counting never affects a request.
143 */
144const record = (
145 $: EngineInterface,
146 ctx: Ctx,
147 backend: Backend,
148 change: () => Promise<(s: Stats) => Stats>,
149 optional = false,
150): boolean => {
151 // Only an optional write (a begin) may be dropped past the cap. An end or cancel always runs,
152 // and a dropped begin drops its end too (the caller checks the return value), so in-flight can't stick.
153 if (optional && ctx.pendingWrites >= MAX_PENDING_WRITES) return false
154 ctx.pendingWrites += 1
155 ctx.statsQueue = ctx.statsQueue
156 .then(async () => {
157 const apply = await change()
158 const write =
159 backend === 'ollama' ? update($, statsOllama, s => apply(asStats(s))) : update($, statsLmstudio, s => apply(asStats(s)))
160 const stop = new AbortController()
161 const timer = $.clock.sleep(WRITE_TIMEOUT_MS, { signal: stop.signal }).catch(() => {})
162 try {
163 await Promise.race([write, timer])
164 } finally {
165 stop.abort()
166 }
167 })
168 .catch(() => {})
169 .finally(() => {
170 ctx.pendingWrites -= 1
171 })
172 return true
173}
174
175/** The end of a counted tool call: a refusal takes the begin back, anything else is a sample. */
176const recordToolEnd = ($: EngineInterface, ctx: Ctx, b: Backend, t0: Promise<number | undefined>, r: ToolCallResult | undefined): void => {
177 const at = stamp($)
178 const outcome = toolOutcome(r)
179 record($, ctx, b, async () => {
180 if (outcome === 'refused') return cancel
181 const [start, endAt] = await Promise.all([t0, at])
182 const s0 = start ?? 0
183 const sample = { kind: 'tool' as const, ok: outcome === 'ok', latencyMs: (endAt ?? s0) - s0, at: endAt ?? s0 }
184 return s => end(s, sample)
185 })
186}
187
188/** A clock read started now and awaited later (inside `record`), so the request never waits on it. */
189const stamp = ($: EngineInterface): Promise<number | undefined> => $.clock.now().then(
190 t => t,
191 () => undefined,
192)
193
194/** Whether an inventory result would change nothing in `s` (so it isn't written and redrawn). */
195const unchanged = (s: BackendStatus, r: InventoryResult): boolean =>
196 'models' in r && s.modelsError === undefined && s.modelsAt !== undefined && sameInventory(s, r.models)
197
198/** Writes an inventory result unless it is what state already holds. Static analysis wants each atom named. */
199const writeInventory = async ($: EngineInterface, ctx: Ctx, backend: Backend, r: InventoryResult, at: number): Promise<void> => {
200 if ('models' in r) ctx.models[backend] = r.models.models
201 if (backend === 'ollama') {
202 if (!unchanged(await read($, ollamaStatus), r)) await update($, ollamaStatus, prev => applyInventory(prev, r, at))
203 } else if (!unchanged(await read($, lmstudioStatus), r)) {
204 await update($, lmstudioStatus, prev => applyInventory(prev, r, at))
205 }
206}
207
208/** FR-5.1: `/api/tags` + `/api/ps` when due. Two more requests per modelsRefreshSec, not per poll. */
209const refreshOllamaModels = async ($: EngineInterface, ctx: Ctx, cameUp: boolean): Promise<void> => {
210 const now = await $.clock.now()
211 if (!inventoryDue(await read($, ollamaStatus), ctx.inventoryAt, now, ctx.config.modelsRefreshMs, cameUp, ctx.forceInventory)) return
212 ctx.inventoryAt = now
213 ctx.forceInventory = false
214 const tagsUrl = buildUrl(ctx.config.ollamaUrl, '/api/tags')
215 const psUrl = buildUrl(ctx.config.ollamaUrl, '/api/ps')
216 if (tagsUrl === undefined || psUrl === undefined) return
217 const tags = await probe($, ctx, 'ollama', tagsUrl)
218 // A failed or unreadable /api/tags decides; skip /api/ps then.
219 const tagsReadable = tags.kind === 'response' && tags.status >= 200 && tags.status <= 299 && parseTags(parseJson(tags.text)) !== undefined
220 const ps: ProbeOutcome = tagsReadable ? await probe($, ctx, 'ollama', psUrl) : { kind: 'timeout' }
221 await writeInventory($, ctx, 'ollama', ollamaInventory(tags, ps), await $.clock.now())
222}
223
224/** Shows the toast a probe result calls for, if any (FR-3.4). */
225const announce = ($: EngineInterface, ctx: Ctx, backend: Backend, health: Health): void => {
226 const r = transitionToast(backend === 'ollama' ? 'Ollama' : 'LM Studio', ctx.toasts[backend], health)
227 ctx.toasts[backend] = r.memory
228 if (r.toast !== undefined) $.ui.toast(r.toast)
229}
230
231/**
232 * One loopback GET raced against the probe timeout (`$.http.fetch` has no timeout of its own).
233 * A late answer is ignored. A rejection's message is passed on only to be classified, never kept.
234 */
235const probe = async (
236 $: EngineInterface,
237 ctx: Ctx,
238 backend: Backend,
239 url: string,
240 headers?: Record<string, string>,
241): Promise<ProbeOutcome> => {
242 const t0 = await $.clock.now()
243 const stopTimer = new AbortController()
244 const timeout = $.clock.sleep(ctx.config.probeTimeoutMs, { signal: stopTimer.signal }).then(
245 (): ProbeOutcome => ({ kind: 'timeout' }),
246 (): ProbeOutcome => ({ kind: 'timeout' }),
247 )
248 ctx.inflight[backend] += 1
249 // Through a then, so even a synchronous throw from fetch lands in `.finally` and frees the slot.
250 const request = Promise.resolve()
251 .then(() => $.http.fetch(url, headers === undefined ? undefined : { headers }))
252 .then(
253 async (r): Promise<ProbeOutcome> => ({ kind: 'response', status: r.status, text: r.text, ms: (await $.clock.now()) - t0 }),
254 (err: unknown): ProbeOutcome => ({ kind: 'rejected', message: err instanceof Error ? err.message : String(err) }),
255 )
256 .finally(() => {
257 ctx.inflight[backend] -= 1
258 // The race is over once the fetch settles: stop the timeout's sleep instead of leaving it pending.
259 stopTimer.abort()
260 })
261 return Promise.race([request, timeout])
262}
263
264const probeOllama = async ($: EngineInterface, ctx: Ctx): Promise<void> => {
265 const url = buildUrl(ctx.config.ollamaUrl, '/api/version')
266 const r: ProbeResult = url === undefined ? { health: 'degraded', error: 'config' } : ollamaProbe(await probe($, ctx, 'ollama', url))
267 const at = await $.clock.now()
268 const wasUp = (await read($, ollamaStatus)).health === 'up'
269 await update($, ollamaStatus, s => applyProbe(s, r, at))
270 announce($, ctx, 'ollama', r.health)
271 if (r.health === 'up') await refreshOllamaModels($, ctx, !wasUp)
272}
273
274const askLmStudio = async ($: EngineInterface, ctx: Ctx, endpoint: Exclude<LmStudioEndpoint, null>, withToken: boolean): Promise<LmStudioStep> => {
275 const url = buildUrl(ctx.config.lmstudioUrl, endpoint === 'native' ? '/api/v1/models' : '/v1/models')
276 if (url === undefined) return { result: { health: 'degraded', error: 'config' }, next: 'done' }
277 // The token is sent only to the guarded lmstudioUrl, never to Ollama or anywhere else.
278 const headers = withToken && ctx.token !== undefined ? { Authorization: `Bearer ${ctx.token}` } : undefined
279 return lmstudioProbe(await probe($, ctx, 'lmstudio', url, headers), endpoint, headers !== undefined)
280}
281
282const probeLmStudio = async ($: EngineInterface, ctx: Ctx): Promise<void> => {
283 // A status never checked means fresh state: a new session after /clear, /resume or /branch
284 // (which reset $.state and fire no session.start). That session must be asked for auth again.
285 if ((await read($, lmstudioStatus)).checkedAt === undefined) ctx.lmAskedForAuth = false
286 const remembered = await read($, lmstudioEndpoint)
287 let endpoint = remembered ?? 'native'
288 // At most two requests per tick (NFR-3): a fallback and a token retry in the same tick would be three.
289 let requests = 1
290 let step = await askLmStudio($, ctx, endpoint, ctx.lmAskedForAuth)
291 if (step.next === 'fallback') {
292 endpoint = 'openai'
293 requests += 1
294 step = await askLmStudio($, ctx, endpoint, ctx.lmAskedForAuth)
295 }
296 if (step.next === 'retryWithToken' && ctx.token !== undefined) {
297 ctx.lmAskedForAuth = true
298 // After a fallback, the endpoint is remembered below and the next tick asks with the token.
299 if (requests < 2) step = await askLmStudio($, ctx, endpoint, true)
300 }
301 if ((step.result.health === 'up' || step.result.health === 'auth') && endpoint !== remembered) {
302 const found = endpoint
303 await update($, lmstudioEndpoint, () => found)
304 }
305 // Whatever answers on the port next must ask for auth again before it gets the token.
306 if (step.result.health === 'down') ctx.lmAskedForAuth = false
307 const at = await $.clock.now()
308 const r = step.result
309 await update($, lmstudioStatus, s => applyProbe(s, r, at))
310 announce($, ctx, 'lmstudio', r.health)
311 // FR-5.1: the probe's own response is the inventory: no extra request.
312 if (step.models !== undefined) await writeInventory($, ctx, 'lmstudio', { models: step.models }, at)
313}
314
315/** Probes one backend unless its previous probe is still pending. rmod's own failures never escape. */
316const runProbe = async ($: EngineInterface, ctx: Ctx, backend: Backend): Promise<void> => {
317 // NFR-3: one open request per backend. A fetch that never settles (it can't be cancelled) holds
318 // the backend at "no answer in time" until it does, or until the mod reloads.
319 if (ctx.pending[backend] || ctx.inflight[backend] > 0) return
320 ctx.pending[backend] = true
321 try {
322 if (backend === 'ollama') await probeOllama($, ctx)
323 else await probeLmStudio($, ctx)
324 } catch {
325 // Keep the last known state and try again next tick.
326 } finally {
327 ctx.pending[backend] = false
328 }
329}
330
331/**
332 * Keeps a disabled backend showing off (FR-1.4, FR-2.4), also after /clear has reset $.state to
333 * `unknown` without a session.start. Reads first, so an unchanged value is never rewritten.
334 */
335const keepOff = async ($: EngineInterface, ctx: Ctx): Promise<void> => {
336 try {
337 if (!ctx.config.enableOllama && (await read($, ollamaStatus)).health !== 'off') await update($, ollamaStatus, s => onOff(s, false))
338 if (!ctx.config.enableLmStudio && (await read($, lmstudioStatus)).health !== 'off') await update($, lmstudioStatus, s => onOff(s, false))
339 // FR-3.6: after /clear the atom is back at its default; put the person's choice back.
340 const hidden = ctx.bandHidden
341 if ((await read($, bandHidden)) !== hidden) await update($, bandHidden, () => hidden)
342 } catch {
343 // Try again next tick.
344 }
345}
346
347/** One poll: each enabled backend, in parallel. */
348const tick = ($: EngineInterface, ctx: Ctx): void => {
349 void keepOff($, ctx)
350 if (ctx.config.enableOllama) void runProbe($, ctx, 'ollama')
351 if (ctx.config.enableLmStudio) void runProbe($, ctx, 'lmstudio')
352 // The pane's "checked N s ago" and "last request N s ago" must age even when no probe writes
353 // state (one still pending, both backends off). The engine throttles this; it does no I/O.
354 try {
355 $.ui.invalidate('ui.render')
356 } catch {
357 // The next state write redraws.
358 }
359}
360
361/**
362 * (Re)starts the poll timer and applies on/off (FR-1.4, FR-2.4). A second session.start replaces
363 * the timer instead of adding one. The first probe runs from a timer, where `$` calls outside a
364 * hook are documented to run.
365 */
366const startPolling = async ($: EngineInterface, ctx: Ctx): Promise<void> => {
367 try {
368 ctx.poll?.cancel()
369 ctx.poll = $.clock.every(ctx.config.pollMs, () => tick($, ctx))
370 $.clock.after(0, () => tick($, ctx))
371 } catch {
372 $.ui.log('rmod: could not start polling')
373 }
374 // Not awaited: session.start must never wait on bookkeeping (a host whose state writes hang
375 // would hold the session's start). On/off is applied at read time anyway, so this only tidies state.
376 void update($, ollamaStatus, s => onOff(s, ctx.config.enableOllama)).catch(() => {})
377 void update($, lmstudioStatus, s => onOff(s, ctx.config.enableLmStudio)).catch(() => {})
378}
379
380/** The /rmod-status text (FR-3.7), also the answer to /rmod-refresh. */
381const statusOf = async ($: EngineInterface, ctx: Ctx): Promise<string> =>
382 statusText({
383 ollama: onOff(await read($, ollamaStatus), ctx.config.enableOllama),
384 lmstudio: onOff(await read($, lmstudioStatus), ctx.config.enableLmStudio),
385 ollamaUrl: ctx.config.ollamaUrl,
386 lmstudioUrl: ctx.config.lmstudioUrl,
387 ollamaExposed: ctx.config.ollamaExposed,
388 lmstudioEndpoint: await read($, lmstudioEndpoint),
389 statsOllama: asStats(await read($, statsOllama)),
390 statsLmstudio: asStats(await read($, statsLmstudio)),
391 })
392
393/** FR-1.4 / FR-2.4: a disabled backend shows off; one enabled again after an earlier load starts over. */
394const onOff = (s: BackendStatus, enabled: boolean): BackendStatus =>
395 enabled ? (s.health === 'off' ? { ...s, health: 'unknown' } : s) : { ...s, health: 'off' }
396
397export const register: Register = (on, options) => {
398 // Re-derived on every load: a change in /config reloads the module with new options.
399 // OLLAMA_HOST needs `$`, so session.start parses again with it.
400 const ctx: Ctx = {
401 config: defaultConfig(),
402 token: undefined,
403 poll: undefined,
404 pending: { ollama: false, lmstudio: false },
405 inflight: { ollama: 0, lmstudio: 0 },
406 lmAskedForAuth: false,
407 toasts: { ollama: freshMemory(), lmstudio: freshMemory() },
408 inventoryAt: undefined,
409 forceInventory: false,
410 baseUrl: undefined,
411 models: { ollama: [], lmstudio: [] },
412 statsQueue: Promise.resolve(),
413 pendingWrites: 0,
414 bandHidden: false,
415 bandQueue: Promise.resolve(),
416 headless: false,
417 refreshScheduled: false,
418 }
419 try {
420 const parsed = parseConfig(options, undefined)
421 ctx.config = parsed.config
422 ctx.token = parsed.token
423 } catch {
424 // Keep the defaults.
425 }
426
427 on('session.start', async ($, e, next) => {
428 ctx.headless = !e.isInteractive || e.surface === null
429 try {
430 const parsed = parseConfig(options, await $.env.get('OLLAMA_HOST'))
431 ctx.config = parsed.config
432 ctx.token = parsed.token
433 // FR-6.3: one line per invalid field. Warnings never carry the rejected value or the token.
434 for (const line of parsed.warnings) $.ui.log(line)
435 } catch {
436 $.ui.log('rmod: could not read OLLAMA_HOST; using configured values')
437 }
438 try {
439 ctx.baseUrl = await $.env.get('ANTHROPIC_BASE_URL')
440 } catch {
441 ctx.baseUrl = undefined
442 }
443 // FR-3.6: a saved choice wins; showBand is only the default. Anything but a boolean is ignored.
444 ctx.bandHidden = !ctx.config.showBand
445 try {
446 const saved = await $.store.get(STORE_BAND_HIDDEN)
447 if (typeof saved === 'boolean') ctx.bandHidden = saved
448 } catch {
449 // Keep the default.
450 }
451 try {
452 await $.command.register({ name: 'rmod', description: 'rmod: open the local models pane', immediate: true })
453 } catch {
454 $.ui.log('rmod: could not register /rmod')
455 }
456 // A load or reload can't have requests of its own in flight: clear any count a lost one left behind.
457 record($, ctx, 'ollama', async () => s => ({ ...s, inFlight: 0 }))
458 record($, ctx, 'lmstudio', async () => s => ({ ...s, inFlight: 0 }))
459
460 try {
461 await $.command.register({
462 name: 'rmod-status',
463 description: 'rmod: local Ollama / LM Studio status as plain text',
464 immediate: true,
465 })
466 } catch {
467 // Fixed text only: never echo the error, which could carry user data.
468 $.ui.log('rmod: could not register /rmod-status')
469 }
470 try {
471 await $.command.register({
472 name: 'rmod-reset',
473 description: 'rmod: reset this session\'s local request counters',
474 immediate: true,
475 })
476 } catch {
477 $.ui.log('rmod: could not register /rmod-reset')
478 }
479 try {
480 await $.command.register({
481 name: 'rmod-refresh',
482 description: 'rmod: re-read the local model lists now',
483 immediate: true,
484 })
485 } catch {
486 $.ui.log('rmod: could not register /rmod-refresh')
487 }
488
489 await startPolling($, ctx)
490 return next(e)
491 })
492
493 // FR-4.1a, FR-4.2: model requests to a local backend. An observer: every chunk is passed on as
494 // it came and the result is returned unchanged. A step that isn't local is handed straight on.
495 // Nothing on the request path awaits rmod: clock reads are started here and resolved in `record`.
496 on('turn.step', async function* ($, e, next) {
497 let backend: Backend | null = null
498 try {
499 backend = backendForModelStep(ctx.baseUrl, e.model, targetsOf(ctx.config), ctx.models)
500 } catch {
501 backend = null
502 }
503 if (backend === null) return yield* next(e)
504 const b = backend
505 const t0 = stamp($)
506 let first: Promise<number | undefined> | undefined
507 let result: TurnStepResult | undefined
508 let done = false
509 // Once only: from finally, or from the dispatch's abort if the stream is abandoned (FR-4.3).
510 const finish = () => {
511 if (done || !counted) return
512 done = true
513 const at = stamp($)
514 const r = result
515 record($, ctx, b, async () => {
516 const [start, firstAt, endAt] = await Promise.all([t0, first, at])
517 const s0 = start ?? 0
518 const u = usageOf(r?.usage)
519 const sample = {
520 kind: 'model' as const,
521 ok: stepSucceeded(r),
522 latencyMs: (endAt ?? s0) - s0,
523 at: endAt ?? s0,
524 ttftMs: firstAt === undefined ? undefined : firstAt - s0,
525 outTokens: u.outTokens,
526 model: u.model ?? e.model,
527 }
528 return s => end(s, sample)
529 })
530 }
531 const counted = record($, ctx, b, async () => begin, true)
532 try {
533 next.signal.addEventListener('abort', finish, { once: true })
534 } catch {
535 // No abort signal: finally still runs.
536 }
537 try {
538 const stream = next(e)
539 // Closing early makes `result` reject: that's an interrupt, already counted below.
540 stream.result.catch(() => {})
541 for await (const chunk of stream) {
542 try {
543 if (first === undefined && isFirstOutput(chunk)) first = stamp($)
544 } catch {
545 // No first-token time for this step.
546 }
547 yield chunk
548 }
549 result = await stream.result
550 return result
551 } finally {
552 try {
553 finish()
554 } catch {
555 // Counting never affects the step.
556 }
557 }
558 })
559
560 // FR-4.1b, FR-4.2: tool calls that reach a local backend. An observer: the result is returned
561 // unchanged, and only the backend's name is kept, never the command, URL or input (FR-4.6).
562 // A refused call never reached the backend, so it isn't a request (FR-4.1).
563 on('tool.call', async ($, e, next) => {
564 let backend: Backend | null = null
565 try {
566 const input = e as unknown as Readonly<Record<string, unknown>>
567 const tool = String(e.tool)
568 if (!isBackgroundBash(tool, input)) backend = classifyToolCall(tool, input, targetsOf(ctx.config), ctx.config.mcpServerNames)
569 } catch {
570 backend = null
571 }
572 if (backend === null) return next(e)
573 const b = backend
574 const t0 = stamp($)
575 const counted = record($, ctx, b, async () => begin, true)
576 let r: ToolCallResult | undefined
577 try {
578 r = await next(e)
579 return r
580 } finally {
581 try {
582 if (counted) recordToolEnd($, ctx, b, t0, r)
583 } catch {
584 // Counting never affects the call.
585 }
586 }
587 })
588
589 // FR-3.1: the band above the prompt. Reads only; it never fetches or writes (NFR-3).
590 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
591 try {
592 // Subscribes the band to the setting; the value drawn is ctx's (FR-3.6).
593 await read($, bandHidden)
594 if (e.props.hasSurvey || ctx.bandHidden) return next(e)
595 const input = {
596 // On/off is applied as it is read, so a disabled backend shows off at once after /clear.
597 ollama: onOff(await read($, ollamaStatus), ctx.config.enableOllama),
598 lmstudio: onOff(await read($, lmstudioStatus), ctx.config.enableLmStudio),
599 statsOllama: asStats(await read($, statsOllama)),
600 statsLmstudio: asStats(await read($, statsLmstudio)),
601 }
602 // Nothing to watch and nothing counted: leave the row to the engine.
603 if (!ctx.config.enableOllama && !ctx.config.enableLmStudio && summarize(input.statsOllama, input.statsLmstudio).total === 0) return next(e)
604 const ours = bandTree($.ui.resolve(e), bandLine(input, e.props.bodyColumns))
605 const { Box } = $.ui.resolve(e)
606 // A tree here replaces what the mods after rmod draw: keep theirs, below rmod's row.
607 // `next` is called once, here; a failure past this point still returns what it gave.
608 const theirs = await next(e).catch(() => undefined)
609 return theirs === undefined || theirs === null ? (
610 ours
611 ) : (
612 <Box flexDirection="column">
613 {ours}
614 {theirs}
615 </Box>
616 )
617 } catch {
618 return next(e)
619 }
620 })
621
622 // FR-3.5: the pane. Reads only; the buttons' handlers do the work (NFR-3).
623 on('ui.render', { component: 'Pane', requestId: PANE_ID }, async ($, e, next) => {
624 try {
625 await read($, bandHidden)
626 const now = await $.clock.now()
627 const sections = [
628 paneSection('Ollama', onOff(await read($, ollamaStatus), ctx.config.enableOllama), asStats(await read($, statsOllama)), now),
629 paneSection('LM Studio', onOff(await read($, lmstudioStatus), ctx.config.enableLmStudio), asStats(await read($, statsLmstudio)), now),
630 ]
631 return paneTree($.ui.resolve(e), sections, ctx.bandHidden, {
632 onRefresh: () => requestRefresh($, ctx),
633 onReset: () => resetCounters($, ctx),
634 onToggleBand: () => toggleBand($, ctx),
635 })
636 } catch {
637 return next(e)
638 }
639 })
640
641 on('command.run', { command: 'rmod' }, async $ => {
642 try {
643 // FR-3.7: where nothing draws, answer with the same information as text. `claude -p` places
644 // every pane without drawing it, so it is told apart by the session, not by isPlaced.
645 if (ctx.headless) return { text: await statusOf($, ctx) }
646 const placed = await $.ui.open({ id: PANE_ID, title: 'rmod · local models', focus: true, closeOnEscape: true })
647 // Nothing for the transcript (or the model) when the pane opens.
648 return placed.isPlaced ? {} : { text: await statusOf($, ctx) }
649 } catch {
650 return { text: 'rmod: could not open the pane; try /rmod-status' }
651 }
652 })
653
654 on('command.run', { command: 'rmod-status' }, async $ => {
655 try {
656 return { text: await statusOf($, ctx) }
657 } catch {
658 return { text: 'rmod: status unavailable' }
659 }
660 })
661
662 // FR-4.5: counters reset on demand. A request still running when this lands ends without going negative.
663 on('command.run', { command: 'rmod-reset' }, async $ => {
664 try {
665 resetCounters($, ctx)
666 const stop = new AbortController()
667 try {
668 await Promise.race([ctx.statsQueue, $.clock.sleep(1000, { signal: stop.signal }).catch(() => {})])
669 } finally {
670 stop.abort()
671 }
672 return { text: 'rmod: local request counters reset' }
673 } catch {
674 return { text: 'rmod: reset failed' }
675 }
676 })
677
678 // FR-5.1: inventories on demand. The probes run from a timer, not in this hook: their timeout
679 // sleeps would count against the hook's budget (three in a row at probeTimeoutMs = 10 s overruns it).
680 on('command.run', { command: 'rmod-refresh' }, async $ => {
681 try {
682 const busy = ctx.pending.ollama || ctx.pending.lmstudio || ctx.inflight.ollama > 0 || ctx.inflight.lmstudio > 0
683 requestRefresh($, ctx)
684 return {
685 text: busy
686 ? 'rmod: a check is already running; the model lists refresh with it. Run /rmod-status in a moment.'
687 : 'rmod: refreshing the model lists now. Run /rmod-status in a moment.',
688 }
689 } catch {
690 return { text: 'rmod: refresh failed' }
691 }
692 })
693}
694hooks/lib/attribution.ts 181 lines1// Which local backend a model request or tool call belongs to (FR-4.1). Pure: no `$`.
2//
3// Every function returns a BackendId or null and keeps nothing: the command text, URL or tool input
4// it reads is never stored, returned or logged (FR-4.6). Backends are matched by port, so every
5// loopback spelling (127.0.0.1, localhost, [::1]) counts as the same host. Ties go to Ollama
6// (same port) or to the first match in a command; a model id in both inventories counts for neither.
7
8import type { BackendId, ModelInfo } from '../../types'
9import type { Config } from './config'
10import { inlineScript, programOf, simpleCommands } from './shell'
11import { isLoopbackUrl, normalizeOllamaHost, portOf } from './url-guard'
12
13/** The ports of the enabled backends, and of disabled ones (which get nothing attributed). */
14export type Targets = { ollama?: number; lmstudio?: number; offPorts?: readonly number[] }
15
16const BACKENDS: readonly BackendId[] = ['ollama', 'lmstudio']
17/** Only this much of a command is read, to keep each tool call's bookkeeping well under NFR-2's 5 ms. */
18const MAX_SCAN_CHARS = 10_000
19const MAX_TOOL_NAME = 128
20const MAX_SCRIPT_DEPTH = 2
21
22/** Programs whose loopback host:port arguments are requests. */
23const CLIENTS: ReadonlySet<string> = new Set(['curl', 'wget', 'http', 'https', 'xh', 'httpie', 'aria2c'])
24/** Programs that take host and port as two arguments. */
25const SOCKET_CLIENTS: ReadonlySet<string> = new Set(['nc', 'ncat', 'netcat', 'telnet'])
26/** Interpreters whose inline code may hold the URL anywhere (`python3 -c "...localhost:11434..."`). */
27const INTERPRETERS: ReadonlySet<string> = new Set(['python', 'python3', 'node', 'deno', 'bun', 'ruby', 'perl'])
28/** Runs on another machine: out of scope (SPEC §5). */
29const REMOTE: ReadonlySet<string> = new Set(['ssh', 'mosh'])
30
31const HOST = String.raw`(?:127\.0\.0\.1|localhost|\[::1\]|0\.0\.0\.0)`
32/** A whole argument that is a loopback URL or host:port: `http://localhost:11434/api`, `127.0.0.1:1234`. */
33const URL_ARG = new RegExp(String.raw`^(?:https?://)?${HOST}:(\d{1,5})(?:[/?#]|$)`, 'i')
34/** A loopback host:port anywhere in a code string, not glued to a longer name or number. */
35const IN_CODE = new RegExp(String.raw`(?<![\w.-])${HOST}(?![\w.-]):(\d{1,5})(?!\d)`, 'gi')
36const LOOPBACK_HOST = new RegExp(`^${HOST}$`, 'i')
37
38export const targetsOf = (c: Pick<Config, 'enableOllama' | 'enableLmStudio' | 'ollamaUrl' | 'lmstudioUrl'>): Targets => {
39 const t: Targets = {}
40 const off: number[] = []
41 const o = portOf(c.ollamaUrl)
42 const l = portOf(c.lmstudioUrl)
43 if (o !== undefined) {
44 if (c.enableOllama) t.ollama = o
45 else off.push(o)
46 }
47 if (l !== undefined) {
48 if (c.enableLmStudio) t.lmstudio = l
49 else off.push(l)
50 }
51 if (off.length > 0) t.offPorts = off
52 return t
53}
54
55const byPort = (targets: Targets, port: number): BackendId | null => BACKENDS.find(b => targets[b] === port) ?? null
56
57const backendOfArgs = (program: string, args: readonly string[], targets: Targets): BackendId | null => {
58 if (CLIENTS.has(program)) {
59 for (const a of args) {
60 const m = URL_ARG.exec(a)
61 const b = m === null ? null : byPort(targets, Number(m[1]))
62 if (b !== null) return b
63 }
64 } else if (SOCKET_CLIENTS.has(program)) {
65 for (let i = 0; i + 1 < args.length; i += 1) {
66 if (LOOPBACK_HOST.test(args[i] ?? '')) {
67 const b = byPort(targets, Number(args[i + 1]))
68 if (b !== null) return b
69 }
70 }
71 } else if (INTERPRETERS.has(program)) {
72 for (const a of args) {
73 for (const m of a.matchAll(IN_CODE)) {
74 const b = byPort(targets, Number(m[1]))
75 if (b !== null) return b
76 }
77 }
78 }
79 return null
80}
81
82const backendOfCommand = (command: string, targets: Targets, depth: number): BackendId | null => {
83 for (const words of simpleCommands(command.slice(0, MAX_SCAN_CHARS))) {
84 const p = programOf(words)
85 if (p === undefined || REMOTE.has(p.program)) continue
86 const script = inlineScript(p.program, p.args)
87 if (script !== undefined) {
88 const b = depth < MAX_SCRIPT_DEPTH ? backendOfCommand(script, targets, depth + 1) : null
89 if (b !== null) return b
90 continue
91 }
92 const verb = p.args.find(a => !a.startsWith('-'))
93 // `OLLAMA_HOST=1.2.3.4 ollama run …` talks to another machine: out of scope (SPEC §5).
94 const host = p.env.get('OLLAMA_HOST')
95 const remote = host !== undefined && host.trim() !== '' && normalizeOllamaHost(host) === undefined
96 if (p.program === 'ollama' && verb === 'run' && targets.ollama !== undefined && !remote) return 'ollama'
97 if (p.program === 'lms' && verb === 'chat' && targets.lmstudio !== undefined) return 'lmstudio'
98 const b = backendOfArgs(p.program, p.args, targets)
99 if (b !== null) return b
100 }
101 return null
102}
103
104const backendOfUrl = (url: string, targets: Targets): BackendId | null => {
105 const port = portOf(url)
106 return port === undefined ? null : byPort(targets, port)
107}
108
109/** The server part of `mcp__<server>__<tool>`, lower-cased; undefined when there is no tool part. */
110const mcpServerOf = (tool: string): string | undefined => {
111 const rest = tool.slice('mcp__'.length)
112 const cut = rest.indexOf('__')
113 return cut <= 0 ? undefined : rest.slice(0, cut).toLowerCase()
114}
115
116/**
117 * The backend an MCP tool belongs to: the one its server name mentions, else the only enabled
118 * backend, else none. Called only for a server already found local.
119 */
120export const backendOfMcpTool = (tool: string, targets: Targets): BackendId | null => {
121 const server = tool.slice('mcp__'.length).split('__')[0] ?? ''
122 const named: BackendId | null = /ollama/i.test(server) ? 'ollama' : /lm[-_]?studio/i.test(server) ? 'lmstudio' : null
123 if (named !== null) return targets[named] !== undefined ? named : null
124 const enabled = BACKENDS.filter(b => targets[b] !== undefined)
125 return enabled.length === 1 ? (enabled[0] ?? null) : null
126}
127
128/**
129 * FR-4.1b: a Bash command that runs `ollama run` / `lms chat`, or sends a network client to a
130 * backend's loopback port (not one that only mentions it in quotes, comments or heredocs); WebFetch to
131 * a backend URL; MCP tools whose server name contains one of `mcpServerNames`. `input` is the tool
132 * call's arguments (`e` itself).
133 */
134export const classifyToolCall = (
135 tool: string,
136 input: Readonly<Record<string, unknown>>,
137 targets: Targets,
138 mcpServerNames: readonly string[],
139): BackendId | null => {
140 if (tool === 'Bash') {
141 const command = input['command']
142 return typeof command === 'string' ? backendOfCommand(command, targets, 0) : null
143 }
144 if (tool === 'WebFetch') {
145 const url = input['url']
146 return typeof url === 'string' ? backendOfUrl(url, targets) : null
147 }
148 if (tool.startsWith('mcp__') && tool.length <= MAX_TOOL_NAME) {
149 const server = mcpServerOf(tool)
150 if (server !== undefined && mcpServerNames.some(w => server.includes(w))) return backendOfMcpTool(tool, targets)
151 }
152 return null
153}
154
155/** Ollama treats `name` and `name:latest` as the same model. */
156const sameModel = (b: BackendId, a: string, z: string): boolean =>
157 a === z || (b === 'ollama' && (a.includes(':') ? a : `${a}:latest`) === (z.includes(':') ? z : `${z}:latest`))
158
159/**
160 * FR-4.1a: a loopback `ANTHROPIC_BASE_URL` on an enabled backend's port attributes every step there.
161 * The model id decides only when there is no base URL, or it is loopback on a port that is no
162 * backend's (a local proxy): then an id found in exactly one enabled inventory picks that backend.
163 * A remote base URL, or one on a disabled backend's port, attributes nothing.
164 */
165export const backendForModelStep = (
166 baseUrl: string | undefined,
167 model: string,
168 targets: Targets,
169 inventories: { readonly [B in BackendId]: readonly ModelInfo[] },
170): BackendId | null => {
171 if (baseUrl !== undefined && baseUrl.trim() !== '') {
172 if (!isLoopbackUrl(baseUrl)) return null
173 const port = portOf(baseUrl)
174 if (port === undefined || targets.offPorts?.includes(port) === true) return null
175 const b = byPort(targets, port)
176 if (b !== null) return b
177 }
178 const owners = BACKENDS.filter(b => targets[b] !== undefined && inventories[b].some(m => sameModel(b, m.id, model)))
179 return owners.length === 1 ? (owners[0] ?? null) : null
180}
181hooks/lib/band.ts 103 lines1// The status band's content (FR-3.1, FR-3.2, FR-3.8, FR-5.2, NFR-5). Pure: no `$`; the drawing is in hooks/ui/band.tsx.
2//
3// Every state is a glyph plus a word; `tone` only adds colour. The line is built from the most
4// detailed variant down until one fits `bodyColumns`. Below COMPACT_BELOW columns it collapses
5// to icons and numbers.
6
7import type { BackendStatus, Health, Stats } from '../../types'
8import { summarize } from './stats'
9
10export const COMPACT_BELOW = 60
11
12export type Tone = 'ok' | 'bad' | 'warn' | 'dim' | 'plain'
13/** `shrink`: the part that is cut first if the row is still too wide (wide glyphs in some terminals). */
14export type Segment = { text: string; tone: Tone; shrink?: true }
15
16export type BandInput = {
17 ollama: BackendStatus
18 lmstudio: BackendStatus
19 statsOllama: Stats
20 statsLmstudio: Stats
21}
22
23const GLYPH: { readonly [H in Health]: string } = { up: '✓', auth: '⚠', degraded: '⚠', down: '✗', unknown: '…', off: '–' }
24const WORD: { readonly [H in Health]: string } = { up: '', auth: ' auth', degraded: ' degraded', down: '', unknown: '', off: ' off' }
25export const TONE: { readonly [H in Health]: Tone } = { up: 'ok', auth: 'warn', degraded: 'warn', down: 'bad', unknown: 'dim', off: 'dim' }
26
27export const formatMs = (ms: number): string =>
28 ms < 1000 ? `${Math.round(ms)}ms` : ms < 10_000 ? `${(ms / 1000).toFixed(1)}s` : `${Math.round(ms / 1000)}s`
29
30type Detail = { version: boolean; models: boolean; avg: boolean; p95: boolean; tokPerSec: boolean }
31
32const FULL: Detail = { version: true, models: true, avg: true, p95: true, tokPerSec: true }
33/** Detail dropped in this order until the line fits. */
34const VARIANTS: readonly Detail[] = [
35 FULL,
36 { ...FULL, tokPerSec: false },
37 { ...FULL, tokPerSec: false, p95: false },
38 { ...FULL, tokPerSec: false, p95: false, version: false },
39 { version: false, models: false, avg: true, p95: false, tokPerSec: false },
40 { version: false, models: false, avg: false, p95: false, tokPerSec: false },
41]
42
43/** `N models (M loaded)` once the inventory is known; just `N models` when loaded state is unknown. */
44export const modelCount = (s: BackendStatus): string | undefined => {
45 if (s.modelsAt === undefined || s.health === 'off' || s.health === 'down') return undefined
46 const n = s.models.length + s.modelsTruncated
47 const unknown = s.models.some(m => m.loaded === null)
48 const loaded = s.models.filter(m => m.loaded === true).length
49 const label = `${n} model${n === 1 ? '' : 's'}`
50 return unknown ? label : `${label} (${loaded} loaded)`
51}
52
53const backend = (name: string, s: BackendStatus, d: Detail): Segment => {
54 let text = `${GLYPH[s.health]} ${name}${WORD[s.health]}`
55 if (d.version && s.health === 'up' && s.version !== undefined) text += ` ${s.version}`
56 // Only while up: under auth, degraded or down the last list may be stale (FR-5.4).
57 const models = d.models && s.health === 'up' ? modelCount(s) : undefined
58 if (models !== undefined) text += ` · ${models}${s.modelsError === undefined ? '' : ' ⚠'}`
59 return { text, tone: TONE[s.health] }
60}
61
62const summary = (i: BandInput, d: Detail): string | undefined => {
63 const sum = summarize(i.statsOllama, i.statsLmstudio)
64 if (sum.total === 0 && sum.inFlight === 0) return undefined
65 let text = `local reqs ${sum.total}${sum.inFlight > 0 ? ` (⟳${sum.inFlight})` : ''}`
66 if (d.avg && sum.avgMs !== undefined) text += ` · avg ${formatMs(sum.avgMs)}`
67 if (d.p95 && sum.p95Ms !== undefined) text += ` · p95 ${formatMs(sum.p95Ms)}`
68 if (d.tokPerSec && sum.tokPerSec !== undefined) text += ` · ${Math.round(sum.tokPerSec)} tok/s`
69 return text
70}
71
72const full = (i: BandInput, d: Detail): Segment[] => {
73 const line: Segment[] = [backend('Ollama', i.ollama, d), { text: ' ', tone: 'plain' }, backend('LM Studio', i.lmstudio, d)]
74 const sum = summary(i, d)
75 if (sum !== undefined) line.push({ text: ` │ ${sum}`, tone: 'plain', shrink: true })
76 return line
77}
78
79const compact = (i: BandInput, columns: number): Segment[] => {
80 const head: Segment[] = [
81 { text: `${GLYPH[i.ollama.health]}O`, tone: TONE[i.ollama.health] },
82 { text: ' ', tone: 'plain' },
83 { text: `${GLYPH[i.lmstudio.health]}L`, tone: TONE[i.lmstudio.health] },
84 ]
85 const sum = summarize(i.statsOllama, i.statsLmstudio)
86 const tails = sum.total === 0 && sum.inFlight === 0 ? [] : [` · ${sum.total} req${sum.avgMs === undefined ? '' : ` · ${formatMs(sum.avgMs)}`}`, ` · ${sum.total} req`]
87 for (const tail of tails) if (5 + tail.length <= columns) return [...head, { text: tail, tone: 'plain', shrink: true }]
88 return head
89}
90
91const width = (line: readonly Segment[]): number => line.reduce((n, s) => n + s.text.length, 0)
92
93/** The band's segments for a body `columns` wide. */
94export const bandLine = (i: BandInput, columns: number): Segment[] => {
95 if (columns >= COMPACT_BELOW) {
96 for (const d of VARIANTS) {
97 const line = full(i, d)
98 if (width(line) <= columns) return line
99 }
100 }
101 return compact(i, columns)
102}
103hooks/lib/inventory.ts 53 lines1// Model inventories in state (FR-5.1, FR-5.4). Pure: no `$`; register.tsx does the fetches.
2
3import type { BackendStatus, ProbeError } from '../../types'
4import type { ModelList } from './models'
5import { parseJson } from './models'
6import { mergeOllama } from './parse-ollama'
7import { classifyRejection } from './probe'
8import type { ProbeOutcome } from './probe'
9
10export type InventoryResult = { models: ModelList } | { error: ProbeError }
11
12/** Why a body could not be read, or the body when it can be (2xx). */
13const bodyOf = (o: ProbeOutcome): { text: string } | { error: ProbeError } => {
14 if (o.kind === 'timeout') return { error: 'timeout' }
15 if (o.kind === 'rejected') return { error: classifyRejection(o.message) }
16 return o.status >= 200 && o.status <= 299 ? { text: o.text } : { error: 'http' }
17}
18
19/** `/api/tags` decides; `/api/ps` failing only makes loaded state unknown. */
20export const ollamaInventory = (tags: ProbeOutcome, ps: ProbeOutcome): InventoryResult => {
21 const t = bodyOf(tags)
22 if ('error' in t) return t
23 const p = bodyOf(ps)
24 const list = mergeOllama(parseJson(t.text), 'text' in p ? parseJson(p.text) : undefined)
25 return list === undefined ? { error: 'parse' } : { models: list }
26}
27
28/** The status with a new inventory, or with the error recorded and the last good list kept (FR-5.4). */
29export const applyInventory = (prev: BackendStatus, r: InventoryResult, at: number): BackendStatus => {
30 if ('error' in r) return { ...prev, modelsError: r.error }
31 const { modelsError: _dropped, ...rest } = prev
32 return { ...rest, models: r.models.models, modelsTruncated: r.models.truncated, modelsAt: at }
33}
34
35/**
36 * Whether to fetch the inventory now: asked for, just came up, never tried, state fresh after
37 * /clear (no list and no error), or `refreshMs` since the last attempt. After a failed attempt it
38 * waits for the interval rather than retrying on every probe.
39 */
40export const inventoryDue = (
41 s: BackendStatus,
42 lastAttempt: number | undefined,
43 now: number,
44 refreshMs: number,
45 cameUp: boolean,
46 force: boolean,
47): boolean =>
48 force || cameUp || lastAttempt === undefined || (s.modelsAt === undefined && s.modelsError === undefined) || now - lastAttempt >= refreshMs
49
50/** True when `list` is what the status already holds, so state isn't rewritten (and redrawn) for nothing. */
51export const sameInventory = (s: BackendStatus, list: ModelList): boolean =>
52 s.modelsTruncated === list.truncated && JSON.stringify(s.models) === JSON.stringify(list.models)
53hooks/lib/config.ts 179 lines1// Parses and validates the mod's userConfig options (FR-6.2, FR-6.3). Pure: no `$`.
2//
3// Every invalid field falls back to its default and adds one warning line that
4// names the field and the rule, never the rejected value (it could be a secret
5// pasted into the wrong field). The LM Studio token is returned apart from
6// Config, so Config can be logged or compared without leaking it.
7
8import { guardBaseUrl, normalizeOllamaHost, portOf } from './url-guard'
9
10export const DEFAULT_OLLAMA_URL = 'http://127.0.0.1:11434'
11export const DEFAULT_LMSTUDIO_URL = 'http://127.0.0.1:1234'
12/**
13 * Words that mark an MCP server as local when its name (the part between `mcp__` and the next `__`)
14 * contains one, in any letter case. Literal substrings, not a regular expression: matching costs
15 * linear time whatever the person configures, so a tool call can never be held up (NFR-2).
16 */
17export const DEFAULT_MCP_SERVER_NAMES = 'ollama,lmstudio,lm-studio,lm_studio'
18
19const MAX_TOKEN_LENGTH = 4096
20const MAX_SERVER_NAMES = 10
21/** Letters, digits, `-` and `_`: what an MCP server name can hold. */
22const SERVER_WORD = /^[a-z0-9_-]{1,40}$/
23
24export type Config = {
25 enableOllama: boolean
26 enableLmStudio: boolean
27 showBand: boolean
28 ollamaUrl: string
29 /** Where ollamaUrl came from: the user's setting, OLLAMA_HOST, or the built-in default. */
30 ollamaSource: 'config' | 'env' | 'default'
31 /** OLLAMA_HOST binds all interfaces (0.0.0.0); rmod still probes 127.0.0.1. */
32 ollamaExposed: boolean
33 lmstudioUrl: string
34 pollMs: number
35 probeTimeoutMs: number
36 modelsRefreshMs: number
37 /** Lower-case words; an MCP server whose name contains one is local (FR-4.1b). */
38 mcpServerNames: readonly string[]
39}
40
41export type ParsedConfig = {
42 config: Config
43 /** The LM Studio API token, trimmed; undefined when unset or unusable. */
44 token: string | undefined
45 warnings: string[]
46}
47
48type Options = Readonly<Record<string, unknown>>
49
50const bool = (o: Options, key: string, fallback: boolean, warn: (m: string) => void): boolean => {
51 const v = o[key]
52 if (v === undefined) return fallback
53 if (typeof v === 'boolean') return v
54 warn(`rmod: ${key} must be true or false; using ${String(fallback)}`)
55 return fallback
56}
57
58const num = (o: Options, key: string, min: number, max: number, fallback: number, warn: (m: string) => void): number => {
59 const v = o[key]
60 if (v === undefined) return fallback
61 if (typeof v === 'number' && Number.isFinite(v) && v >= min && v <= max) return v
62 warn(`rmod: ${key} must be a number from ${min} to ${max}; using ${fallback}`)
63 return fallback
64}
65
66/** A configured string, trimmed; undefined when unset or blank (blank is never an error). */
67const str = (o: Options, key: string, warn: (m: string) => void): string | undefined => {
68 const v = o[key]
69 if (v === undefined) return undefined
70 if (typeof v !== 'string') {
71 warn(`rmod: ${key} must be text; using the default`)
72 return undefined
73 }
74 const t = v.trim()
75 return t === '' ? undefined : t
76}
77
78/** A configured base URL: `base` when valid, `rejected` when set but not loopback; neither when blank. */
79const urlField = (o: Options, key: string, fallback: string, warn: (m: string) => void): { base?: string; rejected: boolean } => {
80 const raw = str(o, key, warn)
81 if (raw === undefined) return { rejected: false }
82 const g = guardBaseUrl(raw)
83 if (g.ok) return { base: g.base, rejected: false }
84 warn(`rmod: ${key} rejected (${g.reason}); using ${fallback}`)
85 return { rejected: true }
86}
87
88const serverNames = (o: Options, warn: (m: string) => void): readonly string[] => {
89 const fallback = DEFAULT_MCP_SERVER_NAMES.split(',')
90 const raw = str(o, 'localMcpServerNames', warn)
91 if (raw === undefined) return fallback
92 const words = raw
93 .toLowerCase()
94 .split(',')
95 .map(w => w.trim())
96 .filter(w => w !== '')
97 if (words.length === 0 || words.length > MAX_SERVER_NAMES || !words.every(w => SERVER_WORD.test(w))) {
98 warn(`rmod: localMcpServerNames must be up to ${MAX_SERVER_NAMES} comma-separated words of letters, digits, - or _; using the default`)
99 return fallback
100 }
101 return words
102}
103
104const tokenOf = (o: Options, warn: (m: string) => void): string | undefined => {
105 const v = o['lmstudioApiToken']
106 if (typeof v !== 'string') return undefined
107 const t = v.trim()
108 if (t === '') return undefined
109 if (t.length > MAX_TOKEN_LENGTH || !/^[\x21-\x7e]+$/.test(t)) {
110 warn('rmod: lmstudioApiToken is not usable (too long, or not printable ASCII); sending no token')
111 return undefined
112 }
113 return t
114}
115
116/** The config rmod uses when nothing else can be read: the manifest's defaults. */
117export const defaultConfig = (): Config => parseConfig({}, undefined).config
118
119/**
120 * @param options what register() received (manifest defaults filled in by the loader)
121 * @param ollamaHostEnv the OLLAMA_HOST environment value, if any
122 */
123export const parseConfig = (options: Options, ollamaHostEnv: string | undefined): ParsedConfig => {
124 const warnings: string[] = []
125 const warn = (m: string) => {
126 warnings.push(m)
127 }
128
129 const enableOllama = bool(options, 'enableOllama', true, warn)
130 const enableLmStudio = bool(options, 'enableLmStudio', true, warn)
131 const showBand = bool(options, 'showBand', true, warn)
132
133 let ollamaUrl = DEFAULT_OLLAMA_URL
134 let ollamaSource: Config['ollamaSource'] = 'default'
135 let ollamaExposed = false
136 const ollamaField = urlField(options, 'ollamaUrl', DEFAULT_OLLAMA_URL, warn)
137 if (ollamaField.base !== undefined) {
138 ollamaUrl = ollamaField.base
139 ollamaSource = 'config'
140 } else if (!ollamaField.rejected && ollamaHostEnv !== undefined && ollamaHostEnv.trim() !== '') {
141 const env = normalizeOllamaHost(ollamaHostEnv)
142 if (env !== undefined) {
143 ollamaUrl = env.url
144 ollamaSource = 'env'
145 ollamaExposed = env.exposed
146 } else {
147 warn(`rmod: OLLAMA_HOST is not a loopback address; probing ${DEFAULT_OLLAMA_URL}`)
148 }
149 }
150
151 const lmField = urlField(options, 'lmstudioUrl', DEFAULT_LMSTUDIO_URL, warn)
152 const lmstudioUrl = lmField.base ?? DEFAULT_LMSTUDIO_URL
153 if (enableOllama && enableLmStudio && portOf(ollamaUrl) === portOf(lmstudioUrl)) {
154 warn('rmod: ollamaUrl and lmstudioUrl use the same port; requests to it count as Ollama')
155 }
156
157 const config: Config = {
158 enableOllama,
159 enableLmStudio,
160 showBand,
161 ollamaUrl,
162 ollamaSource,
163 ollamaExposed,
164 lmstudioUrl,
165 pollMs: num(options, 'pollIntervalSec', 2, 300, 5, warn) * 1000,
166 probeTimeoutMs: num(options, 'probeTimeoutMs', 200, 10000, 1500, warn),
167 modelsRefreshMs: num(options, 'modelsRefreshSec', 10, 3600, 30, warn) * 1000,
168 mcpServerNames: serverNames(options, warn),
169 }
170
171 let token = tokenOf(options, warn)
172 if (token !== undefined && lmField.rejected) {
173 // The token was meant for the server the user named, not whatever answers on the default port.
174 warn('rmod: lmstudioApiToken not sent because lmstudioUrl was rejected')
175 token = undefined
176 }
177 return { config, token, warnings }
178}
179hooks/lib/probe.ts 75 lines1// Classifies health probes (FR-1, FR-2). Pure: no `$`; register.tsx does the fetch and the timeout race.
2//
3// Server text and rejection messages are untrusted and may hold URLs: they are only classified,
4// never kept. What reaches state is rmod's own Health and ProbeError words.
5
6import type { BackendStatus, Health, LmStudioEndpoint, ProbeError } from '../../types'
7import { parseJson } from './models'
8import type { ModelList } from './models'
9import { parseVersion } from './parse-ollama'
10import { parseOpenAiModels, parseV1Models } from './parse-lmstudio'
11
12/** What one probe came to: a response (any status), a rejected fetch, or no answer in time. */
13export type ProbeOutcome =
14 | { kind: 'response'; status: number; text: string; ms: number }
15 | { kind: 'rejected'; message: string }
16 | { kind: 'timeout' }
17
18export type ProbeResult = { health: Health; error?: ProbeError; version?: string; probeMs?: number }
19
20export const initialStatus = (): BackendStatus => ({ health: 'unknown', models: [], modelsTruncated: 0 })
21
22/** A rejected fetch: an organization's web-fetch policy, or nothing reachable there. */
23export const classifyRejection = (message: string): ProbeError => (/polic/i.test(message) ? 'policy' : 'refused')
24
25const notListening = (o: Exclude<ProbeOutcome, { kind: 'response' }>): ProbeResult =>
26 o.kind === 'timeout' ? { health: 'down', error: 'timeout' } : { health: 'down', error: classifyRejection(o.message) }
27
28/** Any HTTP answer means listening (SPEC §2): 2xx up, 401/403 auth required, anything else degraded. */
29const byStatus = (status: number, probeMs: number): ProbeResult | undefined => {
30 if (status === 401 || status === 403) return { health: 'auth', probeMs }
31 if (status < 200 || status > 299) return { health: 'degraded', error: 'http', probeMs }
32 return undefined
33}
34
35/** `GET /api/version` (FR-1.2, FR-1.3). */
36export const ollamaProbe = (o: ProbeOutcome): ProbeResult => {
37 if (o.kind !== 'response') return notListening(o)
38 const bad = byStatus(o.status, o.ms)
39 if (bad !== undefined) return bad
40 const version = parseVersion(parseJson(o.text))
41 return version === undefined ? { health: 'degraded', error: 'parse', probeMs: o.ms } : { health: 'up', version, probeMs: o.ms }
42}
43
44export type LmStudioStep = {
45 result: ProbeResult
46 /** `fallback`: try `/v1/models` (FR-2.2). `retryWithToken`: ask again with the token, if one is set (FR-2.3). */
47 next: 'done' | 'fallback' | 'retryWithToken'
48 /** The models the response listed, when it was up and readable (FR-5.1: the probe is the inventory). */
49 models?: ModelList
50}
51
52/** `GET /api/v1/models` (native) or `GET /v1/models` (fallback) (FR-2). */
53export const lmstudioProbe = (o: ProbeOutcome, endpoint: Exclude<LmStudioEndpoint, null>, tokenSent: boolean): LmStudioStep => {
54 if (o.kind !== 'response') return { result: notListening(o), next: 'done' }
55 if (o.status === 404 && endpoint === 'native') return { result: { health: 'degraded', error: 'http', probeMs: o.ms }, next: 'fallback' }
56 const bad = byStatus(o.status, o.ms)
57 if (bad !== undefined) return { result: bad, next: bad.health === 'auth' && !tokenSent ? 'retryWithToken' : 'done' }
58 const body = parseJson(o.text)
59 const list = endpoint === 'native' ? parseV1Models(body) : parseOpenAiModels(body)
60 return list === undefined
61 ? { result: { health: 'degraded', error: 'parse', probeMs: o.ms }, next: 'done' }
62 : { result: { health: 'up', probeMs: o.ms }, next: 'done', models: list }
63}
64
65/** The new status after a probe: health fields replaced (stale version/error dropped), models kept. */
66export const applyProbe = (prev: BackendStatus, r: ProbeResult, at: number): BackendStatus => {
67 const next: BackendStatus = { health: r.health, checkedAt: at, models: prev.models, modelsTruncated: prev.modelsTruncated }
68 if (r.version !== undefined) next.version = r.version
69 if (r.probeMs !== undefined) next.probeMs = r.probeMs
70 if (r.error !== undefined) next.error = r.error
71 if (prev.modelsAt !== undefined) next.modelsAt = prev.modelsAt
72 if (prev.modelsError !== undefined) next.modelsError = prev.modelsError
73 return next
74}
75hooks/lib/requests.ts 34 lines1// Reading model steps and tool results for the request counter (FR-4.2, FR-4.6). Pure: no `$`.
2// Only kinds, counts and the model id are read. Never text, input or output.
3
4/** A chunk that is the model's own output (TTFT ends here), not the engine's envelope or a stop marker. */
5export const isFirstOutput = (chunk: { readonly kind: string }): boolean =>
6 chunk.kind === 'text' || chunk.kind === 'thinking' || chunk.kind === 'tool' || chunk.kind === 'input'
7
8/** Output tokens and the answering model from a step's `usage`, when present and sane. */
9export const usageOf = (usage: unknown): { outTokens?: number; model?: string } => {
10 if (typeof usage !== 'object' || usage === null) return {}
11 const u = usage as Readonly<Record<string, unknown>>
12 const out: { outTokens?: number; model?: string } = {}
13 const t = u['output_tokens']
14 if (typeof t === 'number' && Number.isSafeInteger(t) && t >= 0) out.outTokens = t
15 const m = u['model']
16 if (typeof m === 'string' && m !== '') out.model = m
17 return out
18}
19
20/** What a tool call came to: refused (never reached the backend: not a request), failed, or ok. */
21export const toolOutcome = (r: unknown): 'refused' | 'failed' | 'ok' => {
22 if (typeof r !== 'object' || r === null) return 'failed'
23 const o = r as Readonly<Record<string, unknown>>
24 if (typeof o['deny'] === 'string') return 'refused'
25 return o['isError'] === true ? 'failed' : 'ok'
26}
27
28/** A step with no stop reason never got a response (failed or interrupted before a message). */
29export const stepSucceeded = (r: { readonly stopReason: unknown } | undefined): boolean => r !== undefined && r.stopReason !== null
30
31/** Bash run in the background returns at once: its latency isn't the request's, so it isn't counted. */
32export const isBackgroundBash = (tool: string, input: Readonly<Record<string, unknown>>): boolean =>
33 tool === 'Bash' && input['run_in_background'] === true
34hooks/lib/stats.ts 218 lines1// Request statistics per backend (FR-4.2, FR-4.3, FR-4.6, NFR-3). Pure: no `$`. Every update returns a new Stats.
2//
3// Speed figures (latency, TTFT, tok/s) come from successful requests only (SPEC FR-4.2): a backend
4// that fails fast must not look quick. Every request still counts in total, ok/failed and lastAt.
5
6import type { Stats } from '../../types'
7import { cleanId } from './models'
8
9export const MAX_SAMPLES = 200
10export const MAX_BY_MODEL = 50
11const MAX_LATENCY_MS = 24 * 60 * 60 * 1000
12const MAX_TOKENS = 10_000_000
13/** Below this, the time after the first token is too short to measure a rate from. */
14const MIN_GEN_MS = 50
15
16/** One finished request: metadata only, never prompt, output or command text (FR-4.6). */
17export type Sample = {
18 kind: 'model' | 'tool'
19 /** False for an error, an abort or a tool error. */
20 ok: boolean
21 latencyMs: number
22 /** Clock time when it finished. */
23 at: number
24 /** Time to the first streamed chunk (model requests only). */
25 ttftMs?: number
26 /** Output tokens from `usage` (model requests only). */
27 outTokens?: number
28 model?: string
29}
30
31export type Summary = {
32 total: number
33 model: number
34 tool: number
35 inFlight: number
36 ok: number
37 failed: number
38 avgMs?: number
39 p50Ms?: number
40 p95Ms?: number
41 lastMs?: number
42 avgTtftMs?: number
43 tokPerSec?: number
44 lastAt?: number
45}
46
47export const emptyStats = (): Stats => ({
48 total: 0,
49 model: 0,
50 tool: 0,
51 inFlight: 0,
52 ok: 0,
53 failed: 0,
54 latMs: [],
55 ttftMs: [],
56 outTokens: 0,
57 rateTokens: 0,
58 genMs: 0,
59 byModel: {},
60 byModelOther: 0,
61})
62
63const count = (v: unknown): number => (typeof v === 'number' && Number.isSafeInteger(v) && v >= 0 ? v : 0)
64
65const samples = (v: unknown): number[] =>
66 Array.isArray(v)
67 ? v.filter((n): n is number => typeof n === 'number' && Number.isFinite(n) && n >= 0).slice(-MAX_SAMPLES)
68 : []
69
70/**
71 * A full, bounded Stats from whatever `$.state` holds: after a hot reload it may be a value from an
72 * older contract, or partial. Missing or malformed fields become their empty values.
73 */
74export const asStats = (v: unknown): Stats => {
75 if (typeof v !== 'object' || v === null) return emptyStats()
76 const o = v as Readonly<Record<string, unknown>>
77 const counts = new Map<string, number>()
78 const rawByModel = o['byModel']
79 if (typeof rawByModel === 'object' && rawByModel !== null) {
80 for (const [k, n] of Object.entries(rawByModel)) {
81 if (counts.size >= MAX_BY_MODEL) break
82 if (count(n) > 0) counts.set(k, count(n))
83 }
84 }
85 const s: Stats = {
86 total: count(o['total']),
87 model: count(o['model']),
88 tool: count(o['tool']),
89 inFlight: count(o['inFlight']),
90 ok: count(o['ok']),
91 failed: count(o['failed']),
92 latMs: samples(o['latMs']),
93 ttftMs: samples(o['ttftMs']),
94 outTokens: count(o['outTokens']),
95 rateTokens: count(o['rateTokens']),
96 genMs: count(o['genMs']),
97 byModel: Object.fromEntries(counts),
98 byModelOther: count(o['byModelOther']),
99 }
100 const lastAt = o['lastAt']
101 if (typeof lastAt === 'number' && Number.isFinite(lastAt)) s.lastAt = lastAt
102 return s
103}
104
105export const begin = (s: Stats): Stats => ({ ...s, inFlight: s.inFlight + 1 })
106
107/** Takes back a `begin` for something that turned out not to be a request (a refused tool call). */
108export const cancel = (s: Stats): Stats => ({ ...s, inFlight: Math.max(0, s.inFlight - 1) })
109
110/** Clears the counters but keeps requests still running, so the band doesn't hide live work. */
111export const reset = (s: Stats): Stats => ({ ...emptyStats(), inFlight: s.inFlight })
112
113const ms = (v: unknown): number =>
114 typeof v === 'number' && Number.isFinite(v) ? Math.min(Math.max(v, 0), MAX_LATENCY_MS) : 0
115
116const tokens = (v: unknown): number | undefined =>
117 typeof v === 'number' && Number.isSafeInteger(v) && v >= 0 && v <= MAX_TOKENS ? v : undefined
118
119const push = (list: readonly number[], v: number): number[] => [...list.slice(-(MAX_SAMPLES - 1)), v]
120
121/** How many requests this session sent for `model` (0 when unknown or counted in byModelOther). */
122export const requestsFor = (s: Stats, model: string): number =>
123 Object.hasOwn(s.byModel, model) ? (s.byModel[model] ?? 0) : 0
124
125const countModel = (s: Stats, raw: unknown): Pick<Stats, 'byModel' | 'byModelOther'> => {
126 const id = cleanId(raw)
127 if (id === undefined) return { byModel: s.byModel, byModelOther: s.byModelOther }
128 // A Map, then Object.fromEntries: untrusted ids such as `__proto__` stay plain own keys.
129 const counts = new Map(Object.entries(s.byModel))
130 if (!counts.has(id) && counts.size >= MAX_BY_MODEL) return { byModel: s.byModel, byModelOther: s.byModelOther + 1 }
131 counts.set(id, (counts.get(id) ?? 0) + 1)
132 return { byModel: Object.fromEntries(counts), byModelOther: s.byModelOther }
133}
134
135/** The part of a successful model request that feeds tok/s, or nothing when it can't be measured. */
136const rate = (latency: number, ttft: number | undefined, out: number | undefined): { tokens: number; ms: number } | undefined => {
137 if (out === undefined) return undefined
138 const afterFirst = ttft === undefined ? latency : latency - ttft
139 // One chunk, or a first token at the very end: measure over the whole request instead.
140 const span = afterFirst >= Math.max(MIN_GEN_MS, latency * 0.1) ? afterFirst : latency
141 return span > 0 ? { tokens: out, ms: span } : undefined
142}
143
144export const end = (s: Stats, sample: Sample): Stats => {
145 const isModel = sample.kind === 'model'
146 const latency = ms(sample.latencyMs)
147 const t = isModel ? sample.ttftMs : undefined
148 // Ignored unless it is a real time that fits inside the request: a first token can't come after the end.
149 const ttft = typeof t === 'number' && Number.isFinite(t) && t >= 0 && t <= latency ? t : undefined
150 const out = isModel ? tokens(sample.outTokens) : undefined
151 const measured = sample.ok ? rate(latency, ttft, out) : undefined
152
153 // Only known fields are copied: nothing else on `sample` (text of any kind) reaches state (FR-4.6).
154 const next: Stats = {
155 ...s,
156 total: s.total + 1,
157 model: s.model + (isModel ? 1 : 0),
158 tool: s.tool + (isModel ? 0 : 1),
159 inFlight: Math.max(0, s.inFlight - 1),
160 ok: s.ok + (sample.ok ? 1 : 0),
161 failed: s.failed + (sample.ok ? 0 : 1),
162 latMs: sample.ok ? push(s.latMs, latency) : s.latMs,
163 ttftMs: sample.ok && ttft !== undefined ? push(s.ttftMs, ttft) : s.ttftMs,
164 outTokens: s.outTokens + (out ?? 0),
165 rateTokens: s.rateTokens + (measured?.tokens ?? 0),
166 genMs: s.genMs + (measured?.ms ?? 0),
167 ...(sample.model === undefined ? {} : countModel(s, sample.model)),
168 }
169 if (typeof sample.at === 'number' && Number.isFinite(sample.at)) next.lastAt = sample.at
170 return next
171}
172
173const nearestRank = (sorted: readonly number[], p: number): number | undefined => {
174 if (sorted.length === 0) return undefined
175 return sorted[Math.min(sorted.length, Math.max(1, Math.ceil((p / 100) * sorted.length))) - 1]
176}
177
178/** Nearest-rank percentile; undefined for no samples. */
179export const percentile = (list: readonly number[], p: number): number | undefined =>
180 nearestRank([...list].sort((a, b) => a - b), p)
181
182const mean = (list: readonly number[]): number | undefined =>
183 list.length === 0 ? undefined : list.reduce((a, b) => a + b, 0) / list.length
184
185const sum = (all: readonly Stats[], pick: (s: Stats) => number): number => all.reduce((n, s) => n + pick(s), 0)
186
187/** Display figures for one backend, or for several combined (the band's total). Unknowns are left out. */
188export const summarize = (...all: readonly Stats[]): Summary => {
189 const lat = all.flatMap(s => s.latMs).sort((a, b) => a - b)
190 const genMs = sum(all, s => s.genMs)
191 // The backend that finished a request most recently gives lastAt; lastMs is its newest successful latency.
192 const newest = all.reduce<Stats | undefined>(
193 (best, s) => (s.lastAt !== undefined && (best?.lastAt === undefined || s.lastAt > best.lastAt) ? s : best),
194 undefined,
195 )
196 const lastMs = newest?.latMs.at(-1) ?? [...all].reverse().find(s => s.latMs.length > 0)?.latMs.at(-1)
197
198 const out: Summary = {
199 total: sum(all, s => s.total),
200 model: sum(all, s => s.model),
201 tool: sum(all, s => s.tool),
202 inFlight: sum(all, s => s.inFlight),
203 ok: sum(all, s => s.ok),
204 failed: sum(all, s => s.failed),
205 }
206 const set = <K extends keyof Summary>(k: K, v: Summary[K] | undefined) => {
207 if (v !== undefined) out[k] = v
208 }
209 set('avgMs', mean(lat))
210 set('p50Ms', nearestRank(lat, 50))
211 set('p95Ms', nearestRank(lat, 95))
212 set('lastMs', lastMs)
213 set('avgTtftMs', mean(all.flatMap(s => s.ttftMs)))
214 set('tokPerSec', genMs > 0 ? (sum(all, s => s.rateTokens) * 1000) / genMs : undefined)
215 set('lastAt', newest?.lastAt)
216 return out
217}
218hooks/lib/status-text.ts 67 lines1// The plain-text status that /rmod-status returns (FR-3.7, NFR-5). Pure: no `$`.
2// Glyph + word for every state, so the meaning never depends on colour.
3
4import type { BackendStatus, Health, LmStudioEndpoint, ProbeError, Stats } from '../../types'
5import { formatMs, modelCount } from './band'
6import { summarize } from './stats'
7
8export const HEALTH_TEXT: { readonly [H in Health]: string } = {
9 up: '✓ up',
10 auth: '⚠ up, auth required',
11 degraded: '⚠ degraded',
12 down: '✗ not listening',
13 unknown: '… checking',
14 off: '– off',
15}
16
17export const ERROR_TEXT: { readonly [E in ProbeError]: string } = {
18 refused: 'connection refused',
19 timeout: 'no answer in time',
20 policy: 'blocked by policy',
21 http: 'unexpected HTTP status',
22 parse: 'unreadable response',
23 config: 'invalid URL',
24}
25
26const line = (name: string, s: BackendStatus, url: string, extra: readonly string[]): string => {
27 const parts = [`${name}: ${HEALTH_TEXT[s.health]}${s.error === undefined || s.health === 'off' ? '' : ` (${ERROR_TEXT[s.error]})`}`]
28 if (s.health === 'up' && s.version !== undefined) parts.push(`v${s.version}`)
29 if (s.health !== 'down' && s.health !== 'off' && s.probeMs !== undefined) parts.push(`${Math.round(s.probeMs)} ms`)
30 const models = modelCount(s)
31 if (models !== undefined) parts.push(models)
32 if (s.modelsError !== undefined && s.health !== 'down' && s.health !== 'off') parts.push('⚠ model list: unreadable response')
33 if (s.health !== 'off') parts.push(url)
34 return [...parts, ...extra].join(' · ')
35}
36
37export type StatusInput = {
38 ollama: BackendStatus
39 lmstudio: BackendStatus
40 ollamaUrl: string
41 lmstudioUrl: string
42 ollamaExposed: boolean
43 lmstudioEndpoint: LmStudioEndpoint
44 statsOllama: Stats
45 statsLmstudio: Stats
46}
47
48/** The band's session summary as a line, or nothing before the first local request. */
49const sessionLine = (i: StatusInput): string[] => {
50 const sum = summarize(i.statsOllama, i.statsLmstudio)
51 if (sum.total === 0 && sum.inFlight === 0) return []
52 const parts = [`Session: ${sum.total} local requests (${sum.model} model, ${sum.tool} tool)`, `${sum.ok} ok, ${sum.failed} failed`]
53 if (sum.inFlight > 0) parts.push(`${sum.inFlight} in flight`)
54 if (sum.avgMs !== undefined) parts.push(`avg ${formatMs(sum.avgMs)}`)
55 if (sum.p95Ms !== undefined) parts.push(`p95 ${formatMs(sum.p95Ms)}`)
56 if (sum.tokPerSec !== undefined) parts.push(`${Math.round(sum.tokPerSec)} tok/s`)
57 return [parts.join(' · ')]
58}
59
60export const statusText = (i: StatusInput): string =>
61 [
62 'rmod',
63 line('Ollama', i.ollama, i.ollamaUrl, i.ollamaExposed ? ['OLLAMA_HOST listens on all interfaces'] : []),
64 line('LM Studio', i.lmstudio, i.lmstudioUrl, i.lmstudioEndpoint === 'openai' ? ['OpenAI-compatible API (older LM Studio)'] : []),
65 ...sessionLine(i),
66 ].join('\n')
67hooks/lib/transitions.ts 38 lines1// Which health changes deserve a toast (FR-3.4). Pure: no `$`.
2//
3// A toast says what the person was last told has changed. "Stopped listening" needs two probes in
4// a row that found the backend down, so one slow answer (a timeout while a model loads) toasts
5// nothing. "Listening again" follows only a backend the person knows to be down. The first probe
6// sets what the person is told without a toast.
7
8import type { Health } from '../../types'
9
10const listening = (h: Health): boolean => h === 'up' || h === 'auth' || h === 'degraded'
11
12/** Probes in a row that must find a backend down before "stopped listening" is shown. */
13export const DOWN_CONFIRMATIONS = 2
14
15export type ToastMemory = {
16 /** What the person was last told (or saw at the first probe); undefined before any probe. */
17 shown: 'listening' | 'down' | undefined
18 /** Probes in a row that found the backend down. */
19 downStreak: number
20}
21
22export const freshMemory = (): ToastMemory => ({ shown: undefined, downStreak: 0 })
23
24/** The memory after a probe found `health`, and the toast to show for it, if any. */
25export const transitionToast = (name: string, m: ToastMemory, health: Health): { memory: ToastMemory; toast?: string } => {
26 if (health === 'unknown' || health === 'off') return { memory: m }
27 if (listening(health)) {
28 const memory: ToastMemory = { shown: 'listening', downStreak: 0 }
29 return m.shown === 'down' ? { memory, toast: `✓ ${name} is listening again` } : { memory }
30 }
31 const downStreak = m.downStreak + 1
32 if (m.shown === undefined) return { memory: { shown: 'down', downStreak } }
33 if (m.shown === 'listening' && downStreak >= DOWN_CONFIRMATIONS) {
34 return { memory: { shown: 'down', downStreak }, toast: `✗ ${name} stopped listening` }
35 }
36 return { memory: { shown: m.shown, downStreak } }
37}
38hooks/lib/models.ts 100 lines1// Shared helpers for model inventories from untrusted servers (FR-5, NFR-1, NFR-3). Pure: no `$`.
2
3import type { ModelInfo } from '../../types'
4
5export const MAX_MODELS = 200
6/** How many raw entries a parser looks at; the rest are only counted. */
7export const MAX_SCAN = 5000
8const MAX_BODY_CHARS = 4_000_000
9const MAX_ID_CHARS = 200
10const MAX_SHORT_CHARS = 32
11
12export type ModelList = { models: ModelInfo[]; truncated: number }
13
14/** JSON.parse that never throws; undefined for oversized or malformed bodies. */
15export const parseJson = (text: string): unknown => {
16 if (text.length > MAX_BODY_CHARS) return undefined
17 try {
18 return JSON.parse(text) as unknown
19 } catch {
20 return undefined
21 }
22}
23
24export const isRecord = (v: unknown): v is Readonly<Record<string, unknown>> =>
25 typeof v === 'object' && v !== null && !Array.isArray(v)
26
27/**
28 * Characters that could repaint, reorder or hide text: controls (Cc), format characters such as
29 * bidi overrides, zero-width and tag characters (Cf), line/paragraph separators (Zl, Zp),
30 * private-use (Co) and lone surrogates (Cs).
31 */
32const UNSAFE = /[\p{Cc}\p{Cf}\p{Zl}\p{Zp}\p{Co}\p{Cs}]/gu
33
34/** A display-safe string: unsafe characters removed, trimmed, cut to `max` code points; undefined when empty. */
35export const cleanText = (v: unknown, max: number): string | undefined => {
36 if (typeof v !== 'string') return undefined
37 // Bound the work first (a code point is at most 2 UTF-16 units), then cut on code points.
38 const t = [...v.slice(0, max * 4).replace(UNSAFE, '').trim()].slice(0, max).join('').trim()
39 return t === '' ? undefined : t
40}
41
42export const cleanId = (v: unknown): string | undefined => cleanText(v, MAX_ID_CHARS)
43
44/** A short label (parameters, quantization). Too long means garbage, so it is dropped, not cut. */
45export const cleanShort = (v: unknown): string | undefined => {
46 if (typeof v !== 'string' || v.length > MAX_SHORT_CHARS * 2) return undefined
47 const t = cleanText(v, MAX_SHORT_CHARS + 1)
48 return t === undefined || t.length > MAX_SHORT_CHARS ? undefined : t
49}
50
51export const cleanBytes = (v: unknown): number | undefined =>
52 typeof v === 'number' && Number.isSafeInteger(v) && v >= 0 ? v : undefined
53
54/** Builds a ModelInfo without undefined-valued keys, so results compare and serialise cleanly. */
55export const modelInfo = (id: string, loaded: boolean | null, extra: Omit<ModelInfo, 'id' | 'loaded'>): ModelInfo => {
56 const m: ModelInfo = { id, loaded }
57 if (extra.sizeBytes !== undefined) m.sizeBytes = extra.sizeBytes
58 if (extra.params !== undefined) m.params = extra.params
59 if (extra.quant !== undefined) m.quant = extra.quant
60 if (extra.kind !== undefined) m.kind = extra.kind
61 if (extra.vramBytes !== undefined) m.vramBytes = extra.vramBytes
62 if (extra.expiresAt !== undefined) m.expiresAt = extra.expiresAt
63 return m
64}
65
66/** An ISO 8601 timestamp as epoch ms; undefined unless it parses to a finite time. */
67export const cleanTime = (v: unknown): number | undefined => {
68 if (typeof v !== 'string' || v.length > 64) return undefined
69 // Ollama sends nanoseconds (`.000000000Z`); Date.parse is only specified for milliseconds.
70 const t = Date.parse(v.replace(/(\.\d{3})\d+/, '$1'))
71 return Number.isFinite(t) ? t : undefined
72}
73
74const rank = (m: ModelInfo): number => (m.loaded === true ? 0 : m.loaded === null ? 1 : 2)
75
76/**
77 * Dedupes by id (the most-loaded copy wins), sorts loaded first then by id, and caps at MAX_MODELS.
78 * `unscanned` counts raw entries a parser never looked at; they count as truncated, so past
79 * MAX_SCAN entries the `+N more` figure is an upper bound (it includes invalid and duplicate entries).
80 */
81export const finalize = (entries: Iterable<ModelInfo>, unscanned = 0): ModelList => {
82 const byId = new Map<string, ModelInfo>()
83 for (const m of entries) {
84 // On a duplicate id keep the entry with the better loaded state (loaded > unknown > unloaded).
85 const seen = byId.get(m.id)
86 if (seen === undefined || rank(m) < rank(seen)) byId.set(m.id, m)
87 }
88 const all = [...byId.values()].sort((a, b) => rank(a) - rank(b) || (a.id < b.id ? -1 : a.id > b.id ? 1 : 0))
89 return { models: all.slice(0, MAX_MODELS), truncated: Math.max(0, all.length - MAX_MODELS) + unscanned }
90}
91
92/** The array under `key` of a record body, or undefined when the body is not of that shape. */
93export const listAt = (body: unknown, key: string): { items: readonly unknown[]; unscanned: number } | undefined => {
94 if (!isRecord(body)) return undefined
95 const list = body[key]
96 if (!Array.isArray(list)) return undefined
97 const items: readonly unknown[] = list
98 return { items: items.slice(0, MAX_SCAN), unscanned: Math.max(0, items.length - MAX_SCAN) }
99}
100hooks/lib/parse-ollama.ts 104 lines1// Ollama response parsers (FR-1.3, FR-5.1, FR-5.4). Pure: no `$`. Input is untrusted JSON.
2// Shapes: docs/LOCAL_SERVERS.md (verified against Ollama 0.40.2).
3
4import type { ModelInfo } from '../../types'
5import { cleanBytes, cleanId, cleanShort, cleanTime, finalize, isRecord, listAt, MAX_MODELS, modelInfo } from './models'
6import type { ModelList } from './models'
7
8/** `GET /api/version` → `{ version }`. Undefined unless it is a plain version string of at most 32 characters. */
9export const parseVersion = (body: unknown): string | undefined => {
10 if (!isRecord(body)) return undefined
11 const v = body['version']
12 return typeof v === 'string' && /^[0-9A-Za-z.+_-]{1,32}$/.test(v) ? v : undefined
13}
14
15const kindOf = (capabilities: unknown): ModelInfo['kind'] => {
16 if (!Array.isArray(capabilities)) return undefined
17 const caps: readonly unknown[] = capabilities
18 return caps.includes('embedding') ? 'embedding' : 'llm'
19}
20
21/** The tag entries before dedupe and cap, so a merge with `/api/ps` caps only once. */
22const tagEntries = (body: unknown): { models: ModelInfo[]; unscanned: number } | undefined => {
23 const list = listAt(body, 'models')
24 if (list === undefined) return undefined
25 const models: ModelInfo[] = []
26 for (const entry of list.items) {
27 if (!isRecord(entry)) continue
28 // Cloud models (remote_host set) are not local: out of scope (SPEC §5). The host is never stored.
29 if (typeof entry['remote_host'] === 'string') continue
30 const id = cleanId(entry['name'])
31 if (id === undefined) continue
32 const details = isRecord(entry['details']) ? entry['details'] : {}
33 models.push(
34 modelInfo(id, false, {
35 sizeBytes: cleanBytes(entry['size']),
36 // MLX models report these as "": cleanShort turns that into unknown.
37 params: cleanShort(details['parameter_size']),
38 quant: cleanShort(details['quantization_level']),
39 kind: kindOf(entry['capabilities']),
40 }),
41 )
42 }
43 return { models, unscanned: list.unscanned }
44}
45
46/** `GET /api/tags` → available models, all `loaded: false`. */
47export const parseTags = (body: unknown): ModelList | undefined => {
48 const t = tagEntries(body)
49 return t === undefined ? undefined : finalize(t.models, t.unscanned)
50}
51
52type Loaded = { sizeBytes?: number; vramBytes?: number; expiresAt?: number }
53
54/** `/api/ps` entries by name (at most MAX_MODELS), with their memory figures. */
55const psEntries = (body: unknown): Map<string, Loaded> | undefined => {
56 const list = listAt(body, 'models')
57 if (list === undefined) return undefined
58 const loaded = new Map<string, Loaded>()
59 for (const entry of list.items) {
60 if (loaded.size >= MAX_MODELS) break
61 if (!isRecord(entry)) continue
62 const id = cleanId(entry['name'])
63 if (id === undefined || loaded.has(id)) continue
64 loaded.set(id, {
65 sizeBytes: cleanBytes(entry['size']),
66 vramBytes: cleanBytes(entry['size_vram']),
67 expiresAt: cleanTime(entry['expires_at']),
68 })
69 }
70 return loaded
71}
72
73/** `GET /api/ps` → names of the models loaded in memory (at most MAX_MODELS). */
74export const parsePs = (body: unknown): string[] | undefined => {
75 const loaded = psEntries(body)
76 return loaded === undefined ? undefined : [...loaded.keys()]
77}
78
79/**
80 * The Ollama inventory from the `/api/tags` and `/api/ps` bodies. Undefined when the tags are
81 * unreadable (FR-5.4). An unreadable `/api/ps` makes every loaded state unknown. Loaded models
82 * missing from the tags are added. Loaded models sort first, so the cap never drops one.
83 */
84export const mergeOllama = (tagsBody: unknown, psBody: unknown): ModelList | undefined => {
85 const tags = tagEntries(tagsBody)
86 if (tags === undefined) return undefined
87 const loaded = psEntries(psBody)
88 if (loaded === undefined) return finalize(tags.models.map(m => ({ ...m, loaded: null })), tags.unscanned)
89 const merged: ModelInfo[] = tags.models.map(m => {
90 const ps = loaded.get(m.id)
91 // Keep the tags' disk size; /api/ps `size` is the in-memory size and is used only for ps-only models.
92 return ps === undefined ? { ...m, loaded: false } : modelInfo(m.id, true, { ...m, vramBytes: ps.vramBytes, expiresAt: ps.expiresAt })
93 })
94 const known = new Set(tags.models.map(m => m.id))
95 let psOnly = 0
96 for (const [id, ps] of loaded) {
97 if (known.has(id)) continue
98 merged.push(modelInfo(id, true, ps))
99 psOnly += 1
100 }
101 // A ps-only model may sit in the unscanned tail of the tags: don't count it twice.
102 return finalize(merged, Math.max(0, tags.unscanned - psOnly))
103}
104