SLOPSHOPPER

rmod

Local LLM monitor: shows whether Ollama and LM Studio are listening, counts this session's local requests, and lists local models.

newpanebandguardcommandtoast
v0.0.1MITupdated 2026-10-10beerbox/rmod
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · rmod
│ ┃ rmod · local models ✕ › fix the failing auth test and add an audit log call │ ┃ r: Refresh z: Reset h: Hide band │ ┃ ⏺ Read(src/auth.ts) │ ┃ Ollama: … checking ⎿ Read 6 lines │ ┃ no local requests yet ⏺ Update(src/auth.ts) │ ┃ ⎿ Added 2 lines, removed 1 line │ ┃ LM Studio: … checking ⏺ Bash(bun test) │ ┃ no local requests yet ⎿ 3 pass, 1 fail │ ┃ │ ┃ ● loaded · ○ available · ◌ unknown · ×N used ● Done. refresh now rejects expired claims and logs an audit event. │ │ ✻ Worked for 42s · done 4:20 PM │ │ › /rmod │ │ … Ollama … LM Studio ⟨Claude Code's own drawing⟩ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Band
… Ollama … LM Studio ⟨Claude Code's own drawing⟩
Pane · rmod · local models
r: Refresh z: Reset h: Hide band Ollama: … checking no local requests yet LM Studio: … checking no local requests yet ● loaded · ○ available · ◌ unknown · ×N used
README

rmod: local LLM monitor for Claude Code

A Claude Code mod that shows, right above your prompt:

  • whether Ollama and LM Studio are listening on this Mac (green ✓ / red ✗)
  • which local models are available, and which are loaded
  • a live counter of the requests this session sends to local models, with avg, p95, time to first token and tokens per second

It looks like this above the prompt (green, red and yellow follow your theme; every state also has a word):

✓ Ollama 0.40.2 · 3 models (1 loaded)   ✗ LM Studio   │ local reqs 12 (⟳1) · avg 2.3s · p95 5.1s · 41 tok/s

Below 60 columns it collapses to ✓O ✗L · 12 req · 2.3s.

Status: feature-complete for v0.1 (milestones M0–M7). See docs/PLAN.md.

Requirements

  • macOS (Apple Silicon MacBook)
  • Claude Code v2.1.287 or later (claude --version; update with claude update)
  • Optional: Ollama and/or LM Studio with its local server started
  • For development: Node.js (TypeScript type-check only) and brew install gitleaks

Develop

git clone <repo-url> rmod && cd rmod
git config core.hooksPath .githooks   # secret scanning before every commit
npm ci --ignore-scripts               # dev-only: pinned TypeScript for the type-check
claude                                # dev session; CLAUDE.md + skills load automatically

In a second terminal, run the mod with hot reload (from outside the repo, with the clone's path):

claude --plugin-dir /path/to/rmod --debug

Useful slash commands in the dev session:

CommandWhat it does
/docs-checkrefresh against the latest Claude Code, mods, Ollama and LM Studio docs
/implement-milestone [id]build the next PLAN item test-first
/verifyvalidate, type-check, run tests, scan for secrets, check the call allowlist
/probe-local-serversinspect the real local servers and capture sanitized fixtures
`/security-review [diff\all]`independent security review
`/release [patch\minor\major]`cut a release (you trigger it, never Claude)

Install

Terminal, for one session (also how you try a change):

claude --plugin-dir /path/to/rmod

Every session, including the Desktop app's Code tab: add the folder to env in ~/.claude/settings.json:

{ "env": { "CLAUDE_CODE_PLUGIN_DIRS": "/path/to/rmod" } }

A marketplace install line is added at release (.claude-plugin/marketplace.json).

Use

CommandWhat it does
/rmodopens the pane: status, request stats and model lists per backend. Keys: r Refresh, z Reset, h Hide/Show band, Esc closes. In claude -p it prints the text below instead
/rmod-statusthe same information as plain text (works in claude -p and VS Code)
/rmod-refreshre-reads the model lists now (it answers at once; run /rmod-status a moment later)
/rmod-resetresets this session's counters (requests still running are kept)

What counts as a local request:

  • Model requests: steps whose ANTHROPIC_BASE_URL is a local backend's port, or whose model id is in exactly one backend's model list.
  • Tool calls: Bash commands that run ollama run / lms chat or point a client (curl, wget, nc, inline python/node) at a backend port; WebFetch to a backend URL; and MCP tools whose server name contains "ollama" or "lm studio".

A command that only mentions one in quotes, a comment or a heredoc doesn't count, and rmod's own health checks never count.

Settings (/plugin → rmod → configure, or /config): server URLs (loopback only), enable or disable each backend, poll interval (2–300 s), probe timeout (200–10000 ms), model refresh interval, the band's default visibility, the MCP server names that count as local (literal words, default ollama,lmstudio,lm-studio,lm_studio; narrow it if, say, an ollama-docs server isn't local), and an optional LM Studio API token (stored in the macOS Keychain). A bad URL or name list falls back to the default with one log line. An out-of-range number is refused by Claude Code itself when it loads the mod.

Privacy and trust

  • Loopback only. rmod sends HTTP only to 127.0.0.1, ::1 or localhost on the configured ports. Any other host in the config is rejected. There's no telemetry.
  • Metadata only. It keeps counts, timings, model names and status codes. It never stores, logs or shows your prompts, model output, Bash commands, URLs or file contents.
  • Observer, never a gate. rmod's tool.call and turn.step hooks pass every result through unchanged. claude plugin validate . lists the tool.call hook as gating hook without .catch: tool.call, because any hook there could refuse a call. rmod's never does, and if it ever failed, Claude Code would skip it and run the call.
  • The token is sent only as Authorization to your LM Studio URL, and only after LM Studio has asked for it with a 401/403. It's never logged or shown.
  • One limit you should know. Claude Code's $.http.fetch follows HTTP redirects and offers no way to turn that off. A program listening on Ollama's or LM Studio's port could redirect rmod's health check elsewhere, even to the internet. That request carries no token and no data of yours. The token is dropped on any cross-origin redirect.
  • Like every mod, it runs with your permissions inside Claude Code. See what it hooks and calls with claude plugin validate .. Full model: docs/SECURITY.md.

Operational advice:

  • Ollama listens on 127.0.0.1 by default. OLLAMA_HOST=0.0.0.0 exposes it to your network, and rmod notes that in /rmod-status.
  • LM Studio's "Serve on Local Network" toggle does the same. Keep it off unless you need it, and turn on its API token if you do.
  • Install mods only from sources you trust, and keep Claude Code up to date (claude update).

Troubleshooting

  • A backend shows ✗ not listening (no answer in time) although it's running again. A request to it never finished, and Claude Code can't cancel one. rmod keeps one request open per backend at most, so it waits. Run /reload-plugins, or restart the session.
  • /rmod-status shows ⚠ model list: unreadable response. The server answered with something rmod couldn't read. The last good list is kept, and /rmod-refresh tries again.

Tested with

Claude Code 2.1.296, Ollama 0.40.2 (live), LM Studio (response shapes from its docs; live check pending), macOS on Apple Silicon. Final versions are recorded in CHANGELOG.md at each release.

License

MIT © 2026 Radoslaw Szpila

Docs

Spec · Architecture · Plan · Testing · Security · Server APIs · References

Source 20 files
hooks/register.tsx 694 lines
1import type { EngineInterface, Register, Timer, ToolCallResult, TurnStepResult } from 'claude-code'
2import { atom, read, update } from 'claude-code'
3
4import { backendForModelStep, classifyToolCall, targetsOf } from './lib/attribution'
5import { bandLine } from './lib/band'
6import { applyInventory, inventoryDue, ollamaInventory, sameInventory } from './lib/inventory'
7import type { InventoryResult } from './lib/inventory'
8import { defaultConfig, parseConfig } from './lib/config'
9import type { Config } from './lib/config'
10import { applyProbe, initialStatus, lmstudioProbe, ollamaProbe } from './lib/probe'
11import type { LmStudioStep, ProbeOutcome, ProbeResult } from './lib/probe'
12import { isBackgroundBash, isFirstOutput, stepSucceeded, toolOutcome, usageOf } from './lib/requests'
13import { asStats, begin, cancel, emptyStats, end, reset, summarize } from './lib/stats'
14import { statusText } from './lib/status-text'
15import { freshMemory, transitionToast } from './lib/transitions'
16import type { ToastMemory } from './lib/transitions'
17import { parseJson } from './lib/models'
18import { parseTags } from './lib/parse-ollama'
19import { paneSection } from './lib/pane'
20import { buildUrl } from './lib/url-guard'
21import { bandTree } from './ui/band'
22import { paneTree } from './ui/pane'
23import type { BackendStatus, Health, LmStudioEndpoint, ModelInfo, Stats } from '../types'
24
25// Static analysis lets `$` reach only functions declared at the top of this file, and lets `read` /
26// `update` take only atoms defined in this file. So the work and the atoms live here, and the
27// module's mutable state is passed as `ctx`.
28
29const ollamaStatus = atom({ plugin: 'rmod', key: 'ollama' } as const, initialStatus())
30const lmstudioStatus = atom({ plugin: 'rmod', key: 'lmstudio' } as const, initialStatus())
31const statsOllama = atom({ plugin: 'rmod', key: 'statsOllama' } as const, emptyStats())
32const statsLmstudio = atom({ plugin: 'rmod', key: 'statsLmstudio' } as const, emptyStats())
33/** FR-3.6: drawn from ctx.bandHidden; this atom only makes the band and pane redraw when it changes. */
34const bandHidden = atom({ plugin: 'rmod', key: 'bandHidden' } as const, false)
35/** Which LM Studio endpoint answered this session (FR-2.2); null until known. */
36const lmstudioEndpoint = atom({ plugin: 'rmod', key: 'lmstudioEndpoint' } as const, null)
37
38type Backend = 'ollama' | 'lmstudio'
39
40type Ctx = {
41  config: Config
42  /** Kept out of Config so config can be logged or compared without leaking it. */
43  token: string | undefined
44  poll: Timer | undefined
45  /**
46   * No overlapping probes per backend (NFR-3). `pending`: a probe is running. `inflight`: fetches
47   * still running, including ones whose probe already timed out, since `$.http.fetch` can't be
48   * cancelled. A tick is skipped while either is set, so a hung server holds one request at most.
49   */
50  pending: { ollama: boolean; lmstudio: boolean }
51  inflight: { ollama: number; lmstudio: number }
52  /** FR-2.3: the token goes out only after a token-less probe to lmstudioUrl asked for auth. */
53  lmAskedForAuth: boolean
54  /** FR-3.4: what the person was last told about each backend. */
55  toasts: { ollama: ToastMemory; lmstudio: ToastMemory }
56  /** FR-5.1: when the Ollama inventory was last fetched, and whether /rmod-refresh asked for it now. */
57  inventoryAt: number | undefined
58  forceInventory: boolean
59  /** FR-4.1a: ANTHROPIC_BASE_URL, read at session.start. */
60  baseUrl: string | undefined
61  /** FR-4.1a: model ids per backend, cached as inventories are written so a model step reads no state. */
62  models: { ollama: readonly ModelInfo[]; lmstudio: readonly ModelInfo[] }
63  /**
64   * Stats writes in order, never awaited by a tool or model hook (NFR-2): begin and end land in the
65   * order they were made without delaying the request.
66   */
67  statsQueue: Promise<void>
68  /** Writes queued and not yet settled; past MAX_PENDING_WRITES new ones are dropped. */
69  pendingWrites: number
70  /**
71   * FR-3.6: the person's band choice, loaded from $.store at session.start (showBand is only the
72   * default). Kept here because /clear resets $.state: the band reads this, the tick re-seeds the atom.
73   */
74  bandHidden: boolean
75  /** Band-setting writes in order, so a double press leaves $.store with the last choice. */
76  bandQueue: Promise<void>
77  /** FR-3.7: `claude -p` or the SDK: panes are "placed" there but nothing draws them. */
78  headless: boolean
79  /** A refresh is scheduled or ran less than MIN_REFRESH_MS ago. */
80  refreshScheduled: boolean
81}
82
83const PANE_ID = 'rmod'
84/** The shortest gap between on-demand refreshes, the same floor as the poll interval (SECURITY.md §5). */
85const MIN_REFRESH_MS = 2000
86const STORE_BAND_HIDDEN = 'bandHidden'
87
88/** FR-5.1: re-read the inventories now, from a timer (the probes' sleeps must not run in a hook). */
89const requestRefresh = ($: EngineInterface, ctx: Ctx): void => {
90  ctx.forceInventory = true
91  // NFR-3: at most one on-demand refresh per MIN_REFRESH_MS, however fast the key repeats.
92  if (ctx.refreshScheduled) return
93  ctx.refreshScheduled = true
94  try {
95    $.clock.after(0, () => {
96      tick($, ctx)
97      $.clock.after(MIN_REFRESH_MS, () => {
98        ctx.refreshScheduled = false
99      })
100    })
101  } catch {
102    ctx.refreshScheduled = false
103  }
104}
105
106/** FR-4.5: clear the counters, keeping requests still running. */
107const resetCounters = ($: EngineInterface, ctx: Ctx): void => {
108  record($, ctx, 'ollama', async () => reset)
109  record($, ctx, 'lmstudio', async () => reset)
110}
111
112/** FR-3.6: flip the band, remember it across sessions, redraw the band and the pane. */
113const toggleBand = ($: EngineInterface, ctx: Ctx): void => {
114  ctx.bandHidden = !ctx.bandHidden
115  const hidden = ctx.bandHidden
116  ctx.bandQueue = ctx.bandQueue.then(async () => {
117    try {
118      await update($, bandHidden, () => hidden)
119    } catch {
120      // The band still reads ctx; the next tick re-seeds the atom.
121    }
122    // An equal-value write (right after /clear) may not redraw: ask for one.
123    try {
124      $.ui.invalidate('ui.render')
125    } catch {
126      // The next state write redraws.
127    }
128    try {
129      await $.store.set(STORE_BAND_HIDDEN, hidden)
130    } catch {
131      $.ui.log('rmod: could not save the band setting')
132    }
133  }).catch(() => {})
134}
135
136const MAX_PENDING_WRITES = 100
137const WRITE_TIMEOUT_MS = 2000
138
139/**
140 * Queues one stats change for a backend. `change` may await the clock reads the request hook
141 * started without awaiting, so no request waits on rmod. Each write is raced against a timer, so
142 * one that never settles can't stall the queue. Failures are dropped: counting never affects a request.
143 */
144const record = (
145  $: EngineInterface,
146  ctx: Ctx,
147  backend: Backend,
148  change: () => Promise<(s: Stats) => Stats>,
149  optional = false,
150): boolean => {
151  // Only an optional write (a begin) may be dropped past the cap. An end or cancel always runs,
152  // and a dropped begin drops its end too (the caller checks the return value), so in-flight can't stick.
153  if (optional && ctx.pendingWrites >= MAX_PENDING_WRITES) return false
154  ctx.pendingWrites += 1
155  ctx.statsQueue = ctx.statsQueue
156    .then(async () => {
157      const apply = await change()
158      const write =
159        backend === 'ollama' ? update($, statsOllama, s => apply(asStats(s))) : update($, statsLmstudio, s => apply(asStats(s)))
160      const stop = new AbortController()
161      const timer = $.clock.sleep(WRITE_TIMEOUT_MS, { signal: stop.signal }).catch(() => {})
162      try {
163        await Promise.race([write, timer])
164      } finally {
165        stop.abort()
166      }
167    })
168    .catch(() => {})
169    .finally(() => {
170      ctx.pendingWrites -= 1
171    })
172  return true
173}
174
175/** The end of a counted tool call: a refusal takes the begin back, anything else is a sample. */
176const recordToolEnd = ($: EngineInterface, ctx: Ctx, b: Backend, t0: Promise<number | undefined>, r: ToolCallResult | undefined): void => {
177  const at = stamp($)
178  const outcome = toolOutcome(r)
179  record($, ctx, b, async () => {
180    if (outcome === 'refused') return cancel
181    const [start, endAt] = await Promise.all([t0, at])
182    const s0 = start ?? 0
183    const sample = { kind: 'tool' as const, ok: outcome === 'ok', latencyMs: (endAt ?? s0) - s0, at: endAt ?? s0 }
184    return s => end(s, sample)
185  })
186}
187
188/** A clock read started now and awaited later (inside `record`), so the request never waits on it. */
189const stamp = ($: EngineInterface): Promise<number | undefined> => $.clock.now().then(
190  t => t,
191  () => undefined,
192)
193
194/** Whether an inventory result would change nothing in `s` (so it isn't written and redrawn). */
195const unchanged = (s: BackendStatus, r: InventoryResult): boolean =>
196  'models' in r && s.modelsError === undefined && s.modelsAt !== undefined && sameInventory(s, r.models)
197
198/** Writes an inventory result unless it is what state already holds. Static analysis wants each atom named. */
199const writeInventory = async ($: EngineInterface, ctx: Ctx, backend: Backend, r: InventoryResult, at: number): Promise<void> => {
200  if ('models' in r) ctx.models[backend] = r.models.models
201  if (backend === 'ollama') {
202    if (!unchanged(await read($, ollamaStatus), r)) await update($, ollamaStatus, prev => applyInventory(prev, r, at))
203  } else if (!unchanged(await read($, lmstudioStatus), r)) {
204    await update($, lmstudioStatus, prev => applyInventory(prev, r, at))
205  }
206}
207
208/** FR-5.1: `/api/tags` + `/api/ps` when due. Two more requests per modelsRefreshSec, not per poll. */
209const refreshOllamaModels = async ($: EngineInterface, ctx: Ctx, cameUp: boolean): Promise<void> => {
210  const now = await $.clock.now()
211  if (!inventoryDue(await read($, ollamaStatus), ctx.inventoryAt, now, ctx.config.modelsRefreshMs, cameUp, ctx.forceInventory)) return
212  ctx.inventoryAt = now
213  ctx.forceInventory = false
214  const tagsUrl = buildUrl(ctx.config.ollamaUrl, '/api/tags')
215  const psUrl = buildUrl(ctx.config.ollamaUrl, '/api/ps')
216  if (tagsUrl === undefined || psUrl === undefined) return
217  const tags = await probe($, ctx, 'ollama', tagsUrl)
218  // A failed or unreadable /api/tags decides; skip /api/ps then.
219  const tagsReadable = tags.kind === 'response' && tags.status >= 200 && tags.status <= 299 && parseTags(parseJson(tags.text)) !== undefined
220  const ps: ProbeOutcome = tagsReadable ? await probe($, ctx, 'ollama', psUrl) : { kind: 'timeout' }
221  await writeInventory($, ctx, 'ollama', ollamaInventory(tags, ps), await $.clock.now())
222}
223
224/** Shows the toast a probe result calls for, if any (FR-3.4). */
225const announce = ($: EngineInterface, ctx: Ctx, backend: Backend, health: Health): void => {
226  const r = transitionToast(backend === 'ollama' ? 'Ollama' : 'LM Studio', ctx.toasts[backend], health)
227  ctx.toasts[backend] = r.memory
228  if (r.toast !== undefined) $.ui.toast(r.toast)
229}
230
231/**
232 * One loopback GET raced against the probe timeout (`$.http.fetch` has no timeout of its own).
233 * A late answer is ignored. A rejection's message is passed on only to be classified, never kept.
234 */
235const probe = async (
236  $: EngineInterface,
237  ctx: Ctx,
238  backend: Backend,
239  url: string,
240  headers?: Record<string, string>,
241): Promise<ProbeOutcome> => {
242  const t0 = await $.clock.now()
243  const stopTimer = new AbortController()
244  const timeout = $.clock.sleep(ctx.config.probeTimeoutMs, { signal: stopTimer.signal }).then(
245    (): ProbeOutcome => ({ kind: 'timeout' }),
246    (): ProbeOutcome => ({ kind: 'timeout' }),
247  )
248  ctx.inflight[backend] += 1
249  // Through a then, so even a synchronous throw from fetch lands in `.finally` and frees the slot.
250  const request = Promise.resolve()
251    .then(() => $.http.fetch(url, headers === undefined ? undefined : { headers }))
252    .then(
253      async (r): Promise<ProbeOutcome> => ({ kind: 'response', status: r.status, text: r.text, ms: (await $.clock.now()) - t0 }),
254      (err: unknown): ProbeOutcome => ({ kind: 'rejected', message: err instanceof Error ? err.message : String(err) }),
255    )
256    .finally(() => {
257      ctx.inflight[backend] -= 1
258      // The race is over once the fetch settles: stop the timeout's sleep instead of leaving it pending.
259      stopTimer.abort()
260    })
261  return Promise.race([request, timeout])
262}
263
264const probeOllama = async ($: EngineInterface, ctx: Ctx): Promise<void> => {
265  const url = buildUrl(ctx.config.ollamaUrl, '/api/version')
266  const r: ProbeResult = url === undefined ? { health: 'degraded', error: 'config' } : ollamaProbe(await probe($, ctx, 'ollama', url))
267  const at = await $.clock.now()
268  const wasUp = (await read($, ollamaStatus)).health === 'up'
269  await update($, ollamaStatus, s => applyProbe(s, r, at))
270  announce($, ctx, 'ollama', r.health)
271  if (r.health === 'up') await refreshOllamaModels($, ctx, !wasUp)
272}
273
274const askLmStudio = async ($: EngineInterface, ctx: Ctx, endpoint: Exclude<LmStudioEndpoint, null>, withToken: boolean): Promise<LmStudioStep> => {
275  const url = buildUrl(ctx.config.lmstudioUrl, endpoint === 'native' ? '/api/v1/models' : '/v1/models')
276  if (url === undefined) return { result: { health: 'degraded', error: 'config' }, next: 'done' }
277  // The token is sent only to the guarded lmstudioUrl, never to Ollama or anywhere else.
278  const headers = withToken && ctx.token !== undefined ? { Authorization: `Bearer ${ctx.token}` } : undefined
279  return lmstudioProbe(await probe($, ctx, 'lmstudio', url, headers), endpoint, headers !== undefined)
280}
281
282const probeLmStudio = async ($: EngineInterface, ctx: Ctx): Promise<void> => {
283  // A status never checked means fresh state: a new session after /clear, /resume or /branch
284  // (which reset $.state and fire no session.start). That session must be asked for auth again.
285  if ((await read($, lmstudioStatus)).checkedAt === undefined) ctx.lmAskedForAuth = false
286  const remembered = await read($, lmstudioEndpoint)
287  let endpoint = remembered ?? 'native'
288  // At most two requests per tick (NFR-3): a fallback and a token retry in the same tick would be three.
289  let requests = 1
290  let step = await askLmStudio($, ctx, endpoint, ctx.lmAskedForAuth)
291  if (step.next === 'fallback') {
292    endpoint = 'openai'
293    requests += 1
294    step = await askLmStudio($, ctx, endpoint, ctx.lmAskedForAuth)
295  }
296  if (step.next === 'retryWithToken' && ctx.token !== undefined) {
297    ctx.lmAskedForAuth = true
298    // After a fallback, the endpoint is remembered below and the next tick asks with the token.
299    if (requests < 2) step = await askLmStudio($, ctx, endpoint, true)
300  }
301  if ((step.result.health === 'up' || step.result.health === 'auth') && endpoint !== remembered) {
302    const found = endpoint
303    await update($, lmstudioEndpoint, () => found)
304  }
305  // Whatever answers on the port next must ask for auth again before it gets the token.
306  if (step.result.health === 'down') ctx.lmAskedForAuth = false
307  const at = await $.clock.now()
308  const r = step.result
309  await update($, lmstudioStatus, s => applyProbe(s, r, at))
310  announce($, ctx, 'lmstudio', r.health)
311  // FR-5.1: the probe's own response is the inventory: no extra request.
312  if (step.models !== undefined) await writeInventory($, ctx, 'lmstudio', { models: step.models }, at)
313}
314
315/** Probes one backend unless its previous probe is still pending. rmod's own failures never escape. */
316const runProbe = async ($: EngineInterface, ctx: Ctx, backend: Backend): Promise<void> => {
317  // NFR-3: one open request per backend. A fetch that never settles (it can't be cancelled) holds
318  // the backend at "no answer in time" until it does, or until the mod reloads.
319  if (ctx.pending[backend] || ctx.inflight[backend] > 0) return
320  ctx.pending[backend] = true
321  try {
322    if (backend === 'ollama') await probeOllama($, ctx)
323    else await probeLmStudio($, ctx)
324  } catch {
325    // Keep the last known state and try again next tick.
326  } finally {
327    ctx.pending[backend] = false
328  }
329}
330
331/**
332 * Keeps a disabled backend showing off (FR-1.4, FR-2.4), also after /clear has reset $.state to
333 * `unknown` without a session.start. Reads first, so an unchanged value is never rewritten.
334 */
335const keepOff = async ($: EngineInterface, ctx: Ctx): Promise<void> => {
336  try {
337    if (!ctx.config.enableOllama && (await read($, ollamaStatus)).health !== 'off') await update($, ollamaStatus, s => onOff(s, false))
338    if (!ctx.config.enableLmStudio && (await read($, lmstudioStatus)).health !== 'off') await update($, lmstudioStatus, s => onOff(s, false))
339    // FR-3.6: after /clear the atom is back at its default; put the person's choice back.
340    const hidden = ctx.bandHidden
341    if ((await read($, bandHidden)) !== hidden) await update($, bandHidden, () => hidden)
342  } catch {
343    // Try again next tick.
344  }
345}
346
347/** One poll: each enabled backend, in parallel. */
348const tick = ($: EngineInterface, ctx: Ctx): void => {
349  void keepOff($, ctx)
350  if (ctx.config.enableOllama) void runProbe($, ctx, 'ollama')
351  if (ctx.config.enableLmStudio) void runProbe($, ctx, 'lmstudio')
352  // The pane's "checked N s ago" and "last request N s ago" must age even when no probe writes
353  // state (one still pending, both backends off). The engine throttles this; it does no I/O.
354  try {
355    $.ui.invalidate('ui.render')
356  } catch {
357    // The next state write redraws.
358  }
359}
360
361/**
362 * (Re)starts the poll timer and applies on/off (FR-1.4, FR-2.4). A second session.start replaces
363 * the timer instead of adding one. The first probe runs from a timer, where `$` calls outside a
364 * hook are documented to run.
365 */
366const startPolling = async ($: EngineInterface, ctx: Ctx): Promise<void> => {
367  try {
368    ctx.poll?.cancel()
369    ctx.poll = $.clock.every(ctx.config.pollMs, () => tick($, ctx))
370    $.clock.after(0, () => tick($, ctx))
371  } catch {
372    $.ui.log('rmod: could not start polling')
373  }
374  // Not awaited: session.start must never wait on bookkeeping (a host whose state writes hang
375  // would hold the session's start). On/off is applied at read time anyway, so this only tidies state.
376  void update($, ollamaStatus, s => onOff(s, ctx.config.enableOllama)).catch(() => {})
377  void update($, lmstudioStatus, s => onOff(s, ctx.config.enableLmStudio)).catch(() => {})
378}
379
380/** The /rmod-status text (FR-3.7), also the answer to /rmod-refresh. */
381const statusOf = async ($: EngineInterface, ctx: Ctx): Promise<string> =>
382  statusText({
383    ollama: onOff(await read($, ollamaStatus), ctx.config.enableOllama),
384    lmstudio: onOff(await read($, lmstudioStatus), ctx.config.enableLmStudio),
385    ollamaUrl: ctx.config.ollamaUrl,
386    lmstudioUrl: ctx.config.lmstudioUrl,
387    ollamaExposed: ctx.config.ollamaExposed,
388    lmstudioEndpoint: await read($, lmstudioEndpoint),
389    statsOllama: asStats(await read($, statsOllama)),
390    statsLmstudio: asStats(await read($, statsLmstudio)),
391  })
392
393/** FR-1.4 / FR-2.4: a disabled backend shows off; one enabled again after an earlier load starts over. */
394const onOff = (s: BackendStatus, enabled: boolean): BackendStatus =>
395  enabled ? (s.health === 'off' ? { ...s, health: 'unknown' } : s) : { ...s, health: 'off' }
396
397export const register: Register = (on, options) => {
398  // Re-derived on every load: a change in /config reloads the module with new options.
399  // OLLAMA_HOST needs `$`, so session.start parses again with it.
400  const ctx: Ctx = {
401    config: defaultConfig(),
402    token: undefined,
403    poll: undefined,
404    pending: { ollama: false, lmstudio: false },
405    inflight: { ollama: 0, lmstudio: 0 },
406    lmAskedForAuth: false,
407    toasts: { ollama: freshMemory(), lmstudio: freshMemory() },
408    inventoryAt: undefined,
409    forceInventory: false,
410    baseUrl: undefined,
411    models: { ollama: [], lmstudio: [] },
412    statsQueue: Promise.resolve(),
413    pendingWrites: 0,
414    bandHidden: false,
415    bandQueue: Promise.resolve(),
416    headless: false,
417    refreshScheduled: false,
418  }
419  try {
420    const parsed = parseConfig(options, undefined)
421    ctx.config = parsed.config
422    ctx.token = parsed.token
423  } catch {
424    // Keep the defaults.
425  }
426
427  on('session.start', async ($, e, next) => {
428    ctx.headless = !e.isInteractive || e.surface === null
429    try {
430      const parsed = parseConfig(options, await $.env.get('OLLAMA_HOST'))
431      ctx.config = parsed.config
432      ctx.token = parsed.token
433      // FR-6.3: one line per invalid field. Warnings never carry the rejected value or the token.
434      for (const line of parsed.warnings) $.ui.log(line)
435    } catch {
436      $.ui.log('rmod: could not read OLLAMA_HOST; using configured values')
437    }
438    try {
439      ctx.baseUrl = await $.env.get('ANTHROPIC_BASE_URL')
440    } catch {
441      ctx.baseUrl = undefined
442    }
443    // FR-3.6: a saved choice wins; showBand is only the default. Anything but a boolean is ignored.
444    ctx.bandHidden = !ctx.config.showBand
445    try {
446      const saved = await $.store.get(STORE_BAND_HIDDEN)
447      if (typeof saved === 'boolean') ctx.bandHidden = saved
448    } catch {
449      // Keep the default.
450    }
451    try {
452      await $.command.register({ name: 'rmod', description: 'rmod: open the local models pane', immediate: true })
453    } catch {
454      $.ui.log('rmod: could not register /rmod')
455    }
456    // A load or reload can't have requests of its own in flight: clear any count a lost one left behind.
457    record($, ctx, 'ollama', async () => s => ({ ...s, inFlight: 0 }))
458    record($, ctx, 'lmstudio', async () => s => ({ ...s, inFlight: 0 }))
459
460    try {
461      await $.command.register({
462        name: 'rmod-status',
463        description: 'rmod: local Ollama / LM Studio status as plain text',
464        immediate: true,
465      })
466    } catch {
467      // Fixed text only: never echo the error, which could carry user data.
468      $.ui.log('rmod: could not register /rmod-status')
469    }
470    try {
471      await $.command.register({
472        name: 'rmod-reset',
473        description: 'rmod: reset this session\'s local request counters',
474        immediate: true,
475      })
476    } catch {
477      $.ui.log('rmod: could not register /rmod-reset')
478    }
479    try {
480      await $.command.register({
481        name: 'rmod-refresh',
482        description: 'rmod: re-read the local model lists now',
483        immediate: true,
484      })
485    } catch {
486      $.ui.log('rmod: could not register /rmod-refresh')
487    }
488
489    await startPolling($, ctx)
490    return next(e)
491  })
492
493  // FR-4.1a, FR-4.2: model requests to a local backend. An observer: every chunk is passed on as
494  // it came and the result is returned unchanged. A step that isn't local is handed straight on.
495  // Nothing on the request path awaits rmod: clock reads are started here and resolved in `record`.
496  on('turn.step', async function* ($, e, next) {
497    let backend: Backend | null = null
498    try {
499      backend = backendForModelStep(ctx.baseUrl, e.model, targetsOf(ctx.config), ctx.models)
500    } catch {
501      backend = null
502    }
503    if (backend === null) return yield* next(e)
504    const b = backend
505    const t0 = stamp($)
506    let first: Promise<number | undefined> | undefined
507    let result: TurnStepResult | undefined
508    let done = false
509    // Once only: from finally, or from the dispatch's abort if the stream is abandoned (FR-4.3).
510    const finish = () => {
511      if (done || !counted) return
512      done = true
513      const at = stamp($)
514      const r = result
515      record($, ctx, b, async () => {
516        const [start, firstAt, endAt] = await Promise.all([t0, first, at])
517        const s0 = start ?? 0
518        const u = usageOf(r?.usage)
519        const sample = {
520          kind: 'model' as const,
521          ok: stepSucceeded(r),
522          latencyMs: (endAt ?? s0) - s0,
523          at: endAt ?? s0,
524          ttftMs: firstAt === undefined ? undefined : firstAt - s0,
525          outTokens: u.outTokens,
526          model: u.model ?? e.model,
527        }
528        return s => end(s, sample)
529      })
530    }
531    const counted = record($, ctx, b, async () => begin, true)
532    try {
533      next.signal.addEventListener('abort', finish, { once: true })
534    } catch {
535      // No abort signal: finally still runs.
536    }
537    try {
538      const stream = next(e)
539      // Closing early makes `result` reject: that's an interrupt, already counted below.
540      stream.result.catch(() => {})
541      for await (const chunk of stream) {
542        try {
543          if (first === undefined && isFirstOutput(chunk)) first = stamp($)
544        } catch {
545          // No first-token time for this step.
546        }
547        yield chunk
548      }
549      result = await stream.result
550      return result
551    } finally {
552      try {
553        finish()
554      } catch {
555        // Counting never affects the step.
556      }
557    }
558  })
559
560  // FR-4.1b, FR-4.2: tool calls that reach a local backend. An observer: the result is returned
561  // unchanged, and only the backend's name is kept, never the command, URL or input (FR-4.6).
562  // A refused call never reached the backend, so it isn't a request (FR-4.1).
563  on('tool.call', async ($, e, next) => {
564    let backend: Backend | null = null
565    try {
566      const input = e as unknown as Readonly<Record<string, unknown>>
567      const tool = String(e.tool)
568      if (!isBackgroundBash(tool, input)) backend = classifyToolCall(tool, input, targetsOf(ctx.config), ctx.config.mcpServerNames)
569    } catch {
570      backend = null
571    }
572    if (backend === null) return next(e)
573    const b = backend
574    const t0 = stamp($)
575    const counted = record($, ctx, b, async () => begin, true)
576    let r: ToolCallResult | undefined
577    try {
578      r = await next(e)
579      return r
580    } finally {
581      try {
582        if (counted) recordToolEnd($, ctx, b, t0, r)
583      } catch {
584        // Counting never affects the call.
585      }
586    }
587  })
588
589  // FR-3.1: the band above the prompt. Reads only; it never fetches or writes (NFR-3).
590  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
591    try {
592      // Subscribes the band to the setting; the value drawn is ctx's (FR-3.6).
593      await read($, bandHidden)
594      if (e.props.hasSurvey || ctx.bandHidden) return next(e)
595      const input = {
596        // On/off is applied as it is read, so a disabled backend shows off at once after /clear.
597        ollama: onOff(await read($, ollamaStatus), ctx.config.enableOllama),
598        lmstudio: onOff(await read($, lmstudioStatus), ctx.config.enableLmStudio),
599        statsOllama: asStats(await read($, statsOllama)),
600        statsLmstudio: asStats(await read($, statsLmstudio)),
601      }
602      // Nothing to watch and nothing counted: leave the row to the engine.
603      if (!ctx.config.enableOllama && !ctx.config.enableLmStudio && summarize(input.statsOllama, input.statsLmstudio).total === 0) return next(e)
604      const ours = bandTree($.ui.resolve(e), bandLine(input, e.props.bodyColumns))
605      const { Box } = $.ui.resolve(e)
606      // A tree here replaces what the mods after rmod draw: keep theirs, below rmod's row.
607      // `next` is called once, here; a failure past this point still returns what it gave.
608      const theirs = await next(e).catch(() => undefined)
609      return theirs === undefined || theirs === null ? (
610        ours
611      ) : (
612        <Box flexDirection="column">
613          {ours}
614          {theirs}
615        </Box>
616      )
617    } catch {
618      return next(e)
619    }
620  })
621
622  // FR-3.5: the pane. Reads only; the buttons' handlers do the work (NFR-3).
623  on('ui.render', { component: 'Pane', requestId: PANE_ID }, async ($, e, next) => {
624    try {
625      await read($, bandHidden)
626      const now = await $.clock.now()
627      const sections = [
628        paneSection('Ollama', onOff(await read($, ollamaStatus), ctx.config.enableOllama), asStats(await read($, statsOllama)), now),
629        paneSection('LM Studio', onOff(await read($, lmstudioStatus), ctx.config.enableLmStudio), asStats(await read($, statsLmstudio)), now),
630      ]
631      return paneTree($.ui.resolve(e), sections, ctx.bandHidden, {
632        onRefresh: () => requestRefresh($, ctx),
633        onReset: () => resetCounters($, ctx),
634        onToggleBand: () => toggleBand($, ctx),
635      })
636    } catch {
637      return next(e)
638    }
639  })
640
641  on('command.run', { command: 'rmod' }, async $ => {
642    try {
643      // FR-3.7: where nothing draws, answer with the same information as text. `claude -p` places
644      // every pane without drawing it, so it is told apart by the session, not by isPlaced.
645      if (ctx.headless) return { text: await statusOf($, ctx) }
646      const placed = await $.ui.open({ id: PANE_ID, title: 'rmod · local models', focus: true, closeOnEscape: true })
647      // Nothing for the transcript (or the model) when the pane opens.
648      return placed.isPlaced ? {} : { text: await statusOf($, ctx) }
649    } catch {
650      return { text: 'rmod: could not open the pane; try /rmod-status' }
651    }
652  })
653
654  on('command.run', { command: 'rmod-status' }, async $ => {
655    try {
656      return { text: await statusOf($, ctx) }
657    } catch {
658      return { text: 'rmod: status unavailable' }
659    }
660  })
661
662  // FR-4.5: counters reset on demand. A request still running when this lands ends without going negative.
663  on('command.run', { command: 'rmod-reset' }, async $ => {
664    try {
665      resetCounters($, ctx)
666      const stop = new AbortController()
667      try {
668        await Promise.race([ctx.statsQueue, $.clock.sleep(1000, { signal: stop.signal }).catch(() => {})])
669      } finally {
670        stop.abort()
671      }
672      return { text: 'rmod: local request counters reset' }
673    } catch {
674      return { text: 'rmod: reset failed' }
675    }
676  })
677
678  // FR-5.1: inventories on demand. The probes run from a timer, not in this hook: their timeout
679  // sleeps would count against the hook's budget (three in a row at probeTimeoutMs = 10 s overruns it).
680  on('command.run', { command: 'rmod-refresh' }, async $ => {
681    try {
682      const busy = ctx.pending.ollama || ctx.pending.lmstudio || ctx.inflight.ollama > 0 || ctx.inflight.lmstudio > 0
683      requestRefresh($, ctx)
684      return {
685        text: busy
686          ? 'rmod: a check is already running; the model lists refresh with it. Run /rmod-status in a moment.'
687          : 'rmod: refreshing the model lists now. Run /rmod-status in a moment.',
688      }
689    } catch {
690      return { text: 'rmod: refresh failed' }
691    }
692  })
693}
694
hooks/lib/attribution.ts 181 lines
1// Which local backend a model request or tool call belongs to (FR-4.1). Pure: no `$`.
2//
3// Every function returns a BackendId or null and keeps nothing: the command text, URL or tool input
4// it reads is never stored, returned or logged (FR-4.6). Backends are matched by port, so every
5// loopback spelling (127.0.0.1, localhost, [::1]) counts as the same host. Ties go to Ollama
6// (same port) or to the first match in a command; a model id in both inventories counts for neither.
7
8import type { BackendId, ModelInfo } from '../../types'
9import type { Config } from './config'
10import { inlineScript, programOf, simpleCommands } from './shell'
11import { isLoopbackUrl, normalizeOllamaHost, portOf } from './url-guard'
12
13/** The ports of the enabled backends, and of disabled ones (which get nothing attributed). */
14export type Targets = { ollama?: number; lmstudio?: number; offPorts?: readonly number[] }
15
16const BACKENDS: readonly BackendId[] = ['ollama', 'lmstudio']
17/** Only this much of a command is read, to keep each tool call's bookkeeping well under NFR-2's 5 ms. */
18const MAX_SCAN_CHARS = 10_000
19const MAX_TOOL_NAME = 128
20const MAX_SCRIPT_DEPTH = 2
21
22/** Programs whose loopback host:port arguments are requests. */
23const CLIENTS: ReadonlySet<string> = new Set(['curl', 'wget', 'http', 'https', 'xh', 'httpie', 'aria2c'])
24/** Programs that take host and port as two arguments. */
25const SOCKET_CLIENTS: ReadonlySet<string> = new Set(['nc', 'ncat', 'netcat', 'telnet'])
26/** Interpreters whose inline code may hold the URL anywhere (`python3 -c "...localhost:11434..."`). */
27const INTERPRETERS: ReadonlySet<string> = new Set(['python', 'python3', 'node', 'deno', 'bun', 'ruby', 'perl'])
28/** Runs on another machine: out of scope (SPEC §5). */
29const REMOTE: ReadonlySet<string> = new Set(['ssh', 'mosh'])
30
31const HOST = String.raw`(?:127\.0\.0\.1|localhost|\[::1\]|0\.0\.0\.0)`
32/** A whole argument that is a loopback URL or host:port: `http://localhost:11434/api`, `127.0.0.1:1234`. */
33const URL_ARG = new RegExp(String.raw`^(?:https?://)?${HOST}:(\d{1,5})(?:[/?#]|$)`, 'i')
34/** A loopback host:port anywhere in a code string, not glued to a longer name or number. */
35const IN_CODE = new RegExp(String.raw`(?<![\w.-])${HOST}(?![\w.-]):(\d{1,5})(?!\d)`, 'gi')
36const LOOPBACK_HOST = new RegExp(`^${HOST}$`, 'i')
37
38export const targetsOf = (c: Pick<Config, 'enableOllama' | 'enableLmStudio' | 'ollamaUrl' | 'lmstudioUrl'>): Targets => {
39  const t: Targets = {}
40  const off: number[] = []
41  const o = portOf(c.ollamaUrl)
42  const l = portOf(c.lmstudioUrl)
43  if (o !== undefined) {
44    if (c.enableOllama) t.ollama = o
45    else off.push(o)
46  }
47  if (l !== undefined) {
48    if (c.enableLmStudio) t.lmstudio = l
49    else off.push(l)
50  }
51  if (off.length > 0) t.offPorts = off
52  return t
53}
54
55const byPort = (targets: Targets, port: number): BackendId | null => BACKENDS.find(b => targets[b] === port) ?? null
56
57const backendOfArgs = (program: string, args: readonly string[], targets: Targets): BackendId | null => {
58  if (CLIENTS.has(program)) {
59    for (const a of args) {
60      const m = URL_ARG.exec(a)
61      const b = m === null ? null : byPort(targets, Number(m[1]))
62      if (b !== null) return b
63    }
64  } else if (SOCKET_CLIENTS.has(program)) {
65    for (let i = 0; i + 1 < args.length; i += 1) {
66      if (LOOPBACK_HOST.test(args[i] ?? '')) {
67        const b = byPort(targets, Number(args[i + 1]))
68        if (b !== null) return b
69      }
70    }
71  } else if (INTERPRETERS.has(program)) {
72    for (const a of args) {
73      for (const m of a.matchAll(IN_CODE)) {
74        const b = byPort(targets, Number(m[1]))
75        if (b !== null) return b
76      }
77    }
78  }
79  return null
80}
81
82const backendOfCommand = (command: string, targets: Targets, depth: number): BackendId | null => {
83  for (const words of simpleCommands(command.slice(0, MAX_SCAN_CHARS))) {
84    const p = programOf(words)
85    if (p === undefined || REMOTE.has(p.program)) continue
86    const script = inlineScript(p.program, p.args)
87    if (script !== undefined) {
88      const b = depth < MAX_SCRIPT_DEPTH ? backendOfCommand(script, targets, depth + 1) : null
89      if (b !== null) return b
90      continue
91    }
92    const verb = p.args.find(a => !a.startsWith('-'))
93    // `OLLAMA_HOST=1.2.3.4 ollama run …` talks to another machine: out of scope (SPEC §5).
94    const host = p.env.get('OLLAMA_HOST')
95    const remote = host !== undefined && host.trim() !== '' && normalizeOllamaHost(host) === undefined
96    if (p.program === 'ollama' && verb === 'run' && targets.ollama !== undefined && !remote) return 'ollama'
97    if (p.program === 'lms' && verb === 'chat' && targets.lmstudio !== undefined) return 'lmstudio'
98    const b = backendOfArgs(p.program, p.args, targets)
99    if (b !== null) return b
100  }
101  return null
102}
103
104const backendOfUrl = (url: string, targets: Targets): BackendId | null => {
105  const port = portOf(url)
106  return port === undefined ? null : byPort(targets, port)
107}
108
109/** The server part of `mcp__<server>__<tool>`, lower-cased; undefined when there is no tool part. */
110const mcpServerOf = (tool: string): string | undefined => {
111  const rest = tool.slice('mcp__'.length)
112  const cut = rest.indexOf('__')
113  return cut <= 0 ? undefined : rest.slice(0, cut).toLowerCase()
114}
115
116/**
117 * The backend an MCP tool belongs to: the one its server name mentions, else the only enabled
118 * backend, else none. Called only for a server already found local.
119 */
120export const backendOfMcpTool = (tool: string, targets: Targets): BackendId | null => {
121  const server = tool.slice('mcp__'.length).split('__')[0] ?? ''
122  const named: BackendId | null = /ollama/i.test(server) ? 'ollama' : /lm[-_]?studio/i.test(server) ? 'lmstudio' : null
123  if (named !== null) return targets[named] !== undefined ? named : null
124  const enabled = BACKENDS.filter(b => targets[b] !== undefined)
125  return enabled.length === 1 ? (enabled[0] ?? null) : null
126}
127
128/**
129 * FR-4.1b: a Bash command that runs `ollama run` / `lms chat`, or sends a network client to a
130 * backend's loopback port (not one that only mentions it in quotes, comments or heredocs); WebFetch to
131 * a backend URL; MCP tools whose server name contains one of `mcpServerNames`. `input` is the tool
132 * call's arguments (`e` itself).
133 */
134export const classifyToolCall = (
135  tool: string,
136  input: Readonly<Record<string, unknown>>,
137  targets: Targets,
138  mcpServerNames: readonly string[],
139): BackendId | null => {
140  if (tool === 'Bash') {
141    const command = input['command']
142    return typeof command === 'string' ? backendOfCommand(command, targets, 0) : null
143  }
144  if (tool === 'WebFetch') {
145    const url = input['url']
146    return typeof url === 'string' ? backendOfUrl(url, targets) : null
147  }
148  if (tool.startsWith('mcp__') && tool.length <= MAX_TOOL_NAME) {
149    const server = mcpServerOf(tool)
150    if (server !== undefined && mcpServerNames.some(w => server.includes(w))) return backendOfMcpTool(tool, targets)
151  }
152  return null
153}
154
155/** Ollama treats `name` and `name:latest` as the same model. */
156const sameModel = (b: BackendId, a: string, z: string): boolean =>
157  a === z || (b === 'ollama' && (a.includes(':') ? a : `${a}:latest`) === (z.includes(':') ? z : `${z}:latest`))
158
159/**
160 * FR-4.1a: a loopback `ANTHROPIC_BASE_URL` on an enabled backend's port attributes every step there.
161 * The model id decides only when there is no base URL, or it is loopback on a port that is no
162 * backend's (a local proxy): then an id found in exactly one enabled inventory picks that backend.
163 * A remote base URL, or one on a disabled backend's port, attributes nothing.
164 */
165export const backendForModelStep = (
166  baseUrl: string | undefined,
167  model: string,
168  targets: Targets,
169  inventories: { readonly [B in BackendId]: readonly ModelInfo[] },
170): BackendId | null => {
171  if (baseUrl !== undefined && baseUrl.trim() !== '') {
172    if (!isLoopbackUrl(baseUrl)) return null
173    const port = portOf(baseUrl)
174    if (port === undefined || targets.offPorts?.includes(port) === true) return null
175    const b = byPort(targets, port)
176    if (b !== null) return b
177  }
178  const owners = BACKENDS.filter(b => targets[b] !== undefined && inventories[b].some(m => sameModel(b, m.id, model)))
179  return owners.length === 1 ? (owners[0] ?? null) : null
180}
181
hooks/lib/band.ts 103 lines
1// The status band's content (FR-3.1, FR-3.2, FR-3.8, FR-5.2, NFR-5). Pure: no `$`; the drawing is in hooks/ui/band.tsx.
2//
3// Every state is a glyph plus a word; `tone` only adds colour. The line is built from the most
4// detailed variant down until one fits `bodyColumns`. Below COMPACT_BELOW columns it collapses
5// to icons and numbers.
6
7import type { BackendStatus, Health, Stats } from '../../types'
8import { summarize } from './stats'
9
10export const COMPACT_BELOW = 60
11
12export type Tone = 'ok' | 'bad' | 'warn' | 'dim' | 'plain'
13/** `shrink`: the part that is cut first if the row is still too wide (wide glyphs in some terminals). */
14export type Segment = { text: string; tone: Tone; shrink?: true }
15
16export type BandInput = {
17  ollama: BackendStatus
18  lmstudio: BackendStatus
19  statsOllama: Stats
20  statsLmstudio: Stats
21}
22
23const GLYPH: { readonly [H in Health]: string } = { up: '✓', auth: '⚠', degraded: '⚠', down: '✗', unknown: '…', off: '–' }
24const WORD: { readonly [H in Health]: string } = { up: '', auth: ' auth', degraded: ' degraded', down: '', unknown: '', off: ' off' }
25export const TONE: { readonly [H in Health]: Tone } = { up: 'ok', auth: 'warn', degraded: 'warn', down: 'bad', unknown: 'dim', off: 'dim' }
26
27export const formatMs = (ms: number): string =>
28  ms < 1000 ? `${Math.round(ms)}ms` : ms < 10_000 ? `${(ms / 1000).toFixed(1)}s` : `${Math.round(ms / 1000)}s`
29
30type Detail = { version: boolean; models: boolean; avg: boolean; p95: boolean; tokPerSec: boolean }
31
32const FULL: Detail = { version: true, models: true, avg: true, p95: true, tokPerSec: true }
33/** Detail dropped in this order until the line fits. */
34const VARIANTS: readonly Detail[] = [
35  FULL,
36  { ...FULL, tokPerSec: false },
37  { ...FULL, tokPerSec: false, p95: false },
38  { ...FULL, tokPerSec: false, p95: false, version: false },
39  { version: false, models: false, avg: true, p95: false, tokPerSec: false },
40  { version: false, models: false, avg: false, p95: false, tokPerSec: false },
41]
42
43/** `N models (M loaded)` once the inventory is known; just `N models` when loaded state is unknown. */
44export const modelCount = (s: BackendStatus): string | undefined => {
45  if (s.modelsAt === undefined || s.health === 'off' || s.health === 'down') return undefined
46  const n = s.models.length + s.modelsTruncated
47  const unknown = s.models.some(m => m.loaded === null)
48  const loaded = s.models.filter(m => m.loaded === true).length
49  const label = `${n} model${n === 1 ? '' : 's'}`
50  return unknown ? label : `${label} (${loaded} loaded)`
51}
52
53const backend = (name: string, s: BackendStatus, d: Detail): Segment => {
54  let text = `${GLYPH[s.health]} ${name}${WORD[s.health]}`
55  if (d.version && s.health === 'up' && s.version !== undefined) text += ` ${s.version}`
56  // Only while up: under auth, degraded or down the last list may be stale (FR-5.4).
57  const models = d.models && s.health === 'up' ? modelCount(s) : undefined
58  if (models !== undefined) text += ` · ${models}${s.modelsError === undefined ? '' : ' ⚠'}`
59  return { text, tone: TONE[s.health] }
60}
61
62const summary = (i: BandInput, d: Detail): string | undefined => {
63  const sum = summarize(i.statsOllama, i.statsLmstudio)
64  if (sum.total === 0 && sum.inFlight === 0) return undefined
65  let text = `local reqs ${sum.total}${sum.inFlight > 0 ? ` (⟳${sum.inFlight})` : ''}`
66  if (d.avg && sum.avgMs !== undefined) text += ` · avg ${formatMs(sum.avgMs)}`
67  if (d.p95 && sum.p95Ms !== undefined) text += ` · p95 ${formatMs(sum.p95Ms)}`
68  if (d.tokPerSec && sum.tokPerSec !== undefined) text += ` · ${Math.round(sum.tokPerSec)} tok/s`
69  return text
70}
71
72const full = (i: BandInput, d: Detail): Segment[] => {
73  const line: Segment[] = [backend('Ollama', i.ollama, d), { text: '   ', tone: 'plain' }, backend('LM Studio', i.lmstudio, d)]
74  const sum = summary(i, d)
75  if (sum !== undefined) line.push({ text: `   │ ${sum}`, tone: 'plain', shrink: true })
76  return line
77}
78
79const compact = (i: BandInput, columns: number): Segment[] => {
80  const head: Segment[] = [
81    { text: `${GLYPH[i.ollama.health]}O`, tone: TONE[i.ollama.health] },
82    { text: ' ', tone: 'plain' },
83    { text: `${GLYPH[i.lmstudio.health]}L`, tone: TONE[i.lmstudio.health] },
84  ]
85  const sum = summarize(i.statsOllama, i.statsLmstudio)
86  const tails = sum.total === 0 && sum.inFlight === 0 ? [] : [` · ${sum.total} req${sum.avgMs === undefined ? '' : ` · ${formatMs(sum.avgMs)}`}`, ` · ${sum.total} req`]
87  for (const tail of tails) if (5 + tail.length <= columns) return [...head, { text: tail, tone: 'plain', shrink: true }]
88  return head
89}
90
91const width = (line: readonly Segment[]): number => line.reduce((n, s) => n + s.text.length, 0)
92
93/** The band's segments for a body `columns` wide. */
94export const bandLine = (i: BandInput, columns: number): Segment[] => {
95  if (columns >= COMPACT_BELOW) {
96    for (const d of VARIANTS) {
97      const line = full(i, d)
98      if (width(line) <= columns) return line
99    }
100  }
101  return compact(i, columns)
102}
103
hooks/lib/inventory.ts 53 lines
1// Model inventories in state (FR-5.1, FR-5.4). Pure: no `$`; register.tsx does the fetches.
2
3import type { BackendStatus, ProbeError } from '../../types'
4import type { ModelList } from './models'
5import { parseJson } from './models'
6import { mergeOllama } from './parse-ollama'
7import { classifyRejection } from './probe'
8import type { ProbeOutcome } from './probe'
9
10export type InventoryResult = { models: ModelList } | { error: ProbeError }
11
12/** Why a body could not be read, or the body when it can be (2xx). */
13const bodyOf = (o: ProbeOutcome): { text: string } | { error: ProbeError } => {
14  if (o.kind === 'timeout') return { error: 'timeout' }
15  if (o.kind === 'rejected') return { error: classifyRejection(o.message) }
16  return o.status >= 200 && o.status <= 299 ? { text: o.text } : { error: 'http' }
17}
18
19/** `/api/tags` decides; `/api/ps` failing only makes loaded state unknown. */
20export const ollamaInventory = (tags: ProbeOutcome, ps: ProbeOutcome): InventoryResult => {
21  const t = bodyOf(tags)
22  if ('error' in t) return t
23  const p = bodyOf(ps)
24  const list = mergeOllama(parseJson(t.text), 'text' in p ? parseJson(p.text) : undefined)
25  return list === undefined ? { error: 'parse' } : { models: list }
26}
27
28/** The status with a new inventory, or with the error recorded and the last good list kept (FR-5.4). */
29export const applyInventory = (prev: BackendStatus, r: InventoryResult, at: number): BackendStatus => {
30  if ('error' in r) return { ...prev, modelsError: r.error }
31  const { modelsError: _dropped, ...rest } = prev
32  return { ...rest, models: r.models.models, modelsTruncated: r.models.truncated, modelsAt: at }
33}
34
35/**
36 * Whether to fetch the inventory now: asked for, just came up, never tried, state fresh after
37 * /clear (no list and no error), or `refreshMs` since the last attempt. After a failed attempt it
38 * waits for the interval rather than retrying on every probe.
39 */
40export const inventoryDue = (
41  s: BackendStatus,
42  lastAttempt: number | undefined,
43  now: number,
44  refreshMs: number,
45  cameUp: boolean,
46  force: boolean,
47): boolean =>
48  force || cameUp || lastAttempt === undefined || (s.modelsAt === undefined && s.modelsError === undefined) || now - lastAttempt >= refreshMs
49
50/** True when `list` is what the status already holds, so state isn't rewritten (and redrawn) for nothing. */
51export const sameInventory = (s: BackendStatus, list: ModelList): boolean =>
52  s.modelsTruncated === list.truncated && JSON.stringify(s.models) === JSON.stringify(list.models)
53
hooks/lib/config.ts 179 lines
1// Parses and validates the mod's userConfig options (FR-6.2, FR-6.3). Pure: no `$`.
2//
3// Every invalid field falls back to its default and adds one warning line that
4// names the field and the rule, never the rejected value (it could be a secret
5// pasted into the wrong field). The LM Studio token is returned apart from
6// Config, so Config can be logged or compared without leaking it.
7
8import { guardBaseUrl, normalizeOllamaHost, portOf } from './url-guard'
9
10export const DEFAULT_OLLAMA_URL = 'http://127.0.0.1:11434'
11export const DEFAULT_LMSTUDIO_URL = 'http://127.0.0.1:1234'
12/**
13 * Words that mark an MCP server as local when its name (the part between `mcp__` and the next `__`)
14 * contains one, in any letter case. Literal substrings, not a regular expression: matching costs
15 * linear time whatever the person configures, so a tool call can never be held up (NFR-2).
16 */
17export const DEFAULT_MCP_SERVER_NAMES = 'ollama,lmstudio,lm-studio,lm_studio'
18
19const MAX_TOKEN_LENGTH = 4096
20const MAX_SERVER_NAMES = 10
21/** Letters, digits, `-` and `_`: what an MCP server name can hold. */
22const SERVER_WORD = /^[a-z0-9_-]{1,40}$/
23
24export type Config = {
25  enableOllama: boolean
26  enableLmStudio: boolean
27  showBand: boolean
28  ollamaUrl: string
29  /** Where ollamaUrl came from: the user's setting, OLLAMA_HOST, or the built-in default. */
30  ollamaSource: 'config' | 'env' | 'default'
31  /** OLLAMA_HOST binds all interfaces (0.0.0.0); rmod still probes 127.0.0.1. */
32  ollamaExposed: boolean
33  lmstudioUrl: string
34  pollMs: number
35  probeTimeoutMs: number
36  modelsRefreshMs: number
37  /** Lower-case words; an MCP server whose name contains one is local (FR-4.1b). */
38  mcpServerNames: readonly string[]
39}
40
41export type ParsedConfig = {
42  config: Config
43  /** The LM Studio API token, trimmed; undefined when unset or unusable. */
44  token: string | undefined
45  warnings: string[]
46}
47
48type Options = Readonly<Record<string, unknown>>
49
50const bool = (o: Options, key: string, fallback: boolean, warn: (m: string) => void): boolean => {
51  const v = o[key]
52  if (v === undefined) return fallback
53  if (typeof v === 'boolean') return v
54  warn(`rmod: ${key} must be true or false; using ${String(fallback)}`)
55  return fallback
56}
57
58const num = (o: Options, key: string, min: number, max: number, fallback: number, warn: (m: string) => void): number => {
59  const v = o[key]
60  if (v === undefined) return fallback
61  if (typeof v === 'number' && Number.isFinite(v) && v >= min && v <= max) return v
62  warn(`rmod: ${key} must be a number from ${min} to ${max}; using ${fallback}`)
63  return fallback
64}
65
66/** A configured string, trimmed; undefined when unset or blank (blank is never an error). */
67const str = (o: Options, key: string, warn: (m: string) => void): string | undefined => {
68  const v = o[key]
69  if (v === undefined) return undefined
70  if (typeof v !== 'string') {
71    warn(`rmod: ${key} must be text; using the default`)
72    return undefined
73  }
74  const t = v.trim()
75  return t === '' ? undefined : t
76}
77
78/** A configured base URL: `base` when valid, `rejected` when set but not loopback; neither when blank. */
79const urlField = (o: Options, key: string, fallback: string, warn: (m: string) => void): { base?: string; rejected: boolean } => {
80  const raw = str(o, key, warn)
81  if (raw === undefined) return { rejected: false }
82  const g = guardBaseUrl(raw)
83  if (g.ok) return { base: g.base, rejected: false }
84  warn(`rmod: ${key} rejected (${g.reason}); using ${fallback}`)
85  return { rejected: true }
86}
87
88const serverNames = (o: Options, warn: (m: string) => void): readonly string[] => {
89  const fallback = DEFAULT_MCP_SERVER_NAMES.split(',')
90  const raw = str(o, 'localMcpServerNames', warn)
91  if (raw === undefined) return fallback
92  const words = raw
93    .toLowerCase()
94    .split(',')
95    .map(w => w.trim())
96    .filter(w => w !== '')
97  if (words.length === 0 || words.length > MAX_SERVER_NAMES || !words.every(w => SERVER_WORD.test(w))) {
98    warn(`rmod: localMcpServerNames must be up to ${MAX_SERVER_NAMES} comma-separated words of letters, digits, - or _; using the default`)
99    return fallback
100  }
101  return words
102}
103
104const tokenOf = (o: Options, warn: (m: string) => void): string | undefined => {
105  const v = o['lmstudioApiToken']
106  if (typeof v !== 'string') return undefined
107  const t = v.trim()
108  if (t === '') return undefined
109  if (t.length > MAX_TOKEN_LENGTH || !/^[\x21-\x7e]+$/.test(t)) {
110    warn('rmod: lmstudioApiToken is not usable (too long, or not printable ASCII); sending no token')
111    return undefined
112  }
113  return t
114}
115
116/** The config rmod uses when nothing else can be read: the manifest's defaults. */
117export const defaultConfig = (): Config => parseConfig({}, undefined).config
118
119/**
120 * @param options what register() received (manifest defaults filled in by the loader)
121 * @param ollamaHostEnv the OLLAMA_HOST environment value, if any
122 */
123export const parseConfig = (options: Options, ollamaHostEnv: string | undefined): ParsedConfig => {
124  const warnings: string[] = []
125  const warn = (m: string) => {
126    warnings.push(m)
127  }
128
129  const enableOllama = bool(options, 'enableOllama', true, warn)
130  const enableLmStudio = bool(options, 'enableLmStudio', true, warn)
131  const showBand = bool(options, 'showBand', true, warn)
132
133  let ollamaUrl = DEFAULT_OLLAMA_URL
134  let ollamaSource: Config['ollamaSource'] = 'default'
135  let ollamaExposed = false
136  const ollamaField = urlField(options, 'ollamaUrl', DEFAULT_OLLAMA_URL, warn)
137  if (ollamaField.base !== undefined) {
138    ollamaUrl = ollamaField.base
139    ollamaSource = 'config'
140  } else if (!ollamaField.rejected && ollamaHostEnv !== undefined && ollamaHostEnv.trim() !== '') {
141    const env = normalizeOllamaHost(ollamaHostEnv)
142    if (env !== undefined) {
143      ollamaUrl = env.url
144      ollamaSource = 'env'
145      ollamaExposed = env.exposed
146    } else {
147      warn(`rmod: OLLAMA_HOST is not a loopback address; probing ${DEFAULT_OLLAMA_URL}`)
148    }
149  }
150
151  const lmField = urlField(options, 'lmstudioUrl', DEFAULT_LMSTUDIO_URL, warn)
152  const lmstudioUrl = lmField.base ?? DEFAULT_LMSTUDIO_URL
153  if (enableOllama && enableLmStudio && portOf(ollamaUrl) === portOf(lmstudioUrl)) {
154    warn('rmod: ollamaUrl and lmstudioUrl use the same port; requests to it count as Ollama')
155  }
156
157  const config: Config = {
158    enableOllama,
159    enableLmStudio,
160    showBand,
161    ollamaUrl,
162    ollamaSource,
163    ollamaExposed,
164    lmstudioUrl,
165    pollMs: num(options, 'pollIntervalSec', 2, 300, 5, warn) * 1000,
166    probeTimeoutMs: num(options, 'probeTimeoutMs', 200, 10000, 1500, warn),
167    modelsRefreshMs: num(options, 'modelsRefreshSec', 10, 3600, 30, warn) * 1000,
168    mcpServerNames: serverNames(options, warn),
169  }
170
171  let token = tokenOf(options, warn)
172  if (token !== undefined && lmField.rejected) {
173    // The token was meant for the server the user named, not whatever answers on the default port.
174    warn('rmod: lmstudioApiToken not sent because lmstudioUrl was rejected')
175    token = undefined
176  }
177  return { config, token, warnings }
178}
179
hooks/lib/probe.ts 75 lines
1// Classifies health probes (FR-1, FR-2). Pure: no `$`; register.tsx does the fetch and the timeout race.
2//
3// Server text and rejection messages are untrusted and may hold URLs: they are only classified,
4// never kept. What reaches state is rmod's own Health and ProbeError words.
5
6import type { BackendStatus, Health, LmStudioEndpoint, ProbeError } from '../../types'
7import { parseJson } from './models'
8import type { ModelList } from './models'
9import { parseVersion } from './parse-ollama'
10import { parseOpenAiModels, parseV1Models } from './parse-lmstudio'
11
12/** What one probe came to: a response (any status), a rejected fetch, or no answer in time. */
13export type ProbeOutcome =
14  | { kind: 'response'; status: number; text: string; ms: number }
15  | { kind: 'rejected'; message: string }
16  | { kind: 'timeout' }
17
18export type ProbeResult = { health: Health; error?: ProbeError; version?: string; probeMs?: number }
19
20export const initialStatus = (): BackendStatus => ({ health: 'unknown', models: [], modelsTruncated: 0 })
21
22/** A rejected fetch: an organization's web-fetch policy, or nothing reachable there. */
23export const classifyRejection = (message: string): ProbeError => (/polic/i.test(message) ? 'policy' : 'refused')
24
25const notListening = (o: Exclude<ProbeOutcome, { kind: 'response' }>): ProbeResult =>
26  o.kind === 'timeout' ? { health: 'down', error: 'timeout' } : { health: 'down', error: classifyRejection(o.message) }
27
28/** Any HTTP answer means listening (SPEC §2): 2xx up, 401/403 auth required, anything else degraded. */
29const byStatus = (status: number, probeMs: number): ProbeResult | undefined => {
30  if (status === 401 || status === 403) return { health: 'auth', probeMs }
31  if (status < 200 || status > 299) return { health: 'degraded', error: 'http', probeMs }
32  return undefined
33}
34
35/** `GET /api/version` (FR-1.2, FR-1.3). */
36export const ollamaProbe = (o: ProbeOutcome): ProbeResult => {
37  if (o.kind !== 'response') return notListening(o)
38  const bad = byStatus(o.status, o.ms)
39  if (bad !== undefined) return bad
40  const version = parseVersion(parseJson(o.text))
41  return version === undefined ? { health: 'degraded', error: 'parse', probeMs: o.ms } : { health: 'up', version, probeMs: o.ms }
42}
43
44export type LmStudioStep = {
45  result: ProbeResult
46  /** `fallback`: try `/v1/models` (FR-2.2). `retryWithToken`: ask again with the token, if one is set (FR-2.3). */
47  next: 'done' | 'fallback' | 'retryWithToken'
48  /** The models the response listed, when it was up and readable (FR-5.1: the probe is the inventory). */
49  models?: ModelList
50}
51
52/** `GET /api/v1/models` (native) or `GET /v1/models` (fallback) (FR-2). */
53export const lmstudioProbe = (o: ProbeOutcome, endpoint: Exclude<LmStudioEndpoint, null>, tokenSent: boolean): LmStudioStep => {
54  if (o.kind !== 'response') return { result: notListening(o), next: 'done' }
55  if (o.status === 404 && endpoint === 'native') return { result: { health: 'degraded', error: 'http', probeMs: o.ms }, next: 'fallback' }
56  const bad = byStatus(o.status, o.ms)
57  if (bad !== undefined) return { result: bad, next: bad.health === 'auth' && !tokenSent ? 'retryWithToken' : 'done' }
58  const body = parseJson(o.text)
59  const list = endpoint === 'native' ? parseV1Models(body) : parseOpenAiModels(body)
60  return list === undefined
61    ? { result: { health: 'degraded', error: 'parse', probeMs: o.ms }, next: 'done' }
62    : { result: { health: 'up', probeMs: o.ms }, next: 'done', models: list }
63}
64
65/** The new status after a probe: health fields replaced (stale version/error dropped), models kept. */
66export const applyProbe = (prev: BackendStatus, r: ProbeResult, at: number): BackendStatus => {
67  const next: BackendStatus = { health: r.health, checkedAt: at, models: prev.models, modelsTruncated: prev.modelsTruncated }
68  if (r.version !== undefined) next.version = r.version
69  if (r.probeMs !== undefined) next.probeMs = r.probeMs
70  if (r.error !== undefined) next.error = r.error
71  if (prev.modelsAt !== undefined) next.modelsAt = prev.modelsAt
72  if (prev.modelsError !== undefined) next.modelsError = prev.modelsError
73  return next
74}
75
hooks/lib/requests.ts 34 lines
1// Reading model steps and tool results for the request counter (FR-4.2, FR-4.6). Pure: no `$`.
2// Only kinds, counts and the model id are read. Never text, input or output.
3
4/** A chunk that is the model's own output (TTFT ends here), not the engine's envelope or a stop marker. */
5export const isFirstOutput = (chunk: { readonly kind: string }): boolean =>
6  chunk.kind === 'text' || chunk.kind === 'thinking' || chunk.kind === 'tool' || chunk.kind === 'input'
7
8/** Output tokens and the answering model from a step's `usage`, when present and sane. */
9export const usageOf = (usage: unknown): { outTokens?: number; model?: string } => {
10  if (typeof usage !== 'object' || usage === null) return {}
11  const u = usage as Readonly<Record<string, unknown>>
12  const out: { outTokens?: number; model?: string } = {}
13  const t = u['output_tokens']
14  if (typeof t === 'number' && Number.isSafeInteger(t) && t >= 0) out.outTokens = t
15  const m = u['model']
16  if (typeof m === 'string' && m !== '') out.model = m
17  return out
18}
19
20/** What a tool call came to: refused (never reached the backend: not a request), failed, or ok. */
21export const toolOutcome = (r: unknown): 'refused' | 'failed' | 'ok' => {
22  if (typeof r !== 'object' || r === null) return 'failed'
23  const o = r as Readonly<Record<string, unknown>>
24  if (typeof o['deny'] === 'string') return 'refused'
25  return o['isError'] === true ? 'failed' : 'ok'
26}
27
28/** A step with no stop reason never got a response (failed or interrupted before a message). */
29export const stepSucceeded = (r: { readonly stopReason: unknown } | undefined): boolean => r !== undefined && r.stopReason !== null
30
31/** Bash run in the background returns at once: its latency isn't the request's, so it isn't counted. */
32export const isBackgroundBash = (tool: string, input: Readonly<Record<string, unknown>>): boolean =>
33  tool === 'Bash' && input['run_in_background'] === true
34
hooks/lib/stats.ts 218 lines
1// Request statistics per backend (FR-4.2, FR-4.3, FR-4.6, NFR-3). Pure: no `$`. Every update returns a new Stats.
2//
3// Speed figures (latency, TTFT, tok/s) come from successful requests only (SPEC FR-4.2): a backend
4// that fails fast must not look quick. Every request still counts in total, ok/failed and lastAt.
5
6import type { Stats } from '../../types'
7import { cleanId } from './models'
8
9export const MAX_SAMPLES = 200
10export const MAX_BY_MODEL = 50
11const MAX_LATENCY_MS = 24 * 60 * 60 * 1000
12const MAX_TOKENS = 10_000_000
13/** Below this, the time after the first token is too short to measure a rate from. */
14const MIN_GEN_MS = 50
15
16/** One finished request: metadata only, never prompt, output or command text (FR-4.6). */
17export type Sample = {
18  kind: 'model' | 'tool'
19  /** False for an error, an abort or a tool error. */
20  ok: boolean
21  latencyMs: number
22  /** Clock time when it finished. */
23  at: number
24  /** Time to the first streamed chunk (model requests only). */
25  ttftMs?: number
26  /** Output tokens from `usage` (model requests only). */
27  outTokens?: number
28  model?: string
29}
30
31export type Summary = {
32  total: number
33  model: number
34  tool: number
35  inFlight: number
36  ok: number
37  failed: number
38  avgMs?: number
39  p50Ms?: number
40  p95Ms?: number
41  lastMs?: number
42  avgTtftMs?: number
43  tokPerSec?: number
44  lastAt?: number
45}
46
47export const emptyStats = (): Stats => ({
48  total: 0,
49  model: 0,
50  tool: 0,
51  inFlight: 0,
52  ok: 0,
53  failed: 0,
54  latMs: [],
55  ttftMs: [],
56  outTokens: 0,
57  rateTokens: 0,
58  genMs: 0,
59  byModel: {},
60  byModelOther: 0,
61})
62
63const count = (v: unknown): number => (typeof v === 'number' && Number.isSafeInteger(v) && v >= 0 ? v : 0)
64
65const samples = (v: unknown): number[] =>
66  Array.isArray(v)
67    ? v.filter((n): n is number => typeof n === 'number' && Number.isFinite(n) && n >= 0).slice(-MAX_SAMPLES)
68    : []
69
70/**
71 * A full, bounded Stats from whatever `$.state` holds: after a hot reload it may be a value from an
72 * older contract, or partial. Missing or malformed fields become their empty values.
73 */
74export const asStats = (v: unknown): Stats => {
75  if (typeof v !== 'object' || v === null) return emptyStats()
76  const o = v as Readonly<Record<string, unknown>>
77  const counts = new Map<string, number>()
78  const rawByModel = o['byModel']
79  if (typeof rawByModel === 'object' && rawByModel !== null) {
80    for (const [k, n] of Object.entries(rawByModel)) {
81      if (counts.size >= MAX_BY_MODEL) break
82      if (count(n) > 0) counts.set(k, count(n))
83    }
84  }
85  const s: Stats = {
86    total: count(o['total']),
87    model: count(o['model']),
88    tool: count(o['tool']),
89    inFlight: count(o['inFlight']),
90    ok: count(o['ok']),
91    failed: count(o['failed']),
92    latMs: samples(o['latMs']),
93    ttftMs: samples(o['ttftMs']),
94    outTokens: count(o['outTokens']),
95    rateTokens: count(o['rateTokens']),
96    genMs: count(o['genMs']),
97    byModel: Object.fromEntries(counts),
98    byModelOther: count(o['byModelOther']),
99  }
100  const lastAt = o['lastAt']
101  if (typeof lastAt === 'number' && Number.isFinite(lastAt)) s.lastAt = lastAt
102  return s
103}
104
105export const begin = (s: Stats): Stats => ({ ...s, inFlight: s.inFlight + 1 })
106
107/** Takes back a `begin` for something that turned out not to be a request (a refused tool call). */
108export const cancel = (s: Stats): Stats => ({ ...s, inFlight: Math.max(0, s.inFlight - 1) })
109
110/** Clears the counters but keeps requests still running, so the band doesn't hide live work. */
111export const reset = (s: Stats): Stats => ({ ...emptyStats(), inFlight: s.inFlight })
112
113const ms = (v: unknown): number =>
114  typeof v === 'number' && Number.isFinite(v) ? Math.min(Math.max(v, 0), MAX_LATENCY_MS) : 0
115
116const tokens = (v: unknown): number | undefined =>
117  typeof v === 'number' && Number.isSafeInteger(v) && v >= 0 && v <= MAX_TOKENS ? v : undefined
118
119const push = (list: readonly number[], v: number): number[] => [...list.slice(-(MAX_SAMPLES - 1)), v]
120
121/** How many requests this session sent for `model` (0 when unknown or counted in byModelOther). */
122export const requestsFor = (s: Stats, model: string): number =>
123  Object.hasOwn(s.byModel, model) ? (s.byModel[model] ?? 0) : 0
124
125const countModel = (s: Stats, raw: unknown): Pick<Stats, 'byModel' | 'byModelOther'> => {
126  const id = cleanId(raw)
127  if (id === undefined) return { byModel: s.byModel, byModelOther: s.byModelOther }
128  // A Map, then Object.fromEntries: untrusted ids such as `__proto__` stay plain own keys.
129  const counts = new Map(Object.entries(s.byModel))
130  if (!counts.has(id) && counts.size >= MAX_BY_MODEL) return { byModel: s.byModel, byModelOther: s.byModelOther + 1 }
131  counts.set(id, (counts.get(id) ?? 0) + 1)
132  return { byModel: Object.fromEntries(counts), byModelOther: s.byModelOther }
133}
134
135/** The part of a successful model request that feeds tok/s, or nothing when it can't be measured. */
136const rate = (latency: number, ttft: number | undefined, out: number | undefined): { tokens: number; ms: number } | undefined => {
137  if (out === undefined) return undefined
138  const afterFirst = ttft === undefined ? latency : latency - ttft
139  // One chunk, or a first token at the very end: measure over the whole request instead.
140  const span = afterFirst >= Math.max(MIN_GEN_MS, latency * 0.1) ? afterFirst : latency
141  return span > 0 ? { tokens: out, ms: span } : undefined
142}
143
144export const end = (s: Stats, sample: Sample): Stats => {
145  const isModel = sample.kind === 'model'
146  const latency = ms(sample.latencyMs)
147  const t = isModel ? sample.ttftMs : undefined
148  // Ignored unless it is a real time that fits inside the request: a first token can't come after the end.
149  const ttft = typeof t === 'number' && Number.isFinite(t) && t >= 0 && t <= latency ? t : undefined
150  const out = isModel ? tokens(sample.outTokens) : undefined
151  const measured = sample.ok ? rate(latency, ttft, out) : undefined
152
153  // Only known fields are copied: nothing else on `sample` (text of any kind) reaches state (FR-4.6).
154  const next: Stats = {
155    ...s,
156    total: s.total + 1,
157    model: s.model + (isModel ? 1 : 0),
158    tool: s.tool + (isModel ? 0 : 1),
159    inFlight: Math.max(0, s.inFlight - 1),
160    ok: s.ok + (sample.ok ? 1 : 0),
161    failed: s.failed + (sample.ok ? 0 : 1),
162    latMs: sample.ok ? push(s.latMs, latency) : s.latMs,
163    ttftMs: sample.ok && ttft !== undefined ? push(s.ttftMs, ttft) : s.ttftMs,
164    outTokens: s.outTokens + (out ?? 0),
165    rateTokens: s.rateTokens + (measured?.tokens ?? 0),
166    genMs: s.genMs + (measured?.ms ?? 0),
167    ...(sample.model === undefined ? {} : countModel(s, sample.model)),
168  }
169  if (typeof sample.at === 'number' && Number.isFinite(sample.at)) next.lastAt = sample.at
170  return next
171}
172
173const nearestRank = (sorted: readonly number[], p: number): number | undefined => {
174  if (sorted.length === 0) return undefined
175  return sorted[Math.min(sorted.length, Math.max(1, Math.ceil((p / 100) * sorted.length))) - 1]
176}
177
178/** Nearest-rank percentile; undefined for no samples. */
179export const percentile = (list: readonly number[], p: number): number | undefined =>
180  nearestRank([...list].sort((a, b) => a - b), p)
181
182const mean = (list: readonly number[]): number | undefined =>
183  list.length === 0 ? undefined : list.reduce((a, b) => a + b, 0) / list.length
184
185const sum = (all: readonly Stats[], pick: (s: Stats) => number): number => all.reduce((n, s) => n + pick(s), 0)
186
187/** Display figures for one backend, or for several combined (the band's total). Unknowns are left out. */
188export const summarize = (...all: readonly Stats[]): Summary => {
189  const lat = all.flatMap(s => s.latMs).sort((a, b) => a - b)
190  const genMs = sum(all, s => s.genMs)
191  // The backend that finished a request most recently gives lastAt; lastMs is its newest successful latency.
192  const newest = all.reduce<Stats | undefined>(
193    (best, s) => (s.lastAt !== undefined && (best?.lastAt === undefined || s.lastAt > best.lastAt) ? s : best),
194    undefined,
195  )
196  const lastMs = newest?.latMs.at(-1) ?? [...all].reverse().find(s => s.latMs.length > 0)?.latMs.at(-1)
197
198  const out: Summary = {
199    total: sum(all, s => s.total),
200    model: sum(all, s => s.model),
201    tool: sum(all, s => s.tool),
202    inFlight: sum(all, s => s.inFlight),
203    ok: sum(all, s => s.ok),
204    failed: sum(all, s => s.failed),
205  }
206  const set = <K extends keyof Summary>(k: K, v: Summary[K] | undefined) => {
207    if (v !== undefined) out[k] = v
208  }
209  set('avgMs', mean(lat))
210  set('p50Ms', nearestRank(lat, 50))
211  set('p95Ms', nearestRank(lat, 95))
212  set('lastMs', lastMs)
213  set('avgTtftMs', mean(all.flatMap(s => s.ttftMs)))
214  set('tokPerSec', genMs > 0 ? (sum(all, s => s.rateTokens) * 1000) / genMs : undefined)
215  set('lastAt', newest?.lastAt)
216  return out
217}
218
hooks/lib/status-text.ts 67 lines
1// The plain-text status that /rmod-status returns (FR-3.7, NFR-5). Pure: no `$`.
2// Glyph + word for every state, so the meaning never depends on colour.
3
4import type { BackendStatus, Health, LmStudioEndpoint, ProbeError, Stats } from '../../types'
5import { formatMs, modelCount } from './band'
6import { summarize } from './stats'
7
8export const HEALTH_TEXT: { readonly [H in Health]: string } = {
9  up: '✓ up',
10  auth: '⚠ up, auth required',
11  degraded: '⚠ degraded',
12  down: '✗ not listening',
13  unknown: '… checking',
14  off: '– off',
15}
16
17export const ERROR_TEXT: { readonly [E in ProbeError]: string } = {
18  refused: 'connection refused',
19  timeout: 'no answer in time',
20  policy: 'blocked by policy',
21  http: 'unexpected HTTP status',
22  parse: 'unreadable response',
23  config: 'invalid URL',
24}
25
26const line = (name: string, s: BackendStatus, url: string, extra: readonly string[]): string => {
27  const parts = [`${name}: ${HEALTH_TEXT[s.health]}${s.error === undefined || s.health === 'off' ? '' : ` (${ERROR_TEXT[s.error]})`}`]
28  if (s.health === 'up' && s.version !== undefined) parts.push(`v${s.version}`)
29  if (s.health !== 'down' && s.health !== 'off' && s.probeMs !== undefined) parts.push(`${Math.round(s.probeMs)} ms`)
30  const models = modelCount(s)
31  if (models !== undefined) parts.push(models)
32  if (s.modelsError !== undefined && s.health !== 'down' && s.health !== 'off') parts.push('⚠ model list: unreadable response')
33  if (s.health !== 'off') parts.push(url)
34  return [...parts, ...extra].join(' · ')
35}
36
37export type StatusInput = {
38  ollama: BackendStatus
39  lmstudio: BackendStatus
40  ollamaUrl: string
41  lmstudioUrl: string
42  ollamaExposed: boolean
43  lmstudioEndpoint: LmStudioEndpoint
44  statsOllama: Stats
45  statsLmstudio: Stats
46}
47
48/** The band's session summary as a line, or nothing before the first local request. */
49const sessionLine = (i: StatusInput): string[] => {
50  const sum = summarize(i.statsOllama, i.statsLmstudio)
51  if (sum.total === 0 && sum.inFlight === 0) return []
52  const parts = [`Session: ${sum.total} local requests (${sum.model} model, ${sum.tool} tool)`, `${sum.ok} ok, ${sum.failed} failed`]
53  if (sum.inFlight > 0) parts.push(`${sum.inFlight} in flight`)
54  if (sum.avgMs !== undefined) parts.push(`avg ${formatMs(sum.avgMs)}`)
55  if (sum.p95Ms !== undefined) parts.push(`p95 ${formatMs(sum.p95Ms)}`)
56  if (sum.tokPerSec !== undefined) parts.push(`${Math.round(sum.tokPerSec)} tok/s`)
57  return [parts.join(' · ')]
58}
59
60export const statusText = (i: StatusInput): string =>
61  [
62    'rmod',
63    line('Ollama', i.ollama, i.ollamaUrl, i.ollamaExposed ? ['OLLAMA_HOST listens on all interfaces'] : []),
64    line('LM Studio', i.lmstudio, i.lmstudioUrl, i.lmstudioEndpoint === 'openai' ? ['OpenAI-compatible API (older LM Studio)'] : []),
65    ...sessionLine(i),
66  ].join('\n')
67
hooks/lib/transitions.ts 38 lines
1// Which health changes deserve a toast (FR-3.4). Pure: no `$`.
2//
3// A toast says what the person was last told has changed. "Stopped listening" needs two probes in
4// a row that found the backend down, so one slow answer (a timeout while a model loads) toasts
5// nothing. "Listening again" follows only a backend the person knows to be down. The first probe
6// sets what the person is told without a toast.
7
8import type { Health } from '../../types'
9
10const listening = (h: Health): boolean => h === 'up' || h === 'auth' || h === 'degraded'
11
12/** Probes in a row that must find a backend down before "stopped listening" is shown. */
13export const DOWN_CONFIRMATIONS = 2
14
15export type ToastMemory = {
16  /** What the person was last told (or saw at the first probe); undefined before any probe. */
17  shown: 'listening' | 'down' | undefined
18  /** Probes in a row that found the backend down. */
19  downStreak: number
20}
21
22export const freshMemory = (): ToastMemory => ({ shown: undefined, downStreak: 0 })
23
24/** The memory after a probe found `health`, and the toast to show for it, if any. */
25export const transitionToast = (name: string, m: ToastMemory, health: Health): { memory: ToastMemory; toast?: string } => {
26  if (health === 'unknown' || health === 'off') return { memory: m }
27  if (listening(health)) {
28    const memory: ToastMemory = { shown: 'listening', downStreak: 0 }
29    return m.shown === 'down' ? { memory, toast: `✓ ${name} is listening again` } : { memory }
30  }
31  const downStreak = m.downStreak + 1
32  if (m.shown === undefined) return { memory: { shown: 'down', downStreak } }
33  if (m.shown === 'listening' && downStreak >= DOWN_CONFIRMATIONS) {
34    return { memory: { shown: 'down', downStreak }, toast: `✗ ${name} stopped listening` }
35  }
36  return { memory: { shown: m.shown, downStreak } }
37}
38
hooks/lib/models.ts 100 lines
1// Shared helpers for model inventories from untrusted servers (FR-5, NFR-1, NFR-3). Pure: no `$`.
2
3import type { ModelInfo } from '../../types'
4
5export const MAX_MODELS = 200
6/** How many raw entries a parser looks at; the rest are only counted. */
7export const MAX_SCAN = 5000
8const MAX_BODY_CHARS = 4_000_000
9const MAX_ID_CHARS = 200
10const MAX_SHORT_CHARS = 32
11
12export type ModelList = { models: ModelInfo[]; truncated: number }
13
14/** JSON.parse that never throws; undefined for oversized or malformed bodies. */
15export const parseJson = (text: string): unknown => {
16  if (text.length > MAX_BODY_CHARS) return undefined
17  try {
18    return JSON.parse(text) as unknown
19  } catch {
20    return undefined
21  }
22}
23
24export const isRecord = (v: unknown): v is Readonly<Record<string, unknown>> =>
25  typeof v === 'object' && v !== null && !Array.isArray(v)
26
27/**
28 * Characters that could repaint, reorder or hide text: controls (Cc), format characters such as
29 * bidi overrides, zero-width and tag characters (Cf), line/paragraph separators (Zl, Zp),
30 * private-use (Co) and lone surrogates (Cs).
31 */
32const UNSAFE = /[\p{Cc}\p{Cf}\p{Zl}\p{Zp}\p{Co}\p{Cs}]/gu
33
34/** A display-safe string: unsafe characters removed, trimmed, cut to `max` code points; undefined when empty. */
35export const cleanText = (v: unknown, max: number): string | undefined => {
36  if (typeof v !== 'string') return undefined
37  // Bound the work first (a code point is at most 2 UTF-16 units), then cut on code points.
38  const t = [...v.slice(0, max * 4).replace(UNSAFE, '').trim()].slice(0, max).join('').trim()
39  return t === '' ? undefined : t
40}
41
42export const cleanId = (v: unknown): string | undefined => cleanText(v, MAX_ID_CHARS)
43
44/** A short label (parameters, quantization). Too long means garbage, so it is dropped, not cut. */
45export const cleanShort = (v: unknown): string | undefined => {
46  if (typeof v !== 'string' || v.length > MAX_SHORT_CHARS * 2) return undefined
47  const t = cleanText(v, MAX_SHORT_CHARS + 1)
48  return t === undefined || t.length > MAX_SHORT_CHARS ? undefined : t
49}
50
51export const cleanBytes = (v: unknown): number | undefined =>
52  typeof v === 'number' && Number.isSafeInteger(v) && v >= 0 ? v : undefined
53
54/** Builds a ModelInfo without undefined-valued keys, so results compare and serialise cleanly. */
55export const modelInfo = (id: string, loaded: boolean | null, extra: Omit<ModelInfo, 'id' | 'loaded'>): ModelInfo => {
56  const m: ModelInfo = { id, loaded }
57  if (extra.sizeBytes !== undefined) m.sizeBytes = extra.sizeBytes
58  if (extra.params !== undefined) m.params = extra.params
59  if (extra.quant !== undefined) m.quant = extra.quant
60  if (extra.kind !== undefined) m.kind = extra.kind
61  if (extra.vramBytes !== undefined) m.vramBytes = extra.vramBytes
62  if (extra.expiresAt !== undefined) m.expiresAt = extra.expiresAt
63  return m
64}
65
66/** An ISO 8601 timestamp as epoch ms; undefined unless it parses to a finite time. */
67export const cleanTime = (v: unknown): number | undefined => {
68  if (typeof v !== 'string' || v.length > 64) return undefined
69  // Ollama sends nanoseconds (`.000000000Z`); Date.parse is only specified for milliseconds.
70  const t = Date.parse(v.replace(/(\.\d{3})\d+/, '$1'))
71  return Number.isFinite(t) ? t : undefined
72}
73
74const rank = (m: ModelInfo): number => (m.loaded === true ? 0 : m.loaded === null ? 1 : 2)
75
76/**
77 * Dedupes by id (the most-loaded copy wins), sorts loaded first then by id, and caps at MAX_MODELS.
78 * `unscanned` counts raw entries a parser never looked at; they count as truncated, so past
79 * MAX_SCAN entries the `+N more` figure is an upper bound (it includes invalid and duplicate entries).
80 */
81export const finalize = (entries: Iterable<ModelInfo>, unscanned = 0): ModelList => {
82  const byId = new Map<string, ModelInfo>()
83  for (const m of entries) {
84    // On a duplicate id keep the entry with the better loaded state (loaded > unknown > unloaded).
85    const seen = byId.get(m.id)
86    if (seen === undefined || rank(m) < rank(seen)) byId.set(m.id, m)
87  }
88  const all = [...byId.values()].sort((a, b) => rank(a) - rank(b) || (a.id < b.id ? -1 : a.id > b.id ? 1 : 0))
89  return { models: all.slice(0, MAX_MODELS), truncated: Math.max(0, all.length - MAX_MODELS) + unscanned }
90}
91
92/** The array under `key` of a record body, or undefined when the body is not of that shape. */
93export const listAt = (body: unknown, key: string): { items: readonly unknown[]; unscanned: number } | undefined => {
94  if (!isRecord(body)) return undefined
95  const list = body[key]
96  if (!Array.isArray(list)) return undefined
97  const items: readonly unknown[] = list
98  return { items: items.slice(0, MAX_SCAN), unscanned: Math.max(0, items.length - MAX_SCAN) }
99}
100
hooks/lib/parse-ollama.ts 104 lines
1// Ollama response parsers (FR-1.3, FR-5.1, FR-5.4). Pure: no `$`. Input is untrusted JSON.
2// Shapes: docs/LOCAL_SERVERS.md (verified against Ollama 0.40.2).
3
4import type { ModelInfo } from '../../types'
5import { cleanBytes, cleanId, cleanShort, cleanTime, finalize, isRecord, listAt, MAX_MODELS, modelInfo } from './models'
6import type { ModelList } from './models'
7
8/** `GET /api/version` → `{ version }`. Undefined unless it is a plain version string of at most 32 characters. */
9export const parseVersion = (body: unknown): string | undefined => {
10  if (!isRecord(body)) return undefined
11  const v = body['version']
12  return typeof v === 'string' && /^[0-9A-Za-z.+_-]{1,32}$/.test(v) ? v : undefined
13}
14
15const kindOf = (capabilities: unknown): ModelInfo['kind'] => {
16  if (!Array.isArray(capabilities)) return undefined
17  const caps: readonly unknown[] = capabilities
18  return caps.includes('embedding') ? 'embedding' : 'llm'
19}
20
21/** The tag entries before dedupe and cap, so a merge with `/api/ps` caps only once. */
22const tagEntries = (body: unknown): { models: ModelInfo[]; unscanned: number } | undefined => {
23  const list = listAt(body, 'models')
24  if (list === undefined) return undefined
25  const models: ModelInfo[] = []
26  for (const entry of list.items) {
27    if (!isRecord(entry)) continue
28    // Cloud models (remote_host set) are not local: out of scope (SPEC §5). The host is never stored.
29    if (typeof entry['remote_host'] === 'string') continue
30    const id = cleanId(entry['name'])
31    if (id === undefined) continue
32    const details = isRecord(entry['details']) ? entry['details'] : {}
33    models.push(
34      modelInfo(id, false, {
35        sizeBytes: cleanBytes(entry['size']),
36        // MLX models report these as "": cleanShort turns that into unknown.
37        params: cleanShort(details['parameter_size']),
38        quant: cleanShort(details['quantization_level']),
39        kind: kindOf(entry['capabilities']),
40      }),
41    )
42  }
43  return { models, unscanned: list.unscanned }
44}
45
46/** `GET /api/tags` → available models, all `loaded: false`. */
47export const parseTags = (body: unknown): ModelList | undefined => {
48  const t = tagEntries(body)
49  return t === undefined ? undefined : finalize(t.models, t.unscanned)
50}
51
52type Loaded = { sizeBytes?: number; vramBytes?: number; expiresAt?: number }
53
54/** `/api/ps` entries by name (at most MAX_MODELS), with their memory figures. */
55const psEntries = (body: unknown): Map<string, Loaded> | undefined => {
56  const list = listAt(body, 'models')
57  if (list === undefined) return undefined
58  const loaded = new Map<string, Loaded>()
59  for (const entry of list.items) {
60    if (loaded.size >= MAX_MODELS) break
61    if (!isRecord(entry)) continue
62    const id = cleanId(entry['name'])
63    if (id === undefined || loaded.has(id)) continue
64    loaded.set(id, {
65      sizeBytes: cleanBytes(entry['size']),
66      vramBytes: cleanBytes(entry['size_vram']),
67      expiresAt: cleanTime(entry['expires_at']),
68    })
69  }
70  return loaded
71}
72
73/** `GET /api/ps` → names of the models loaded in memory (at most MAX_MODELS). */
74export const parsePs = (body: unknown): string[] | undefined => {
75  const loaded = psEntries(body)
76  return loaded === undefined ? undefined : [...loaded.keys()]
77}
78
79/**
80 * The Ollama inventory from the `/api/tags` and `/api/ps` bodies. Undefined when the tags are
81 * unreadable (FR-5.4). An unreadable `/api/ps` makes every loaded state unknown. Loaded models
82 * missing from the tags are added. Loaded models sort first, so the cap never drops one.
83 */
84export const mergeOllama = (tagsBody: unknown, psBody: unknown): ModelList | undefined => {
85  const tags = tagEntries(tagsBody)
86  if (tags === undefined) return undefined
87  const loaded = psEntries(psBody)
88  if (loaded === undefined) return finalize(tags.models.map(m => ({ ...m, loaded: null })), tags.unscanned)
89  const merged: ModelInfo[] = tags.models.map(m => {
90    const ps = loaded.get(m.id)
91    // Keep the tags' disk size; /api/ps `size` is the in-memory size and is used only for ps-only models.
92    return ps === undefined ? { ...m, loaded: false } : modelInfo(m.id, true, { ...m, vramBytes: ps.vramBytes, expiresAt: ps.expiresAt })
93  })
94  const known = new Set(tags.models.map(m => m.id))
95  let psOnly = 0
96  for (const [id, ps] of loaded) {
97    if (known.has(id)) continue
98    merged.push(modelInfo(id, true, ps))
99    psOnly += 1
100  }
101  // A ps-only model may sit in the unscanned tail of the tags: don't count it twice.
102  return finalize(merged, Math.max(0, tags.unscanned - psOnly))
103}
104