SLOPSHOPPER

token-speed

主控与活跃子代理各一行实时估算速率,各代理独立按请求模型累计 API 速率

newbandcommandprocesstimeragents
★ 1v0.3.2MITupdated 2026-10-09liuyejinghong/Claude-workflow/mods/token-speed
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · token-speed
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /tok-speed ⎿ token-speed: token-speed 0.3.2 · 各代理独立累计 · 主控 first ⎿ token-speed: 各代理同一行显示模型、effort、Ctx、Live/Last/Avg(tok/s 只标一次);工作区最后一行。effort 来源:request 为 turn.step 请求,applied 为 classic 实际档 ⎿ token-speed: context registry v1 (load failed: confirmed overrides only) · 优先实际 session 窗口、明确 override、官方 default(未给默认时明确 c ⎿ token-speed: main · 主控 · ended · claude-opus-5-5 · effort — (unknown) · Ctx —/? (unknown) · Live — · idle ⎿ token-speed: Last — · Avg — · 尚无本会话统计 ⎿ token-speed: Live:最近 3 秒收到的文本、thinking、工具 JSON 参数估算;CJK 每 Unicode code point 1 token,其余 0.25。分母 min(3 秒, 请求已耗时),不足 250ms 为 ⚡ main · claude-opus-5-5 · — · Ctx —/? · Live — · Last — · Avg — tok/s ⌂ app on feat/auth-refresh ⟨Claude Code's own drawing⟩ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Band
⚡ main · claude-opus-5-5 · — · Ctx —/? · Live — · Last — · Avg — tok/s ⌂ app on feat/auth-refresh ⟨Claude Code's own drawing⟩
README

token-speed 0.3.2

适用于 Claude Code 2.1.294 的 function-hooks mod。在提示框上方,主控与每个活跃子代理各用一行显示模型、effort、上下文与输出速率,工作区放在最后一行。只观察原有请求,不调用额外模型、不引入 tokenizer 或网络请求;分支通过引擎 process API 调用本机 git。

⚡ main · gpt-6.1-sol · high · Ctx ██████▊    67%/272k · Live ~42.1 · Last 31.2 · Avg 28.6 tok/s · streaming
↳ abcdefg1 · glm-5.3-flash · max · Ctx █▍         14%/1.0M · Live ~35.2 · Last 29.1 · Avg 30.4 tok/s · streaming
⌂ Claude-workflow on main

主控始终第一,子代理按首次观察顺序每个一行,切换 view 不过滤。短 id 自动延长到可区分同会话代理;工具间隙保留行,Live 为 —,明确终态后隐藏,历史通过 /tok-speed 查询。带 survey 或 maxRows=0 时隐藏,其它插件和引擎的内容原样保留。只有主控时为代理行加工作区行,共两行;行数预算不足先丢工作区。

各列按终端显示宽度对齐,包括双宽的 ⚡、CJK 与常见 emoji。速率数字右对齐,单位 tok/s 每行只显示一次。动态宽度不足时依次省略 status、Last、缩短进度条,始终保留 effort、上下文数值、Live 与 Avg;特别窄时可缩短模型显示。渲染前计算可见内容宽度,Text 的 truncate-end 只作最后保护。

进度条是固定 10 个终端 cell 的短矩形,轨道使用 theme rate_limit_empty,整格填充 █,尾部使用 ▏▎▍▌▋▊▉ 表示 1/8 cell。共 80 个视觉档位(10 格 × 8 档),因此 0–100 的整数百分点不是每个都有不同的条形:相邻百分点可落在同一档(例如 67% 与 68% 的条形相同,但数字 67%、68% 不同)。宽屏也不扩展为长条。窄屏可省略条形,数值始终保持整数 1% 精度。填充固定使用 theme rate_limit_fill,不随占用变化,没有横向渐变;Ctx 数值用 text,未知占用用 inactive。未知占用只画轨道空条,并以 —% 或 ? 明确标为未知,不能解释为 0%。

主控上下文来自 $.session.usage(),与状态栏同口径,最近响应的 uncached + cache-read + cache-written 输入 tokens 除以窗口;首次响应或成功压缩后无读数显示 —%/窗口。已知读数不会被打回未知:缺少有效 tokens 的部分读数(窗口相同)与主控模型切换都会沿用上一次已知读数,只有成功压缩(主控)或整会话重置才清空,因此进度条不会在两次轮询之间闪回 —%。子代理输入量来自自己最近一次响应的 CLI usage,不累计多次请求。子代理分母由启动时读取一次的 data/model-contexts.json 提供,无联网刷新。优先顺序为实际运行窗口、明确 CLI override、官方默认、未知;官网仅提供容量时明确标为官方 capacity。输入始终取自己最近一次 CLI response,主控仍以 $.session.usage() 为准。数据中的官方来源、核验日期和 registry 版本可通过 /tok-speed 查看。读取或严格解析失败时,官方表为空,保留已确认的 Sol 272000、GLM-5.3/Flash 1000000 后备,不中断请求。

官方记录覆盖当前前沿型号,共 24 条(核验日期 2026-10-09)。Anthropic:claude-haiku-5-5、claude-sonnet-5-5、claude-opus-5-5、claude-fable-5-1、claude-fable-5、claude-mythos-5-1、claude-mythos-5、claude-opus-5、claude-sonnet-5、claude-opus-4-8、claude-opus-4-7、claude-opus-4-6、claude-sonnet-4-6 为共享输入输出上下文 1M,claude-haiku-4-5、claude-sonnet-4-5 为 200k。OpenAI:gpt-6.1-sol、gpt-6-astra、gpt-6-sol、gpt-6-luna 为 Codex 官方默认 272000 / max 872000,其中前两者另带 API 总上下文 1050000 与 max input 922000 元数据。glm-5.3、glm-5.3-flash、deepseek-v4.1-flash(aliases deepseek-flash、deepseek-v4-flash)、deepseek-v4-pro、gemini-3.8-flash 官网只给出容量且未声明默认/最大,按官方 capacity 记录;Gemini 的 1,048,576 是输入上限,输出上限 65,536 单独存在,不是共享窗口。用户 override 独立存放,不混作官方事实。匹配只用去渠道的 canonical 尾名及精确 aliases,不匹配子串,不把未知邻近型号映射到已知型号;不使用动态 haiku/opus/sonnet 别名,不因 [1m] 选择官方最大值。来源分别记录为 cli-input-window-config、cli-input-official-default、cli-input-official-capacity。

新增 provider/model 时,只添加具有官方 URL、核验日期和精确型号的记录;未核实的型号保持 unknown,或单独添加有理由的 local override。模型最大上下文不保证路由实际 cap,registry 不把最大值当默认。不复制主控百分比,数值超过窗口可显示超过 100%,只限制条形填充。成功安装的压缩清对应行输入读数,保留已知窗口;precompute 和 skip 不清。

Effort 来源保存在状态并可由命令查看:request 为 turn.step.effort,applied 为 classic PostToolUse/Stop 的 effort.level,configured 为主控当前模型的 modelSettings[model].effortLevel、全局 effortLevel 或明确模型后缀。reload 后主控立即读取配置后备,无数据为 —;不硬编码 high,不扫描 agent 定义。配置档位不保证外部 provider 已应用。只选取 settings 的这些 effort 字段,不保存或打印其它 settings。

工作区为仓库根目录名(非 git 使用会话根目录名),附当前分支;detached HEAD 不显示分支,约每 5 秒刷新。显示使用 theme tokens:主控标签为 claude、子代理标签与行次要信息为 inactive;模型名为 text(普通文本,不着色);有效 Live/Avg/Last 数字为 text,缺值为 inactive,速率不以 success 表示;effort 按档位着色(数值为 subtle);状态 idle/waiting/streaming 为 inactive,aborted 为 warning,error 为 error,usage unavailable 为 warning。工作区行为 inactive。颜色跟随 CLI 主题。

安装

/plugin marketplace add liuyejinghong/Claude-workflow
/plugin install token-speed@claude-workflow-mods

也可用固定 tag token-speed-v0.3.2 获取源文件,并从 checkout 加载:

git clone --branch token-speed-v0.3.2 https://github.com/liuyejinghong/Claude-workflow.git
claude --plugin-dir ./Claude-workflow/mods/token-speed

安装后的插件由 Claude Code 加载。开发时修改自己的 checkout;请勿直接修改 plugins/cache 副本。

命令

  • /tok-speed:列出各代理的短 id 角色、任务描述、当前模型、实时估算和各模型历史累计,以及统计口径。
  • /tok-speed reset:清空本会话统计。任意代理请求仍 active 时拒绝;工具间隙可 reset,保留活跃代理元数据和行。
  • /clear:清空统计、代理行与实时请求;晚返回的旧流或 spawn 元数据不复活统计,新 turn.step 重启计时器。

状态包括 waiting、streaming、idle、aborted、error、usage unavailable。新采纳的代理没有可观察 stream 时 Live 为 —,模型为 unknown,直到 step 或有明确 resolved model 的 spawn。流内停顿仍为 streaming,Live 随窗口衰减。

模型名称与独立统计

每个代理独立按 请求 turn.step.model 累计;两个相同模型代理不共享统计桶。usage.model 仅是响应事实,不改变标签或桶。真实请求模型切换建立新桶。

统计桶与命令保留已有完整渠道模型名,UI 仅隐藏最后 / 之前的渠道前缀;其它合法冒号部分不会删除。

0.1 的 host state 会迁移到主控行,其 response-model 均值保留为 [legacy v0.1 response model] 桶,在 /tok-speed 可查。由于旧桶缺乏请求渠道证据,它们不会随意合并或混入 0.2 的请求模型均值;UI Avg 只取新请求桶。

统计含义

Live 是估算值,始终带 ~。 收集最近 3 秒实际观察到的 text、thinking 和工具调用 JSON 参数字符;tool、engine、stop chunk 不增加计数。按 Unicode code point 计数:CJK Unified Ideographs(U+3400–4DBF、U+4E00–9FFF)、兼容汉字(U+F900–FAFF)及补充汉字(U+20000–323AF)每 code point 估为 1 token,其余为 0.25。不是 tokenizer 精确计数。

分母为 min(3000ms, 当前时间 - 请求开始时间)。不足 250ms 为 —;请求开始至首个 chunk 的等待也计时。仅计近 3 秒字符,停流超过 3 秒变为 ~0.0 tok/s。每 250ms 共用一次 host clock 时间和一次原子 snapshot update,批量发布所有活跃 loop;每个字符 chunk 调用一次 clock.now,不按 chunk 写 host state。

Last / Avg 是 CLI usage 的 API 全程输出速率。 每个 turn.step 从 hook 开始计时至最终结果返回,包含 TTFT、thinking、网络、请求内停顿和内置 API 工具耗时;排除两次请求之间的工具或用户空闲,非纯解码速度。

  • Last:当前请求模型最近请求的 output_tokens / 秒数。失败、中断、无有效 usage 或耗时不大于 0 为 —。
  • Avg:该代理、该请求模型最近 24 小时滚动窗口内的 有效 output_tokens 总和 / 有效请求耗时总和,不取速度算术平均,按请求完成时间戳(clock epoch ms)判窗,出窗样本即弃——窗口是相对时间,与本地时区(含 UTC+8)无关。例如窗口内 100/2s + 300/3s = 80 tok/s。迁移自 v0.2 的无时间戳桶在首个新请求前退回全程累计口径。
  • usage null、count 缺失/NaN/Infinity/负数、耗时不大于 0 均不入平均;有限非负 0 是有效样本。turn.complete 的汇总 usage 不重复累计。

CPA 上游缺失 token count 可能被 CLI 规范化为 0。mod 无法区分实际 0 和上游遗漏后的 0;API 统计表示 CLI usage 口径,整段 usage null 和仍可观察的无效 count 不计入。

生命周期与边界

统计在本 session 的 version 2 host PluginState,每 loop 单独 row,顺序稳定;本地 Map 只保存流累加器。reload 保留统计,丢弃无法接管的 stale active,以 $.agent.list() 重建在跑子代理元数据。该 API 不含模型;未观察到请求的代理显示 unknown/Live —。每 1 秒轻量 reconcile 与上下文用量轮询、每 5 秒工作区刷新在无流时也运行,值未变化不写状态;渲染不查询或写这些来源,只读 host state。每请求的完成样本保留 24 小时滚动日志支撑 Avg。失败 list 保留行;生命周期 revision guard 防止旧 list 返回覆盖刚开始/完成的新 turn。只采纳 running;waiting 和工具间隙不判完成;仅 completed/failed/killed 明确终态收尾,不根据 absence 结束新 spawn。

agent.list() 只涉及本 session 的子代理/teammate,不扫描其它 CLI 或远程 session。API 不列 workflow agent,也无法观察远程 workflow loop;有本地 turn.step 时可单独统计。无可观察 stream 的行永远不推算速度。重复 session.start 不重复创建 timer;session.end 取消 timer 清空统计,不跨会话持久化。

流 hook 只调用一次 next(e),原样 yield 所有 chunk/ref,原样返回 HookStream.result;不更改请求、顺序或工具 JSON。仪表自身异常被捕获,真实下游错误继续抛出,不重复请求。每个 loop 仅清理自己的 active/live。

开发检查

claude plugin validate <mod-dir> --strict
tsc -p <mod-dir>
claude plugin test <mod-dir>

本目录 tsconfig 开启 strict/noUncheckedIndexedAccess,包含 hooks、tests、自有 types 和生成的 .claude-plugin/types。测试使用实际 claude-code/testing 和 mock.clock 合成 stream,不联网或调用模型。

测试保留流透传、单次 next、测量/Unicode/加权平均、失败清理、reset/clear/reload、并发独立桶与生命周期竞态合同,并核验统一代理行、配置 effort 后备、主控/子代理上下文隔离、成功压缩、10 格块状条 80 档视觉映射(0–100 为 81 种可见状态,67/68 同形但数字不同)、配置子代理窗口、80/120/240 列预算、显示宽对齐、UI 隐藏渠道和命令完整渠道。kit 只验证引擎 tree 与合同,不代替终端字体的实际截图验收。

Source 3 files
hooks/register.tsx 714 lines
1import { atom, read, update } from 'claude-code'
2import { fallbackRegistry, lookupWindow, registryLoader } from './context-registry'
3import type { AgentInfo, EngineInterface, Register, ThemeKey, Timer, TurnStepChunk, TurnStepResult } from 'claude-code'
4import type { TokenSpeedActive, TokenSpeedContextInfo, TokenSpeedEffort, TokenSpeedModelStats, TokenSpeedRow,
5  TokenSpeedSample, TokenSpeedSessionInfo, TokenSpeedSnapshot, TokenSpeedStatus, TokenSpeedWorkspace } from '../types'
6
7const WINDOW_MS = 3000
8const TICK_MS = 250
9// Avg rolls over the last day; sample timestamps are clock epoch ms, so no timezone applies.
10const AVG_WINDOW_MS = 24 * 60 * 60 * 1000
11const loadRegistry = registryLoader()
12let registry = fallbackRegistry()
13function childContext(model: string | null, tokens: number | null = null): TokenSpeedContextInfo {
14  const reading = lookupWindow(registry, model === null ? null : canonical(model))
15  const window = reading.window
16  return { tokens, window, percent: tokens === null || window === 0 ? null : Math.round(tokens * 100 / window),
17    source: reading.source === 'configured' ? 'cli-input-window-config'
18      : reading.source === 'official-default' ? 'cli-input-official-default'
19        : reading.source === 'official-capacity' ? 'cli-input-official-capacity' : tokens === null ? 'unknown' : 'cli-input' }
20}
21const emptyContext = (): TokenSpeedContextInfo => ({ tokens: null, window: 0, percent: null, source: 'unknown' })
22const row = (id: string): TokenSpeedRow => ({ id, description: '', running: false, currentModel: null,
23  models: [], active: null, live: null, status: 'idle', seen: false, turnId: null, revision: 0,
24  effort: null, effortSource: 'unknown', context: emptyContext() })
25const emptySession = (): TokenSpeedSessionInfo => ({ context: emptyContext(), workspace: null })
26const empty = (): TokenSpeedSnapshot => ({ version: 2, rows: [row('main')], session: emptySession() })
27const snapshot = atom({ plugin: 'token-speed', key: 'snapshot' } as const, empty())
28
29type LiveRequest = TokenSpeedActive & {
30  loopId: string
31  epoch: number
32  chunks: { at: number; tokens: number }[]
33  hasContent: boolean
34}
35
36// Only stream accumulators and lifecycle guards are local. Drawings read host state.
37const liveRequests = new Map<string, LiveRequest>()
38let timer: Timer | null = null
39let tickBusy = false
40let listBusy = false
41let usageBusy = false
42let repoBusy = false
43let ticks = 0
44let sequence = 0
45let revision = 0
46let epoch = 0
47
48async function observe(work: () => Promise<unknown>): Promise<void> {
49  try { await work() } catch { /* Meter failures never interrupt or repeat the actual request. */ }
50}
51
52const object = (value: unknown): value is Record<string, unknown> => typeof value === 'object' && value !== null
53const finite = (value: unknown): value is number => typeof value === 'number' && Number.isFinite(value)
54const status = (value: unknown): TokenSpeedStatus => value === 'waiting' || value === 'streaming'
55  || value === 'aborted' || value === 'error' || value === 'usage unavailable' ? value : 'idle'
56function requestLog(value: unknown): TokenSpeedSample[] {
57  if (!Array.isArray(value)) return []
58  const log: TokenSpeedSample[] = []
59  for (const item of value) {
60    if (!object(item) || !finite(item.at) || !finite(item.tokens) || item.tokens < 0 || !finite(item.ms) || item.ms <= 0) continue
61    log.push({ at: item.at, tokens: item.tokens, ms: item.ms })
62  }
63  return log
64}
65function stats(value: unknown, legacy = false): TokenSpeedModelStats[] {
66  if (!Array.isArray(value)) return []
67  const buckets: TokenSpeedModelStats[] = []
68  for (const item of value) {
69    if (!object(item) || typeof item.model !== 'string' || !finite(item.outputTokens) || item.outputTokens < 0
70      || !finite(item.durationMs) || item.durationMs < 0 || !finite(item.samples) || item.samples < 0) continue
71    // A v0.2 bucket carries no log; its average falls back to the aggregates until a new request.
72    buckets.push({ model: item.model, outputTokens: item.outputTokens, durationMs: item.durationMs,
73      samples: item.samples, lastApi: finite(item.lastApi) && item.lastApi >= 0 ? item.lastApi : null,
74      ...(legacy || item.legacy === true ? { legacy: true as const }
75        : Array.isArray(item.log) ? { log: requestLog(item.log) } : {}) })
76  }
77  return buckets
78}
79function activeValue(value: unknown): TokenSpeedActive | null {
80  return object(value) && typeof value.id === 'string' && typeof value.turnId === 'string'
81    && typeof value.model === 'string' && finite(value.startedAt)
82    ? { id: value.id, turnId: value.turnId, model: value.model, startedAt: value.startedAt } : null
83}
84function effortValue(value: unknown): TokenSpeedEffort | null {
85  if (typeof value === 'string' && ['low', 'medium', 'high', 'xhigh', 'max'].includes(value)) return value as TokenSpeedEffort
86  return finite(value) && value >= 0 ? value : null
87}
88function contextInfo(value: unknown): TokenSpeedContextInfo {
89  const held = object(value) ? value : {}
90  return { tokens: finite(held.tokens) && held.tokens >= 0 ? held.tokens : null,
91    window: finite(held.window) && held.window > 0 ? held.window : 0,
92    percent: finite(held.percent) && held.percent >= 0 ? held.percent : null,
93    source: held.source === 'session' || held.source === 'cli-input' || held.source === 'cli-input-window-config'
94      || held.source === 'cli-input-official-default' || held.source === 'cli-input-official-capacity' ? held.source : 'unknown' }
95}
96function workspaceValue(value: unknown): TokenSpeedWorkspace | null {
97  if (!object(value) || typeof value.name !== 'string' || !value.name) return null
98  return { name: value.name, branch: typeof value.branch === 'string' && value.branch ? value.branch : null }
99}
100function sessionInfo(value: unknown): TokenSpeedSessionInfo {
101  const held = object(value) ? value : {}
102  return { context: contextInfo(held.context), workspace: workspaceValue(held.workspace) }
103}
104// Runtime validation also admits the v0.1 shape; response-model averages stay legacy buckets.
105function normalize(value: unknown): TokenSpeedSnapshot {
106  if (!object(value)) return empty()
107  if (value.version !== 2 || !Array.isArray(value.rows)) {
108    const main = row('main')
109    main.models = stats(value.models, true)
110    main.currentModel = typeof value.currentModel === 'string' ? value.currentModel : null
111    main.active = activeValue(value.active)
112    main.live = finite(value.live) ? value.live : null
113    main.status = status(value.status)
114    main.seen = value.seen === true || main.models.length > 0 || main.active !== null
115    main.turnId = main.active?.turnId ?? null
116    return { version: 2, rows: [main], session: sessionInfo(value.session) }
117  }
118  const rows: TokenSpeedRow[] = []
119  for (const item of value.rows) {
120    if (!object(item) || typeof item.id !== 'string' || rows.some(r => r.id === item.id)) continue
121    rows.push({ id: item.id, description: typeof item.description === 'string' ? item.description : '',
122      running: item.running === true, currentModel: typeof item.currentModel === 'string' ? item.currentModel : null,
123      models: stats(item.models), active: activeValue(item.active), live: finite(item.live) ? item.live : null,
124      status: status(item.status), seen: item.seen === true,
125      turnId: typeof item.turnId === 'string' ? item.turnId : null,
126      revision: finite(item.revision) ? item.revision : 0,
127      effort: effortValue(item.effort),
128      effortSource: item.effortSource === 'configured' || item.effortSource === 'request' || item.effortSource === 'applied'
129        ? item.effortSource : 'unknown',
130      context: item.id === 'main' ? contextInfo(item.context)
131        : childContext(typeof item.currentModel === 'string' ? item.currentModel : null, contextInfo(item.context).tokens) })
132  }
133  return { version: 2, rows: [rows.find(r => r.id === 'main') ?? row('main'), ...rows.filter(r => r.id !== 'main')],
134    session: sessionInfo(value.session) }
135}
136async function mutate($: EngineInterface, heldEpoch: number, change: (state: TokenSpeedSnapshot) => TokenSpeedSnapshot): Promise<void> {
137  await update($, snapshot, raw => heldEpoch === epoch ? change(normalize(raw)) : normalize(raw))
138}
139const readState = async ($: EngineInterface): Promise<TokenSpeedSnapshot> => normalize(await read($, snapshot))
140function changeRow(state: TokenSpeedSnapshot, id: string, change: (held: TokenSpeedRow) => TokenSpeedRow): TokenSpeedSnapshot {
141  const present = state.rows.some(r => r.id === id)
142  return { version: 2, session: state.session,
143    rows: present ? state.rows.map(r => r.id === id ? change(r) : r) : [...state.rows, change(row(id))] }
144}
145
146// Only known configuration suffixes are removed; an existing provider prefix is evidence.
147function canonical(model: string, previous: string | null = null): string {
148  let normalized = model.trim()
149  for (;;) {
150    const trimmed = normalized.replace(/(?:\[1m\]|\((?:low|medium|high|xhigh|max)\)|:(?:low|medium|high|xhigh|max))$/i, '').trim()
151    if (trimmed === normalized) break
152    normalized = trimmed
153  }
154  if (!normalized) return 'unknown'
155  if (!normalized.includes('/') && previous?.includes('/') && previous.slice(previous.lastIndexOf('/') + 1) === normalized) return previous
156  return normalized
157}
158function suffixEffort(model: string): TokenSpeedEffort | null {
159  // Only an explicit effort suffix is evidence; [1m] says nothing about effort/window.
160  const match = model.trim().replace(/\[1m\]$/i, '').match(/(?:\((low|medium|high|xhigh|max)\)|:(low|medium|high|xhigh|max))$/i)
161  return effortValue((match?.[1] ?? match?.[2])?.toLowerCase())
162}
163async function configuredEffort($: EngineInterface, model: string): Promise<TokenSpeedEffort | null> {
164  try {
165    const settings = await $.settings.read()
166    const models = object(settings.modelSettings) ? settings.modelSettings : {}
167    const selected = models[model] ?? models[canonical(model)]
168    return effortValue(object(selected) ? selected.effortLevel : undefined)
169      ?? effortValue(settings.effortLevel) ?? suffixEffort(model)
170  } catch { return suffixEffort(model) }
171}
172const sameContext = (a: TokenSpeedContextInfo, b: TokenSpeedContextInfo): boolean =>
173  a.tokens === b.tokens && a.window === b.window && a.percent === b.percent && a.source === b.source
174async function refreshMain($: EngineInterface, heldEpoch: number): Promise<void> {
175  const before = (await readState($)).rows.find(r => r.id === 'main') ?? row('main')
176  let fullModel = before.currentModel
177  try { fullModel = await $.session.model() } catch { /* Keep the observed request model when the host has none. */ }
178  if (!fullModel || heldEpoch !== epoch) return
179  const model = canonical(fullModel, before.currentModel)
180  const effort = await configuredEffort($, fullModel)
181    ?? (model === before.currentModel && before.effortSource === 'configured' ? before.effort : null)
182  const expected = before.currentModel === model && (before.effortSource === 'request' || before.effortSource === 'applied')
183    ? { effort: before.effort, source: before.effortSource } : { effort, source: effort === null ? 'unknown' : 'configured' }
184  if (before.currentModel === model && before.effort === expected.effort && before.effortSource === expected.source && before.seen) return
185  await mutate($, heldEpoch, state => changeRow(state, 'main', r => {
186    if (r.revision !== before.revision) return r
187    const changedModel = r.currentModel !== model
188    const retained = !changedModel && (r.effortSource === 'request' || r.effortSource === 'applied')
189    const value = retained ? r.effort : effort
190    const source = retained ? r.effortSource : value === null ? 'unknown' : 'configured'
191    if (!changedModel && r.effort === value && r.effortSource === source && r.seen) return r
192    return { ...r, currentModel: model, effort: value, effortSource: source, seen: true }
193  }))
194}
195const owns = (active: LiveRequest): boolean => active.epoch === epoch && liveRequests.get(active.loopId) === active
196function stopTimer(): void {
197  try { timer?.cancel() } catch { /* Disposal is observational too. */ }
198  timer = null
199  ticks = 0
200}
201
202async function reconcile($: EngineInterface, adoptExisting = false): Promise<void> {
203  if (listBusy) return
204  listBusy = true
205  const heldEpoch = epoch
206  try {
207    const before = await readState($)
208    const guards = new Map(before.rows.map(r => [r.id, r.revision]))
209    const agents = await $.agent.list() // A rejected list must not finish any row.
210    if (heldEpoch !== epoch) return
211    const ended: { id: string; requestId: string | undefined }[] = []
212    await mutate($, heldEpoch, state => {
213      ended.length = 0 // update may retry after a version collision.
214      let changed = state
215      for (const agent of agents) {
216        const held = changed.rows.find(r => r.id === agent.id)
217        if (held && held.revision !== guards.get(agent.id)) continue
218        if (agent.status === 'running') {
219          changed = changeRow(changed, agent.id, r => ({ ...r, description: agent.description,
220            running: held ? (adoptExisting ? true : r.running) : true,
221            seen: true }))
222        } else if (held && (agent.status === 'completed' || agent.status === 'failed' || agent.status === 'killed')) {
223          ended.push({ id: agent.id, requestId: held.active?.id })
224          changed = changeRow(changed, agent.id, r => ({ ...r, description: agent.description, running: false,
225            active: null, live: null, status: agent.status === 'failed' ? 'error' : agent.status === 'killed' ? 'aborted' : 'idle' }))
226        }
227      }
228      // Absence never finishes a newly spawned loop: only explicit terminal entries do.
229      return changed
230    })
231    for (const item of ended) {
232      if (liveRequests.get(item.id)?.id === item.requestId) liveRequests.delete(item.id)
233    }
234  } finally { listBusy = false }
235}
236function ensureTimer($: EngineInterface): void {
237  if (timer) return
238  timer = $.clock.every(TICK_MS, () => {
239    ticks++
240    if (ticks % 4 === 0) void observe(() => reconcile($))
241    if (ticks % 4 === 2) void observe(() => pollUsage($))
242    if (ticks % 20 === 10) void observe(() => refreshWorkspace($))
243    if (tickBusy || liveRequests.size === 0) return
244    tickBusy = true
245    void observe(async () => {
246      const active = [...liveRequests.values()]
247      const heldEpoch = epoch
248      const now = await $.clock.now() // One timestamp for every active loop in this batch.
249      const values = new Map<string, { id: string; live: number | null; status: TokenSpeedStatus }>()
250      for (const request of active) {
251        if (!owns(request)) continue
252        request.chunks = request.chunks.filter(chunk => chunk.at > now - WINDOW_MS)
253        const elapsed = Math.min(WINDOW_MS, now - request.startedAt)
254        values.set(request.loopId, { id: request.id,
255          live: elapsed < TICK_MS ? null : request.chunks.reduce((sum, chunk) => sum + chunk.tokens, 0) * 1000 / elapsed,
256          status: request.hasContent ? 'streaming' : 'waiting' })
257      }
258      await mutate($, heldEpoch, state => ({ version: 2, session: state.session, rows: state.rows.map(r => {
259        const value = values.get(r.id)
260        return value && r.active?.id === value.id ? { ...r, live: value.live, status: value.status } : r
261      }) }))
262    }).finally(() => { tickBusy = false })
263  })
264}
265async function pollUsage($: EngineInterface): Promise<void> {
266  if (usageBusy) return
267  usageBusy = true
268  const heldEpoch = epoch
269  try {
270    await observe(() => refreshMain($, heldEpoch))
271    const held = (await $.session.usage()).context
272    if (heldEpoch !== epoch || !object(held) || !finite(held.window) || held.window <= 0) return
273    const tokens = finite(held.tokens) && held.tokens >= 0 ? held.tokens : null
274    const measured = tokens !== null
275    const read: TokenSpeedContextInfo = { tokens, window: held.window,
276      percent: finite(held.percent) && held.percent >= 0 ? held.percent : null, source: 'session' }
277    const before = await readState($)
278    const current = (before.rows.find(r => r.id === 'main') ?? row('main')).context
279    // A partial session reading must not erase a known measurement for the same window.
280    const next = !measured && current.tokens !== null && current.window === held.window ? current : read
281    if (sameContext(before.session.context, next) && sameContext(current, next)) return
282    await mutate($, heldEpoch, state => {
283      const changed = changeRow(state, 'main', r => sameContext(r.context, next) ? r : { ...r, context: next })
284      return sameContext(state.session.context, next) ? changed : { ...changed, session: { ...state.session, context: next } }
285    })
286  } finally { usageBusy = false }
287}
288async function refreshWorkspace($: EngineInterface): Promise<void> {
289  if (repoBusy) return
290  repoBusy = true
291  const heldEpoch = epoch
292  try {
293    let name: string
294    let branch: string | null = null
295    const repo = await $.session.repo()
296    if (repo && repo.root) {
297      name = folderName(repo.root)
298      try {
299        const git = await $.process.run(['git', 'branch', '--show-current'], { cwd: repo.root, timeoutMs: 2000 })
300        if (git.exitCode === 0) branch = git.stdout.trim() || null // Empty output is a detached HEAD.
301      } catch { /* A missing git never hides the workspace name. */ }
302    } else {
303      name = folderName(await $.session.root())
304    }
305    if (heldEpoch !== epoch) return
306    const next: TokenSpeedWorkspace = { name, branch }
307    await mutate($, heldEpoch, state => {
308      const held = state.session.workspace
309      return held?.name === next.name && held.branch === next.branch ? state
310        : { version: 2, rows: state.rows, session: { ...state.session, workspace: next } }
311    })
312  } finally { repoBusy = false }
313}
314
315// Unicode code points, never UTF-16 units. Deliberately not a tokenizer.
316function estimate(text: string): number {
317  let tokens = 0
318  for (const char of text) {
319    const code = char.codePointAt(0)!
320    const cjk = (code >= 0x3400 && code <= 0x4dbf) || (code >= 0x4e00 && code <= 0x9fff)
321      || (code >= 0xf900 && code <= 0xfaff) || (code >= 0x20000 && code <= 0x323af)
322    tokens += cjk ? 1 : 0.25
323  }
324  return tokens
325}
326async function observeChunk($: EngineInterface, active: LiveRequest, chunk: TurnStepChunk): Promise<void> {
327  const text = chunk.kind === 'text' || chunk.kind === 'thinking' ? chunk.text : chunk.kind === 'input' ? chunk.json : ''
328  if (!text || !owns(active)) return
329  const at = await $.clock.now()
330  if (!owns(active)) return
331  active.hasContent = true
332  active.chunks = active.chunks.filter(item => item.at > at - WINDOW_MS)
333  active.chunks.push({ at, tokens: estimate(text) })
334}
335function bucketFor(r: TokenSpeedRow, model: string): TokenSpeedModelStats {
336  return r.models.find(bucket => bucket.model === model && !bucket.legacy)
337    ?? { model, outputTokens: 0, durationMs: 0, samples: 0, lastApi: null }
338}
339async function finish($: EngineInterface, active: LiveRequest, result: TurnStepResult): Promise<void> {
340  const endedAt = await $.clock.now()
341  const durationMs = endedAt - active.startedAt
342  const tokens = result.usage?.output_tokens
343  const validUsage = typeof tokens === 'number' && Number.isFinite(tokens) && tokens >= 0
344  const completed = result.stopReason !== null
345  const valid = completed && validUsage && Number.isFinite(durationMs) && durationMs > 0
346  const inputs = result.usage && [result.usage.input_tokens, result.usage.cache_read_input_tokens,
347    result.usage.cache_creation_input_tokens]
348  const inputTokens = completed && inputs && inputs.every(n => finite(n) && n >= 0)
349    ? inputs.reduce((sum, n) => sum + n, 0) : null
350  await mutate($, active.epoch, state => changeRow(state, active.loopId, r => {
351    if (r.active?.id !== active.id) return r
352    const old = bucketFor(r, active.model)
353    // The log prunes to the rolling day window; the aggregates stay all-time.
354    const bucket = valid
355      ? { ...old, log: withinWindow([...(old.log ?? []), { at: endedAt, tokens, ms: durationMs }], endedAt),
356          outputTokens: old.outputTokens + tokens, durationMs: old.durationMs + durationMs, samples: old.samples + 1,
357          lastApi: tokens * 1000 / durationMs }
358      : { ...old, lastApi: null }
359    return { ...r, models: [...r.models.filter(item => item.legacy || item.model !== active.model), bucket],
360      context: active.loopId === 'main' ? r.context
361        : childContext(active.model, inputTokens !== null && finite(inputTokens) ? inputTokens : null),
362      active: null, live: null, status: valid ? 'idle' : 'usage unavailable' }
363  }))
364}
365async function failed($: EngineInterface, active: LiveRequest, failure: TokenSpeedStatus): Promise<void> {
366  await mutate($, active.epoch, state => changeRow(state, active.loopId, r => {
367    if (r.active?.id !== active.id) return r
368    const old = bucketFor(r, active.model)
369    return { ...r, active: null, live: null, status: failure,
370      models: [...r.models.filter(bucket => bucket.legacy || bucket.model !== active.model), { ...old, lastApi: null }] }
371  }))
372}
373const rate = (value: number | null | undefined): string => value == null ? '—' : `${value.toFixed(1)} tok/s`
374const average = (tokens: number, ms: number): number | null => ms > 0 ? tokens * 1000 / ms : null
375function withinWindow(log: TokenSpeedSample[], now: number): TokenSpeedSample[] {
376  return log.filter(sample => sample.at > now - AVG_WINDOW_MS)
377}
378function windowedAverage(bucket: TokenSpeedModelStats, now: number): number | null {
379  // A migrated v0.2 bucket has no timestamps; its aggregates stand in until a new request.
380  if (bucket.log === undefined) return average(bucket.outputTokens, bucket.durationMs)
381  const held = withinWindow(bucket.log, now)
382  const tokens = held.reduce((sum, sample) => sum + sample.tokens, 0)
383  const ms = held.reduce((sum, sample) => sum + sample.ms, 0)
384  return ms > 0 ? tokens * 1000 / ms : null
385}
386async function safeNow($: EngineInterface): Promise<number> {
387  try { return await $.clock.now() } catch { return Number.MAX_SAFE_INTEGER }
388}
389const folderName = (path: string): string => {
390  const trimmed = path.replace(/\/+$/, '')
391  return trimmed.slice(trimmed.lastIndexOf('/') + 1) || path
392}
393function shortId(id: string, rows: TokenSpeedRow[]): string {
394  let length = Math.min(8, id.length)
395  while (length < id.length && rows.some(r => r.id !== id && r.id.slice(0, length) === id.slice(0, length))) length++
396  return id.slice(0, length)
397}
398const role = (r: TokenSpeedRow, rows: TokenSpeedRow[]): string => r.id === 'main' ? 'main' : `agent:${shortId(r.id, rows)}`
399const statusColors: Record<TokenSpeedStatus, ThemeKey> = {
400  waiting: 'inactive', streaming: 'inactive', idle: 'inactive',
401  aborted: 'warning', error: 'error', 'usage unavailable': 'warning',
402}
403
404async function appliedEffort($: EngineInterface, heldEpoch: number, id: string, value: unknown): Promise<void> {
405  const effort = effortValue(value)
406  if (effort === null || heldEpoch !== epoch) return
407  await mutate($, heldEpoch, state => changeRow(state, id, r => ({ ...r, effort, effortSource: 'applied' })))
408}
409const humanWindow = (window: number): string =>
410  window >= 1e6 ? `${(window / 1e6).toFixed(1)}M` : window >= 1e3 ? `${Math.round(window / 1e3)}k` : `${window}`
411const contextLabel = (context: TokenSpeedContextInfo): string => context.window > 0
412  ? `${context.percent === null ? '—' : Math.round(context.percent)}%/${humanWindow(context.window)}`
413  : context.tokens === null ? '—/?' : `${(context.tokens / 1000).toFixed(1)}k/?`
414const displayModel = (model: string | null): string => model?.slice(model.lastIndexOf('/') + 1) ?? 'unknown'
415// Terminal cells, not UTF-16 length. Combining marks/ZWJ consume no extra cell;
416// the common East Asian and emoji ranges are wide (including ⚡).
417function cellWidth(text: string): number {
418  let width = 0
419  let joined = false
420  for (const char of text) {
421    const n = char.codePointAt(0)!
422    if (n === 0x200d) { joined = true; continue }
423    if (/\p{Mark}/u.test(char) || n === 0xfe0f || n < 32 || n === 127) continue
424    if (joined) { joined = false; continue }
425    width += n >= 0x1100 && (n <= 0x115f || n === 0x2329 || n === 0x232a
426      || (n >= 0x2e80 && n <= 0xa4cf && n !== 0x303f) || (n >= 0xac00 && n <= 0xd7a3)
427      || (n >= 0xf900 && n <= 0xfaff) || (n >= 0xfe10 && n <= 0xfe19)
428      || (n >= 0xfe30 && n <= 0xfe6f) || (n >= 0xff00 && n <= 0xff60)
429      || (n >= 0xffe0 && n <= 0xffe6) || (n >= 0x1f000 && n <= 0x1faff)
430      || (n >= 0x20000 && n <= 0x3fffd)) || n === 0x26a1 ? 2 : 1
431  }
432  return width
433}
434const padEndTo = (text: string, width: number): string => text + ' '.repeat(Math.max(0, width - cellWidth(text)))
435const padStartTo = (text: string, width: number): string => ' '.repeat(Math.max(0, width - cellWidth(text))) + text
436function fitText(text: string, width: number): string {
437  if (cellWidth(text) <= width) return text
438  if (width <= 0) return ''
439  let fitted = ''
440  for (const char of text) {
441    if (cellWidth(fitted + char) > width - 1) break
442    fitted += char
443  }
444  return fitted + '…'
445}
446
447export const register: Register = on => {
448  on('session.start', async ($, e, next) => {
449    const rest = await next(e)
450    await observe(async () => {
451      registry = await loadRegistry(() => $.fs.read(`${$.plugin.root}/data/model-contexts.json`))
452      await $.command.register({ name: 'tok-speed', description: '各代理独立按请求模型统计;主控 first,reset 清空本会话统计', argumentHint: '[reset]' }).catch(() => {})
453      await mutate($, epoch, state => ({ version: 2, session: state.session, rows: state.rows.map(r => {
454        revision = Math.max(revision, r.revision)
455        return r.active && r.active.id !== liveRequests.get(r.id)?.id ? { ...r, active: null, live: null, status: 'aborted' } : r
456      }) }))
457      ensureTimer($)
458      await reconcile($, true)
459      await observe(() => refreshMain($, epoch))
460      await observe(() => refreshWorkspace($))
461    })
462    return rest
463  })
464  on('session.end', async ($, e, next) => {
465    epoch++
466    liveRequests.clear()
467    stopTimer()
468    await observe(() => update($, snapshot, empty))
469    return next(e)
470  })
471  on('agent.spawn', async ($, e, next) => {
472    const heldEpoch = epoch
473    const rest = await next(e) // Once only; metadata failure must not repeat a spawn.
474    await observe(async () => {
475      if (rest.deny !== undefined || !rest.agentId || heldEpoch !== epoch) return
476      const id = rest.agentId
477      // Remote workflow agents have no local list/loop. Do not invent rows for them.
478      const agents: AgentInfo[] = await $.agent.list()
479      const local = agents.find(agent => agent.id === id)
480      await mutate($, heldEpoch, state => {
481        const existing = state.rows.find(r => r.id === id)
482        if (!existing && local?.status !== 'running') return state
483        return changeRow(state, id, r => ({ ...r, description: e.description,
484          currentModel: r.currentModel ?? canonical(rest.model),
485          context: r.currentModel === null ? childContext(rest.model) : r.context,
486          effort: r.currentModel === null ? suffixEffort(rest.model) : r.effort,
487          effortSource: r.currentModel === null ? suffixEffort(rest.model) === null ? 'unknown' : 'configured' : r.effortSource,
488          running: existing ? r.running : true, seen: true }))
489      })
490      ensureTimer($)
491    })
492    return rest
493  })
494  on('turn.step', async function* ($, e, next) {
495    const heldEpoch = epoch
496    const loopId = e.agentId ?? 'main'
497    const startedRevision = ++revision
498    let active: LiveRequest | null = null
499    let completed = false
500    let failureStatus: TokenSpeedStatus = 'aborted'
501    await observe(async () => {
502      // Even when clock.now fails, retain loop metadata, but never invent Live.
503      let model = canonical(e.model)
504      const requested = effortValue(e.effort)
505      const fallback = requested === null
506        ? loopId === 'main' ? await configuredEffort($, e.model) : suffixEffort(e.model) : null
507      await mutate($, heldEpoch, state => changeRow(state, loopId, r => {
508        model = canonical(e.model, r.currentModel)
509        const effort = requested ?? fallback
510        return { ...r, running: true, currentModel: model, seen: true, turnId: e.turnId,
511          context: loopId === 'main' || model === r.currentModel ? r.context : childContext(model),
512          revision: startedRevision, active: null, live: null, status: 'waiting', effort,
513          effortSource: requested !== null ? 'request' : effort !== null ? 'configured' : 'unknown' }
514      }))
515      const startedAt = await $.clock.now()
516      if (heldEpoch !== epoch) return
517      const request: LiveRequest = { id: `${e.turnId}:${e.index}:${++sequence}`, turnId: e.turnId,
518        model, startedAt, loopId, epoch: heldEpoch, chunks: [], hasContent: false }
519      active = request
520      liveRequests.set(loopId, request)
521      await mutate($, heldEpoch, state => changeRow(state, loopId, r => r.revision !== startedRevision ? r : { ...r,
522        active: { id: request.id, turnId: e.turnId, model: request.model, startedAt } }))
523      ensureTimer($)
524    })
525    try {
526      const stream = next(e)
527      for await (const chunk of stream) {
528        if (active) await observe(() => observeChunk($, active!, chunk))
529        yield chunk
530      }
531      const result = await stream.result
532      if (active) {
533        if (next.signal.aborted) await observe(() => failed($, active!, 'aborted'))
534        else {
535          try { await finish($, active, result) }
536          catch { await observe(() => failed($, active!, 'usage unavailable')) }
537        }
538      }
539      completed = true
540      return result
541    } catch (error) {
542      failureStatus = next.signal.aborted || (error instanceof Error && error.name === 'AbortError') ? 'aborted' : 'error'
543      throw error
544    } finally {
545      if (active) {
546        if (owns(active)) liveRequests.delete(loopId)
547        if (!completed) await observe(() => failed($, active!, failureStatus))
548      }
549    }
550  })
551  on('turn.complete', async ($, e, next) => {
552    const heldEpoch = epoch
553    const rest = await next(e)
554    await observe(async () => {
555      const id = e.agentId ?? 'main'
556      const active = liveRequests.get(id)
557      if (active?.turnId === e.turnId) liveRequests.delete(id)
558      const endedRevision = ++revision
559      await mutate($, heldEpoch, state => {
560        const held = state.rows.find(r => r.id === id)
561        if (!held || (held.turnId !== e.turnId && !(id !== 'main' && held.turnId === null && held.running))) return state
562        return changeRow(state, id, r => ({ ...r, running: false, active: null, live: null,
563          turnId: e.turnId, revision: endedRevision,
564          status: e.reason === 'error' ? 'error' : e.reason === 'aborted' ? 'aborted' : 'idle' }))
565      })
566    })
567    // Turn usage is a turn total, never another request sample.
568    return rest
569  })
570  on('classic.PostToolUse', async ($, e, next) => {
571    const heldEpoch = epoch
572    const rest = await next(e)
573    await observe(() => appliedEffort($, heldEpoch, e.agent_id ?? 'main', e.effort?.level))
574    return rest
575  })
576  on('classic.Stop', async ($, e, next) => {
577    const heldEpoch = epoch
578    const rest = await next(e)
579    await observe(() => appliedEffort($, heldEpoch, e.agent_id ?? 'main', e.effort?.level))
580    return rest
581  })
582  on('session.compact', async ($, e, next) => {
583    const heldEpoch = epoch
584    const rest = await next(e)
585    if (e.trigger !== 'precompute' && !('skip' in rest)) await observe(() => mutate($, heldEpoch, state => {
586      const changed = changeRow(state, e.agentId ?? 'main', r => ({ ...r,
587        context: e.agentId === undefined ? emptyContext() : childContext(r.currentModel) }))
588      return e.agentId === undefined ? { ...changed, session: { ...state.session, context: emptyContext() } } : changed
589    }))
590    return rest
591  })
592  on('command.run', { command: 'tok-speed' }, async ($, e) => {
593    const args = e.args.trim()
594    if (args && args !== 'reset') return { text: '用法:/tok-speed 或 /tok-speed reset' }
595    if (args === 'reset') {
596      let refused = liveRequests.size > 0
597      await mutate($, epoch, state => {
598        refused = refused || state.rows.some(r => r.active !== null)
599        if (refused) return state
600        return { version: 2, session: state.session, rows: state.rows.filter(r => r.id === 'main' || r.running).map(r => ({ ...r,
601          models: [], active: null, live: null, seen: r.running, status: r.running ? 'waiting' : 'idle' })) }
602      })
603      return { text: refused ? '仍有代理请求在进行,无法 reset;请在请求结束后重试。' : '已重置本会话 token-speed 统计。' }
604    }
605    const state = await readState($)
606    const now = await safeNow($)
607    const provenance = state.rows.filter(r => r.id !== 'main').flatMap(r => {
608      const reading = lookupWindow(registry, r.currentModel)
609      return reading.sourceURL ? [`${role(r, state.rows)} · context source ${reading.sourceURL} · checked ${reading.date}`] : []
610    })
611    const lines = state.rows.flatMap(r => [
612      `${role(r, state.rows)} · ${r.description || (r.id === 'main' ? '主控' : '代理')} · ${r.running ? 'running' : 'ended'} · ${r.currentModel ?? 'unknown'} · effort ${r.effort ?? '—'} (${r.effortSource}) · Ctx ${contextLabel(r.context)} (${r.context.source}) · Live ${r.live === null ? '—' : `~${rate(r.live)}`} · ${r.status}`,
613      ...r.models.map(bucket => {
614        const held = bucket.log === undefined ? null : withinWindow(bucket.log, now)
615        const scope = held === null
616          ? `累计 ${bucket.samples} 个有效请求 · ${bucket.outputTokens} tokens / ${(bucket.durationMs / 1000).toFixed(3)}s`
617          : `24h ${held.length} 个请求 · ${held.reduce((sum, sample) => sum + sample.tokens, 0)} tokens / ${(held.reduce((sum, sample) => sum + sample.ms, 0) / 1000).toFixed(3)}s`
618        return `${role(r, state.rows)} · ${bucket.model}${bucket.legacy ? ' [legacy v0.1 response model]' : ''} · Last ${rate(bucket.lastApi)} · Avg ${rate(windowedAverage(bucket, now))} · ${scope}`
619      }),
620    ])
621    return { text: [
622      'token-speed 0.3.2 · 各代理独立累计 · 主控 first',
623      '各代理同一行显示模型、effort、Ctx、Live/Last/Avg(tok/s 只标一次);工作区最后一行。effort 来源:request 为 turn.step 请求,applied 为 classic 实际档位,configured 为主控配置或明确模型后缀,unknown 为 —。主控 Ctx 来自 session.usage;子代理输入来自最近 CLI uncached+cache读写总量,窗口从 data/model-contexts.json 精确匹配模型尾名/aliases,明确 override 优先官方 default,来源分别 cli-input-window-config、cli-input-official-default 与 cli-input-official-capacity,均不冒充 API 实测窗口;未匹配模型显示 k/?。暗底块状条固定 10 格,每格 8 档共 80 视觉档位(相邻整数百分比可能落在同一档,数字仍为整数 1%),填充固定使用 theme `rate_limit_fill`,轨道为 `rate_limit_empty`,窄屏可省条;无占用读数显示 —%。',
624      `context registry v${registry.version}${registry.fallback ? ' (load failed: confirmed overrides only)' : ''} · 优先实际 session 窗口、明确 override、官方 default(未给默认时明确 capacity)、未知;官方 max 仅元数据,不因 [1m] 自动选择。`,
625      ...provenance,
626      ...lines,
627      ...(state.rows.some(r => r.models.length) ? [] : ['Last — · Avg — · 尚无本会话统计']),
628      'Live:最近 3 秒收到的文本、thinking、工具 JSON 参数估算;CJK 每 Unicode code point 1 token,其余 0.25。分母 min(3 秒, 请求已耗时),不足 250ms 为 —;停流 3 秒后为 0。~ 非 tokenizer 精确计数。无法观察的代理流为 —,不推算速度。',
629      'Last/Avg:CLI usage.output_tokens / 请求全程耗时,包含 TTFT、thinking、网络与请求内停顿;排除请求之间的工具/用户空闲,非纯解码速度。Last 为该代理该模型最近一次有效请求;Avg 为最近 24 小时滚动窗口内该模型的有效 token 总数 / 有效耗时总和,按请求完成时间戳(clock epoch ms)判窗,与本地时区无关,出窗即弃。迁移自 v0.2 的无时间戳桶在首个新请求前退回全程累计口径;失败/中断、无有效 usage、耗时不大于 0 不入统计;有效 0 计入。CPA 缺字段可能被 CLI 归零,无法确认其来源;turn.complete usage 不重复累计。',
630      '模型名:以请求 model 为准,去除末尾 [1m]、(high) 或 :high 等配置后缀;统计桶与命令保留已知渠道,UI 仅隐藏最后 / 之前的渠道,不推断渠道。同代理 bare 模型与已知渠道模型尾名相同时沿用渠道;usage.model 不改变标签或桶。旧 v0.1 response-model 均值标为 legacy,保留且不混入新的请求模型均值。',
631      'UI 主控始终第一,活跃子代理按首次观察顺序每个一行;切换 view 不过滤。工具间隙保留行,Live —;结束立即隐藏,命令可查看历史。reload 保留统计、清除无法接管的旧流,以本 session agent.list 同步元数据;/clear 清空,reset 任意 active 请求时拒绝、工具间隙保留活跃代理元数据。',
632    ].join('\n') }
633  })
634  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
635    const rest = await next(e)
636    const state = await readState($)
637    if (e.props.hasSurvey || e.props.maxRows <= 0 || !state.rows.some(r => r.seen || r.running)) return rest
638    const visible = state.rows.filter(r => r.id === 'main' || r.running)
639    const { Box, Text } = $.ui.resolve(e)
640    const now = await safeNow($)
641    const budget = e.props.maxRows - visible.length
642    const effortColors: Record<Exclude<TokenSpeedEffort, number>, ThemeKey> =
643      { low: 'success', medium: 'planMode', high: 'warning', xhigh: 'error', max: 'error' }
644    const workspace = state.session.workspace
645    const workspaceLine = budget >= 1 && workspace
646      ? <Text key="token-speed-workspace" color="inactive" wrap="truncate-end">
647          {fitText(`⌂ ${workspace.name}${workspace.branch ? ` on ${workspace.branch}` : ''}`, e.props.bodyColumns)}
648        </Text>
649      : null
650    const labels = visible.map(r => r.id === 'main' ? '⚡ main' : `↳ ${shortId(r.id, state.rows)}`)
651    const models = visible.map(r => displayModel(r.currentModel))
652    const efforts = visible.map(r => r.effort === null ? '—' : String(r.effort))
653    const contexts = visible.map(r => contextLabel(r.context))
654    const buckets = visible.map(r => r.models.find(item => item.model === r.currentModel && !item.legacy))
655    const averages = buckets.map(bucket => bucket ? windowedAverage(bucket, now) : null)
656    const number = (n: number | null | undefined): string => n == null ? '—' : n.toFixed(1)
657    const lives = visible.map(r => r.live === null ? '—' : `~${number(r.live)}`)
658    const lasts = buckets.map(bucket => number(bucket?.lastApi))
659    const avgs = averages.map(number)
660    const maxWidth = (values: string[], minimum = 0): number => Math.max(minimum, ...values.map(cellWidth))
661    const multi = visible.length > 1
662    const labelWidth = maxWidth(labels)
663    let modelWidth = maxWidth(models)
664    const effortWidth = maxWidth(efforts)
665    const contextWidth = maxWidth(contexts)
666    const liveWidth = maxWidth(lives, 5)
667    const lastWidth = maxWidth(lasts, 5)
668    const avgWidth = maxWidth(avgs, 5)
669    const statusWidth = maxWidth(visible.map(r => r.status))
670    const columns = Math.max(1, e.props.bodyColumns)
671    let showStatus = true
672    let showLast = true
673    // Fixed 10-cell bar: 80 eighth-cell steps; no widening on wide terminals.
674    let cells = 10
675    const width = (bar: number): number => labelWidth + 3 + modelWidth + 3 + effortWidth
676      + 7 + contextWidth + (bar ? bar + 1 : 0)
677      + 8 + liveWidth + (showLast ? 8 + lastWidth : 0) + 7 + avgWidth + 6
678      + (showStatus ? 3 + statusWidth : 0)
679    // Keep a usable compact meter first; discard status, then Last, before the bar.
680    if (width(cells) > columns) showStatus = false
681    if (width(cells) > columns) showLast = false
682    if (width(cells) > columns) cells = 0
683    if (width(cells) > columns) modelWidth = Math.max(1, modelWidth - (width(cells) - columns))
684    // At extremely small widths preserve the numeric context/effort and rates;
685    // model alone can be abbreviated after optional columns have been exhausted.
686    const aligned = (text: string, size: number): string => multi ? padEndTo(text, size) : text
687    return <Box flexDirection="column">
688      {visible.map((r, index) => {
689        const percent = r.context.percent
690        const known = r.context.window > 0 && percent !== null
691        const eighths = known ? Math.max(0, Math.min(cells * 8, Math.round(percent * cells * 8 / 100))) : 0
692        const full = Math.floor(eighths / 8)
693        const partial = ['', '▏', '▎', '▍', '▌', '▋', '▊', '▉'][eighths % 8] ?? ''
694        const fill = '█'.repeat(full) + partial
695        const track = ' '.repeat(cells - full - (partial ? 1 : 0))
696        const effortColor = r.effort === null ? 'inactive' : typeof r.effort === 'number' ? 'subtle' : effortColors[r.effort]
697        return <Text key={`token-speed-${r.id}`} color="inactive" wrap="truncate-end">
698          <Text color={r.id === 'main' ? 'claude' : 'inactive'} bold={r.id === 'main'}>{aligned(labels[index] ?? '', labelWidth)}</Text>
699          {' · '}<Text color="text">{aligned(fitText(models[index] ?? '', modelWidth), modelWidth)}</Text>
700          {' · '}<Text color={effortColor}>{aligned(efforts[index] ?? '—', effortWidth)}</Text>
701          {' · Ctx '}{cells === 0 ? null : <Text><Text key={`context-meter-${r.id}`} backgroundColor="rate_limit_empty"><Text color="rate_limit_fill">{fill}</Text>{track}</Text>{' '}</Text>}
702          <Text color={known ? 'text' : 'inactive'}>{aligned(contexts[index] ?? '—/?', contextWidth)}</Text>
703          {' · Live '}<Text color={r.live === null ? 'inactive' : 'text'} bold={r.live !== null}>{padStartTo(lives[index] ?? '—', liveWidth)}</Text>
704          {showLast ? <Text>{' · Last '}<Text color={buckets[index]?.lastApi == null ? 'inactive' : 'text'}>{padStartTo(lasts[index] ?? '—', lastWidth)}</Text></Text> : null}
705          {' · Avg '}<Text color={averages[index] == null ? 'inactive' : 'text'}>{padStartTo(avgs[index] ?? '—', avgWidth)}</Text>{' tok/s'}
706          {showStatus ? <Text>{' · '}<Text color={statusColors[r.status]}>{r.status}</Text></Text> : null}
707        </Text>
708      })}
709      {workspaceLine}
710      {rest}
711    </Box>
712  })
713}
714
hooks/context-registry.ts 105 lines
1export type OfficialContext = {
2  id: string
3  aliases: string[]
4  defaultWindow: number | null
5  maxWindow?: number
6  contextWindow?: number
7  sourceScope?: string
8  additionalSources?: string[]
9  apiContextWindow?: number
10  maxInputTokens?: number
11  sourceValue?: string
12  sourceURL: string
13  date: string
14}
15export type ContextOverride = { id: string; aliases: string[]; window: number; reason: string }
16export type ContextRegistry = { version: 1; models: OfficialContext[]; overrides: ContextOverride[]; fallback?: true }
17export type WindowReading = { window: number; source: 'configured' | 'official-default' | 'official-capacity' | 'unknown'; sourceURL?: string; date?: string }
18
19// Keep the already-confirmed user windows if the packaged registry cannot load.
20const FALLBACK_OVERRIDES: ContextOverride[] = [
21  { id: 'gpt-6.1-sol', aliases: [], window: 272000, reason: 'User-confirmed CLI window' },
22  { id: 'glm-5.3', aliases: [], window: 1000000, reason: 'User-confirmed CLI window' },
23  { id: 'glm-5.3-flash', aliases: [], window: 1000000, reason: 'User-confirmed CLI window' },
24]
25const object = (v: unknown): v is Record<string, unknown> => typeof v === 'object' && v !== null && !Array.isArray(v)
26const integer = (v: unknown): v is number => typeof v === 'number' && Number.isSafeInteger(v) && v > 0
27const name = (v: unknown): v is string => typeof v === 'string' && /^[a-z0-9][a-z0-9._:-]*$/i.test(v)
28const sourceURL = (v: unknown): v is string => typeof v === 'string' && /^https:\/\/[a-z0-9][a-z0-9.-]*(?:\/[\x21-\x7e]*)?$/i.test(v)
29const fail = (): never => { throw new Error('Invalid model context registry') }
30function aliases(v: unknown): string[] {
31  if (!Array.isArray(v) || !v.every(name)) return fail()
32  return v.map(s => s.toLowerCase())
33}
34function fields(v: Record<string, unknown>, keys: string[]): void {
35  if (Object.keys(v).some(k => !keys.includes(k))) fail()
36}
37function date(v: unknown): v is string {
38  if (typeof v !== 'string' || !/^\d{4}-\d{2}-\d{2}$/.test(v)) return false
39  const parsed = new Date(v + 'T00:00:00Z')
40  return Number.isFinite(parsed.getTime()) && parsed.toISOString().slice(0, 10) === v
41}
42function unique(entries: { id: string; aliases: string[] }[]): void {
43  const seen = new Set<string>()
44  for (const entry of entries) for (const alias of [entry.id, ...entry.aliases]) {
45    if (seen.has(alias)) fail()
46    seen.add(alias)
47  }
48}
49export function parseRegistry(text: string): ContextRegistry {
50  const value: unknown = JSON.parse(text)
51  if (!object(value) || value.version !== 1 || !Array.isArray(value.models) || !Array.isArray(value.overrides)) return fail()
52  fields(value, ['version', 'models', 'overrides'])
53  const models: OfficialContext[] = value.models.map(v => {
54    if (!object(v) || !name(v.id) || (v.defaultWindow !== null && !integer(v.defaultWindow)) || !date(v.date)
55      || !sourceURL(v.sourceURL) || (v.defaultWindow === null && !integer(v.contextWindow))) return fail()
56    for (const key of ['maxWindow', 'contextWindow', 'apiContextWindow', 'maxInputTokens']) {
57      if (v[key] !== undefined && !integer(v[key])) fail()
58    }
59    if (v.maxWindow !== undefined && v.defaultWindow !== null && (v.maxWindow as number) < (v.defaultWindow as number)) fail()
60    if (v.additionalSources !== undefined && (!Array.isArray(v.additionalSources) || !v.additionalSources.every(sourceURL))) fail()
61    if ((v.sourceScope !== undefined && (typeof v.sourceScope !== 'string' || !v.sourceScope.trim()))
62      || (v.sourceValue !== undefined && (typeof v.sourceValue !== 'string' || !v.sourceValue.trim()))) fail()
63    fields(v, ['id', 'aliases', 'defaultWindow', 'maxWindow', 'contextWindow', 'sourceScope', 'additionalSources',
64      'apiContextWindow', 'maxInputTokens', 'sourceValue', 'sourceURL', 'date'])
65    return { id: v.id.toLowerCase(), aliases: aliases(v.aliases), defaultWindow: v.defaultWindow as number | null,
66      ...(v.maxWindow === undefined ? {} : { maxWindow: v.maxWindow as number }),
67      ...(v.contextWindow === undefined ? {} : { contextWindow: v.contextWindow as number }),
68      ...(v.apiContextWindow === undefined ? {} : { apiContextWindow: v.apiContextWindow as number }),
69      ...(v.maxInputTokens === undefined ? {} : { maxInputTokens: v.maxInputTokens as number }),
70      ...(v.sourceScope === undefined ? {} : { sourceScope: v.sourceScope as string }),
71      ...(v.sourceValue === undefined ? {} : { sourceValue: v.sourceValue as string }),
72      ...(v.additionalSources === undefined ? {} : { additionalSources: v.additionalSources as string[] }),
73      sourceURL: v.sourceURL, date: v.date }
74  })
75  const overrides: ContextOverride[] = value.overrides.map(v => {
76    if (!object(v) || !name(v.id) || !integer(v.window) || typeof v.reason !== 'string' || !v.reason.trim()) return fail()
77    fields(v, ['id', 'aliases', 'window', 'reason'])
78    return { id: v.id.toLowerCase(), aliases: aliases(v.aliases), window: v.window, reason: v.reason }
79  })
80  // A deliberate override may share an official id; ambiguity within either tier is refused.
81  unique(models)
82  unique(overrides)
83  return { version: 1, models, overrides }
84}
85export function fallbackRegistry(): ContextRegistry {
86  return { version: 1, models: [], overrides: FALLBACK_OVERRIDES.map(v => ({ ...v, aliases: [...v.aliases] })), fallback: true }
87}
88export function lookupWindow(registry: ContextRegistry, canonicalModel: string | null): WindowReading {
89  const tail = canonicalModel?.slice(canonicalModel.lastIndexOf('/') + 1).toLowerCase() ?? ''
90  const matches = (entry: { id: string; aliases: string[] }) => entry.id === tail || entry.aliases.includes(tail)
91  const override = registry.overrides.find(matches)
92  if (override) return { window: override.window, source: 'configured' }
93  const official = registry.models.find(matches)
94  return official ? { window: official.defaultWindow ?? official.contextWindow ?? 0,
95    source: official.defaultWindow === null ? 'official-capacity' : 'official-default', sourceURL: official.sourceURL, date: official.date }
96    : { window: 0, source: 'unknown' }
97}
98export function registryLoader(): (read: () => Promise<string>) => Promise<ContextRegistry> {
99  let pending: Promise<ContextRegistry> | null = null
100  return read => pending ??= (async () => {
101    try { return parseRegistry(await read()) }
102    catch { return fallbackRegistry() }
103  })()
104}
105
types/index.d.ts 85 lines
1export type TokenSpeedStatus = 'idle' | 'waiting' | 'streaming' | 'aborted' | 'error' | 'usage unavailable'
2
3export type TokenSpeedEffort = 'low' | 'medium' | 'high' | 'xhigh' | 'max' | number
4
5export type TokenSpeedSample = {
6  /** $.clock.now() epoch ms when the request finished. */
7  at: number
8  tokens: number
9  ms: number
10}
11
12export type TokenSpeedModelStats = {
13  model: string
14  outputTokens: number
15  durationMs: number
16  samples: number
17  lastApi: number | null
18  /** v0.1 used response models, so its averages remain separately identified. */
19  legacy?: true
20  /**
21   * v0.3+ per-request log backing the rolling 24h Avg. Absent on v0.1 legacy
22   * buckets and on v0.2 state migrated before its first new request.
23   */
24  log?: TokenSpeedSample[]
25}
26
27export type TokenSpeedActive = {
28  id: string
29  turnId: string
30  model: string
31  startedAt: number
32}
33
34export type TokenSpeedRow = {
35  /** main, or the session-local agentId. Array order is first observation order. */
36  id: string
37  description: string
38  running: boolean
39  currentModel: string | null
40  models: TokenSpeedModelStats[]
41  active: TokenSpeedActive | null
42  live: number | null
43  status: TokenSpeedStatus
44  seen: boolean
45  turnId: string | null
46  /** Lifecycle guard for asynchronously returned roster snapshots. */
47  revision: number
48  /** Observed request/applied effort, or an explicit configuration fallback. */
49  effort: TokenSpeedEffort | null
50  effortSource: 'configured' | 'request' | 'applied' | 'unknown'
51  /** Each loop owns its latest reading; child windows are explicit configuration. */
52  context: TokenSpeedContextInfo
53}
54
55export type TokenSpeedContextInfo = {
56  /** Last response's uncached + cache-read + cache-written inputs, never a turn sum. */
57  tokens: number | null
58  /** Zero means unknown, never a guessed model/context cap. */
59  window: number
60  percent: number | null
61  source: 'session' | 'cli-input' | 'cli-input-window-config' | 'cli-input-official-default' | 'cli-input-official-capacity' | 'unknown'
62}
63
64export type TokenSpeedWorkspace = { name: string; branch: string | null }
65
66export type TokenSpeedSessionInfo = {
67  context: TokenSpeedContextInfo
68  workspace: TokenSpeedWorkspace | null
69}
70
71export type TokenSpeedSnapshot = {
72  version: 2
73  rows: TokenSpeedRow[]
74  session: TokenSpeedSessionInfo
75}
76
77/** Read only compatibility shape; all writes use version 2. */
78export type TokenSpeedLegacySnapshot = Pick<TokenSpeedRow, 'models' | 'currentModel' | 'active' | 'live' | 'status' | 'seen'>
79
80declare module 'claude-code' {
81  interface PluginState {
82    'token-speed': { snapshot: TokenSpeedSnapshot }
83  }
84}
85