主控与活跃子代理各一行实时估算速率,各代理独立按请求模型累计 API 速率

适用于 Claude Code 2.1.294 的 function-hooks mod。在提示框上方,主控与每个活跃子代理各用一行显示模型、effort、上下文与输出速率,工作区放在最后一行。只观察原有请求,不调用额外模型、不引入 tokenizer 或网络请求;分支通过引擎 process API 调用本机 git。
⚡ main · gpt-6.1-sol · high · Ctx ██████▊ 67%/272k · Live ~42.1 · Last 31.2 · Avg 28.6 tok/s · streaming
↳ abcdefg1 · glm-5.3-flash · max · Ctx █▍ 14%/1.0M · Live ~35.2 · Last 29.1 · Avg 30.4 tok/s · streaming
⌂ Claude-workflow on main
主控始终第一,子代理按首次观察顺序每个一行,切换 view 不过滤。短 id 自动延长到可区分同会话代理;工具间隙保留行,Live 为 —,明确终态后隐藏,历史通过 /tok-speed 查询。带 survey 或 maxRows=0 时隐藏,其它插件和引擎的内容原样保留。只有主控时为代理行加工作区行,共两行;行数预算不足先丢工作区。
各列按终端显示宽度对齐,包括双宽的 ⚡、CJK 与常见 emoji。速率数字右对齐,单位 tok/s 每行只显示一次。动态宽度不足时依次省略 status、Last、缩短进度条,始终保留 effort、上下文数值、Live 与 Avg;特别窄时可缩短模型显示。渲染前计算可见内容宽度,Text 的 truncate-end 只作最后保护。
进度条是固定 10 个终端 cell 的短矩形,轨道使用 theme rate_limit_empty,整格填充 █,尾部使用 ▏▎▍▌▋▊▉ 表示 1/8 cell。共 80 个视觉档位(10 格 × 8 档),因此 0–100 的整数百分点不是每个都有不同的条形:相邻百分点可落在同一档(例如 67% 与 68% 的条形相同,但数字 67%、68% 不同)。宽屏也不扩展为长条。窄屏可省略条形,数值始终保持整数 1% 精度。填充固定使用 theme rate_limit_fill,不随占用变化,没有横向渐变;Ctx 数值用 text,未知占用用 inactive。未知占用只画轨道空条,并以 —% 或 ? 明确标为未知,不能解释为 0%。
主控上下文来自 $.session.usage(),与状态栏同口径,最近响应的 uncached + cache-read + cache-written 输入 tokens 除以窗口;首次响应或成功压缩后无读数显示 —%/窗口。已知读数不会被打回未知:缺少有效 tokens 的部分读数(窗口相同)与主控模型切换都会沿用上一次已知读数,只有成功压缩(主控)或整会话重置才清空,因此进度条不会在两次轮询之间闪回 —%。子代理输入量来自自己最近一次响应的 CLI usage,不累计多次请求。子代理分母由启动时读取一次的 data/model-contexts.json 提供,无联网刷新。优先顺序为实际运行窗口、明确 CLI override、官方默认、未知;官网仅提供容量时明确标为官方 capacity。输入始终取自己最近一次 CLI response,主控仍以 $.session.usage() 为准。数据中的官方来源、核验日期和 registry 版本可通过 /tok-speed 查看。读取或严格解析失败时,官方表为空,保留已确认的 Sol 272000、GLM-5.3/Flash 1000000 后备,不中断请求。
官方记录覆盖当前前沿型号,共 24 条(核验日期 2026-10-09)。Anthropic:claude-haiku-5-5、claude-sonnet-5-5、claude-opus-5-5、claude-fable-5-1、claude-fable-5、claude-mythos-5-1、claude-mythos-5、claude-opus-5、claude-sonnet-5、claude-opus-4-8、claude-opus-4-7、claude-opus-4-6、claude-sonnet-4-6 为共享输入输出上下文 1M,claude-haiku-4-5、claude-sonnet-4-5 为 200k。OpenAI:gpt-6.1-sol、gpt-6-astra、gpt-6-sol、gpt-6-luna 为 Codex 官方默认 272000 / max 872000,其中前两者另带 API 总上下文 1050000 与 max input 922000 元数据。glm-5.3、glm-5.3-flash、deepseek-v4.1-flash(aliases deepseek-flash、deepseek-v4-flash)、deepseek-v4-pro、gemini-3.8-flash 官网只给出容量且未声明默认/最大,按官方 capacity 记录;Gemini 的 1,048,576 是输入上限,输出上限 65,536 单独存在,不是共享窗口。用户 override 独立存放,不混作官方事实。匹配只用去渠道的 canonical 尾名及精确 aliases,不匹配子串,不把未知邻近型号映射到已知型号;不使用动态 haiku/opus/sonnet 别名,不因 [1m] 选择官方最大值。来源分别记录为 cli-input-window-config、cli-input-official-default、cli-input-official-capacity。
新增 provider/model 时,只添加具有官方 URL、核验日期和精确型号的记录;未核实的型号保持 unknown,或单独添加有理由的 local override。模型最大上下文不保证路由实际 cap,registry 不把最大值当默认。不复制主控百分比,数值超过窗口可显示超过 100%,只限制条形填充。成功安装的压缩清对应行输入读数,保留已知窗口;precompute 和 skip 不清。
Effort 来源保存在状态并可由命令查看:request 为 turn.step.effort,applied 为 classic PostToolUse/Stop 的 effort.level,configured 为主控当前模型的 modelSettings[model].effortLevel、全局 effortLevel 或明确模型后缀。reload 后主控立即读取配置后备,无数据为 —;不硬编码 high,不扫描 agent 定义。配置档位不保证外部 provider 已应用。只选取 settings 的这些 effort 字段,不保存或打印其它 settings。
工作区为仓库根目录名(非 git 使用会话根目录名),附当前分支;detached HEAD 不显示分支,约每 5 秒刷新。显示使用 theme tokens:主控标签为 claude、子代理标签与行次要信息为 inactive;模型名为 text(普通文本,不着色);有效 Live/Avg/Last 数字为 text,缺值为 inactive,速率不以 success 表示;effort 按档位着色(数值为 subtle);状态 idle/waiting/streaming 为 inactive,aborted 为 warning,error 为 error,usage unavailable 为 warning。工作区行为 inactive。颜色跟随 CLI 主题。
/plugin marketplace add liuyejinghong/Claude-workflow
/plugin install token-speed@claude-workflow-mods
也可用固定 tag token-speed-v0.3.2 获取源文件,并从 checkout 加载:
git clone --branch token-speed-v0.3.2 https://github.com/liuyejinghong/Claude-workflow.git
claude --plugin-dir ./Claude-workflow/mods/token-speed
安装后的插件由 Claude Code 加载。开发时修改自己的 checkout;请勿直接修改 plugins/cache 副本。
/tok-speed:列出各代理的短 id 角色、任务描述、当前模型、实时估算和各模型历史累计,以及统计口径。/tok-speed reset:清空本会话统计。任意代理请求仍 active 时拒绝;工具间隙可 reset,保留活跃代理元数据和行。/clear:清空统计、代理行与实时请求;晚返回的旧流或 spawn 元数据不复活统计,新 turn.step 重启计时器。状态包括 waiting、streaming、idle、aborted、error、usage unavailable。新采纳的代理没有可观察 stream 时 Live 为 —,模型为 unknown,直到 step 或有明确 resolved model 的 spawn。流内停顿仍为 streaming,Live 随窗口衰减。
每个代理独立按 请求 turn.step.model 累计;两个相同模型代理不共享统计桶。usage.model 仅是响应事实,不改变标签或桶。真实请求模型切换建立新桶。
统计桶与命令保留已有完整渠道模型名,UI 仅隐藏最后 / 之前的渠道前缀;其它合法冒号部分不会删除。
0.1 的 host state 会迁移到主控行,其 response-model 均值保留为 [legacy v0.1 response model] 桶,在 /tok-speed 可查。由于旧桶缺乏请求渠道证据,它们不会随意合并或混入 0.2 的请求模型均值;UI Avg 只取新请求桶。
Live 是估算值,始终带 ~。 收集最近 3 秒实际观察到的 text、thinking 和工具调用 JSON 参数字符;tool、engine、stop chunk 不增加计数。按 Unicode code point 计数:CJK Unified Ideographs(U+3400–4DBF、U+4E00–9FFF)、兼容汉字(U+F900–FAFF)及补充汉字(U+20000–323AF)每 code point 估为 1 token,其余为 0.25。不是 tokenizer 精确计数。
分母为 min(3000ms, 当前时间 - 请求开始时间)。不足 250ms 为 —;请求开始至首个 chunk 的等待也计时。仅计近 3 秒字符,停流超过 3 秒变为 ~0.0 tok/s。每 250ms 共用一次 host clock 时间和一次原子 snapshot update,批量发布所有活跃 loop;每个字符 chunk 调用一次 clock.now,不按 chunk 写 host state。
Last / Avg 是 CLI usage 的 API 全程输出速率。 每个 turn.step 从 hook 开始计时至最终结果返回,包含 TTFT、thinking、网络、请求内停顿和内置 API 工具耗时;排除两次请求之间的工具或用户空闲,非纯解码速度。
output_tokens / 秒数。失败、中断、无有效 usage 或耗时不大于 0 为 —。有效 output_tokens 总和 / 有效请求耗时总和,不取速度算术平均,按请求完成时间戳(clock epoch ms)判窗,出窗样本即弃——窗口是相对时间,与本地时区(含 UTC+8)无关。例如窗口内 100/2s + 300/3s = 80 tok/s。迁移自 v0.2 的无时间戳桶在首个新请求前退回全程累计口径。CPA 上游缺失 token count 可能被 CLI 规范化为 0。mod 无法区分实际 0 和上游遗漏后的 0;API 统计表示 CLI usage 口径,整段 usage null 和仍可观察的无效 count 不计入。
统计在本 session 的 version 2 host PluginState,每 loop 单独 row,顺序稳定;本地 Map 只保存流累加器。reload 保留统计,丢弃无法接管的 stale active,以 $.agent.list() 重建在跑子代理元数据。该 API 不含模型;未观察到请求的代理显示 unknown/Live —。每 1 秒轻量 reconcile 与上下文用量轮询、每 5 秒工作区刷新在无流时也运行,值未变化不写状态;渲染不查询或写这些来源,只读 host state。每请求的完成样本保留 24 小时滚动日志支撑 Avg。失败 list 保留行;生命周期 revision guard 防止旧 list 返回覆盖刚开始/完成的新 turn。只采纳 running;waiting 和工具间隙不判完成;仅 completed/failed/killed 明确终态收尾,不根据 absence 结束新 spawn。
agent.list() 只涉及本 session 的子代理/teammate,不扫描其它 CLI 或远程 session。API 不列 workflow agent,也无法观察远程 workflow loop;有本地 turn.step 时可单独统计。无可观察 stream 的行永远不推算速度。重复 session.start 不重复创建 timer;session.end 取消 timer 清空统计,不跨会话持久化。
流 hook 只调用一次 next(e),原样 yield 所有 chunk/ref,原样返回 HookStream.result;不更改请求、顺序或工具 JSON。仪表自身异常被捕获,真实下游错误继续抛出,不重复请求。每个 loop 仅清理自己的 active/live。
claude plugin validate <mod-dir> --strict
tsc -p <mod-dir>
claude plugin test <mod-dir>
本目录 tsconfig 开启 strict/noUncheckedIndexedAccess,包含 hooks、tests、自有 types 和生成的 .claude-plugin/types。测试使用实际 claude-code/testing 和 mock.clock 合成 stream,不联网或调用模型。
测试保留流透传、单次 next、测量/Unicode/加权平均、失败清理、reset/clear/reload、并发独立桶与生命周期竞态合同,并核验统一代理行、配置 effort 后备、主控/子代理上下文隔离、成功压缩、10 格块状条 80 档视觉映射(0–100 为 81 种可见状态,67/68 同形但数字不同)、配置子代理窗口、80/120/240 列预算、显示宽对齐、UI 隐藏渠道和命令完整渠道。kit 只验证引擎 tree 与合同,不代替终端字体的实际截图验收。
hooks/register.tsx 714 lines1import { atom, read, update } from 'claude-code'
2import { fallbackRegistry, lookupWindow, registryLoader } from './context-registry'
3import type { AgentInfo, EngineInterface, Register, ThemeKey, Timer, TurnStepChunk, TurnStepResult } from 'claude-code'
4import type { TokenSpeedActive, TokenSpeedContextInfo, TokenSpeedEffort, TokenSpeedModelStats, TokenSpeedRow,
5 TokenSpeedSample, TokenSpeedSessionInfo, TokenSpeedSnapshot, TokenSpeedStatus, TokenSpeedWorkspace } from '../types'
6
7const WINDOW_MS = 3000
8const TICK_MS = 250
9// Avg rolls over the last day; sample timestamps are clock epoch ms, so no timezone applies.
10const AVG_WINDOW_MS = 24 * 60 * 60 * 1000
11const loadRegistry = registryLoader()
12let registry = fallbackRegistry()
13function childContext(model: string | null, tokens: number | null = null): TokenSpeedContextInfo {
14 const reading = lookupWindow(registry, model === null ? null : canonical(model))
15 const window = reading.window
16 return { tokens, window, percent: tokens === null || window === 0 ? null : Math.round(tokens * 100 / window),
17 source: reading.source === 'configured' ? 'cli-input-window-config'
18 : reading.source === 'official-default' ? 'cli-input-official-default'
19 : reading.source === 'official-capacity' ? 'cli-input-official-capacity' : tokens === null ? 'unknown' : 'cli-input' }
20}
21const emptyContext = (): TokenSpeedContextInfo => ({ tokens: null, window: 0, percent: null, source: 'unknown' })
22const row = (id: string): TokenSpeedRow => ({ id, description: '', running: false, currentModel: null,
23 models: [], active: null, live: null, status: 'idle', seen: false, turnId: null, revision: 0,
24 effort: null, effortSource: 'unknown', context: emptyContext() })
25const emptySession = (): TokenSpeedSessionInfo => ({ context: emptyContext(), workspace: null })
26const empty = (): TokenSpeedSnapshot => ({ version: 2, rows: [row('main')], session: emptySession() })
27const snapshot = atom({ plugin: 'token-speed', key: 'snapshot' } as const, empty())
28
29type LiveRequest = TokenSpeedActive & {
30 loopId: string
31 epoch: number
32 chunks: { at: number; tokens: number }[]
33 hasContent: boolean
34}
35
36// Only stream accumulators and lifecycle guards are local. Drawings read host state.
37const liveRequests = new Map<string, LiveRequest>()
38let timer: Timer | null = null
39let tickBusy = false
40let listBusy = false
41let usageBusy = false
42let repoBusy = false
43let ticks = 0
44let sequence = 0
45let revision = 0
46let epoch = 0
47
48async function observe(work: () => Promise<unknown>): Promise<void> {
49 try { await work() } catch { /* Meter failures never interrupt or repeat the actual request. */ }
50}
51
52const object = (value: unknown): value is Record<string, unknown> => typeof value === 'object' && value !== null
53const finite = (value: unknown): value is number => typeof value === 'number' && Number.isFinite(value)
54const status = (value: unknown): TokenSpeedStatus => value === 'waiting' || value === 'streaming'
55 || value === 'aborted' || value === 'error' || value === 'usage unavailable' ? value : 'idle'
56function requestLog(value: unknown): TokenSpeedSample[] {
57 if (!Array.isArray(value)) return []
58 const log: TokenSpeedSample[] = []
59 for (const item of value) {
60 if (!object(item) || !finite(item.at) || !finite(item.tokens) || item.tokens < 0 || !finite(item.ms) || item.ms <= 0) continue
61 log.push({ at: item.at, tokens: item.tokens, ms: item.ms })
62 }
63 return log
64}
65function stats(value: unknown, legacy = false): TokenSpeedModelStats[] {
66 if (!Array.isArray(value)) return []
67 const buckets: TokenSpeedModelStats[] = []
68 for (const item of value) {
69 if (!object(item) || typeof item.model !== 'string' || !finite(item.outputTokens) || item.outputTokens < 0
70 || !finite(item.durationMs) || item.durationMs < 0 || !finite(item.samples) || item.samples < 0) continue
71 // A v0.2 bucket carries no log; its average falls back to the aggregates until a new request.
72 buckets.push({ model: item.model, outputTokens: item.outputTokens, durationMs: item.durationMs,
73 samples: item.samples, lastApi: finite(item.lastApi) && item.lastApi >= 0 ? item.lastApi : null,
74 ...(legacy || item.legacy === true ? { legacy: true as const }
75 : Array.isArray(item.log) ? { log: requestLog(item.log) } : {}) })
76 }
77 return buckets
78}
79function activeValue(value: unknown): TokenSpeedActive | null {
80 return object(value) && typeof value.id === 'string' && typeof value.turnId === 'string'
81 && typeof value.model === 'string' && finite(value.startedAt)
82 ? { id: value.id, turnId: value.turnId, model: value.model, startedAt: value.startedAt } : null
83}
84function effortValue(value: unknown): TokenSpeedEffort | null {
85 if (typeof value === 'string' && ['low', 'medium', 'high', 'xhigh', 'max'].includes(value)) return value as TokenSpeedEffort
86 return finite(value) && value >= 0 ? value : null
87}
88function contextInfo(value: unknown): TokenSpeedContextInfo {
89 const held = object(value) ? value : {}
90 return { tokens: finite(held.tokens) && held.tokens >= 0 ? held.tokens : null,
91 window: finite(held.window) && held.window > 0 ? held.window : 0,
92 percent: finite(held.percent) && held.percent >= 0 ? held.percent : null,
93 source: held.source === 'session' || held.source === 'cli-input' || held.source === 'cli-input-window-config'
94 || held.source === 'cli-input-official-default' || held.source === 'cli-input-official-capacity' ? held.source : 'unknown' }
95}
96function workspaceValue(value: unknown): TokenSpeedWorkspace | null {
97 if (!object(value) || typeof value.name !== 'string' || !value.name) return null
98 return { name: value.name, branch: typeof value.branch === 'string' && value.branch ? value.branch : null }
99}
100function sessionInfo(value: unknown): TokenSpeedSessionInfo {
101 const held = object(value) ? value : {}
102 return { context: contextInfo(held.context), workspace: workspaceValue(held.workspace) }
103}
104// Runtime validation also admits the v0.1 shape; response-model averages stay legacy buckets.
105function normalize(value: unknown): TokenSpeedSnapshot {
106 if (!object(value)) return empty()
107 if (value.version !== 2 || !Array.isArray(value.rows)) {
108 const main = row('main')
109 main.models = stats(value.models, true)
110 main.currentModel = typeof value.currentModel === 'string' ? value.currentModel : null
111 main.active = activeValue(value.active)
112 main.live = finite(value.live) ? value.live : null
113 main.status = status(value.status)
114 main.seen = value.seen === true || main.models.length > 0 || main.active !== null
115 main.turnId = main.active?.turnId ?? null
116 return { version: 2, rows: [main], session: sessionInfo(value.session) }
117 }
118 const rows: TokenSpeedRow[] = []
119 for (const item of value.rows) {
120 if (!object(item) || typeof item.id !== 'string' || rows.some(r => r.id === item.id)) continue
121 rows.push({ id: item.id, description: typeof item.description === 'string' ? item.description : '',
122 running: item.running === true, currentModel: typeof item.currentModel === 'string' ? item.currentModel : null,
123 models: stats(item.models), active: activeValue(item.active), live: finite(item.live) ? item.live : null,
124 status: status(item.status), seen: item.seen === true,
125 turnId: typeof item.turnId === 'string' ? item.turnId : null,
126 revision: finite(item.revision) ? item.revision : 0,
127 effort: effortValue(item.effort),
128 effortSource: item.effortSource === 'configured' || item.effortSource === 'request' || item.effortSource === 'applied'
129 ? item.effortSource : 'unknown',
130 context: item.id === 'main' ? contextInfo(item.context)
131 : childContext(typeof item.currentModel === 'string' ? item.currentModel : null, contextInfo(item.context).tokens) })
132 }
133 return { version: 2, rows: [rows.find(r => r.id === 'main') ?? row('main'), ...rows.filter(r => r.id !== 'main')],
134 session: sessionInfo(value.session) }
135}
136async function mutate($: EngineInterface, heldEpoch: number, change: (state: TokenSpeedSnapshot) => TokenSpeedSnapshot): Promise<void> {
137 await update($, snapshot, raw => heldEpoch === epoch ? change(normalize(raw)) : normalize(raw))
138}
139const readState = async ($: EngineInterface): Promise<TokenSpeedSnapshot> => normalize(await read($, snapshot))
140function changeRow(state: TokenSpeedSnapshot, id: string, change: (held: TokenSpeedRow) => TokenSpeedRow): TokenSpeedSnapshot {
141 const present = state.rows.some(r => r.id === id)
142 return { version: 2, session: state.session,
143 rows: present ? state.rows.map(r => r.id === id ? change(r) : r) : [...state.rows, change(row(id))] }
144}
145
146// Only known configuration suffixes are removed; an existing provider prefix is evidence.
147function canonical(model: string, previous: string | null = null): string {
148 let normalized = model.trim()
149 for (;;) {
150 const trimmed = normalized.replace(/(?:\[1m\]|\((?:low|medium|high|xhigh|max)\)|:(?:low|medium|high|xhigh|max))$/i, '').trim()
151 if (trimmed === normalized) break
152 normalized = trimmed
153 }
154 if (!normalized) return 'unknown'
155 if (!normalized.includes('/') && previous?.includes('/') && previous.slice(previous.lastIndexOf('/') + 1) === normalized) return previous
156 return normalized
157}
158function suffixEffort(model: string): TokenSpeedEffort | null {
159 // Only an explicit effort suffix is evidence; [1m] says nothing about effort/window.
160 const match = model.trim().replace(/\[1m\]$/i, '').match(/(?:\((low|medium|high|xhigh|max)\)|:(low|medium|high|xhigh|max))$/i)
161 return effortValue((match?.[1] ?? match?.[2])?.toLowerCase())
162}
163async function configuredEffort($: EngineInterface, model: string): Promise<TokenSpeedEffort | null> {
164 try {
165 const settings = await $.settings.read()
166 const models = object(settings.modelSettings) ? settings.modelSettings : {}
167 const selected = models[model] ?? models[canonical(model)]
168 return effortValue(object(selected) ? selected.effortLevel : undefined)
169 ?? effortValue(settings.effortLevel) ?? suffixEffort(model)
170 } catch { return suffixEffort(model) }
171}
172const sameContext = (a: TokenSpeedContextInfo, b: TokenSpeedContextInfo): boolean =>
173 a.tokens === b.tokens && a.window === b.window && a.percent === b.percent && a.source === b.source
174async function refreshMain($: EngineInterface, heldEpoch: number): Promise<void> {
175 const before = (await readState($)).rows.find(r => r.id === 'main') ?? row('main')
176 let fullModel = before.currentModel
177 try { fullModel = await $.session.model() } catch { /* Keep the observed request model when the host has none. */ }
178 if (!fullModel || heldEpoch !== epoch) return
179 const model = canonical(fullModel, before.currentModel)
180 const effort = await configuredEffort($, fullModel)
181 ?? (model === before.currentModel && before.effortSource === 'configured' ? before.effort : null)
182 const expected = before.currentModel === model && (before.effortSource === 'request' || before.effortSource === 'applied')
183 ? { effort: before.effort, source: before.effortSource } : { effort, source: effort === null ? 'unknown' : 'configured' }
184 if (before.currentModel === model && before.effort === expected.effort && before.effortSource === expected.source && before.seen) return
185 await mutate($, heldEpoch, state => changeRow(state, 'main', r => {
186 if (r.revision !== before.revision) return r
187 const changedModel = r.currentModel !== model
188 const retained = !changedModel && (r.effortSource === 'request' || r.effortSource === 'applied')
189 const value = retained ? r.effort : effort
190 const source = retained ? r.effortSource : value === null ? 'unknown' : 'configured'
191 if (!changedModel && r.effort === value && r.effortSource === source && r.seen) return r
192 return { ...r, currentModel: model, effort: value, effortSource: source, seen: true }
193 }))
194}
195const owns = (active: LiveRequest): boolean => active.epoch === epoch && liveRequests.get(active.loopId) === active
196function stopTimer(): void {
197 try { timer?.cancel() } catch { /* Disposal is observational too. */ }
198 timer = null
199 ticks = 0
200}
201
202async function reconcile($: EngineInterface, adoptExisting = false): Promise<void> {
203 if (listBusy) return
204 listBusy = true
205 const heldEpoch = epoch
206 try {
207 const before = await readState($)
208 const guards = new Map(before.rows.map(r => [r.id, r.revision]))
209 const agents = await $.agent.list() // A rejected list must not finish any row.
210 if (heldEpoch !== epoch) return
211 const ended: { id: string; requestId: string | undefined }[] = []
212 await mutate($, heldEpoch, state => {
213 ended.length = 0 // update may retry after a version collision.
214 let changed = state
215 for (const agent of agents) {
216 const held = changed.rows.find(r => r.id === agent.id)
217 if (held && held.revision !== guards.get(agent.id)) continue
218 if (agent.status === 'running') {
219 changed = changeRow(changed, agent.id, r => ({ ...r, description: agent.description,
220 running: held ? (adoptExisting ? true : r.running) : true,
221 seen: true }))
222 } else if (held && (agent.status === 'completed' || agent.status === 'failed' || agent.status === 'killed')) {
223 ended.push({ id: agent.id, requestId: held.active?.id })
224 changed = changeRow(changed, agent.id, r => ({ ...r, description: agent.description, running: false,
225 active: null, live: null, status: agent.status === 'failed' ? 'error' : agent.status === 'killed' ? 'aborted' : 'idle' }))
226 }
227 }
228 // Absence never finishes a newly spawned loop: only explicit terminal entries do.
229 return changed
230 })
231 for (const item of ended) {
232 if (liveRequests.get(item.id)?.id === item.requestId) liveRequests.delete(item.id)
233 }
234 } finally { listBusy = false }
235}
236function ensureTimer($: EngineInterface): void {
237 if (timer) return
238 timer = $.clock.every(TICK_MS, () => {
239 ticks++
240 if (ticks % 4 === 0) void observe(() => reconcile($))
241 if (ticks % 4 === 2) void observe(() => pollUsage($))
242 if (ticks % 20 === 10) void observe(() => refreshWorkspace($))
243 if (tickBusy || liveRequests.size === 0) return
244 tickBusy = true
245 void observe(async () => {
246 const active = [...liveRequests.values()]
247 const heldEpoch = epoch
248 const now = await $.clock.now() // One timestamp for every active loop in this batch.
249 const values = new Map<string, { id: string; live: number | null; status: TokenSpeedStatus }>()
250 for (const request of active) {
251 if (!owns(request)) continue
252 request.chunks = request.chunks.filter(chunk => chunk.at > now - WINDOW_MS)
253 const elapsed = Math.min(WINDOW_MS, now - request.startedAt)
254 values.set(request.loopId, { id: request.id,
255 live: elapsed < TICK_MS ? null : request.chunks.reduce((sum, chunk) => sum + chunk.tokens, 0) * 1000 / elapsed,
256 status: request.hasContent ? 'streaming' : 'waiting' })
257 }
258 await mutate($, heldEpoch, state => ({ version: 2, session: state.session, rows: state.rows.map(r => {
259 const value = values.get(r.id)
260 return value && r.active?.id === value.id ? { ...r, live: value.live, status: value.status } : r
261 }) }))
262 }).finally(() => { tickBusy = false })
263 })
264}
265async function pollUsage($: EngineInterface): Promise<void> {
266 if (usageBusy) return
267 usageBusy = true
268 const heldEpoch = epoch
269 try {
270 await observe(() => refreshMain($, heldEpoch))
271 const held = (await $.session.usage()).context
272 if (heldEpoch !== epoch || !object(held) || !finite(held.window) || held.window <= 0) return
273 const tokens = finite(held.tokens) && held.tokens >= 0 ? held.tokens : null
274 const measured = tokens !== null
275 const read: TokenSpeedContextInfo = { tokens, window: held.window,
276 percent: finite(held.percent) && held.percent >= 0 ? held.percent : null, source: 'session' }
277 const before = await readState($)
278 const current = (before.rows.find(r => r.id === 'main') ?? row('main')).context
279 // A partial session reading must not erase a known measurement for the same window.
280 const next = !measured && current.tokens !== null && current.window === held.window ? current : read
281 if (sameContext(before.session.context, next) && sameContext(current, next)) return
282 await mutate($, heldEpoch, state => {
283 const changed = changeRow(state, 'main', r => sameContext(r.context, next) ? r : { ...r, context: next })
284 return sameContext(state.session.context, next) ? changed : { ...changed, session: { ...state.session, context: next } }
285 })
286 } finally { usageBusy = false }
287}
288async function refreshWorkspace($: EngineInterface): Promise<void> {
289 if (repoBusy) return
290 repoBusy = true
291 const heldEpoch = epoch
292 try {
293 let name: string
294 let branch: string | null = null
295 const repo = await $.session.repo()
296 if (repo && repo.root) {
297 name = folderName(repo.root)
298 try {
299 const git = await $.process.run(['git', 'branch', '--show-current'], { cwd: repo.root, timeoutMs: 2000 })
300 if (git.exitCode === 0) branch = git.stdout.trim() || null // Empty output is a detached HEAD.
301 } catch { /* A missing git never hides the workspace name. */ }
302 } else {
303 name = folderName(await $.session.root())
304 }
305 if (heldEpoch !== epoch) return
306 const next: TokenSpeedWorkspace = { name, branch }
307 await mutate($, heldEpoch, state => {
308 const held = state.session.workspace
309 return held?.name === next.name && held.branch === next.branch ? state
310 : { version: 2, rows: state.rows, session: { ...state.session, workspace: next } }
311 })
312 } finally { repoBusy = false }
313}
314
315// Unicode code points, never UTF-16 units. Deliberately not a tokenizer.
316function estimate(text: string): number {
317 let tokens = 0
318 for (const char of text) {
319 const code = char.codePointAt(0)!
320 const cjk = (code >= 0x3400 && code <= 0x4dbf) || (code >= 0x4e00 && code <= 0x9fff)
321 || (code >= 0xf900 && code <= 0xfaff) || (code >= 0x20000 && code <= 0x323af)
322 tokens += cjk ? 1 : 0.25
323 }
324 return tokens
325}
326async function observeChunk($: EngineInterface, active: LiveRequest, chunk: TurnStepChunk): Promise<void> {
327 const text = chunk.kind === 'text' || chunk.kind === 'thinking' ? chunk.text : chunk.kind === 'input' ? chunk.json : ''
328 if (!text || !owns(active)) return
329 const at = await $.clock.now()
330 if (!owns(active)) return
331 active.hasContent = true
332 active.chunks = active.chunks.filter(item => item.at > at - WINDOW_MS)
333 active.chunks.push({ at, tokens: estimate(text) })
334}
335function bucketFor(r: TokenSpeedRow, model: string): TokenSpeedModelStats {
336 return r.models.find(bucket => bucket.model === model && !bucket.legacy)
337 ?? { model, outputTokens: 0, durationMs: 0, samples: 0, lastApi: null }
338}
339async function finish($: EngineInterface, active: LiveRequest, result: TurnStepResult): Promise<void> {
340 const endedAt = await $.clock.now()
341 const durationMs = endedAt - active.startedAt
342 const tokens = result.usage?.output_tokens
343 const validUsage = typeof tokens === 'number' && Number.isFinite(tokens) && tokens >= 0
344 const completed = result.stopReason !== null
345 const valid = completed && validUsage && Number.isFinite(durationMs) && durationMs > 0
346 const inputs = result.usage && [result.usage.input_tokens, result.usage.cache_read_input_tokens,
347 result.usage.cache_creation_input_tokens]
348 const inputTokens = completed && inputs && inputs.every(n => finite(n) && n >= 0)
349 ? inputs.reduce((sum, n) => sum + n, 0) : null
350 await mutate($, active.epoch, state => changeRow(state, active.loopId, r => {
351 if (r.active?.id !== active.id) return r
352 const old = bucketFor(r, active.model)
353 // The log prunes to the rolling day window; the aggregates stay all-time.
354 const bucket = valid
355 ? { ...old, log: withinWindow([...(old.log ?? []), { at: endedAt, tokens, ms: durationMs }], endedAt),
356 outputTokens: old.outputTokens + tokens, durationMs: old.durationMs + durationMs, samples: old.samples + 1,
357 lastApi: tokens * 1000 / durationMs }
358 : { ...old, lastApi: null }
359 return { ...r, models: [...r.models.filter(item => item.legacy || item.model !== active.model), bucket],
360 context: active.loopId === 'main' ? r.context
361 : childContext(active.model, inputTokens !== null && finite(inputTokens) ? inputTokens : null),
362 active: null, live: null, status: valid ? 'idle' : 'usage unavailable' }
363 }))
364}
365async function failed($: EngineInterface, active: LiveRequest, failure: TokenSpeedStatus): Promise<void> {
366 await mutate($, active.epoch, state => changeRow(state, active.loopId, r => {
367 if (r.active?.id !== active.id) return r
368 const old = bucketFor(r, active.model)
369 return { ...r, active: null, live: null, status: failure,
370 models: [...r.models.filter(bucket => bucket.legacy || bucket.model !== active.model), { ...old, lastApi: null }] }
371 }))
372}
373const rate = (value: number | null | undefined): string => value == null ? '—' : `${value.toFixed(1)} tok/s`
374const average = (tokens: number, ms: number): number | null => ms > 0 ? tokens * 1000 / ms : null
375function withinWindow(log: TokenSpeedSample[], now: number): TokenSpeedSample[] {
376 return log.filter(sample => sample.at > now - AVG_WINDOW_MS)
377}
378function windowedAverage(bucket: TokenSpeedModelStats, now: number): number | null {
379 // A migrated v0.2 bucket has no timestamps; its aggregates stand in until a new request.
380 if (bucket.log === undefined) return average(bucket.outputTokens, bucket.durationMs)
381 const held = withinWindow(bucket.log, now)
382 const tokens = held.reduce((sum, sample) => sum + sample.tokens, 0)
383 const ms = held.reduce((sum, sample) => sum + sample.ms, 0)
384 return ms > 0 ? tokens * 1000 / ms : null
385}
386async function safeNow($: EngineInterface): Promise<number> {
387 try { return await $.clock.now() } catch { return Number.MAX_SAFE_INTEGER }
388}
389const folderName = (path: string): string => {
390 const trimmed = path.replace(/\/+$/, '')
391 return trimmed.slice(trimmed.lastIndexOf('/') + 1) || path
392}
393function shortId(id: string, rows: TokenSpeedRow[]): string {
394 let length = Math.min(8, id.length)
395 while (length < id.length && rows.some(r => r.id !== id && r.id.slice(0, length) === id.slice(0, length))) length++
396 return id.slice(0, length)
397}
398const role = (r: TokenSpeedRow, rows: TokenSpeedRow[]): string => r.id === 'main' ? 'main' : `agent:${shortId(r.id, rows)}`
399const statusColors: Record<TokenSpeedStatus, ThemeKey> = {
400 waiting: 'inactive', streaming: 'inactive', idle: 'inactive',
401 aborted: 'warning', error: 'error', 'usage unavailable': 'warning',
402}
403
404async function appliedEffort($: EngineInterface, heldEpoch: number, id: string, value: unknown): Promise<void> {
405 const effort = effortValue(value)
406 if (effort === null || heldEpoch !== epoch) return
407 await mutate($, heldEpoch, state => changeRow(state, id, r => ({ ...r, effort, effortSource: 'applied' })))
408}
409const humanWindow = (window: number): string =>
410 window >= 1e6 ? `${(window / 1e6).toFixed(1)}M` : window >= 1e3 ? `${Math.round(window / 1e3)}k` : `${window}`
411const contextLabel = (context: TokenSpeedContextInfo): string => context.window > 0
412 ? `${context.percent === null ? '—' : Math.round(context.percent)}%/${humanWindow(context.window)}`
413 : context.tokens === null ? '—/?' : `${(context.tokens / 1000).toFixed(1)}k/?`
414const displayModel = (model: string | null): string => model?.slice(model.lastIndexOf('/') + 1) ?? 'unknown'
415// Terminal cells, not UTF-16 length. Combining marks/ZWJ consume no extra cell;
416// the common East Asian and emoji ranges are wide (including ⚡).
417function cellWidth(text: string): number {
418 let width = 0
419 let joined = false
420 for (const char of text) {
421 const n = char.codePointAt(0)!
422 if (n === 0x200d) { joined = true; continue }
423 if (/\p{Mark}/u.test(char) || n === 0xfe0f || n < 32 || n === 127) continue
424 if (joined) { joined = false; continue }
425 width += n >= 0x1100 && (n <= 0x115f || n === 0x2329 || n === 0x232a
426 || (n >= 0x2e80 && n <= 0xa4cf && n !== 0x303f) || (n >= 0xac00 && n <= 0xd7a3)
427 || (n >= 0xf900 && n <= 0xfaff) || (n >= 0xfe10 && n <= 0xfe19)
428 || (n >= 0xfe30 && n <= 0xfe6f) || (n >= 0xff00 && n <= 0xff60)
429 || (n >= 0xffe0 && n <= 0xffe6) || (n >= 0x1f000 && n <= 0x1faff)
430 || (n >= 0x20000 && n <= 0x3fffd)) || n === 0x26a1 ? 2 : 1
431 }
432 return width
433}
434const padEndTo = (text: string, width: number): string => text + ' '.repeat(Math.max(0, width - cellWidth(text)))
435const padStartTo = (text: string, width: number): string => ' '.repeat(Math.max(0, width - cellWidth(text))) + text
436function fitText(text: string, width: number): string {
437 if (cellWidth(text) <= width) return text
438 if (width <= 0) return ''
439 let fitted = ''
440 for (const char of text) {
441 if (cellWidth(fitted + char) > width - 1) break
442 fitted += char
443 }
444 return fitted + '…'
445}
446
447export const register: Register = on => {
448 on('session.start', async ($, e, next) => {
449 const rest = await next(e)
450 await observe(async () => {
451 registry = await loadRegistry(() => $.fs.read(`${$.plugin.root}/data/model-contexts.json`))
452 await $.command.register({ name: 'tok-speed', description: '各代理独立按请求模型统计;主控 first,reset 清空本会话统计', argumentHint: '[reset]' }).catch(() => {})
453 await mutate($, epoch, state => ({ version: 2, session: state.session, rows: state.rows.map(r => {
454 revision = Math.max(revision, r.revision)
455 return r.active && r.active.id !== liveRequests.get(r.id)?.id ? { ...r, active: null, live: null, status: 'aborted' } : r
456 }) }))
457 ensureTimer($)
458 await reconcile($, true)
459 await observe(() => refreshMain($, epoch))
460 await observe(() => refreshWorkspace($))
461 })
462 return rest
463 })
464 on('session.end', async ($, e, next) => {
465 epoch++
466 liveRequests.clear()
467 stopTimer()
468 await observe(() => update($, snapshot, empty))
469 return next(e)
470 })
471 on('agent.spawn', async ($, e, next) => {
472 const heldEpoch = epoch
473 const rest = await next(e) // Once only; metadata failure must not repeat a spawn.
474 await observe(async () => {
475 if (rest.deny !== undefined || !rest.agentId || heldEpoch !== epoch) return
476 const id = rest.agentId
477 // Remote workflow agents have no local list/loop. Do not invent rows for them.
478 const agents: AgentInfo[] = await $.agent.list()
479 const local = agents.find(agent => agent.id === id)
480 await mutate($, heldEpoch, state => {
481 const existing = state.rows.find(r => r.id === id)
482 if (!existing && local?.status !== 'running') return state
483 return changeRow(state, id, r => ({ ...r, description: e.description,
484 currentModel: r.currentModel ?? canonical(rest.model),
485 context: r.currentModel === null ? childContext(rest.model) : r.context,
486 effort: r.currentModel === null ? suffixEffort(rest.model) : r.effort,
487 effortSource: r.currentModel === null ? suffixEffort(rest.model) === null ? 'unknown' : 'configured' : r.effortSource,
488 running: existing ? r.running : true, seen: true }))
489 })
490 ensureTimer($)
491 })
492 return rest
493 })
494 on('turn.step', async function* ($, e, next) {
495 const heldEpoch = epoch
496 const loopId = e.agentId ?? 'main'
497 const startedRevision = ++revision
498 let active: LiveRequest | null = null
499 let completed = false
500 let failureStatus: TokenSpeedStatus = 'aborted'
501 await observe(async () => {
502 // Even when clock.now fails, retain loop metadata, but never invent Live.
503 let model = canonical(e.model)
504 const requested = effortValue(e.effort)
505 const fallback = requested === null
506 ? loopId === 'main' ? await configuredEffort($, e.model) : suffixEffort(e.model) : null
507 await mutate($, heldEpoch, state => changeRow(state, loopId, r => {
508 model = canonical(e.model, r.currentModel)
509 const effort = requested ?? fallback
510 return { ...r, running: true, currentModel: model, seen: true, turnId: e.turnId,
511 context: loopId === 'main' || model === r.currentModel ? r.context : childContext(model),
512 revision: startedRevision, active: null, live: null, status: 'waiting', effort,
513 effortSource: requested !== null ? 'request' : effort !== null ? 'configured' : 'unknown' }
514 }))
515 const startedAt = await $.clock.now()
516 if (heldEpoch !== epoch) return
517 const request: LiveRequest = { id: `${e.turnId}:${e.index}:${++sequence}`, turnId: e.turnId,
518 model, startedAt, loopId, epoch: heldEpoch, chunks: [], hasContent: false }
519 active = request
520 liveRequests.set(loopId, request)
521 await mutate($, heldEpoch, state => changeRow(state, loopId, r => r.revision !== startedRevision ? r : { ...r,
522 active: { id: request.id, turnId: e.turnId, model: request.model, startedAt } }))
523 ensureTimer($)
524 })
525 try {
526 const stream = next(e)
527 for await (const chunk of stream) {
528 if (active) await observe(() => observeChunk($, active!, chunk))
529 yield chunk
530 }
531 const result = await stream.result
532 if (active) {
533 if (next.signal.aborted) await observe(() => failed($, active!, 'aborted'))
534 else {
535 try { await finish($, active, result) }
536 catch { await observe(() => failed($, active!, 'usage unavailable')) }
537 }
538 }
539 completed = true
540 return result
541 } catch (error) {
542 failureStatus = next.signal.aborted || (error instanceof Error && error.name === 'AbortError') ? 'aborted' : 'error'
543 throw error
544 } finally {
545 if (active) {
546 if (owns(active)) liveRequests.delete(loopId)
547 if (!completed) await observe(() => failed($, active!, failureStatus))
548 }
549 }
550 })
551 on('turn.complete', async ($, e, next) => {
552 const heldEpoch = epoch
553 const rest = await next(e)
554 await observe(async () => {
555 const id = e.agentId ?? 'main'
556 const active = liveRequests.get(id)
557 if (active?.turnId === e.turnId) liveRequests.delete(id)
558 const endedRevision = ++revision
559 await mutate($, heldEpoch, state => {
560 const held = state.rows.find(r => r.id === id)
561 if (!held || (held.turnId !== e.turnId && !(id !== 'main' && held.turnId === null && held.running))) return state
562 return changeRow(state, id, r => ({ ...r, running: false, active: null, live: null,
563 turnId: e.turnId, revision: endedRevision,
564 status: e.reason === 'error' ? 'error' : e.reason === 'aborted' ? 'aborted' : 'idle' }))
565 })
566 })
567 // Turn usage is a turn total, never another request sample.
568 return rest
569 })
570 on('classic.PostToolUse', async ($, e, next) => {
571 const heldEpoch = epoch
572 const rest = await next(e)
573 await observe(() => appliedEffort($, heldEpoch, e.agent_id ?? 'main', e.effort?.level))
574 return rest
575 })
576 on('classic.Stop', async ($, e, next) => {
577 const heldEpoch = epoch
578 const rest = await next(e)
579 await observe(() => appliedEffort($, heldEpoch, e.agent_id ?? 'main', e.effort?.level))
580 return rest
581 })
582 on('session.compact', async ($, e, next) => {
583 const heldEpoch = epoch
584 const rest = await next(e)
585 if (e.trigger !== 'precompute' && !('skip' in rest)) await observe(() => mutate($, heldEpoch, state => {
586 const changed = changeRow(state, e.agentId ?? 'main', r => ({ ...r,
587 context: e.agentId === undefined ? emptyContext() : childContext(r.currentModel) }))
588 return e.agentId === undefined ? { ...changed, session: { ...state.session, context: emptyContext() } } : changed
589 }))
590 return rest
591 })
592 on('command.run', { command: 'tok-speed' }, async ($, e) => {
593 const args = e.args.trim()
594 if (args && args !== 'reset') return { text: '用法:/tok-speed 或 /tok-speed reset' }
595 if (args === 'reset') {
596 let refused = liveRequests.size > 0
597 await mutate($, epoch, state => {
598 refused = refused || state.rows.some(r => r.active !== null)
599 if (refused) return state
600 return { version: 2, session: state.session, rows: state.rows.filter(r => r.id === 'main' || r.running).map(r => ({ ...r,
601 models: [], active: null, live: null, seen: r.running, status: r.running ? 'waiting' : 'idle' })) }
602 })
603 return { text: refused ? '仍有代理请求在进行,无法 reset;请在请求结束后重试。' : '已重置本会话 token-speed 统计。' }
604 }
605 const state = await readState($)
606 const now = await safeNow($)
607 const provenance = state.rows.filter(r => r.id !== 'main').flatMap(r => {
608 const reading = lookupWindow(registry, r.currentModel)
609 return reading.sourceURL ? [`${role(r, state.rows)} · context source ${reading.sourceURL} · checked ${reading.date}`] : []
610 })
611 const lines = state.rows.flatMap(r => [
612 `${role(r, state.rows)} · ${r.description || (r.id === 'main' ? '主控' : '代理')} · ${r.running ? 'running' : 'ended'} · ${r.currentModel ?? 'unknown'} · effort ${r.effort ?? '—'} (${r.effortSource}) · Ctx ${contextLabel(r.context)} (${r.context.source}) · Live ${r.live === null ? '—' : `~${rate(r.live)}`} · ${r.status}`,
613 ...r.models.map(bucket => {
614 const held = bucket.log === undefined ? null : withinWindow(bucket.log, now)
615 const scope = held === null
616 ? `累计 ${bucket.samples} 个有效请求 · ${bucket.outputTokens} tokens / ${(bucket.durationMs / 1000).toFixed(3)}s`
617 : `24h ${held.length} 个请求 · ${held.reduce((sum, sample) => sum + sample.tokens, 0)} tokens / ${(held.reduce((sum, sample) => sum + sample.ms, 0) / 1000).toFixed(3)}s`
618 return `${role(r, state.rows)} · ${bucket.model}${bucket.legacy ? ' [legacy v0.1 response model]' : ''} · Last ${rate(bucket.lastApi)} · Avg ${rate(windowedAverage(bucket, now))} · ${scope}`
619 }),
620 ])
621 return { text: [
622 'token-speed 0.3.2 · 各代理独立累计 · 主控 first',
623 '各代理同一行显示模型、effort、Ctx、Live/Last/Avg(tok/s 只标一次);工作区最后一行。effort 来源:request 为 turn.step 请求,applied 为 classic 实际档位,configured 为主控配置或明确模型后缀,unknown 为 —。主控 Ctx 来自 session.usage;子代理输入来自最近 CLI uncached+cache读写总量,窗口从 data/model-contexts.json 精确匹配模型尾名/aliases,明确 override 优先官方 default,来源分别 cli-input-window-config、cli-input-official-default 与 cli-input-official-capacity,均不冒充 API 实测窗口;未匹配模型显示 k/?。暗底块状条固定 10 格,每格 8 档共 80 视觉档位(相邻整数百分比可能落在同一档,数字仍为整数 1%),填充固定使用 theme `rate_limit_fill`,轨道为 `rate_limit_empty`,窄屏可省条;无占用读数显示 —%。',
624 `context registry v${registry.version}${registry.fallback ? ' (load failed: confirmed overrides only)' : ''} · 优先实际 session 窗口、明确 override、官方 default(未给默认时明确 capacity)、未知;官方 max 仅元数据,不因 [1m] 自动选择。`,
625 ...provenance,
626 ...lines,
627 ...(state.rows.some(r => r.models.length) ? [] : ['Last — · Avg — · 尚无本会话统计']),
628 'Live:最近 3 秒收到的文本、thinking、工具 JSON 参数估算;CJK 每 Unicode code point 1 token,其余 0.25。分母 min(3 秒, 请求已耗时),不足 250ms 为 —;停流 3 秒后为 0。~ 非 tokenizer 精确计数。无法观察的代理流为 —,不推算速度。',
629 'Last/Avg:CLI usage.output_tokens / 请求全程耗时,包含 TTFT、thinking、网络与请求内停顿;排除请求之间的工具/用户空闲,非纯解码速度。Last 为该代理该模型最近一次有效请求;Avg 为最近 24 小时滚动窗口内该模型的有效 token 总数 / 有效耗时总和,按请求完成时间戳(clock epoch ms)判窗,与本地时区无关,出窗即弃。迁移自 v0.2 的无时间戳桶在首个新请求前退回全程累计口径;失败/中断、无有效 usage、耗时不大于 0 不入统计;有效 0 计入。CPA 缺字段可能被 CLI 归零,无法确认其来源;turn.complete usage 不重复累计。',
630 '模型名:以请求 model 为准,去除末尾 [1m]、(high) 或 :high 等配置后缀;统计桶与命令保留已知渠道,UI 仅隐藏最后 / 之前的渠道,不推断渠道。同代理 bare 模型与已知渠道模型尾名相同时沿用渠道;usage.model 不改变标签或桶。旧 v0.1 response-model 均值标为 legacy,保留且不混入新的请求模型均值。',
631 'UI 主控始终第一,活跃子代理按首次观察顺序每个一行;切换 view 不过滤。工具间隙保留行,Live —;结束立即隐藏,命令可查看历史。reload 保留统计、清除无法接管的旧流,以本 session agent.list 同步元数据;/clear 清空,reset 任意 active 请求时拒绝、工具间隙保留活跃代理元数据。',
632 ].join('\n') }
633 })
634 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
635 const rest = await next(e)
636 const state = await readState($)
637 if (e.props.hasSurvey || e.props.maxRows <= 0 || !state.rows.some(r => r.seen || r.running)) return rest
638 const visible = state.rows.filter(r => r.id === 'main' || r.running)
639 const { Box, Text } = $.ui.resolve(e)
640 const now = await safeNow($)
641 const budget = e.props.maxRows - visible.length
642 const effortColors: Record<Exclude<TokenSpeedEffort, number>, ThemeKey> =
643 { low: 'success', medium: 'planMode', high: 'warning', xhigh: 'error', max: 'error' }
644 const workspace = state.session.workspace
645 const workspaceLine = budget >= 1 && workspace
646 ? <Text key="token-speed-workspace" color="inactive" wrap="truncate-end">
647 {fitText(`⌂ ${workspace.name}${workspace.branch ? ` on ${workspace.branch}` : ''}`, e.props.bodyColumns)}
648 </Text>
649 : null
650 const labels = visible.map(r => r.id === 'main' ? '⚡ main' : `↳ ${shortId(r.id, state.rows)}`)
651 const models = visible.map(r => displayModel(r.currentModel))
652 const efforts = visible.map(r => r.effort === null ? '—' : String(r.effort))
653 const contexts = visible.map(r => contextLabel(r.context))
654 const buckets = visible.map(r => r.models.find(item => item.model === r.currentModel && !item.legacy))
655 const averages = buckets.map(bucket => bucket ? windowedAverage(bucket, now) : null)
656 const number = (n: number | null | undefined): string => n == null ? '—' : n.toFixed(1)
657 const lives = visible.map(r => r.live === null ? '—' : `~${number(r.live)}`)
658 const lasts = buckets.map(bucket => number(bucket?.lastApi))
659 const avgs = averages.map(number)
660 const maxWidth = (values: string[], minimum = 0): number => Math.max(minimum, ...values.map(cellWidth))
661 const multi = visible.length > 1
662 const labelWidth = maxWidth(labels)
663 let modelWidth = maxWidth(models)
664 const effortWidth = maxWidth(efforts)
665 const contextWidth = maxWidth(contexts)
666 const liveWidth = maxWidth(lives, 5)
667 const lastWidth = maxWidth(lasts, 5)
668 const avgWidth = maxWidth(avgs, 5)
669 const statusWidth = maxWidth(visible.map(r => r.status))
670 const columns = Math.max(1, e.props.bodyColumns)
671 let showStatus = true
672 let showLast = true
673 // Fixed 10-cell bar: 80 eighth-cell steps; no widening on wide terminals.
674 let cells = 10
675 const width = (bar: number): number => labelWidth + 3 + modelWidth + 3 + effortWidth
676 + 7 + contextWidth + (bar ? bar + 1 : 0)
677 + 8 + liveWidth + (showLast ? 8 + lastWidth : 0) + 7 + avgWidth + 6
678 + (showStatus ? 3 + statusWidth : 0)
679 // Keep a usable compact meter first; discard status, then Last, before the bar.
680 if (width(cells) > columns) showStatus = false
681 if (width(cells) > columns) showLast = false
682 if (width(cells) > columns) cells = 0
683 if (width(cells) > columns) modelWidth = Math.max(1, modelWidth - (width(cells) - columns))
684 // At extremely small widths preserve the numeric context/effort and rates;
685 // model alone can be abbreviated after optional columns have been exhausted.
686 const aligned = (text: string, size: number): string => multi ? padEndTo(text, size) : text
687 return <Box flexDirection="column">
688 {visible.map((r, index) => {
689 const percent = r.context.percent
690 const known = r.context.window > 0 && percent !== null
691 const eighths = known ? Math.max(0, Math.min(cells * 8, Math.round(percent * cells * 8 / 100))) : 0
692 const full = Math.floor(eighths / 8)
693 const partial = ['', '▏', '▎', '▍', '▌', '▋', '▊', '▉'][eighths % 8] ?? ''
694 const fill = '█'.repeat(full) + partial
695 const track = ' '.repeat(cells - full - (partial ? 1 : 0))
696 const effortColor = r.effort === null ? 'inactive' : typeof r.effort === 'number' ? 'subtle' : effortColors[r.effort]
697 return <Text key={`token-speed-${r.id}`} color="inactive" wrap="truncate-end">
698 <Text color={r.id === 'main' ? 'claude' : 'inactive'} bold={r.id === 'main'}>{aligned(labels[index] ?? '', labelWidth)}</Text>
699 {' · '}<Text color="text">{aligned(fitText(models[index] ?? '', modelWidth), modelWidth)}</Text>
700 {' · '}<Text color={effortColor}>{aligned(efforts[index] ?? '—', effortWidth)}</Text>
701 {' · Ctx '}{cells === 0 ? null : <Text><Text key={`context-meter-${r.id}`} backgroundColor="rate_limit_empty"><Text color="rate_limit_fill">{fill}</Text>{track}</Text>{' '}</Text>}
702 <Text color={known ? 'text' : 'inactive'}>{aligned(contexts[index] ?? '—/?', contextWidth)}</Text>
703 {' · Live '}<Text color={r.live === null ? 'inactive' : 'text'} bold={r.live !== null}>{padStartTo(lives[index] ?? '—', liveWidth)}</Text>
704 {showLast ? <Text>{' · Last '}<Text color={buckets[index]?.lastApi == null ? 'inactive' : 'text'}>{padStartTo(lasts[index] ?? '—', lastWidth)}</Text></Text> : null}
705 {' · Avg '}<Text color={averages[index] == null ? 'inactive' : 'text'}>{padStartTo(avgs[index] ?? '—', avgWidth)}</Text>{' tok/s'}
706 {showStatus ? <Text>{' · '}<Text color={statusColors[r.status]}>{r.status}</Text></Text> : null}
707 </Text>
708 })}
709 {workspaceLine}
710 {rest}
711 </Box>
712 })
713}
714hooks/context-registry.ts 105 lines1export type OfficialContext = {
2 id: string
3 aliases: string[]
4 defaultWindow: number | null
5 maxWindow?: number
6 contextWindow?: number
7 sourceScope?: string
8 additionalSources?: string[]
9 apiContextWindow?: number
10 maxInputTokens?: number
11 sourceValue?: string
12 sourceURL: string
13 date: string
14}
15export type ContextOverride = { id: string; aliases: string[]; window: number; reason: string }
16export type ContextRegistry = { version: 1; models: OfficialContext[]; overrides: ContextOverride[]; fallback?: true }
17export type WindowReading = { window: number; source: 'configured' | 'official-default' | 'official-capacity' | 'unknown'; sourceURL?: string; date?: string }
18
19// Keep the already-confirmed user windows if the packaged registry cannot load.
20const FALLBACK_OVERRIDES: ContextOverride[] = [
21 { id: 'gpt-6.1-sol', aliases: [], window: 272000, reason: 'User-confirmed CLI window' },
22 { id: 'glm-5.3', aliases: [], window: 1000000, reason: 'User-confirmed CLI window' },
23 { id: 'glm-5.3-flash', aliases: [], window: 1000000, reason: 'User-confirmed CLI window' },
24]
25const object = (v: unknown): v is Record<string, unknown> => typeof v === 'object' && v !== null && !Array.isArray(v)
26const integer = (v: unknown): v is number => typeof v === 'number' && Number.isSafeInteger(v) && v > 0
27const name = (v: unknown): v is string => typeof v === 'string' && /^[a-z0-9][a-z0-9._:-]*$/i.test(v)
28const sourceURL = (v: unknown): v is string => typeof v === 'string' && /^https:\/\/[a-z0-9][a-z0-9.-]*(?:\/[\x21-\x7e]*)?$/i.test(v)
29const fail = (): never => { throw new Error('Invalid model context registry') }
30function aliases(v: unknown): string[] {
31 if (!Array.isArray(v) || !v.every(name)) return fail()
32 return v.map(s => s.toLowerCase())
33}
34function fields(v: Record<string, unknown>, keys: string[]): void {
35 if (Object.keys(v).some(k => !keys.includes(k))) fail()
36}
37function date(v: unknown): v is string {
38 if (typeof v !== 'string' || !/^\d{4}-\d{2}-\d{2}$/.test(v)) return false
39 const parsed = new Date(v + 'T00:00:00Z')
40 return Number.isFinite(parsed.getTime()) && parsed.toISOString().slice(0, 10) === v
41}
42function unique(entries: { id: string; aliases: string[] }[]): void {
43 const seen = new Set<string>()
44 for (const entry of entries) for (const alias of [entry.id, ...entry.aliases]) {
45 if (seen.has(alias)) fail()
46 seen.add(alias)
47 }
48}
49export function parseRegistry(text: string): ContextRegistry {
50 const value: unknown = JSON.parse(text)
51 if (!object(value) || value.version !== 1 || !Array.isArray(value.models) || !Array.isArray(value.overrides)) return fail()
52 fields(value, ['version', 'models', 'overrides'])
53 const models: OfficialContext[] = value.models.map(v => {
54 if (!object(v) || !name(v.id) || (v.defaultWindow !== null && !integer(v.defaultWindow)) || !date(v.date)
55 || !sourceURL(v.sourceURL) || (v.defaultWindow === null && !integer(v.contextWindow))) return fail()
56 for (const key of ['maxWindow', 'contextWindow', 'apiContextWindow', 'maxInputTokens']) {
57 if (v[key] !== undefined && !integer(v[key])) fail()
58 }
59 if (v.maxWindow !== undefined && v.defaultWindow !== null && (v.maxWindow as number) < (v.defaultWindow as number)) fail()
60 if (v.additionalSources !== undefined && (!Array.isArray(v.additionalSources) || !v.additionalSources.every(sourceURL))) fail()
61 if ((v.sourceScope !== undefined && (typeof v.sourceScope !== 'string' || !v.sourceScope.trim()))
62 || (v.sourceValue !== undefined && (typeof v.sourceValue !== 'string' || !v.sourceValue.trim()))) fail()
63 fields(v, ['id', 'aliases', 'defaultWindow', 'maxWindow', 'contextWindow', 'sourceScope', 'additionalSources',
64 'apiContextWindow', 'maxInputTokens', 'sourceValue', 'sourceURL', 'date'])
65 return { id: v.id.toLowerCase(), aliases: aliases(v.aliases), defaultWindow: v.defaultWindow as number | null,
66 ...(v.maxWindow === undefined ? {} : { maxWindow: v.maxWindow as number }),
67 ...(v.contextWindow === undefined ? {} : { contextWindow: v.contextWindow as number }),
68 ...(v.apiContextWindow === undefined ? {} : { apiContextWindow: v.apiContextWindow as number }),
69 ...(v.maxInputTokens === undefined ? {} : { maxInputTokens: v.maxInputTokens as number }),
70 ...(v.sourceScope === undefined ? {} : { sourceScope: v.sourceScope as string }),
71 ...(v.sourceValue === undefined ? {} : { sourceValue: v.sourceValue as string }),
72 ...(v.additionalSources === undefined ? {} : { additionalSources: v.additionalSources as string[] }),
73 sourceURL: v.sourceURL, date: v.date }
74 })
75 const overrides: ContextOverride[] = value.overrides.map(v => {
76 if (!object(v) || !name(v.id) || !integer(v.window) || typeof v.reason !== 'string' || !v.reason.trim()) return fail()
77 fields(v, ['id', 'aliases', 'window', 'reason'])
78 return { id: v.id.toLowerCase(), aliases: aliases(v.aliases), window: v.window, reason: v.reason }
79 })
80 // A deliberate override may share an official id; ambiguity within either tier is refused.
81 unique(models)
82 unique(overrides)
83 return { version: 1, models, overrides }
84}
85export function fallbackRegistry(): ContextRegistry {
86 return { version: 1, models: [], overrides: FALLBACK_OVERRIDES.map(v => ({ ...v, aliases: [...v.aliases] })), fallback: true }
87}
88export function lookupWindow(registry: ContextRegistry, canonicalModel: string | null): WindowReading {
89 const tail = canonicalModel?.slice(canonicalModel.lastIndexOf('/') + 1).toLowerCase() ?? ''
90 const matches = (entry: { id: string; aliases: string[] }) => entry.id === tail || entry.aliases.includes(tail)
91 const override = registry.overrides.find(matches)
92 if (override) return { window: override.window, source: 'configured' }
93 const official = registry.models.find(matches)
94 return official ? { window: official.defaultWindow ?? official.contextWindow ?? 0,
95 source: official.defaultWindow === null ? 'official-capacity' : 'official-default', sourceURL: official.sourceURL, date: official.date }
96 : { window: 0, source: 'unknown' }
97}
98export function registryLoader(): (read: () => Promise<string>) => Promise<ContextRegistry> {
99 let pending: Promise<ContextRegistry> | null = null
100 return read => pending ??= (async () => {
101 try { return parseRegistry(await read()) }
102 catch { return fallbackRegistry() }
103 })()
104}
105types/index.d.ts 85 lines1export type TokenSpeedStatus = 'idle' | 'waiting' | 'streaming' | 'aborted' | 'error' | 'usage unavailable'
2
3export type TokenSpeedEffort = 'low' | 'medium' | 'high' | 'xhigh' | 'max' | number
4
5export type TokenSpeedSample = {
6 /** $.clock.now() epoch ms when the request finished. */
7 at: number
8 tokens: number
9 ms: number
10}
11
12export type TokenSpeedModelStats = {
13 model: string
14 outputTokens: number
15 durationMs: number
16 samples: number
17 lastApi: number | null
18 /** v0.1 used response models, so its averages remain separately identified. */
19 legacy?: true
20 /**
21 * v0.3+ per-request log backing the rolling 24h Avg. Absent on v0.1 legacy
22 * buckets and on v0.2 state migrated before its first new request.
23 */
24 log?: TokenSpeedSample[]
25}
26
27export type TokenSpeedActive = {
28 id: string
29 turnId: string
30 model: string
31 startedAt: number
32}
33
34export type TokenSpeedRow = {
35 /** main, or the session-local agentId. Array order is first observation order. */
36 id: string
37 description: string
38 running: boolean
39 currentModel: string | null
40 models: TokenSpeedModelStats[]
41 active: TokenSpeedActive | null
42 live: number | null
43 status: TokenSpeedStatus
44 seen: boolean
45 turnId: string | null
46 /** Lifecycle guard for asynchronously returned roster snapshots. */
47 revision: number
48 /** Observed request/applied effort, or an explicit configuration fallback. */
49 effort: TokenSpeedEffort | null
50 effortSource: 'configured' | 'request' | 'applied' | 'unknown'
51 /** Each loop owns its latest reading; child windows are explicit configuration. */
52 context: TokenSpeedContextInfo
53}
54
55export type TokenSpeedContextInfo = {
56 /** Last response's uncached + cache-read + cache-written inputs, never a turn sum. */
57 tokens: number | null
58 /** Zero means unknown, never a guessed model/context cap. */
59 window: number
60 percent: number | null
61 source: 'session' | 'cli-input' | 'cli-input-window-config' | 'cli-input-official-default' | 'cli-input-official-capacity' | 'unknown'
62}
63
64export type TokenSpeedWorkspace = { name: string; branch: string | null }
65
66export type TokenSpeedSessionInfo = {
67 context: TokenSpeedContextInfo
68 workspace: TokenSpeedWorkspace | null
69}
70
71export type TokenSpeedSnapshot = {
72 version: 2
73 rows: TokenSpeedRow[]
74 session: TokenSpeedSessionInfo
75}
76
77/** Read only compatibility shape; all writes use version 2. */
78export type TokenSpeedLegacySnapshot = Pick<TokenSpeedRow, 'models' | 'currentModel' | 'active' | 'live' | 'status' | 'seen'>
79
80declare module 'claude-code' {
81 interface PluginState {
82 'token-speed': { snapshot: TokenSpeedSnapshot }
83 }
84}
85