在输入框上方实时显示 token 用量、输出速度、缓存命中率、订阅额度和费用,并检测模型是否被切换或降级

English · Chinese (Simplified)
My own Claude Code mods.
A live token-usage band above the prompt box:

Click the collapse button and the whole band disappears; only a 🪙 is left at the right of the toolbar below the prompt box. Click it to expand again:

271.1k/1.00M (27%) · this turn 3.3k · 113 tok/s · total in 7.92M · total out 38.4k · cache 98% · $8.58 · 5h 99% · 7d 70% ✅ [collapse]
The band's labels in the screenshots are currently Chinese; the line above shows the same fields in English, in the same order.
| Item | Meaning |
|---|---|
271.1k/1.00M (27%) | Current context size / model window |
| this turn | Output tokens generated so far this turn, ticking up while the reply streams |
tok/s | Average output speed: total output ÷ total generation time (first to last token, excluding time to first token) |
| total in / total out | Session totals, including cache reads/writes and subagents |
| cache | Share of total input served from cache; higher is cheaper |
$ | Session cost |
5h / 7d | Share of the 5-hour / 7-day subscription limit used (hidden on API billing) |
✅ / ❌ | Model check: turns ❌ and shows a toast with the reason when the main thread was switched to another model, the model returned by the API differs from the one requested, or thinking effort was silently lowered |
/tokens to toggle between expanded and collapsed.Adds a ⚡ button to the right of the model name. It opens a side panel so you don't have to type /xxx by hand:
/compact, /clear, /context, /cost, /resume, …) run with one click./name into the prompt box so you can add arguments and send it yourself. Nothing fires by accident.⋯ next to a built-in lets you type arguments; when a command's output lists "Available: a, b, c" the mod remembers it and shows those as buttons next time./config settings as click targets. Toggles get On / Off, pick-one settings get one button per option, text and number settings get an input box. The current value has a blue background.language in settings.json or the system language by default; the 🌐 button switches between Auto / Chinese / English by hand. In Chinese, the English command and setting descriptions are translated with haiku and cached. In English, the engine's own descriptions are shown untouched.The ⚡ button, to the right of the model name:

The panel in English (descriptions and setting names are the engine's originals, unchanged):

One line (paste into a terminal):
claude plugin marketplace add crebot51/jack-mods && claude plugin install token-band@jack-mods
To install quick-commands, swap the last name for quick-commands@jack-mods (no need to add the marketplace again if you already have).
Then start a new session. token-band's band shows up on desktop only after the first message; quick-commands' ⚡ sits to the right of the model name.
Update: claude plugin update <name> Uninstall: claude plugin uninstall <name>
The mod API is an early preview and may need updates across Claude Code versions.
hooks/register.tsx 306 lines1import { atom, read, update } from 'claude-code'
2import type { Register, TurnUsage } from 'claude-code'
3
4import type { ModelCheck, TokenStats } from '../types'
5
6const EMPTY: TokenStats = {
7 totalInput: 0,
8 totalCacheRead: 0,
9 cacheInput: 0,
10 output: 0,
11 turnOutput: 0,
12 speedTokens: 0,
13 speedMs: 0,
14 contextTokens: null,
15 contextWindow: null,
16 rateLimits: [],
17 usd: null,
18}
19
20const stats = atom({ plugin: 'token-band', key: 'stats' } as const, EMPTY)
21const isHidden = atom({ plugin: 'token-band', key: 'isHidden' } as const, false)
22const check = atom({ plugin: 'token-band', key: 'check' } as const, {
23 issues: [],
24 effortWanted: null,
25} as ModelCheck)
26
27/** 不换行空格 */
28const NBSP = '\u00a0'
29
30/** 12345 -> "12.3k",1234567 -> "1.23M" */
31export function fmt(n: number): string {
32 if (n >= 1_000_000) return `${(n / 1_000_000).toFixed(2)}M`
33 if (n >= 1_000) return `${(n / 1_000).toFixed(1)}k`
34 return String(n)
35}
36
37/**
38 * 累加一次回复的用量。isMain:主对话的请求,它的输入量就是当前上下文占用;
39 * isTurnStart:主对话一轮的第一次请求,本轮计数从它清零
40 */
41export function addUsage(
42 s: TokenStats,
43 u: TurnUsage,
44 isMain: boolean,
45 isTurnStart: boolean,
46): TokenStats {
47 const input = u.input_tokens + u.cache_read_input_tokens + u.cache_creation_input_tokens
48 return {
49 ...s,
50 totalInput: (s.totalInput ?? 0) + input,
51 totalCacheRead: (s.totalCacheRead ?? 0) + u.cache_read_input_tokens,
52 cacheInput: (s.cacheInput ?? 0) + input,
53 output: s.output + u.output_tokens,
54 turnOutput: (isTurnStart ? 0 : s.turnOutput) + u.output_tokens,
55 contextTokens: isMain ? input : s.contextTokens,
56 }
57}
58
59/** 累加一次回复的测速样本;耗时为 0 的不计 */
60export function addSpeed(s: TokenStats, outputTokens: number, ms: number): TokenStats {
61 if (ms <= 0 || outputTokens <= 0) return s
62 return {
63 ...s,
64 speedTokens: (s.speedTokens ?? 0) + outputTokens,
65 speedMs: (s.speedMs ?? 0) + ms,
66 }
67}
68
69/** 缓存命中率(0~100),还没有输入时为 null */
70export function cacheHitRate(s: TokenStats): number | null {
71 const total = s.cacheInput ?? 0
72 return total > 0 ? ((s.totalCacheRead ?? 0) / total) * 100 : null
73}
74
75const LIMIT_LABEL: Record<string, string> = { five_hour: '5h', seven_day: '7d' }
76
77/** [{ five_hour, 23 }, { seven_day, 41.5 }] -> ['5h 23%', '7d 42%'];别的窗口原样用名字 */
78export function formatLimits(limits: readonly { kind: string; percentUsed: number }[]): string[] {
79 return limits.map(l => `${LIMIT_LABEL[l.kind] ?? l.kind} ${Math.round(l.percentUsed)}%`)
80}
81
82/** 平均输出速度(token/秒),没有样本时为 null */
83export function avgSpeed(s: TokenStats): number | null {
84 const ms = s.speedMs ?? 0
85 return ms > 0 ? ((s.speedTokens ?? 0) / ms) * 1000 : null
86}
87
88/** 'claude-sonnet-4-5-20250929[1m]' -> 'sonnet-4-5' */
89export function shortModel(model: string): string {
90 return model
91 .toLowerCase()
92 .replace(/\[.*\]$/, '')
93 .replace(/-\d{8}$/, '')
94 .replace(/^claude-/, '')
95}
96
97/**
98 * 两个模型名是否指同一个模型。别名(opus、sonnet、default 这类不带版本号的)
99 * 只要求对方属于同一系列,无法判断的一律当作一致,宁可漏报也不误报
100 */
101export function sameModel(a: string, b: string): boolean {
102 const x = shortModel(a)
103 const y = shortModel(b)
104 if (x === y) return true
105 const isAlias = (m: string) => !/\d/.test(m)
106 if (isAlias(x) || isAlias(y)) {
107 const [alias, full] = isAlias(x) ? [x, y] : [y, x]
108 return alias === 'default' || full.includes(alias)
109 }
110 return false
111}
112
113const EFFORT_RANK: Record<string, number> = { low: 1, medium: 2, high: 3, xhigh: 4, max: 5 }
114
115/** 实际强度比请求的低时返回「强度 high→medium」,否则 null */
116export function effortIssue(wanted: string | null, actual: string | undefined): string | null {
117 if (!wanted || !actual) return null
118 const w = EFFORT_RANK[wanted]
119 const a = EFFORT_RANK[actual]
120 if (w === undefined || a === undefined || a >= w) return null
121 return `强度 ${wanted}→${actual}`
122}
123
124/** 把新发现的异常并进本轮 */
125function addIssue(c: ModelCheck, issue: string): ModelCheck {
126 return c.issues.includes(issue) ? c : { ...c, issues: [...c.issues, issue] }
127}
128
129export const register: Register = on => {
130 on('session.start', async ($, e, next) => {
131 await $.command.register({
132 name: 'tokens',
133 description: '展开/收起输入框上方的 token 用量条',
134 })
135 const usage = await $.session.usage()
136 await update($, stats, s => ({
137 ...s,
138 contextTokens: usage.context.tokens ?? s.contextTokens ?? null,
139 contextWindow: usage.context.window,
140 rateLimits: usage.rateLimits.map(l => ({ kind: l.kind, percentUsed: l.percentUsed })),
141 usd: usage.cost?.usd ?? s.usd,
142 }))
143 return next(e)
144 })
145
146 on('command.run', { command: 'tokens' }, async $ => {
147 const hidden = await update($, isHidden, h => !h)
148 return { text: hidden ? 'token 用量条已收起' : 'token 用量条已展开' }
149 })
150
151 // 每次模型请求返回后立刻累加,所以一轮里会实时跳动。
152 // 测速从第一个内容片段到最后一片,不含等待首字的时间
153 on('turn.step', async function* ($, e, next) {
154 // 主对话一轮的第一次请求:新的一轮开始。不用 turn.start,免得子 agent 的轮次把本轮清零
155 const isTurnStart = e.agentId === undefined && e.index === 0
156 if (isTurnStart) {
157 await update($, check, c => ({ ...c, issues: [] }))
158 }
159
160 const stream = next(e)
161 let firstAt: number | null = null
162 for await (const chunk of stream) {
163 if (firstAt === null && chunk.kind !== 'engine') {
164 firstAt = await $.clock.now()
165 }
166 yield chunk
167 }
168 const endAt = await $.clock.now()
169 const result = await stream.result
170 const usage = result.usage
171 if (usage) {
172 const ms = firstAt === null ? 0 : endAt - firstAt
173 await update($, stats, s =>
174 addSpeed(
175 addUsage(s, usage, e.agentId === undefined, isTurnStart),
176 usage.output_tokens,
177 ms,
178 ),
179 )
180 }
181
182 // 模型一致性只看主对话:子 agent 本来就可能用别的模型
183 if (e.agentId === undefined) {
184 const selected = await $.session.model()
185 const found: string[] = []
186 if (!sameModel(selected, e.model)) found.push(`已切到 ${shortModel(e.model)}`)
187 if (usage && !sameModel(e.model, usage.model)) {
188 found.push(`实际回答 ${shortModel(usage.model)}`)
189 }
190 const wanted = typeof e.effort === 'string' ? e.effort : null
191 const before = await read($, check)
192 const after = await update($, check, c =>
193 found.reduce(addIssue, { ...c, effortWanted: wanted }),
194 )
195 if (after.issues.length > before.issues.length) {
196 $.ui.toast(`❌ 模型异常:${after.issues.join(' · ')}`)
197 }
198 }
199 return result
200 })
201
202 // 一轮结束时拿到「静默降级之后」的实际思考强度,和请求时比较
203 on('classic.Stop', async ($, e, next) => {
204 const c = await read($, check)
205 const issue = effortIssue(c.effortWanted, e.effort?.level)
206 if (issue) {
207 await update($, check, x => addIssue(x, issue))
208 $.ui.toast(`❌ 思考强度被调低:${issue.replace('强度 ', '')}`)
209 }
210 return next(e)
211 })
212
213 // 引擎测量后推送:校准上下文、额度和费用
214 on('session.measure', async ($, e, next) => {
215 await update($, stats, s => ({
216 ...s,
217 contextTokens: e.context.tokens ?? s.contextTokens ?? null,
218 contextWindow: e.context.window,
219 rateLimits: e.rateLimits.map(l => ({ kind: l.kind, percentUsed: l.percentUsed })),
220 usd: e.cost?.usd ?? s.usd,
221 }))
222 return next(e)
223 })
224
225 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
226 if (e.props.hasSurvey) {
227 return next(e)
228 }
229
230 // 收起时整栏不画(连桌面端的灰底一起消失),展开按钮在输入框下方工具栏右侧
231 if (await read($, isHidden)) {
232 return next(e)
233 }
234
235 const s = await read($, stats)
236 const { Box, Button, Text } = $.ui.resolve(e)
237
238 const avg = avgSpeed(s)
239 const ctx = !s.contextWindow
240 ? '—'
241 : s.contextTokens == null
242 ? `—/${fmt(s.contextWindow)}`
243 : `${fmt(s.contextTokens)}/${fmt(s.contextWindow)} (${Math.round(
244 (s.contextTokens / s.contextWindow) * 100,
245 )}%)`
246 const parts = [
247 ctx,
248 `本轮 ${fmt(s.turnOutput)}`,
249 `${avg === null ? '—' : Math.round(avg)} tok/s`,
250 `总输入 ${fmt(s.totalInput ?? 0)}`,
251 `总输出 ${fmt(s.output)}`,
252 ]
253 const hit = cacheHitRate(s)
254 if (hit !== null) parts.push(`缓存 ${Math.round(hit)}%`)
255 if (s.usd !== null) parts.push(`$${s.usd.toFixed(2)}`)
256 // 不是订阅账号时没有额度数据,这一段就不显示
257 parts.push(...formatLimits(s.rateLimits ?? []))
258 // 模型检测只用一个符号:正常 ✅,本轮有异常 ❌(细节在弹出的提示里)
259 const health = (await read($, check)).issues.length > 0 ? '❌' : '✅'
260
261 // 文字区可以收缩、自动折行;按钮不收缩,固定在第一行最右边
262 return (
263 <Box flexDirection="row" alignItems="flex-start" columnGap={1}>
264 <Box flexGrow={1} flexShrink={1}>
265 <Text dimColor>
266 {/* 每一项内部换成不换行空格,「·」也粘在前一项后面:只在「· 」之后折行,不会把一项劈成两半 */}
267 {parts.map(p => p.replace(/ /g, NBSP)).join(`${NBSP}· `)} {health}
268 </Text>
269 </Box>
270 <Box flexShrink={0}>
271 <Button
272 key="toggle"
273 label="收起"
274 dimColor
275 onPress={() => update($, isHidden, () => true)}
276 />
277 </Box>
278 </Box>
279 )
280 })
281
282 // 收起时:在输入框下方工具栏右侧的模式标签处加一个小按钮,原有内容保持不动
283 on('ui.render', { component: 'SessionMode' }, async ($, e, next) => {
284 const engine = await next(e)
285 if (!(await read($, isHidden))) {
286 return engine
287 }
288
289 const { Box, Button } = $.ui.resolve(e)
290 const hasIssue = (await read($, check)).issues.length > 0
291
292 return (
293 <Box>
294 {engine}
295 <Button
296 key="expand"
297 label={hasIssue ? '❌' : '🪙'}
298 plain
299 dimColor
300 onPress={() => update($, isHidden, () => false)}
301 />
302 </Box>
303 )
304 })
305}
306types/index.d.ts 36 lines1export type TokenStats = {
2 /** 本会话累计输入 token:新输入 + 缓存读 + 缓存写(含子 agent)。旧状态里可能没有 */
3 totalInput?: number
4 /** 缓存命中率的分子和分母:同一时刻开始统计,避免旧数据把比例拉偏。旧状态里可能没有 */
5 totalCacheRead?: number
6 cacheInput?: number
7 /** 本会话累计输出 token(含子 agent) */
8 output: number
9 /** 本轮已生成的输出 token(含子 agent) */
10 turnOutput: number
11 /** 测速用:计入测速的输出 token 与生成耗时(毫秒)。热重载前的旧状态里可能没有 */
12 speedTokens?: number
13 speedMs?: number
14 /** 当前上下文占用和窗口大小。旧状态里可能没有 */
15 contextTokens?: number | null
16 contextWindow?: number | null
17 /** 订阅额度:five_hour / seven_day 等窗口的已用百分比。旧状态里可能没有 */
18 rateLimits?: { kind: string; percentUsed: number }[]
19 /** 本会话费用(美元) */
20 usd: number | null
21}
22
23/** 模型一致性检查(只看主对话) */
24export type ModelCheck = {
25 /** 最近一轮发现的异常,例如「实际回答 sonnet-4-5」「强度 high→medium」 */
26 issues: string[]
27 /** 最近一次请求的思考强度(静默降级之前) */
28 effortWanted: string | null
29}
30
31declare module 'claude-code' {
32 interface PluginState {
33 'token-band': { stats: TokenStats; isHidden: boolean; check: ModelCheck }
34 }
35}
36