Shows live tokens per second, per minute and per hour in the status line

Version 0.2.0. See CHANGELOG.md.
A Claude Code mod that shows how fast Claude is writing, live, in the status line:
~95 tok/s | 1.2k tok/min | 8.4k tok/hr
~, an estimate from the text so far. When the reply ends it switches to the exact number from the API.It counts output tokens only (what Claude writes, including thinking and tool calls), not what it reads.
A mod is a plugin with code that runs inside Claude Code, so it can draw in the interface and react to events like tool calls and replies. Skills are instructions Claude reads; mods change Claude Code itself. They need Claude Code v2.1.287 or later. See Anthropic's mods overview.
The API only reports the token count when a reply ends, so while Claude writes, the meter counts characters instead. It starts at 3 characters per token, measured on Claude's replies, then learns the real ratio from each finished reply in your session. When Claude's thinking is hidden it costs tokens but shows no text, so the meter skips the live figure for that reply and shows the exact one when it ends.
Installed from this repo's marketplace and run in real sessions of three replies each (a story, a command, a story, a command, a story):
| Result | |
|---|---|
| Exact counts | Matched the API's own totals in every run (for example 1,462, 1,567 and 1,471 tokens) |
Live ~ estimate against the exact speed | Within about 15%: 73 vs 83, 77 vs 90, 84 vs 93, 92 vs 86, 90 vs 80, 104 vs 94 tok/s |
| Before the 0.2.0 fixes | 2.5k tok/s on replies that thought first, and live estimates 30 to 40% low |
6 automated tests cover the exact speed, the live estimate, learning the ratio, hidden thinking, the per-minute window and clearing the status line. Run them with claude plugin test mods/token-meter.
The test sessions ran headless, where the meter writes to the debug log instead of a status line. I couldn't open an interactive terminal on the test machine, so I haven't seen it drawn on screen myself.
In your terminal:
claude plugin marketplace add NotRedFox/NotRedFoxs-Claude-skills
claude plugin install token-meter@notredfox
Or inside a Claude Code session, /plugin marketplace add NotRedFox/NotRedFoxs-Claude-skills then /plugin install token-meter@notredfox, then /reload-plugins.
To check what it does before installing, clone this repo and run claude plugin validate mods/token-meter. It only reads replies and the clock and writes to the status line.
The terminal and the Code tab of the Claude Desktop app. It runs but shows nothing in the VS Code chat panel and claude -p.
To update it: claude plugin update token-meter@notredfox. To remove it: claude plugin uninstall token-meter@notredfox.
hooks/register.ts 106 lines1import type { EngineInterface, Register } from 'claude-code'
2
3// Output tokens are what "speed" means for a model, so the meter counts those.
4// While a reply is still streaming, the API hasn't reported its token count yet,
5// so the live figure is estimated from characters and marked with "~". It starts
6// at 3 characters per token (measured on Claude's replies) and learns the real
7// ratio from each finished reply. Hidden thinking costs tokens but streams no
8// characters, so those replies get no live figure and aren't learned from.
9const CHARS_PER_TOKEN = 3
10const MINUTE = 60_000
11const HOUR = 3_600_000
12// Shorter than this, a reply's rate is mostly noise (one or two chunks).
13const MIN_STREAM_MS = 200
14
15type Sample = { at: number; tokens: number }
16
17export function compact(n: number): string {
18 if (n >= 1_000_000) return `${(n / 1_000_000).toFixed(1)}M`
19 if (n >= 10_000) return `${Math.round(n / 1000)}k`
20 if (n >= 1000) return `${(n / 1000).toFixed(1)}k`
21 return `${Math.round(n)}`
22}
23
24export function meterText(tps: number | undefined, isLive: boolean, perMinute: number, perHour: number): string {
25 const speed = tps === undefined ? '- tok/s' : `${isLive ? '~' : ''}${compact(tps)} tok/s`
26 return `${speed} | ${compact(perMinute)} tok/min | ${compact(perHour)} tok/hr`
27}
28
29type Meter = {
30 samples: Sample[]
31 lastTps: number | undefined
32 live: { start: number; chars: number } | undefined
33 seen: { tokens: number; chars: number }
34}
35
36function tokensPerChar(m: Meter): number {
37 return m.seen.chars > 0 ? m.seen.tokens / m.seen.chars : 1 / CHARS_PER_TOKEN
38}
39
40async function show($: EngineInterface, m: Meter) {
41 const now = await $.clock.now()
42 while (m.samples.length > 0 && m.samples[0].at <= now - HOUR) m.samples.shift()
43 if (m.samples.length === 0 && m.live === undefined) {
44 $.ui.status(undefined)
45 return
46 }
47 const perMinute = m.samples.filter(s => s.at > now - MINUTE).reduce((sum, s) => sum + s.tokens, 0)
48 const perHour = m.samples.reduce((sum, s) => sum + s.tokens, 0)
49 let tps = m.lastTps
50 let isLive = false
51 if (m.live !== undefined && now - m.live.start >= MIN_STREAM_MS) {
52 tps = m.live.chars * tokensPerChar(m) / ((now - m.live.start) / 1000)
53 isLive = true
54 }
55 $.ui.status(meterText(tps, isLive, perMinute, perHour))
56}
57
58export const register: Register = on => {
59 const meter: Meter = { samples: [], lastTps: undefined, live: undefined, seen: { tokens: 0, chars: 0 } }
60
61 on('session.start', async ($, e, next) => {
62 const started = await next(e)
63 // Per-minute and per-hour totals fall as time passes, so refresh while idle too.
64 $.clock.every(5000, () => void show($, meter))
65 return started
66 })
67
68 on('turn.step', async function* ($, e, next) {
69 // Timed from the request, not the first chunk: a reply can think for seconds
70 // and then arrive in one burst, which would make the speed look huge.
71 const start = await $.clock.now()
72 const stream = next(e)
73 let chars = 0
74 let hidden = false
75 let pieces = 0
76 for await (const chunk of stream) {
77 if (chunk.kind === 'text' || chunk.kind === 'thinking' || chunk.kind === 'input') {
78 const length = chunk.kind === 'input' ? chunk.json.length : chunk.text.length
79 if (chunk.kind === 'thinking' && length === 0) hidden = true
80 chars += length
81 pieces += 1
82 if (pieces % 8 === 0 && !hidden) {
83 meter.live = { start, chars }
84 await show($, meter)
85 }
86 } else if (chunk.kind === 'stop') {
87 const end = await $.clock.now()
88 if (chunk.usage) {
89 meter.samples.push({ at: end, tokens: chunk.usage.output_tokens })
90 if (chars > 0 && !hidden) {
91 meter.seen.tokens += chunk.usage.output_tokens
92 meter.seen.chars += chars
93 }
94 if (end - start >= MIN_STREAM_MS) {
95 meter.lastTps = chunk.usage.output_tokens / ((end - start) / 1000)
96 }
97 }
98 meter.live = undefined
99 await show($, meter)
100 }
101 yield chunk
102 }
103 return await stream.result
104 })
105}
106