SLOPSHOPPER

token-meter

Shows live tokens per second, per minute and per hour in the status line

newstatustimer
A shopper browsing a rack in a slop shop
README

Token Meter

Version 0.2.0. See CHANGELOG.md.

A Claude Code mod that shows how fast Claude is writing, live, in the status line:

~95 tok/s | 1.2k tok/min | 8.4k tok/hr
  • tok/s: speed of the current reply, timed from when the request is sent. While Claude is still writing it starts with ~, an estimate from the text so far. When the reply ends it switches to the exact number from the API.
  • tok/min: tokens Claude wrote in the last 60 seconds.
  • tok/hr: tokens Claude wrote in the last hour.

It counts output tokens only (what Claude writes, including thinking and tool calls), not what it reads.

What's a mod?

A mod is a plugin with code that runs inside Claude Code, so it can draw in the interface and react to events like tool calls and replies. Skills are instructions Claude reads; mods change Claude Code itself. They need Claude Code v2.1.287 or later. See Anthropic's mods overview.

How the live estimate works

The API only reports the token count when a reply ends, so while Claude writes, the meter counts characters instead. It starts at 3 characters per token, measured on Claude's replies, then learns the real ratio from each finished reply in your session. When Claude's thinking is hidden it costs tokens but shows no text, so the meter skips the live figure for that reply and shows the exact one when it ends.

Tested

Installed from this repo's marketplace and run in real sessions of three replies each (a story, a command, a story, a command, a story):

Result
Exact countsMatched the API's own totals in every run (for example 1,462, 1,567 and 1,471 tokens)
Live ~ estimate against the exact speedWithin about 15%: 73 vs 83, 77 vs 90, 84 vs 93, 92 vs 86, 90 vs 80, 104 vs 94 tok/s
Before the 0.2.0 fixes2.5k tok/s on replies that thought first, and live estimates 30 to 40% low

6 automated tests cover the exact speed, the live estimate, learning the ratio, hidden thinking, the per-minute window and clearing the status line. Run them with claude plugin test mods/token-meter.

The test sessions ran headless, where the meter writes to the debug log instead of a status line. I couldn't open an interactive terminal on the test machine, so I haven't seen it drawn on screen myself.

Install

In your terminal:

claude plugin marketplace add NotRedFox/NotRedFoxs-Claude-skills
claude plugin install token-meter@notredfox

Or inside a Claude Code session, /plugin marketplace add NotRedFox/NotRedFoxs-Claude-skills then /plugin install token-meter@notredfox, then /reload-plugins.

To check what it does before installing, clone this repo and run claude plugin validate mods/token-meter. It only reads replies and the clock and writes to the status line.

Where it shows

The terminal and the Code tab of the Claude Desktop app. It runs but shows nothing in the VS Code chat panel and claude -p.

To update it: claude plugin update token-meter@notredfox. To remove it: claude plugin uninstall token-meter@notredfox.

Source 1 files
hooks/register.ts 106 lines
1import type { EngineInterface, Register } from 'claude-code'
2
3// Output tokens are what "speed" means for a model, so the meter counts those.
4// While a reply is still streaming, the API hasn't reported its token count yet,
5// so the live figure is estimated from characters and marked with "~". It starts
6// at 3 characters per token (measured on Claude's replies) and learns the real
7// ratio from each finished reply. Hidden thinking costs tokens but streams no
8// characters, so those replies get no live figure and aren't learned from.
9const CHARS_PER_TOKEN = 3
10const MINUTE = 60_000
11const HOUR = 3_600_000
12// Shorter than this, a reply's rate is mostly noise (one or two chunks).
13const MIN_STREAM_MS = 200
14
15type Sample = { at: number; tokens: number }
16
17export function compact(n: number): string {
18  if (n >= 1_000_000) return `${(n / 1_000_000).toFixed(1)}M`
19  if (n >= 10_000) return `${Math.round(n / 1000)}k`
20  if (n >= 1000) return `${(n / 1000).toFixed(1)}k`
21  return `${Math.round(n)}`
22}
23
24export function meterText(tps: number | undefined, isLive: boolean, perMinute: number, perHour: number): string {
25  const speed = tps === undefined ? '- tok/s' : `${isLive ? '~' : ''}${compact(tps)} tok/s`
26  return `${speed} | ${compact(perMinute)} tok/min | ${compact(perHour)} tok/hr`
27}
28
29type Meter = {
30  samples: Sample[]
31  lastTps: number | undefined
32  live: { start: number; chars: number } | undefined
33  seen: { tokens: number; chars: number }
34}
35
36function tokensPerChar(m: Meter): number {
37  return m.seen.chars > 0 ? m.seen.tokens / m.seen.chars : 1 / CHARS_PER_TOKEN
38}
39
40async function show($: EngineInterface, m: Meter) {
41  const now = await $.clock.now()
42  while (m.samples.length > 0 && m.samples[0].at <= now - HOUR) m.samples.shift()
43  if (m.samples.length === 0 && m.live === undefined) {
44    $.ui.status(undefined)
45    return
46  }
47  const perMinute = m.samples.filter(s => s.at > now - MINUTE).reduce((sum, s) => sum + s.tokens, 0)
48  const perHour = m.samples.reduce((sum, s) => sum + s.tokens, 0)
49  let tps = m.lastTps
50  let isLive = false
51  if (m.live !== undefined && now - m.live.start >= MIN_STREAM_MS) {
52    tps = m.live.chars * tokensPerChar(m) / ((now - m.live.start) / 1000)
53    isLive = true
54  }
55  $.ui.status(meterText(tps, isLive, perMinute, perHour))
56}
57
58export const register: Register = on => {
59  const meter: Meter = { samples: [], lastTps: undefined, live: undefined, seen: { tokens: 0, chars: 0 } }
60
61  on('session.start', async ($, e, next) => {
62    const started = await next(e)
63    // Per-minute and per-hour totals fall as time passes, so refresh while idle too.
64    $.clock.every(5000, () => void show($, meter))
65    return started
66  })
67
68  on('turn.step', async function* ($, e, next) {
69    // Timed from the request, not the first chunk: a reply can think for seconds
70    // and then arrive in one burst, which would make the speed look huge.
71    const start = await $.clock.now()
72    const stream = next(e)
73    let chars = 0
74    let hidden = false
75    let pieces = 0
76    for await (const chunk of stream) {
77      if (chunk.kind === 'text' || chunk.kind === 'thinking' || chunk.kind === 'input') {
78        const length = chunk.kind === 'input' ? chunk.json.length : chunk.text.length
79        if (chunk.kind === 'thinking' && length === 0) hidden = true
80        chars += length
81        pieces += 1
82        if (pieces % 8 === 0 && !hidden) {
83          meter.live = { start, chars }
84          await show($, meter)
85        }
86      } else if (chunk.kind === 'stop') {
87        const end = await $.clock.now()
88        if (chunk.usage) {
89          meter.samples.push({ at: end, tokens: chunk.usage.output_tokens })
90          if (chars > 0 && !hidden) {
91            meter.seen.tokens += chunk.usage.output_tokens
92            meter.seen.chars += chars
93          }
94          if (end - start >= MIN_STREAM_MS) {
95            meter.lastTps = chunk.usage.output_tokens / ((end - start) / 1000)
96          }
97        }
98        meter.live = undefined
99        await show($, meter)
100      }
101      yield chunk
102    }
103    return await stream.result
104  })
105}
106