Live scorer numbers inside Claude Code: a /live pane (and /live status) for context, cost, judge activity, gates and wakes.

A portable, provider-agnostic workflow toolkit for AI coding agents. It works the same way in Claude Code, Codex, and Cursor: 48 native skills plus 3 fetched external design packs (impeccable, emil-design-eng, taste-skill), a repo bootstrapper, safety hooks, a bidirectional MCP bridge for multi-agent communication, a cheap-decision judge, a cost/involvement scorer, and token-efficiency tools (the rtk command rewriter and the headroom context compressor).
There is one canonical core: skills, hook logic, MCP servers, bridge, judge, and scorer. Each provider gets a thin adapter on top. Skills describe capabilities ("ask the user", "spawn a subagent", "call an MCP tool"), and each provider maps those to its own tools. See Providers and planning/PROVIDERS.md.
New here? Start with the onboarding guide:
ONBOARDING.md, or the visual walkthrough indocs/onboarding.html(open it in a browser).
Invoking skills: the examples below use
/<name>(Claude Code, Cursor). In Codex, use$<name>. For example,$reviewinstead of/review.
This toolkit supports an end-to-end product workflow where AI sessions replace documents. The goal is to make GitHub issues the single source of truth, instead of piling up local markdown files.
/withInterview
Human in the loop: You're in the hot seat. The agent interviews you. It asks questions, challenges assumptions, and surfaces contradictions while you answer in your own words. The output is a coherent problem statement and set of goals distilled from your raw thinking. You don't have to write polished prose yourself.
/officeHours [feature or problem]
Human in the loop: This is a back-and-forth collaboration. The agent proposes requirements and you push back. It drafts the technical design, and you redirect priorities and flag constraints it doesn't know about. Think of it as a YC office hours session: you leave with decisions made, not just options listed. It also gives structure to multi-team collaboration. Product and engineering can align on vision, scope, and trade-offs in a shared session before anyone writes a line of code. The output lands in ~/.agentic-workflow/<repo>/plans/<feature>/: a canonical plan.md handoff plus per-owner docs:
| File | Owner | Contents |
|---|---|---|
product.md | Product | Problem statement, EARS requirements, acceptance criteria, success metrics, MVP scope |
engineering.md | Engineering | Current state, approach, architecture decisions, open questions |
design-brief.md | Design | Experience goals, key interactions, UX requirements, design language reference |
TASKS.md | Engineering | Atomic task breakdown with domain tags for cross-team visibility |
Each file is a standalone artifact. Everyone leaves the session with a doc they own, not a monolith that nobody owns.
You can optionally pressure-test the outputs before moving on. Run each lens on its own, or use /autoplan to run them all in parallel:
/productReview # Founder/product lens: is this the right thing to build?
/archReview # Engineering lens: is this the right way to build it?
/autoplan # productReview + archReview + planDesignReview + planDevexReview + cso(plan), in parallel
/design-analyze # Extract design tokens from reference sites (web or iOS)
/design-language # Define brand personality and aesthetic direction
/design-shotgun # Optional: 4–6 mockup variants in parallel to pick a direction
/design-mockup # Generate HTML or SwiftUI mockup from design language
/design-refine # Agents self-critique and iterate against the design language
Human in the loop: Once the first mockup exists, agents enter a self-critique loop. They check whether the mockup reflects the design language, find deviations, and refine on their own. You step in at natural breakpoints to review the current state, direct emphasis ("make the data table the focus, not the sidebar"), and decide when the visual spec is ready to lock. You're the final judge of "good enough to build from." You don't take part in every pixel decision.
These produce design-tokens.json, .impeccable.md, and per-screen mockups (HTML or SwiftUI) with screenshot baselines that serve as the visual specification.
Create a multi-phase issue hierarchy directly from the officeHours output:
product.md, and the approach section of engineering.mdTASKS.md, each referencing the epic and embedding relevant contextThe officeHours MD files are ephemeral. Once the context is in GitHub issues, delete or ignore them. The issues become the canonical source of truth: product vision, design language reference, and mockups all in one place, with no local file sprawl.
/specToProvenPR # Turn an approved spec into proven, review-clean PRs
/review # Multi-agent PR code review
/postReview # Publish findings to GitHub as batched comments
/addressReview # Implement fixes with parallel agents
/cso # Pre-ship security check (OWASP Top 10 + STRIDE)
/shipRelease # Sync, test, push, open PR → auto-chains /landAndDeploy → /canary → /syncDocs
/weeklyRetro # Retrospective with shipping streaks
Human in the loop: Shipping is a loop, not a one-shot. The review agents surface issues and publish them to GitHub. You decide what to fix before merge and what to track as follow-ups. /addressReview implements the fixes in parallel, and you review the diff. /shipRelease runs the gate checks, and you approve the PR. The retro closes the loop: what shipped, what slipped, and what to carry into next week.
The toolkit treats Claude Code, Codex, and Cursor as equal hosts. setup.sh installs for every provider CLI it detects, or for the ones you name with --providers.
| Feature | Claude Code | Codex | Cursor |
|---|---|---|---|
| Native skills + external design packs | Yes | Yes | Yes |
| Invocation | /<name> | $<name> | /<name> |
| Skills installed to | ~/.claude/skills/ | ~/.codex/skills/ | ~/.cursor/skills/ |
| Repo instructions | CLAUDE.md (symlink to AGENTS.md) + .claude/rules (symlink) | AGENTS.md (Rules Index) | AGENTS.md + .cursor/rules/*.mdc (symlinks) |
Safety hooks (config/hooks/) | Native | Via adapter (config/hooks/adapters/codex.sh); trust with /hooks | Via adapter (config/hooks/adapters/cursor.sh) |
| MCP servers (bridge, serena, headroom, prism-mcp, …) | claude mcp add --scope user | codex mcp add | merged into ~/.cursor/mcp.json (may need cursor-agent mcp enable <name>) |
| Judge model provider | claude-cli | codex-cli | cursor-cli |
| Scorer transcript source | Yes | Yes | Involvement only (no token/cost data) |
| Statusline + shell integration | Yes | — | — |
| Plugin marketplaces | Yes | — | — |
The per-provider paths, tool names, and hook event names are recorded in planning/PROVIDERS.md. The capability-to-tool map that skills rely on lives in skills/_shared/capabilities.md.
setup.sh creates ~/.agentic-workflow/toolkit as a symlink to this repo. Skills find shared fragments through ~/.agentic-workflow/toolkit/skills/, never through a provider's skills directory, so every provider resolves them the same way. That includes the shared preamble (_preamble.md, and _design-preamble.md for design skills), which each SKILL.md references instead of embedding.~/.agentic-workflow/providers lists each installed provider and its skills directory, one per line.providers/<name>/install.sh (skills, MCP registration, config) and providers/<name>/install-hooks.sh (hook wiring).config/hooks/ speak the Claude Code hook protocol (JSON on stdin, exit 2 = deny). Claude Code runs them directly. For Codex and Cursor, a small adapter translates each provider's hook input and exit codes to that protocol, so there is only one copy of the safety logic. Lever hooks install with scripts/install-*.sh --provider <name>. The event mapping is in config/hooks/adapters/README.md.AGENTS.md and .agents/rules/ are the only copies. scripts/sync-rules.sh symlinks CLAUDE.md → AGENTS.md, .claude/rules → .agents/rules, and each .cursor/rules/<name>.mdc → .agents/rules/<name>.md. It also regenerates the Rules Index table in AGENTS.md for Codex. /bootstrap produces this same layout in any target repo../setup.sh # install for every provider CLI detected on PATH
./setup.sh --providers claude # Claude Code only
./setup.sh --providers codex # Codex only
./setup.sh --providers cursor # Cursor only
./setup.sh --providers claude,codex,cursor # explicit list
Setup is idempotent, so you can re-run it with a different --providers list to add a provider later. Add --dry-run to print every change without writing anything.
After setup, a couple of provider-specific steps remain:
codex, run /hooks, and trust the aw:* entries. Do this again after a reinstall that changes a hook command.cursor-agent mcp enable <name> for each one.claude), Codex (codex), or Cursor CLI (cursor-agent)gh) installed and authenticated (required by review skills)jq installed (required by hooks and the statusline; brew install jq on macOS)rtk: token-compressing CLI proxy (brew install rtk on macOS; installed automatically by setup.sh)headroom: context optimization layer (pip install "headroom-ai[all]"; installed automatically by setup.sh)git clone https://github.com/vitalizecare/agentic-workflow.git ~/repos/agentic-workflow
cd ~/repos/agentic-workflow
./setup.sh # or: ./setup.sh --providers claude,codex,cursor
The setup script:
jq, Docker) and detects which provider CLIs are installed (or uses --providers)~/.agentic-workflow/toolkit symlink and writes the ~/.agentic-workflow/providers registry/bootstrap, and the external design packs (cloned at pinned commits from EXTERNAL_PINS.env) into each selected provider's skills directoryblock-destructive.sh, block-push-main.sh, detect-secrets.sh, rtk-rewrite.sh) and session-context hooks for each provider. Claude Code uses them natively; Codex and Cursor go through their hook adapters.serena-docker wrapper to ~/.local/bin/agentic-bridge, serena, headroom, prism-mcp, and xcodebuildmcp on macOS) with each selected providerprism-mcp (persistent memory, downloaded on first use) with its Mind Palace dashboard at http://localhost:7180 (PRISM_DASHBOARD_PORT). The prism-context.sh session-start hook warns if the dashboard is unreachable; /prismStatus runs a full health checksettings.json, installs the statusline and shell integration, and adds plugin marketplaces and pluginscd mcp-bridge && npm start # Fastify on http://127.0.0.1:3100
| Variable | Default | Description |
|---|---|---|
PORT | 3100 | REST API port |
HOST | 127.0.0.1 | Bind address (loopback only by default) |
DB_PATH | ./bridge.db | SQLite database file path |
ALLOW_REMOTE | unset | Set to 1 to allow non-loopback binding |
48 native skills, installed as symlinks into each provider's skills directory. Every skill uses the same text for every provider: steps name a capability, and the running agent uses its host's tool for it (see skills/_shared/capabilities.md).
| Stage | Skills |
|---|---|
| Ideation & planning | withInterview, enhancePrompt, officeHours, autoplan, productReview, archReview, planDesignReview, planDevexReview |
| Design | design-analyze, design-language, design-evolve, design-shotgun, design-mockup, design-implement, design-refine, design-verify (dispatchers auto-detect web/iOS and route to their -web / -ios sub-skills) |
| Build & verify | specToProvenPR, verify-app (→ verify-web, verify-ios), ui-evidence |
| Review | review, postReview, addressReview, cso |
| Debug & QA | rootCause, bugHunt, bugReport, bugFixOrchestrator, testAudit |
| Ship & operate | shipRelease, landAndDeploy, canary, syncDocs, weeklyRetro, prismStatus, judge |
| Repo setup | bootstrap |
Skills write their artifacts to ~/.agentic-workflow/<repo-slug>/<domain>/, and downstream skills discover them from there.
Run /bootstrap ($bootstrap in Codex) in any repo to generate its documentation:
AGENTS.md (a navigation doc, not a reference manual), with CLAUDE.md as a symlink to it.agents/rules/, then links them for each provider with sync-rules.sh (.claude/rules, .cursor/rules/*.mdc)A TypeScript MCP server for bidirectional multi-agent communication across any mix of Claude Code, Codex, and Cursor sessions. All providers register the same stdio server and share one SQLite database.
MCP Tools:
send_context: send task context + meta-prompt between agentsget_messages: retrieve conversation history by UUIDget_unread: check for unread messages (marks them read on retrieval)assign_task: assign tasks with domain and implementation detailsreport_status: report back with feedback or completionAPI Endpoints:
POST /messages/send: send context between agentsGET /messages/conversation/:id: retrieve conversation historyGET /messages/unread?recipient=: fetch unread messages and mark them readPOST /tasks/assign: assign a task with domain classificationGET /tasks/:id: get a task by IDGET /tasks/conversation/:id: get all tasks for a conversationPOST /tasks/report: report task statusGET /conversations: paginated conversation summariesFeatures:
AppResult<T> patternjudge/ makes cheap, typed decisions: a rules fast path, then a per-content-class model chain. Model calls go through a headless provider CLI: claude-cli (claude -p), codex-cli (codex exec), or cursor-cli (cursor-agent -p). Any one of them on PATH is enough. The order is providers.agentClis in the judge config if set; otherwise claude, codex, cursor, with the current host (AW_PROVIDER) first. Cursor is slow (~8–13s per call), so on Cursor-only machines text questions often time out and fall back to rules.scorer/ produces a daily cost and involvement report from agent transcripts. Each provider has its own transcript source: Claude Code (~/.claude/projects/), Codex (~/.codex/sessions/), and Cursor (~/.cursor/projects/). By default it reads every provider whose directory exists. --provider claude|codex|cursor|all narrows that, and --codex-dir / --cursor-dir override the paths. The report includes a "By provider" section. Cursor data is involvement-only because its transcripts carry no token or cost data. Codex rollouts imported from Claude are skipped so they aren't counted twice. Run scorer --since 7d to write a report to ~/.agentic-workflow/scorer/reports/. Install it with scripts/install-scorer.sh.config/statusline.sh is an adaptive two-line statusline for Claude Code sessions. setup.sh installs it to ~/.claude/statusline.sh and wires it into settings.json when Claude Code is a selected provider.
Columns (left → right, highest priority leftmost):
| Column | Description |
|---|---|
| 5h Usage | 5-hour rate-limit percentage + reset time |
| 7d Usage | 7-day rate-limit percentage + reset day |
| Context | Color-coded bar + percentage of context window used |
| Model | Active model name (trimmed) |
| Branch | Current git branch |
| Cost | Session cost in USD |
| Time | Session duration |
| Cache | Cache read hit rate |
| API | API wait percentage |
| Lines | Lines added/removed |
Adaptive width tiers: columns drop out automatically as the terminal narrows.
| Tier | Min width | Columns shown |
|---|---|---|
| FULL | 116 cols | All columns, branch up to 15 chars |
| MEDIUM | 101 cols | No Lines; branch up to 12 chars |
| NARROW | 78 cols | No Lines/Cache/API; 7d % only; narrow context bar |
| COMPACT | 65 cols | 5h % only; narrow context bar; branch up to 10 chars |
| COMPACT-S | < 65 cols | Same as COMPACT but drops Time column |
The statusline reads the size of the tty its own Claude Code process is attached to (found by walking up the process tree), so every window gets its own width and no state is shared between windows. When that read fails it falls back to ~/.claude/terminal_width.d/<tty>, written by the shell integration from interactive terminals only (never from Claude Code tool shells). AW_STATUSLINE_DEBUG=1 prints the width source to stderr.
setup.sh installs the shell integration to ~/.claude/shell-integration.sh and sources it from ~/.zshrc / ~/.bashrc. It keeps ~/.claude/terminal_width.d/<tty> current and writes ~/.claude/shell_pid.d/<tty>, so the hooks of a session can send SIGWINCH to the shell on its own tty.
cd mcp-bridge && npm test # Vitest, in-memory SQLite
cd scorer && npm test
cd judge && npm test
bash providers/tests/install.test.sh
bash scripts/tests/sync-rules.test.sh
bash config/hooks/tests/codex-adapter.test.sh
bash config/hooks/tests/cursor-adapter.test.sh
bash config/hooks/tests/provider-install-hooks.test.sh
bash config/lib/tests/merge-hook.test.sh
bash config/hooks/tests/probe-log.test.sh
scripts/sync-rules.sh --check
The full list is under Commands in AGENTS.md.
Tests cover unit tests (controllers, services, DB client, schemas, utilities) and integration tests (all REST routes via Fastify inject, plus the MCP tool handlers). /* v8 ignore */ annotations are prohibited; write the test instead.
agentic-workflow/
├── AGENTS.md # Canonical repo instructions (CLAUDE.md is a symlink to it)
├── .agents/rules/ # Glob-scoped rules, the only copy (.claude/rules, .cursor/rules/*.mdc are symlinks)
├── skills/ # 48 native skills (+ _shared/ fragments, incl. capabilities.md)
├── bootstrap/ # /bootstrap — repo documentation generator
├── providers/<name>/ # Per-provider installers: install.sh, install-hooks.sh (claude, codex, cursor)
├── config/ # Settings, MCP config, statusline, hooks (+ hooks/adapters/ for codex, cursor)
├── mcp-bridge/ # MCP bridge + REST API (Fastify, SQLite)
├── judge/ # Cheap typed decisions (rules → model chain via claude/codex/cursor CLI)
├── scorer/ # Cost/involvement report from provider transcripts
├── scripts/ # sync-rules.sh, serena-docker, probe.sh, install-*.sh, refresh-external-pins.sh
├── planning/ # Project documentation (see planning/PROVIDERS.md)
├── Dockerfile.serena* # Serena base image + opt-in C# / Swift extensions
└── setup.sh # One-command setup: ./setup.sh [--providers claude,codex,cursor]hooks/register.tsx 174 lines1// aw-live: the scorer's numbers for THIS session, live inside Claude Code.
2// `/live` toggles a pane (`/live status` answers in text). The headline numbers
3// live in the statusline's Live column (config/statusline.sh), not here.
4// Every number comes from `scorer live --json` (one short Node process, never
5// in-process). A refresh never blocks a hook and every failure is silent: the
6// pane just keeps its last good value.
7
8import { atom, read, update } from 'claude-code'
9import type { EngineInterface, Register } from 'claude-code'
10
11import type { HostUsage, LiveSnapshot, LiveStatus } from '../types'
12import { paneSections, statusText } from './lib/format'
13import type { Tone } from './lib/format'
14import { binCandidates, isDue, liveArgs, runTimeout } from './lib/refresh'
15import { parseSnapshot } from './lib/snapshot'
16
17const PANE = 'aw-live'
18const TITLE = 'Live scorer'
19const PANE_COLUMNS = 72
20
21const snapshot = atom({ plugin: 'aw-live', key: 'snapshot' } as const, null as LiveSnapshot | null)
22const host = atom({ plugin: 'aw-live', key: 'host' } as const, null as HostUsage | null)
23const status = atom({ plugin: 'aw-live', key: 'status' } as const, null as LiveStatus | null)
24
25const TONE_COLOR: Record<Tone, string | undefined> = { ok: '#7ec699', warn: '#d4a054', dim: '#8b949e' }
26
27// Module variables restart on a hot reload; that only costs one extra refresh.
28let lastAt: number | null = null
29let running: Promise<void> | null = null
30
31/** Test seam: forget the refresh bookkeeping so each test starts cold. */
32export const resetForTests = (): void => {
33 lastAt = null
34 running = null
35}
36
37const isTimeout = (err: unknown): boolean => /time/i.test(err instanceof Error ? err.message : String(err))
38
39/** Runs scorer by argv, trying each candidate path; null on any failure. */
40const runScorer = async ($: EngineInterface, args: string[], timeoutMs: number): Promise<string | null> => {
41 const names = binCandidates(await $.env.get('HOME'), 'scorer', await $.env.get('AW_SCORER_BIN'))
42 for (const name of names) {
43 try {
44 const r = await $.process.run([name, ...args], { timeoutMs })
45 if (r.exitCode === 0) return r.stdout
46 return null
47 } catch (err) {
48 // A timeout means the binary exists but is slow: do not stack another wait on it.
49 if (isTimeout(err)) return null
50 // otherwise not runnable at this path: try the next candidate
51 }
52 }
53 return null
54}
55
56const runRefresh = async ($: EngineInterface, force: boolean): Promise<void> => {
57 let now: number | null = null
58 let ok = false
59 try {
60 now = await $.clock.now()
61 if (!force && !isDue(lastAt, now)) return
62 lastAt = now
63 const sessionId = await $.session.id()
64 if (sessionId === '') return
65 const [cwd, usage] = await Promise.all([$.session.cwd(), $.session.usage()])
66 await update($, host, (): HostUsage => ({
67 usd: usage.cost?.usd ?? null,
68 percent: usage.context.percent ?? null,
69 window: usage.context.window,
70 }))
71 const hasSnapshot = (await read($, snapshot)) !== null
72 const out = await runScorer($, liveArgs(sessionId, cwd, usage.context.window), runTimeout(hasSnapshot))
73 const parsed = out === null ? null : parseSnapshot(out)
74 if (parsed !== null) await update($, snapshot, () => parsed)
75 ok = parsed !== null
76 } catch {
77 // fail silent: the pane keeps its last good value
78 } finally {
79 const at = now
80 if (at !== null) await update($, status, (): LiveStatus => ({ ok, at })).catch(() => undefined)
81 }
82}
83
84/** One refresh at a time: a caller that arrives while one runs waits for that one. */
85const refresh = ($: EngineInterface, force: boolean): Promise<void> => {
86 if (running !== null) return running
87 running = runRefresh($, force).finally(() => {
88 running = null
89 })
90 return running
91}
92
93export const register: Register = on => {
94 on('session.start', async ($, e, next) => {
95 try {
96 await $.command.register({ name: 'live', description: 'Toggle the live scorer pane (context, cost, judge, gates)' })
97 } catch {
98 // a host error must not stop session start
99 }
100 return next(e)
101 })
102
103 on('turn.complete', async ($, e, next) => {
104 try {
105 // Only an open pane shows the numbers; /live refreshes on open, so a closed pane costs nothing.
106 if ((await $.ui.panes()).some(pane => pane.id === PANE)) {
107 $.clock.after(0, () => {
108 void refresh($, false)
109 })
110 }
111 } catch {
112 // a host error must not stop the turn
113 }
114 return next(e)
115 })
116
117 on('command.run', { command: 'live' }, async ($, e) => {
118 if (e.args.trim() === 'status') {
119 await refresh($, true)
120 try {
121 return { text: statusText(await read($, snapshot), await read($, host)) }
122 } catch {
123 return { text: 'aw-live: no numbers yet (is scorer installed? scripts/install-scorer.sh)' }
124 }
125 }
126 try {
127 if ((await $.ui.panes()).some(pane => pane.id === PANE)) {
128 await $.ui.close({ id: PANE })
129 return {}
130 }
131 await refresh($, true)
132 await $.ui.open({ id: PANE, title: TITLE, focus: true, closeOnEscape: true, columns: PANE_COLUMNS })
133 } catch {
134 // a host error: leave the pane as it is
135 }
136 return {}
137 })
138
139 on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e, next) => {
140 try {
141 const { Box, Text, Button } = $.ui.resolve(e)
142 const s = await read($, snapshot)
143 if (s === null) {
144 return (
145 <Box flexDirection="column">
146 <Text color={TONE_COLOR.dim}>No numbers yet. Is `scorer` installed? (scripts/install-scorer.sh)</Text>
147 </Box>
148 )
149 }
150 const sections = paneSections(s, await read($, host))
151 return (
152 <Box flexDirection="column">
153 {sections.map(section => (
154 <Box key={section.title} borderStyle="round" borderColor="#3d4450" paddingX={1} flexDirection="column">
155 <Text color="#7eb8da">{section.title}</Text>
156 {section.rows.map(row => (
157 <Text key={row.label} wrap="truncate-end">
158 <Text color={TONE_COLOR.dim}>{row.label.padEnd(14)}</Text>
159 <Text color={TONE_COLOR[row.tone]}>{row.value}</Text>
160 </Text>
161 ))}
162 </Box>
163 ))}
164 <Box flexDirection="row" columnGap={2} paddingX={1}>
165 <Button key="refresh" label="refresh" hotkey="r" plain onPress={() => refresh($, true)} />
166 </Box>
167 </Box>
168 )
169 } catch {
170 return next(e)
171 }
172 })
173}
174hooks/lib/format.ts 107 lines1// Pure formatting for the pane and `/live status`. No `$` in here.
2
3import type { HostUsage, LiveSnapshot } from '../../types'
4
5export type Tone = 'ok' | 'warn' | 'dim'
6export type Row = { label: string; value: string; tone: Tone }
7export type Section = { title: string; rows: Row[] }
8
9export const WARN_PERCENT = 70
10
11export const fmtTokens = (n: number): string => {
12 if (n >= 1_000_000) return `${(n / 1_000_000).toFixed(1)}M`
13 if (n >= 1_000) return `${Math.round(n / 1_000)}k`
14 return String(n)
15}
16
17/** `▓▓▓▓░░░░░░` for a percent; empty cells when unknown. */
18export const bar = (percent: number | null, width: number): string => {
19 const filled = percent === null ? 0 : Math.round((Math.min(100, Math.max(0, percent)) / 100) * width)
20 return '▓'.repeat(filled) + '░'.repeat(width - filled)
21}
22
23/** The host's own context percent when it has one, else the transcript's estimate. */
24export const contextPercent = (s: LiveSnapshot, host: HostUsage | null): number | null =>
25 host?.percent ?? s.context.percent
26
27/** `jev 83% · rules 17%` for the two biggest deciders; empty with none. */
28export const providerShare = (byProvider: Readonly<Record<string, number>>): string => {
29 const entries = Object.entries(byProvider).sort((a, b) => b[1] - a[1] || a[0].localeCompare(b[0]))
30 const total = entries.reduce((sum, [, n]) => sum + n, 0)
31 return entries.slice(0, 2).map(([name, n]) => `${name} ${Math.round((n / total) * 100)}%`).join(' · ')
32}
33
34const GATE_ORDER = ['scope-gate', 'done-gate', 'send-gate'] as const
35
36const decisionList = (byDecision: Readonly<Record<string, number>>): string =>
37 Object.entries(byDecision).sort((a, b) => b[1] - a[1] || a[0].localeCompare(b[0])).map(([d, n]) => `${d} ${n}`).join(' · ')
38
39/** The pane's cards, top to bottom. */
40export const paneSections = (s: LiveSnapshot, host: HostUsage | null): Section[] => {
41 const percent = contextPercent(s, host)
42 const window = host?.window ?? s.context.window
43 const context: Section = {
44 title: 'context',
45 rows: [
46 { label: 'window', value: `${percent === null ? '--' : `${Math.round(percent)}%`} ${bar(percent, 10)}`, tone: percent !== null && percent >= WARN_PERCENT ? 'warn' : 'ok' },
47 { label: 'tokens', value: `${s.context.tokens === null ? '--' : fmtTokens(s.context.tokens)} of ${fmtTokens(window)}`, tone: 'dim' },
48 ],
49 }
50 const session: Section = {
51 title: 'session',
52 rows: [
53 { label: 'cost', value: host?.usd == null ? '--' : `$${host.usd.toFixed(2)}`, tone: 'ok' },
54 { label: 'tokens', value: `${fmtTokens(s.usage.contextTokens)} in · ${fmtTokens(s.usage.outputTokens)} out`, tone: 'dim' },
55 { label: 'calls', value: `${s.usage.calls} (${s.usage.subagentCalls} subagent)`, tone: 'dim' },
56 { label: 'over 200k', value: String(s.usage.callsOver200k), tone: s.usage.callsOver200k > 0 ? 'warn' : 'dim' },
57 ],
58 }
59 if (s.ingest.readErrors > 0) {
60 session.rows.push({ label: 'warning', value: `stale: ${s.ingest.readErrors} transcript read error${s.ingest.readErrors === 1 ? '' : 's'}`, tone: 'warn' })
61 }
62 const sections = [context, session]
63 if (s.judge.state === 'unscoped') {
64 sections.push({ title: 'judge', rows: [{ label: 'sessions', value: 'not linked yet (judge predates decision_details.session_id)', tone: 'dim' }] })
65 }
66 if (s.judge.state === 'ok') {
67 const j = s.judge
68 sections.push({
69 title: 'judge',
70 rows: [
71 { label: 'calls', value: `${j.calls}${providerShare(j.byProvider) === '' ? '' : ` · ${providerShare(j.byProvider)}`}`, tone: 'ok' },
72 { label: 'vs rules', value: `agreed ${j.agreed} · overrode ${j.overrode} · unsure ${j.undecided}`, tone: j.undecided > 0 ? 'warn' : 'dim' },
73 { label: 'latency', value: `p50 ${j.p50LatencyMs} ms · p95 ${j.p95LatencyMs} ms`, tone: 'dim' },
74 ],
75 })
76 sections.push({
77 title: 'gates',
78 rows: GATE_ORDER.map((name): Row => {
79 const g = j.gates[name]
80 return {
81 label: name,
82 value: g.fired === 0 ? 'quiet' : `${g.fired} · ${decisionList(g.byDecision)} · p50 ${g.p50LatencyMs} ms`,
83 tone: g.fired === 0 ? 'dim' : 'ok',
84 }
85 }),
86 })
87 }
88 sections.push({
89 title: 'wakes',
90 rows: [
91 { label: 'queued now', value: String(s.wakes.queuedNow), tone: s.wakes.queuedNow > 0 ? 'warn' : 'dim' },
92 { label: 'context guard', value: `${s.contextGuard.fires} fires`, tone: s.contextGuard.fires > 0 ? 'warn' : 'dim' },
93 ],
94 })
95 return sections
96}
97
98/** The `/live status` answer: every pane row as plain text. */
99export const statusText = (s: LiveSnapshot | null, host: HostUsage | null): string => {
100 if (s === null) return 'aw-live: no numbers yet (is scorer installed? scripts/install-scorer.sh)'
101 const lines = paneSections(s, host).flatMap(section => [
102 `${section.title}`,
103 ...section.rows.map(row => ` ${row.label.padEnd(14)}${row.value}`),
104 ])
105 return lines.join('\n')
106}
107hooks/lib/refresh.ts 24 lines1// Pure pieces of the refresh loop. register.tsx owns every `$` call.
2
3export const MIN_GAP_MS = 2_000
4export const RUN_TIMEOUT_MS = 5_000
5/** The first refresh has no snapshot to fall back on; a cold ingest of a long session needs longer. */
6export const RUN_TIMEOUT_FIRST_MS = 30_000
7
8export const runTimeout = (hasSnapshot: boolean): number => (hasSnapshot ? RUN_TIMEOUT_MS : RUN_TIMEOUT_FIRST_MS)
9
10/** The `scorer live` command line. `window` is the host's context window, when known. */
11export const liveArgs = (sessionId: string, cwd: string, window: number | null): string[] => [
12 'live', '--session', sessionId, '--cwd', cwd, ...(window === null ? [] : ['--window', String(window)]), '--json',
13]
14
15/** Where to look for an aw CLI: an explicit override, the installer's directory, then PATH. */
16export const binCandidates = (home: string | undefined, name: 'scorer' | 'judge', override?: string): string[] => [
17 ...(override === undefined || override === '' ? [] : [override]),
18 ...(home === undefined || home === '' ? [name] : [`${home}/.local/bin/${name}`, name]),
19]
20
21/** A refresh is due when none has run yet or the last one is at least `minGapMs` old. */
22export const isDue = (lastAt: number | null, now: number, minGapMs: number = MIN_GAP_MS): boolean =>
23 lastAt === null || now - lastAt >= minGapMs
24hooks/lib/snapshot.ts 26 lines1// The contract with `scorer live --json`. scorer/src/live-snapshot.ts holds the zod
2// schema for the same shape; scorer/tests/live-contract.test.ts parses this mod's
3// fixture with it, so the two cannot drift apart unnoticed.
4
5import type { LiveSnapshot } from '../../types'
6
7const isObject = (value: unknown): value is Record<string, unknown> =>
8 typeof value === 'object' && value !== null && !Array.isArray(value)
9
10const JUDGE_STATES = ['none', 'unscoped', 'ok']
11
12/** Parses the CLI's stdout; null for anything that is not a version-1 snapshot. */
13export const parseSnapshot = (text: string): LiveSnapshot | null => {
14 let value: unknown
15 try {
16 value = JSON.parse(text)
17 } catch {
18 return null
19 }
20 if (!isObject(value) || value.v !== 1 || typeof value.sessionId !== 'string') return null
21 const { ingest, context, usage, judge, contextGuard, wakes } = value
22 if (!isObject(ingest) || !isObject(context) || !isObject(usage) || !isObject(contextGuard) || !isObject(wakes)) return null
23 if (!isObject(judge) || typeof judge.state !== 'string' || !JUDGE_STATES.includes(judge.state)) return null
24 return value as unknown as LiveSnapshot
25}
26types/index.d.ts 62 lines1export type GateName = 'scope-gate' | 'done-gate' | 'send-gate'
2
3export type GateStats = {
4 question: string
5 fired: number
6 byDecision: Record<string, number>
7 p50LatencyMs: number
8}
9
10export type JudgeLatest = {
11 id: string
12 ts: string
13 question: string
14 decision: string
15 provider: string
16 confidence: number
17 /** The eval_items id for this decision, once `judge label import` has made one. */
18 itemId: string | null
19}
20
21export type JudgeLive =
22 | { state: 'none' }
23 | { state: 'unscoped' }
24 | {
25 state: 'ok'
26 calls: number
27 byProvider: Record<string, number>
28 undecided: number
29 agreed: number
30 overrode: number
31 p50LatencyMs: number
32 p95LatencyMs: number
33 gates: Record<GateName, GateStats>
34 labelable: boolean
35 latest: JudgeLatest | null
36 }
37
38export type LiveSnapshot = {
39 v: 1
40 sessionId: string
41 at: string
42 transcript: { found: boolean; files: number }
43 /** readErrors: transcript files the scorer could not read this pass; the numbers may be stale. */
44 ingest: { newLines: number; ms: number; readErrors: number }
45 context: { tokens: number | null; window: number; percent: number | null }
46 usage: { calls: number; subagentCalls: number; contextTokens: number; outputTokens: number; callsOver200k: number }
47 judge: JudgeLive
48 contextGuard: { fires: number }
49 wakes: { queuedNow: number }
50}
51
52/** What the host itself reports; shown in preference to the transcript's estimate. */
53export type HostUsage = { usd: number | null; percent: number | null; window: number | null }
54
55export type LiveStatus = { ok: boolean; at: number }
56
57declare module 'claude-code' {
58 interface PluginState {
59 'aw-live': { snapshot: LiveSnapshot | null; host: HostUsage | null; status: LiveStatus | null }
60 }
61}
62