SLOPSHOPPER

aw-live

Live scorer numbers inside Claude Code: a /live pane (and /live status) for context, cost, judge activity, gates and wakes.

newpanecommandprocesstimer
★ 4v0.1.0no licenseupdated 2026-10-09joi-fairshare/agentic-workflow/mods/aw-live
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · aw-live
│ ┃ Live scorer ✕ › fix the failing auth test and add an audit log call │ ┃ No numbers yet. Is `scorer` installed? │ ┃ (scripts/install-scorer.sh) ⏺ Read(src/auth.ts) │ ⎿ Read 6 lines │ ⏺ Update(src/auth.ts) │ ⎿ Added 2 lines, removed 1 line │ ⏺ Bash(bun test) │ ⎿ 3 pass, 1 fail │ │ ● Done. refresh now rejects expired claims and logs an audit event. │ │ ✻ Worked for 42s · done 4:20 PM │ │ › /live │ │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Pane · Live scorer
No numbers yet. Is `scorer` installed? (scripts/install-scorer.sh)
README

Agentic Workflow

A portable, provider-agnostic workflow toolkit for AI coding agents. It works the same way in Claude Code, Codex, and Cursor: 48 native skills plus 3 fetched external design packs (impeccable, emil-design-eng, taste-skill), a repo bootstrapper, safety hooks, a bidirectional MCP bridge for multi-agent communication, a cheap-decision judge, a cost/involvement scorer, and token-efficiency tools (the rtk command rewriter and the headroom context compressor).

There is one canonical core: skills, hook logic, MCP servers, bridge, judge, and scorer. Each provider gets a thin adapter on top. Skills describe capabilities ("ask the user", "spawn a subagent", "call an MCP tool"), and each provider maps those to its own tools. See Providers and planning/PROVIDERS.md.

New here? Start with the onboarding guide: ONBOARDING.md, or the visual walkthrough in docs/onboarding.html (open it in a browser).

Invoking skills: the examples below use /<name> (Claude Code, Cursor). In Codex, use $<name>. For example, $review instead of /review.

Workflow: Product Vision → Ship

This toolkit supports an end-to-end product workflow where AI sessions replace documents. The goal is to make GitHub issues the single source of truth, instead of piling up local markdown files.

Stage 1 — Ideation

/withInterview

Human in the loop: You're in the hot seat. The agent interviews you. It asks questions, challenges assumptions, and surfaces contradictions while you answer in your own words. The output is a coherent problem statement and set of goals distilled from your raw thinking. You don't have to write polished prose yourself.

Stage 2 — Spec & Design Doc

/officeHours [feature or problem]

Human in the loop: This is a back-and-forth collaboration. The agent proposes requirements and you push back. It drafts the technical design, and you redirect priorities and flag constraints it doesn't know about. Think of it as a YC office hours session: you leave with decisions made, not just options listed. It also gives structure to multi-team collaboration. Product and engineering can align on vision, scope, and trade-offs in a shared session before anyone writes a line of code. The output lands in ~/.agentic-workflow/<repo>/plans/<feature>/: a canonical plan.md handoff plus per-owner docs:

FileOwnerContents
product.mdProductProblem statement, EARS requirements, acceptance criteria, success metrics, MVP scope
engineering.mdEngineeringCurrent state, approach, architecture decisions, open questions
design-brief.mdDesignExperience goals, key interactions, UX requirements, design language reference
TASKS.mdEngineeringAtomic task breakdown with domain tags for cross-team visibility

Each file is a standalone artifact. Everyone leaves the session with a doc they own, not a monolith that nobody owns.

You can optionally pressure-test the outputs before moving on. Run each lens on its own, or use /autoplan to run them all in parallel:

/productReview    # Founder/product lens: is this the right thing to build?
/archReview       # Engineering lens: is this the right way to build it?
/autoplan         # productReview + archReview + planDesignReview + planDevexReview + cso(plan), in parallel

Stage 3 — Design System & Mockups

/design-analyze   # Extract design tokens from reference sites (web or iOS)
/design-language  # Define brand personality and aesthetic direction
/design-shotgun   # Optional: 4–6 mockup variants in parallel to pick a direction
/design-mockup    # Generate HTML or SwiftUI mockup from design language
/design-refine    # Agents self-critique and iterate against the design language

Human in the loop: Once the first mockup exists, agents enter a self-critique loop. They check whether the mockup reflects the design language, find deviations, and refine on their own. You step in at natural breakpoints to review the current state, direct emphasis ("make the data table the focus, not the sidebar"), and decide when the visual spec is ready to lock. You're the final judge of "good enough to build from." You don't take part in every pixel decision.

These produce design-tokens.json, .impeccable.md, and per-screen mockups (HTML or SwiftUI) with screenshot baselines that serve as the visual specification.

Stage 4 — Engineering Roadmap (GitHub Issues)

Create a multi-phase issue hierarchy directly from the officeHours output:

  1. Epic issue: paste the product vision, product.md, and the approach section of engineering.md
  2. Task issues: one per entry in TASKS.md, each referencing the epic and embedding relevant context
  3. Attach mockups: link or embed the mockup screenshot so the visual spec lives in the issue

The officeHours MD files are ephemeral. Once the context is in GitHub issues, delete or ignore them. The issues become the canonical source of truth: product vision, design language reference, and mockups all in one place, with no local file sprawl.

Stage 5 — Ship

/specToProvenPR   # Turn an approved spec into proven, review-clean PRs
/review           # Multi-agent PR code review
/postReview       # Publish findings to GitHub as batched comments
/addressReview    # Implement fixes with parallel agents
/cso              # Pre-ship security check (OWASP Top 10 + STRIDE)
/shipRelease      # Sync, test, push, open PR → auto-chains /landAndDeploy → /canary → /syncDocs
/weeklyRetro      # Retrospective with shipping streaks

Human in the loop: Shipping is a loop, not a one-shot. The review agents surface issues and publish them to GitHub. You decide what to fix before merge and what to track as follow-ups. /addressReview implements the fixes in parallel, and you review the diff. /shipRelease runs the gate checks, and you approve the PR. The retro closes the loop: what shipped, what slipped, and what to carry into next week.


Providers

The toolkit treats Claude Code, Codex, and Cursor as equal hosts. setup.sh installs for every provider CLI it detects, or for the ones you name with --providers.

Support matrix

FeatureClaude CodeCodexCursor
Native skills + external design packsYesYesYes
Invocation/<name>$<name>/<name>
Skills installed to~/.claude/skills/~/.codex/skills/~/.cursor/skills/
Repo instructionsCLAUDE.md (symlink to AGENTS.md) + .claude/rules (symlink)AGENTS.md (Rules Index)AGENTS.md + .cursor/rules/*.mdc (symlinks)
Safety hooks (config/hooks/)NativeVia adapter (config/hooks/adapters/codex.sh); trust with /hooksVia adapter (config/hooks/adapters/cursor.sh)
MCP servers (bridge, serena, headroom, prism-mcp, …)claude mcp add --scope usercodex mcp addmerged into ~/.cursor/mcp.json (may need cursor-agent mcp enable <name>)
Judge model providerclaude-clicodex-clicursor-cli
Scorer transcript sourceYesYesInvolvement only (no token/cost data)
Statusline + shell integrationYes——
Plugin marketplacesYes——

The per-provider paths, tool names, and hook event names are recorded in planning/PROVIDERS.md. The capability-to-tool map that skills rely on lives in skills/_shared/capabilities.md.

How it fits together

  • Stable toolkit path. setup.sh creates ~/.agentic-workflow/toolkit as a symlink to this repo. Skills find shared fragments through ~/.agentic-workflow/toolkit/skills/, never through a provider's skills directory, so every provider resolves them the same way. That includes the shared preamble (_preamble.md, and _design-preamble.md for design skills), which each SKILL.md references instead of embedding.
  • Provider registry. ~/.agentic-workflow/providers lists each installed provider and its skills directory, one per line.
  • Per-provider installers. Provider-specific logic lives in providers/<name>/install.sh (skills, MCP registration, config) and providers/<name>/install-hooks.sh (hook wiring).
  • Hooks. The canonical hook scripts in config/hooks/ speak the Claude Code hook protocol (JSON on stdin, exit 2 = deny). Claude Code runs them directly. For Codex and Cursor, a small adapter translates each provider's hook input and exit codes to that protocol, so there is only one copy of the safety logic. Lever hooks install with scripts/install-*.sh --provider <name>. The event mapping is in config/hooks/adapters/README.md.
  • Repo instructions. AGENTS.md and .agents/rules/ are the only copies. scripts/sync-rules.sh symlinks CLAUDE.md → AGENTS.md, .claude/rules → .agents/rules, and each .cursor/rules/<name>.mdc → .agents/rules/<name>.md. It also regenerates the Rules Index table in AGENTS.md for Codex. /bootstrap produces this same layout in any target repo.

Per-provider setup

./setup.sh                                  # install for every provider CLI detected on PATH
./setup.sh --providers claude               # Claude Code only
./setup.sh --providers codex                # Codex only
./setup.sh --providers cursor               # Cursor only
./setup.sh --providers claude,codex,cursor  # explicit list

Setup is idempotent, so you can re-run it with a different --providers list to add a provider later. Add --dry-run to print every change without writing anything.

After setup, a couple of provider-specific steps remain:

  • Codex only runs hooks you trust. Open codex, run /hooks, and trust the aw:* entries. Do this again after a reinstall that changes a hook command.
  • Cursor may ask you to approve new MCP servers on first use. Run cursor-agent mcp enable <name> for each one.

Prerequisites

  • At least one agent CLI: Claude Code (claude), Codex (codex), or Cursor CLI (cursor-agent)
  • Node.js >= 20
  • Docker Desktop installed and running (required for Serena LSP)
  • GitHub CLI (gh) installed and authenticated (required by review skills)
  • jq installed (required by hooks and the statusline; brew install jq on macOS)
  • rtk: token-compressing CLI proxy (brew install rtk on macOS; installed automatically by setup.sh)
  • Python 3 + pip, required for headroom
  • headroom: context optimization layer (pip install "headroom-ai[all]"; installed automatically by setup.sh)

Setup

git clone https://github.com/vitalizecare/agentic-workflow.git ~/repos/agentic-workflow
cd ~/repos/agentic-workflow
./setup.sh                 # or: ./setup.sh --providers claude,codex,cursor

The setup script:

  • Checks hard prerequisites (jq, Docker) and detects which provider CLIs are installed (or uses --providers)
  • Creates the ~/.agentic-workflow/toolkit symlink and writes the ~/.agentic-workflow/providers registry
  • Symlinks native skills, /bootstrap, and the external design packs (cloned at pinned commits from EXTERNAL_PINS.env) into each selected provider's skills directory
  • Installs the safety hooks (block-destructive.sh, block-push-main.sh, detect-secrets.sh, rtk-rewrite.sh) and session-context hooks for each provider. Claude Code uses them natively; Codex and Cursor go through their hook adapters.
  • Installs and builds the MCP bridge
  • Builds the Serena Docker images (base TS/Python image; opt-in C# and Swift extensions) and installs the serena-docker wrapper to ~/.local/bin/
  • Registers the MCP servers (agentic-bridge, serena, headroom, prism-mcp, and xcodebuildmcp on macOS) with each selected provider
  • Configures prism-mcp (persistent memory, downloaded on first use) with its Mind Palace dashboard at http://localhost:7180 (PRISM_DASHBOARD_PORT). The prism-context.sh session-start hook warns if the dashboard is unreachable; /prismStatus runs a full health check
  • Installs rtk and headroom
  • Claude Code only: copies settings.json, installs the statusline and shell integration, and adds plugin marketplaces and plugins

Start the bridge

cd mcp-bridge && npm start    # Fastify on http://127.0.0.1:3100

Environment Variables

VariableDefaultDescription
PORT3100REST API port
HOST127.0.0.1Bind address (loopback only by default)
DB_PATH./bridge.dbSQLite database file path
ALLOW_REMOTEunsetSet to 1 to allow non-loopback binding

Contents

1. Skills

48 native skills, installed as symlinks into each provider's skills directory. Every skill uses the same text for every provider: steps name a capability, and the running agent uses its host's tool for it (see skills/_shared/capabilities.md).

StageSkills
Ideation & planningwithInterview, enhancePrompt, officeHours, autoplan, productReview, archReview, planDesignReview, planDevexReview
Designdesign-analyze, design-language, design-evolve, design-shotgun, design-mockup, design-implement, design-refine, design-verify (dispatchers auto-detect web/iOS and route to their -web / -ios sub-skills)
Build & verifyspecToProvenPR, verify-app (→ verify-web, verify-ios), ui-evidence
Reviewreview, postReview, addressReview, cso
Debug & QArootCause, bugHunt, bugReport, bugFixOrchestrator, testAudit
Ship & operateshipRelease, landAndDeploy, canary, syncDocs, weeklyRetro, prismStatus, judge
Repo setupbootstrap

Skills write their artifacts to ~/.agentic-workflow/<repo-slug>/<domain>/, and downstream skills discover them from there.

2. Bootstrap Skill

Run /bootstrap ($bootstrap in Codex) in any repo to generate its documentation:

  • Detects which of 17 Pivot-pattern docs exist (BUSINESS_PLAN, ARCHITECTURE, ERD, etc.)
  • Generates missing docs adapted to the target repo's tech stack
  • Writes a canonical AGENTS.md (a navigation doc, not a reference manual), with CLAUDE.md as a symlink to it
  • Infers glob-scoped rule files from the repo's structure into .agents/rules/, then links them for each provider with sync-rules.sh (.claude/rules, .cursor/rules/*.mdc)
  • Handles bare repos, partially documented repos, and well-documented repos

3. MCP Bridge

A TypeScript MCP server for bidirectional multi-agent communication across any mix of Claude Code, Codex, and Cursor sessions. All providers register the same stdio server and share one SQLite database.

MCP Tools:

  • send_context: send task context + meta-prompt between agents
  • get_messages: retrieve conversation history by UUID
  • get_unread: check for unread messages (marks them read on retrieval)
  • assign_task: assign tasks with domain and implementation details
  • report_status: report back with feedback or completion

API Endpoints:

  • POST /messages/send: send context between agents
  • GET /messages/conversation/:id: retrieve conversation history
  • GET /messages/unread?recipient=: fetch unread messages and mark them read
  • POST /tasks/assign: assign a task with domain classification
  • GET /tasks/:id: get a task by ID
  • GET /tasks/conversation/:id: get all tasks for a conversation
  • POST /tasks/report: report task status
  • GET /conversations: paginated conversation summaries

Features:

  • SQLite store-and-forward (messages queue while the recipient is offline)
  • Conversation continuity via UUID
  • Fastify REST API (port 3100) + MCP stdio server
  • End-to-end type safety with the AppResult<T> pattern
  • Atomic transactions for multi-step operations

4. Judge and Scorer

  • judge/ makes cheap, typed decisions: a rules fast path, then a per-content-class model chain. Model calls go through a headless provider CLI: claude-cli (claude -p), codex-cli (codex exec), or cursor-cli (cursor-agent -p). Any one of them on PATH is enough. The order is providers.agentClis in the judge config if set; otherwise claude, codex, cursor, with the current host (AW_PROVIDER) first. Cursor is slow (~8–13s per call), so on Cursor-only machines text questions often time out and fall back to rules.
  • scorer/ produces a daily cost and involvement report from agent transcripts. Each provider has its own transcript source: Claude Code (~/.claude/projects/), Codex (~/.codex/sessions/), and Cursor (~/.cursor/projects/). By default it reads every provider whose directory exists. --provider claude|codex|cursor|all narrows that, and --codex-dir / --cursor-dir override the paths. The report includes a "By provider" section. Cursor data is involvement-only because its transcripts carry no token or cost data. Codex rollouts imported from Claude are skipped so they aren't counted twice. Run scorer --since 7d to write a report to ~/.agentic-workflow/scorer/reports/. Install it with scripts/install-scorer.sh.

5. Statusline (Claude Code only)

config/statusline.sh is an adaptive two-line statusline for Claude Code sessions. setup.sh installs it to ~/.claude/statusline.sh and wires it into settings.json when Claude Code is a selected provider.

Columns (left → right, highest priority leftmost):

ColumnDescription
5h Usage5-hour rate-limit percentage + reset time
7d Usage7-day rate-limit percentage + reset day
ContextColor-coded bar + percentage of context window used
ModelActive model name (trimmed)
BranchCurrent git branch
CostSession cost in USD
TimeSession duration
CacheCache read hit rate
APIAPI wait percentage
LinesLines added/removed

Adaptive width tiers: columns drop out automatically as the terminal narrows.

TierMin widthColumns shown
FULL116 colsAll columns, branch up to 15 chars
MEDIUM101 colsNo Lines; branch up to 12 chars
NARROW78 colsNo Lines/Cache/API; 7d % only; narrow context bar
COMPACT65 cols5h % only; narrow context bar; branch up to 10 chars
COMPACT-S< 65 colsSame as COMPACT but drops Time column

The statusline reads the size of the tty its own Claude Code process is attached to (found by walking up the process tree), so every window gets its own width and no state is shared between windows. When that read fails it falls back to ~/.claude/terminal_width.d/<tty>, written by the shell integration from interactive terminals only (never from Claude Code tool shells). AW_STATUSLINE_DEBUG=1 prints the width source to stderr.

setup.sh installs the shell integration to ~/.claude/shell-integration.sh and sources it from ~/.zshrc / ~/.bashrc. It keeps ~/.claude/terminal_width.d/<tty> current and writes ~/.claude/shell_pid.d/<tty>, so the hooks of a session can send SIGWINCH to the shell on its own tty.

Testing

cd mcp-bridge && npm test   # Vitest, in-memory SQLite
cd scorer && npm test
cd judge && npm test
bash providers/tests/install.test.sh
bash scripts/tests/sync-rules.test.sh
bash config/hooks/tests/codex-adapter.test.sh
bash config/hooks/tests/cursor-adapter.test.sh
bash config/hooks/tests/provider-install-hooks.test.sh
bash config/lib/tests/merge-hook.test.sh
bash config/hooks/tests/probe-log.test.sh
scripts/sync-rules.sh --check

The full list is under Commands in AGENTS.md.

Tests cover unit tests (controllers, services, DB client, schemas, utilities) and integration tests (all REST routes via Fastify inject, plus the MCP tool handlers). /* v8 ignore */ annotations are prohibited; write the test instead.

Repository Layout

agentic-workflow/
├── AGENTS.md                # Canonical repo instructions (CLAUDE.md is a symlink to it)
├── .agents/rules/           # Glob-scoped rules, the only copy (.claude/rules, .cursor/rules/*.mdc are symlinks)
├── skills/                  # 48 native skills (+ _shared/ fragments, incl. capabilities.md)
├── bootstrap/               # /bootstrap — repo documentation generator
├── providers/<name>/        # Per-provider installers: install.sh, install-hooks.sh (claude, codex, cursor)
├── config/                  # Settings, MCP config, statusline, hooks (+ hooks/adapters/ for codex, cursor)
├── mcp-bridge/              # MCP bridge + REST API (Fastify, SQLite)
├── judge/                   # Cheap typed decisions (rules → model chain via claude/codex/cursor CLI)
├── scorer/                  # Cost/involvement report from provider transcripts
├── scripts/                 # sync-rules.sh, serena-docker, probe.sh, install-*.sh, refresh-external-pins.sh
├── planning/                # Project documentation (see planning/PROVIDERS.md)
├── Dockerfile.serena*       # Serena base image + opt-in C# / Swift extensions
└── setup.sh                 # One-command setup: ./setup.sh [--providers claude,codex,cursor]
Source 5 files
hooks/register.tsx 174 lines
1// aw-live: the scorer's numbers for THIS session, live inside Claude Code.
2// `/live` toggles a pane (`/live status` answers in text). The headline numbers
3// live in the statusline's Live column (config/statusline.sh), not here.
4// Every number comes from `scorer live --json` (one short Node process, never
5// in-process). A refresh never blocks a hook and every failure is silent: the
6// pane just keeps its last good value.
7
8import { atom, read, update } from 'claude-code'
9import type { EngineInterface, Register } from 'claude-code'
10
11import type { HostUsage, LiveSnapshot, LiveStatus } from '../types'
12import { paneSections, statusText } from './lib/format'
13import type { Tone } from './lib/format'
14import { binCandidates, isDue, liveArgs, runTimeout } from './lib/refresh'
15import { parseSnapshot } from './lib/snapshot'
16
17const PANE = 'aw-live'
18const TITLE = 'Live scorer'
19const PANE_COLUMNS = 72
20
21const snapshot = atom({ plugin: 'aw-live', key: 'snapshot' } as const, null as LiveSnapshot | null)
22const host = atom({ plugin: 'aw-live', key: 'host' } as const, null as HostUsage | null)
23const status = atom({ plugin: 'aw-live', key: 'status' } as const, null as LiveStatus | null)
24
25const TONE_COLOR: Record<Tone, string | undefined> = { ok: '#7ec699', warn: '#d4a054', dim: '#8b949e' }
26
27// Module variables restart on a hot reload; that only costs one extra refresh.
28let lastAt: number | null = null
29let running: Promise<void> | null = null
30
31/** Test seam: forget the refresh bookkeeping so each test starts cold. */
32export const resetForTests = (): void => {
33  lastAt = null
34  running = null
35}
36
37const isTimeout = (err: unknown): boolean => /time/i.test(err instanceof Error ? err.message : String(err))
38
39/** Runs scorer by argv, trying each candidate path; null on any failure. */
40const runScorer = async ($: EngineInterface, args: string[], timeoutMs: number): Promise<string | null> => {
41  const names = binCandidates(await $.env.get('HOME'), 'scorer', await $.env.get('AW_SCORER_BIN'))
42  for (const name of names) {
43    try {
44      const r = await $.process.run([name, ...args], { timeoutMs })
45      if (r.exitCode === 0) return r.stdout
46      return null
47    } catch (err) {
48      // A timeout means the binary exists but is slow: do not stack another wait on it.
49      if (isTimeout(err)) return null
50      // otherwise not runnable at this path: try the next candidate
51    }
52  }
53  return null
54}
55
56const runRefresh = async ($: EngineInterface, force: boolean): Promise<void> => {
57  let now: number | null = null
58  let ok = false
59  try {
60    now = await $.clock.now()
61    if (!force && !isDue(lastAt, now)) return
62    lastAt = now
63    const sessionId = await $.session.id()
64    if (sessionId === '') return
65    const [cwd, usage] = await Promise.all([$.session.cwd(), $.session.usage()])
66    await update($, host, (): HostUsage => ({
67      usd: usage.cost?.usd ?? null,
68      percent: usage.context.percent ?? null,
69      window: usage.context.window,
70    }))
71    const hasSnapshot = (await read($, snapshot)) !== null
72    const out = await runScorer($, liveArgs(sessionId, cwd, usage.context.window), runTimeout(hasSnapshot))
73    const parsed = out === null ? null : parseSnapshot(out)
74    if (parsed !== null) await update($, snapshot, () => parsed)
75    ok = parsed !== null
76  } catch {
77    // fail silent: the pane keeps its last good value
78  } finally {
79    const at = now
80    if (at !== null) await update($, status, (): LiveStatus => ({ ok, at })).catch(() => undefined)
81  }
82}
83
84/** One refresh at a time: a caller that arrives while one runs waits for that one. */
85const refresh = ($: EngineInterface, force: boolean): Promise<void> => {
86  if (running !== null) return running
87  running = runRefresh($, force).finally(() => {
88    running = null
89  })
90  return running
91}
92
93export const register: Register = on => {
94  on('session.start', async ($, e, next) => {
95    try {
96      await $.command.register({ name: 'live', description: 'Toggle the live scorer pane (context, cost, judge, gates)' })
97    } catch {
98      // a host error must not stop session start
99    }
100    return next(e)
101  })
102
103  on('turn.complete', async ($, e, next) => {
104    try {
105      // Only an open pane shows the numbers; /live refreshes on open, so a closed pane costs nothing.
106      if ((await $.ui.panes()).some(pane => pane.id === PANE)) {
107        $.clock.after(0, () => {
108          void refresh($, false)
109        })
110      }
111    } catch {
112      // a host error must not stop the turn
113    }
114    return next(e)
115  })
116
117  on('command.run', { command: 'live' }, async ($, e) => {
118    if (e.args.trim() === 'status') {
119      await refresh($, true)
120      try {
121        return { text: statusText(await read($, snapshot), await read($, host)) }
122      } catch {
123        return { text: 'aw-live: no numbers yet (is scorer installed? scripts/install-scorer.sh)' }
124      }
125    }
126    try {
127      if ((await $.ui.panes()).some(pane => pane.id === PANE)) {
128        await $.ui.close({ id: PANE })
129        return {}
130      }
131      await refresh($, true)
132      await $.ui.open({ id: PANE, title: TITLE, focus: true, closeOnEscape: true, columns: PANE_COLUMNS })
133    } catch {
134      // a host error: leave the pane as it is
135    }
136    return {}
137  })
138
139  on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e, next) => {
140    try {
141      const { Box, Text, Button } = $.ui.resolve(e)
142      const s = await read($, snapshot)
143      if (s === null) {
144        return (
145          <Box flexDirection="column">
146            <Text color={TONE_COLOR.dim}>No numbers yet. Is `scorer` installed? (scripts/install-scorer.sh)</Text>
147          </Box>
148        )
149      }
150      const sections = paneSections(s, await read($, host))
151      return (
152        <Box flexDirection="column">
153          {sections.map(section => (
154            <Box key={section.title} borderStyle="round" borderColor="#3d4450" paddingX={1} flexDirection="column">
155              <Text color="#7eb8da">{section.title}</Text>
156              {section.rows.map(row => (
157                <Text key={row.label} wrap="truncate-end">
158                  <Text color={TONE_COLOR.dim}>{row.label.padEnd(14)}</Text>
159                  <Text color={TONE_COLOR[row.tone]}>{row.value}</Text>
160                </Text>
161              ))}
162            </Box>
163          ))}
164          <Box flexDirection="row" columnGap={2} paddingX={1}>
165            <Button key="refresh" label="refresh" hotkey="r" plain onPress={() => refresh($, true)} />
166          </Box>
167        </Box>
168      )
169    } catch {
170      return next(e)
171    }
172  })
173}
174
hooks/lib/format.ts 107 lines
1// Pure formatting for the pane and `/live status`. No `$` in here.
2
3import type { HostUsage, LiveSnapshot } from '../../types'
4
5export type Tone = 'ok' | 'warn' | 'dim'
6export type Row = { label: string; value: string; tone: Tone }
7export type Section = { title: string; rows: Row[] }
8
9export const WARN_PERCENT = 70
10
11export const fmtTokens = (n: number): string => {
12  if (n >= 1_000_000) return `${(n / 1_000_000).toFixed(1)}M`
13  if (n >= 1_000) return `${Math.round(n / 1_000)}k`
14  return String(n)
15}
16
17/** `▓▓▓▓░░░░░░` for a percent; empty cells when unknown. */
18export const bar = (percent: number | null, width: number): string => {
19  const filled = percent === null ? 0 : Math.round((Math.min(100, Math.max(0, percent)) / 100) * width)
20  return '▓'.repeat(filled) + '░'.repeat(width - filled)
21}
22
23/** The host's own context percent when it has one, else the transcript's estimate. */
24export const contextPercent = (s: LiveSnapshot, host: HostUsage | null): number | null =>
25  host?.percent ?? s.context.percent
26
27/** `jev 83% · rules 17%` for the two biggest deciders; empty with none. */
28export const providerShare = (byProvider: Readonly<Record<string, number>>): string => {
29  const entries = Object.entries(byProvider).sort((a, b) => b[1] - a[1] || a[0].localeCompare(b[0]))
30  const total = entries.reduce((sum, [, n]) => sum + n, 0)
31  return entries.slice(0, 2).map(([name, n]) => `${name} ${Math.round((n / total) * 100)}%`).join(' · ')
32}
33
34const GATE_ORDER = ['scope-gate', 'done-gate', 'send-gate'] as const
35
36const decisionList = (byDecision: Readonly<Record<string, number>>): string =>
37  Object.entries(byDecision).sort((a, b) => b[1] - a[1] || a[0].localeCompare(b[0])).map(([d, n]) => `${d} ${n}`).join(' · ')
38
39/** The pane's cards, top to bottom. */
40export const paneSections = (s: LiveSnapshot, host: HostUsage | null): Section[] => {
41  const percent = contextPercent(s, host)
42  const window = host?.window ?? s.context.window
43  const context: Section = {
44    title: 'context',
45    rows: [
46      { label: 'window', value: `${percent === null ? '--' : `${Math.round(percent)}%`} ${bar(percent, 10)}`, tone: percent !== null && percent >= WARN_PERCENT ? 'warn' : 'ok' },
47      { label: 'tokens', value: `${s.context.tokens === null ? '--' : fmtTokens(s.context.tokens)} of ${fmtTokens(window)}`, tone: 'dim' },
48    ],
49  }
50  const session: Section = {
51    title: 'session',
52    rows: [
53      { label: 'cost', value: host?.usd == null ? '--' : `$${host.usd.toFixed(2)}`, tone: 'ok' },
54      { label: 'tokens', value: `${fmtTokens(s.usage.contextTokens)} in · ${fmtTokens(s.usage.outputTokens)} out`, tone: 'dim' },
55      { label: 'calls', value: `${s.usage.calls} (${s.usage.subagentCalls} subagent)`, tone: 'dim' },
56      { label: 'over 200k', value: String(s.usage.callsOver200k), tone: s.usage.callsOver200k > 0 ? 'warn' : 'dim' },
57    ],
58  }
59  if (s.ingest.readErrors > 0) {
60    session.rows.push({ label: 'warning', value: `stale: ${s.ingest.readErrors} transcript read error${s.ingest.readErrors === 1 ? '' : 's'}`, tone: 'warn' })
61  }
62  const sections = [context, session]
63  if (s.judge.state === 'unscoped') {
64    sections.push({ title: 'judge', rows: [{ label: 'sessions', value: 'not linked yet (judge predates decision_details.session_id)', tone: 'dim' }] })
65  }
66  if (s.judge.state === 'ok') {
67    const j = s.judge
68    sections.push({
69      title: 'judge',
70      rows: [
71        { label: 'calls', value: `${j.calls}${providerShare(j.byProvider) === '' ? '' : ` · ${providerShare(j.byProvider)}`}`, tone: 'ok' },
72        { label: 'vs rules', value: `agreed ${j.agreed} · overrode ${j.overrode} · unsure ${j.undecided}`, tone: j.undecided > 0 ? 'warn' : 'dim' },
73        { label: 'latency', value: `p50 ${j.p50LatencyMs} ms · p95 ${j.p95LatencyMs} ms`, tone: 'dim' },
74      ],
75    })
76    sections.push({
77      title: 'gates',
78      rows: GATE_ORDER.map((name): Row => {
79        const g = j.gates[name]
80        return {
81          label: name,
82          value: g.fired === 0 ? 'quiet' : `${g.fired} · ${decisionList(g.byDecision)} · p50 ${g.p50LatencyMs} ms`,
83          tone: g.fired === 0 ? 'dim' : 'ok',
84        }
85      }),
86    })
87  }
88  sections.push({
89    title: 'wakes',
90    rows: [
91      { label: 'queued now', value: String(s.wakes.queuedNow), tone: s.wakes.queuedNow > 0 ? 'warn' : 'dim' },
92      { label: 'context guard', value: `${s.contextGuard.fires} fires`, tone: s.contextGuard.fires > 0 ? 'warn' : 'dim' },
93    ],
94  })
95  return sections
96}
97
98/** The `/live status` answer: every pane row as plain text. */
99export const statusText = (s: LiveSnapshot | null, host: HostUsage | null): string => {
100  if (s === null) return 'aw-live: no numbers yet (is scorer installed? scripts/install-scorer.sh)'
101  const lines = paneSections(s, host).flatMap(section => [
102    `${section.title}`,
103    ...section.rows.map(row => `  ${row.label.padEnd(14)}${row.value}`),
104  ])
105  return lines.join('\n')
106}
107
hooks/lib/refresh.ts 24 lines
1// Pure pieces of the refresh loop. register.tsx owns every `$` call.
2
3export const MIN_GAP_MS = 2_000
4export const RUN_TIMEOUT_MS = 5_000
5/** The first refresh has no snapshot to fall back on; a cold ingest of a long session needs longer. */
6export const RUN_TIMEOUT_FIRST_MS = 30_000
7
8export const runTimeout = (hasSnapshot: boolean): number => (hasSnapshot ? RUN_TIMEOUT_MS : RUN_TIMEOUT_FIRST_MS)
9
10/** The `scorer live` command line. `window` is the host's context window, when known. */
11export const liveArgs = (sessionId: string, cwd: string, window: number | null): string[] => [
12  'live', '--session', sessionId, '--cwd', cwd, ...(window === null ? [] : ['--window', String(window)]), '--json',
13]
14
15/** Where to look for an aw CLI: an explicit override, the installer's directory, then PATH. */
16export const binCandidates = (home: string | undefined, name: 'scorer' | 'judge', override?: string): string[] => [
17  ...(override === undefined || override === '' ? [] : [override]),
18  ...(home === undefined || home === '' ? [name] : [`${home}/.local/bin/${name}`, name]),
19]
20
21/** A refresh is due when none has run yet or the last one is at least `minGapMs` old. */
22export const isDue = (lastAt: number | null, now: number, minGapMs: number = MIN_GAP_MS): boolean =>
23  lastAt === null || now - lastAt >= minGapMs
24
hooks/lib/snapshot.ts 26 lines
1// The contract with `scorer live --json`. scorer/src/live-snapshot.ts holds the zod
2// schema for the same shape; scorer/tests/live-contract.test.ts parses this mod's
3// fixture with it, so the two cannot drift apart unnoticed.
4
5import type { LiveSnapshot } from '../../types'
6
7const isObject = (value: unknown): value is Record<string, unknown> =>
8  typeof value === 'object' && value !== null && !Array.isArray(value)
9
10const JUDGE_STATES = ['none', 'unscoped', 'ok']
11
12/** Parses the CLI's stdout; null for anything that is not a version-1 snapshot. */
13export const parseSnapshot = (text: string): LiveSnapshot | null => {
14  let value: unknown
15  try {
16    value = JSON.parse(text)
17  } catch {
18    return null
19  }
20  if (!isObject(value) || value.v !== 1 || typeof value.sessionId !== 'string') return null
21  const { ingest, context, usage, judge, contextGuard, wakes } = value
22  if (!isObject(ingest) || !isObject(context) || !isObject(usage) || !isObject(contextGuard) || !isObject(wakes)) return null
23  if (!isObject(judge) || typeof judge.state !== 'string' || !JUDGE_STATES.includes(judge.state)) return null
24  return value as unknown as LiveSnapshot
25}
26
types/index.d.ts 62 lines
1export type GateName = 'scope-gate' | 'done-gate' | 'send-gate'
2
3export type GateStats = {
4  question: string
5  fired: number
6  byDecision: Record<string, number>
7  p50LatencyMs: number
8}
9
10export type JudgeLatest = {
11  id: string
12  ts: string
13  question: string
14  decision: string
15  provider: string
16  confidence: number
17  /** The eval_items id for this decision, once `judge label import` has made one. */
18  itemId: string | null
19}
20
21export type JudgeLive =
22  | { state: 'none' }
23  | { state: 'unscoped' }
24  | {
25      state: 'ok'
26      calls: number
27      byProvider: Record<string, number>
28      undecided: number
29      agreed: number
30      overrode: number
31      p50LatencyMs: number
32      p95LatencyMs: number
33      gates: Record<GateName, GateStats>
34      labelable: boolean
35      latest: JudgeLatest | null
36    }
37
38export type LiveSnapshot = {
39  v: 1
40  sessionId: string
41  at: string
42  transcript: { found: boolean; files: number }
43  /** readErrors: transcript files the scorer could not read this pass; the numbers may be stale. */
44  ingest: { newLines: number; ms: number; readErrors: number }
45  context: { tokens: number | null; window: number; percent: number | null }
46  usage: { calls: number; subagentCalls: number; contextTokens: number; outputTokens: number; callsOver200k: number }
47  judge: JudgeLive
48  contextGuard: { fires: number }
49  wakes: { queuedNow: number }
50}
51
52/** What the host itself reports; shown in preference to the transcript's estimate. */
53export type HostUsage = { usd: number | null; percent: number | null; window: number | null }
54
55export type LiveStatus = { ok: boolean; at: number }
56
57declare module 'claude-code' {
58  interface PluginState {
59    'aw-live': { snapshot: LiveSnapshot | null; host: HostUsage | null; status: LiveStatus | null }
60  }
61}
62