SLOPSHOPPER

costclaw-live

COSTCLAW LIVE: per-request cost, cache decay and repeated-read detection while the session runs, and it serves the Nth identical Read of an unchanged file from…

newbandguardcommandstatustimer
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · costclaw-live
› fix the failing auth test and add an audit log call ● costclaw-live: ⟦costclaw-live⟧ armed · serve-after=4 · rate card 2026-09-14 · evidence -> C:/Projects/claude-mods-rnd/prototypes/costclaw-live/evidence/live-preview-session.jsonl ● costclaw-live: ⟦costclaw-live⟧ COSTCLAW LIVE · efficiency 100/100 · repeated reads 0 (~0 tok) · cache health 0% · tool loop none · context 49% · $0.000000 ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /costclaw ⎿ costclaw-live: COSTCLAW LIVE — session preview-session ⎿ costclaw-live: COSTCLAW LIVE · efficiency 100/100 · repeated reads 0 (~0 tok) · cache health 0% · tool loop none · context 49 ⎿ costclaw-live: ⎿ costclaw-live: SCORE 100/100 (rubric: 100 minus cache<=30, repeats<=25, loop<=20, bloat<=15, decay<=10) ⎿ costclaw-live: cache not assessed · repeats -0 · loop -0 · bloat -0 · decay not assessed ⎿ costclaw-live: ● costclaw-live: ⟦costclaw-live⟧ cost cross-check: ours $0.000000 (0 requests, rate card 2026-09-14) vs engine $0.42 · delta -$0.42 · five-hour 31% ╭──────────────────────────────────────────────────────────────────────────────────────────────────────────────────╮ │ COSTCLAW LIVE · efficiency 100/100 · repeated reads 0 (~0 tok) · cache health 0% · tool loop non │ │ not assessed: cache,decay · saved so far ~0 tok ($0.000000) · serve-after 4 · /costclaw for evidence │ ╰──────────────────────────────────────────────────────────────────────────────────────────────────────────────────╯ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts ⚠ costclaw-live: COSTCLAW LIVE · efficiency 100/100 · repeated reads 0 (~0 tok) · cache health 0% · tool loop none

Draws

Band
╭──────────────────────────────────────────────────────────────────────────────────────────────────╮ │ COSTCLAW LIVE · efficiency 100/100 · repeated reads 0 (~0 tok) · cache health 0% · tool loop non │ │ not assessed: cache,decay · saved so far ~0 tok ($0.000000) · serve-after 4 · /costclaw for evi… │ ╰──────────────────────────────────────────────────────────────────────────────────────────────────╯
README

COSTCLAW LIVE

"Claude is wasting tokens RIGHT NOW."

A Claude Code function-hooks module that does live what C:\Projects\costclaw does post-hoc over ~/.claude/projects/.jsonl — and then acts on one finding instead of filing it**.

The action is the point. costclaw's REDUNDANT_TARGET_THRASH can tell you, a week later, that Read hit one file 14 times. This plugin sees the 4th identical Read of a file nothing has written to, answers the call from its own cache, and the tool never runs. The model is told, in a hidden context block, that the result came from cache. That is the only waste signal in the whole costclaw rule set whose fix is free.


What it is

RequirementWhere
Per-request usage, per model, session cost in USDturn.step hook, ported normalizeUsage + rate card
Per-request cache hit rate + early-third/late-third decayturn.step, cacheDecay()
(tool, target) ledger with request index and result sizetool.call hook, ported targetFor / commandTarget
Repeated reads graded high/medium/lowported AgentLens classifyRun
Tool loops, result bloat (> 20k chars)toolLoopStatus(), RESULT_BLOAT_CHARS
Intervention: Nth identical Read served from cachetool.call returns {result, context} without next
HUD above the promptui.render {component: "AbovePrompt", surface: "terminal"}
/costclaw findings with evidence$.command.register + command.run
$.session.usage() cross-checkturn.complete

Ported, not imported

A hooks module has no require and no import of repo code, so every reused piece is a copy, with its origin named in the header of hooks/index.tsx:

Ported fromWhat
C:\Projects\costclaw\packages\engine\src\pricing.tsPRICES_PER_MTOK (snapshot 2026-09-14), canonicalModel/priceFor/isKnownModel/isPremiumModel, readCounter, normalizeUsage (TTL split, invalid-counter rejection, the four conflict warnings), pricingMultiplier, costForNormalizedUsage, uncachedInputExposureForNormalizedUsage, cacheHitRate, roundMoneyHalfUp
C:\Projects\costclaw\packages\engine\src\parser.ts:65-100commandTarget() (the cd X && … hop stripper, dated 2026-09-11) and targetFor()
archaeology\clones\AgentLens\src\repeated-runs.jsclassifyRun() / detectRepeatedRuns() — the high/medium/low confidence taxonomy
C:\Projects\costclaw\packages\engine\src\optimizer.ts:10-37REDUNDANT_TARGET_MIN_FIRES=10, HEAVY_TOOL_FIRES=100, DECAY_DROP=0.25, LATE_FLOOR=0.5, severityForSavings bands

Deliberately NOT ported: costclaw's two share gates (HEAVY_TOOL_MIN_SHARE = 0.6, REDUNDANT_TARGET_MIN_SHARE = 0.25). Their comment records why they exist — "raw counts flagged 52% of real sessions", 2026-09-11 — i.e. they are a false-positive suppressor for a reader that cannot see the request boundary. AgentLens solved the same problem by grading the evidence instead of raising the bar, and turn.step is the request boundary. So the gates are dropped and classifyRun's requestSpread does the work; low-confidence runs (everything inside one model request — a parallel batch) never produce a dollar figure, exactly as AgentLens's header demands.

costclaw's dollar rule, kept

costclaw attaches estimatedMonthlySavingsUsd only where one is derivable, 0 otherwise — "never invented". Kept verbatim in spirit. Every $ this plugin prints comes from measured chars or measured tokens times the dated rate card:

  • repeated read / result bloat → measured result chars ÷ 4 × the primary model's input rate (an upper bound: those tokens might have landed as a cache read on a later request; labelled ~);
  • cache decay → the late third's own uncachedInputExposureForNormalizedUsage, scaled by the shortfall share (costclaw ruleMarathonSession's own arithmetic);
  • tool loops → no dollar figure. HEAVY_TOOL_LOOPS prints (no dollar figure is derivable).

costclaw's 8 rules: LIVE / NEEDS-HISTORY / OBSOLETE, and what is implemented here

#costclaw rule (optimizer.ts)Verdict liveImplemented in this plugin?
1BAD_CACHE_HIT — 3 consecutive sessions under 50% hitNEEDS-HISTORY (cross-session; $.store makes it live from session 4)No. Per-session hit rate is computed and shown; the 3-session streak is not kept.
2HOT_PROJECT — one project ≥60% of spendNEEDS-HISTORY (cross-project)No. $.session.repo() is captured so the key exists; nothing is aggregated.
3HEAVY_TOOL_LOOPS — ≥100 fires AND ≥60% shareLIVE (two counters over tool.call)Yes, HEAVY_TOOL_FIRES=100 kept, share gate dropped, plus AgentLens consecutive-run grading. No dollar figure.
4REDUNDANT_TARGET_THRASH — same (tool,target) ≥10 AND ≥25% shareLIVE, and the cheapest real winYes — and it intervenes. Reported from 2 fires with a confidence grade; REDUNDANT_TARGET_MIN_FIRES=10 upgrades the ruleId. The cache-serve fires at the configurable serveAfter (default 4).
5SHORT_SESSION_BLOAT — <10 min AND ≥$5LIVE (a tripwire, not a postmortem)No. state.startedAt and state.costUsd are both tracked, so it is ~6 lines; out of scope for this build.
6MARATHON_SESSION — ≥40 turns, cache hit drops ≥25pp early→late, late <50%LIVE (turnCache is literally per-request usage)Yes, DECAY_DROP/LATE_FLOOR kept verbatim; MARATHON_MIN_TURNS 40 → MARATHON_MIN_REQUESTS 6 (see caveat below). Does not call $.session.compact().
7MODEL_MISROUTE — premium model on a trivial sessionOBSOLETE as a finding. Every predicate is a property of a finished session; live the rule inverts into a router on agent.spawn / turn.stepNo, by design. isPremiumModel() is ported and each request is tagged premium, so the predicate is available; this plugin never rewrites a model.
8WORKFLOW_COST — delegated spend ≥$5 and its shareLIVE and strictly better (agentId replaces costclaw's directory-nesting inference)Partly. Every request records its agentId, so the per-agent ledger is one groupBy away; the rule itself is not implemented.
—windows.ts computeWindowStats (5-hour / 7-day inference from 15-min cost buckets)OBSOLETE — $.session.usage().rateLimits returns the real numbersN/A. Not ported; the real five_hour / seven_day percentages are read and printed.
—burnClock (day × 6-hour rhythm, ≥7 active days)NEEDS-HISTORYNo.
—RULE_ERROR (per-rule try/catch)LIVEReplaced by the engine: a hook that throws is skipped and the chain continues.

Added here, not in costclaw: RESULT_BLOAT (a single tool result > 20,000 chars) and SERVED_FROM_CACHE (the intervention's own ledger, the only finding with a negative dollar figure — money not spent).

Caveat on #6: costclaw needs 40 turns before it trusts thirds, because it is about to publish a monthly figure. Live, at request 6 (2 per third) the plugin claims only a decay signal with its own numbers printed next to it, never a monthly figure. That is a deliberate loosening, named here.


Architecture

                                  ┌──────────────────────── ENGINE (claude 2.1.273) ────────────────────────┐
  user prompt ──▶ turn.start ──▶  │   turn.step  ×N      tool.call  ×M      turn.complete      ui.render    │
                                  └──┬──────────────────────┬───────────────────┬──────────────────┬───────┘
                                     │                      │                   │                  │
  ══════════════════════════ hooks chain: prepend ▸ USER(this plugin) ▸ append ▸ builtin ▸ core ═══════════════
                                     │                      │                   │                  │
  hooks/index.tsx, registration order (first = outermost within this plugin):
                                     │                      │                   │                  │
  1 on("turn.step")  ◀───────────────┘                      │                   │                  │
      requestIndex += 1          ── THE REQUEST BOUNDARY ──▶ │  (every tool.call below belongs to it)
      for await (c of next(e)) yield c ;  r = await stream.result
      normalizeUsage(r.usage) ▸ costForNormalizedUsage ▸ uncachedInputExposure ▸ cacheHitRate
      push state.requests[] ─────────────────────────────────────────────┐
                                                                         │
  2 on("tool.call") ◀────────────────────────────────────────┘           │
      target = targetFor(e.tool, e)                                      │
      Write/Edit/NotebookEdit  ──▶ drop readCache[target], mark mutated   │
      Read & cached & nth >= serveAfter & !mutated & $.fs.stat unchanged  │
          ──▶ RETURN { result: <first call's result>, context:[ … ] }     │   ◀── tool never runs
      otherwise  r = await next(e)                                        │
          ──▶ targetCount, targetRequests(Set of request idx), toolEvents │
          ──▶ chars > 20000 ? state.bloat                                 │
          ──▶ first successful Read ? readCache[key] = {result,size,mtime}│
                                                                         │
  3 on("turn.complete") ◀──────────────────────────────────┐             │
      $.session.usage() ▸ context % · five_hour % · cost ──┼──▶ cross-check vs our own sum
      $.fs.write evidence/live-<sessionId>.jsonl           │             │
                                                           │             ▼
  4 on("ui.render", {AbovePrompt, terminal}) ◀─────────────┘   efficiencyScore() ▸ buildFindings()
      $.ui.resolve(e) ▸ <Box borderStyle="round" borderColor="cyan"> … </Box>
                                                                         │
  5 on("command.run", {command:"costclaw"}) ──▶ { text: full report } ◀───┘
  6 on("session.start") ──▶ $.command.register ▸ $.store.get ▸ $.clock.every(2000) ▸ $.ui.status

Middleware position. One plugin, user tier. Six registrations on six distinct events, so the registration order above only fixes nesting within this plugin, not against other plugins. If you load it alongside a recorder such as lab/demos/blackbox, put the recorder first: this plugin answers a served Read without calling next, so anything nested beneath it never sees that call (the lesson lab/README.md records from guardian vs blackbox).

Efficiency score — the rubric (deterministic, this plugin's own)

Start at 100 and deduct. A component with no evidence yet is not assessed and deducts nothing — costclaw scoring.ts:28 makeCheck's discipline, because early in a session almost nothing is assessable. The HUD footer names the skipped components rather than implying full marks.

ComponentMaxRuleAssessed once
cache−30max(0, 0.80 − sessionHitRate) × 100 × 0.5, capped≥ 2 model requests
repeats−253 per redundant same-target call, high/medium confidence only≥ 1 tool call
loop−20active −20, emerging −10, none 0≥ 3 tool calls
bloat−155 per tool result over 20,000 chars≥ 1 tool call
decay−10−10 when DECAY_DROP crossed and late third < LATE_FLOOR≥ 6 model requests

Observed: 5 reads of one file, 3 of them real → 4 redundant → −12 → 88/100.


Run it

Validate:

claude plugin validate C:\Projects\claude-mods-rnd\prototypes\costclaw-live --json

Interactive (the HUD):

claude --plugin-dir C:\Projects\claude-mods-rnd\prototypes\costclaw-live

then ask it to read one file several times, and run /costclaw. /costclaw serve 3 lowers the cache-serve threshold; /costclaw serve 10 raises it.

Headless, reproducing the intervention exactly (bash; BATCH_GUARD_LIMIT / REPEAT_GUARD_OFF relax the production harness's own ~/.claude guards, which otherwise deny the 4th consecutive single Read — see "Failure behaviour"):

BATCH_GUARD_LIMIT=99 REPEAT_GUARD_OFF=1 claude \
  --plugin-dir C:/Projects/claude-mods-rnd/prototypes/costclaw-live \
  -p "Do these steps strictly one at a time, never in parallel. Use the Read tool for every Read step; never substitute Bash. The file is C:/Projects/claude-mods-rnd/prototypes/costclaw-live/evidence/probe-target.txt 1) Read the file 2) Run the Bash command: echo s2 # SEQ: probe 3) Read the file again 4) Run the Bash command: echo s4 # SEQ: probe 5) Read the file again 6) Run the Bash command: echo s6 # SEQ: probe 7) Read the file again with the Read tool 8) Read the file again with the Read tool. Then say DONE and quote verbatim any extra note that came with any tool result." \
  --model haiku --allowedTools "Bash,Read" --debug

Interactive via pty (captures the HUD to a file):

BATCH_GUARD_LIMIT=99 REPEAT_GUARD_OFF=1 python C:/Projects/claude-mods-rnd/tools/pty_drive.py \
  C:/Projects/claude-mods-rnd/prototypes/costclaw-live/evidence/pty-hud.txt 25 \
  "C:/Projects/claude-mods-rnd/prototypes/costclaw-live" \
  '<the same prompt>|||WAIT:6|||KEYS:/costclaw\r|||WAIT:10' 90 "" --allowedTools Read,Bash --debug

Every run appends a machine-readable ledger at prototypes/costclaw-live/evidence/live-<sessionId>.jsonl (kind ∈ start | request | tool | cached | served | invalidate | bloat | usage-check | findings). The findings row is written at every main-loop turn.complete and is the only way a headless (-p) run can read the findings, since there is no HUD and no /costclaw there.


What it proves

  1. A hook can answer a tool call with a previous call's result and the model accepts it. Session ad705704: five Read calls on one file; the ledger holds 3 tool rows (requests 1, 3, 5) and 2 served rows (requests 7 and 9, nth 4 and 5). The model's own report: > Read #4 … System-reminders: tool.call hook additional context: costclaw-live: served from cache, file unchanged since first read (saved ~22 tokens)

Engine debug log:

[costclaw-live] $.ui.log: ⟦costclaw-live⟧ SERVED FROM CACHE read #4 of …probe-target.txt saved ~22 tokens ($0.000022) · file unchanged since request #1

  1. The guard refuses when the file changed (this is the check made to fail on purpose; a cache that never declines has been run, not verified). Session f7bb2c37: the same five-read sequence with a Bash append between read 3 and read 4 — which does not trip the in-session Write/Edit flag, so only the $.fs.stat size+mtime comparison can catch it. Result: 5 tool rows, 0 served rows; reads 4 and 5 carry chars: 108 against the cached 87, and the model reported "Read 4: 5 lines … APPENDED-LINE-9999 present". Session 8749421a covers the other branch: an Edit on the path emits {"kind":"invalidate","tool":"Edit",…,"requestId":6}.
  1. Per-request cost accounting works and disagrees with the engine by a stated amount. Session ced1cacd, 9 requests. The ledger's usage-check row at turn.complete: {"oursUsd":0.1034,"engineUsd":0.1351664,"deltaUsd":-0.0318,"contextPercent":34,"contextTokens":67650,"contextWindow":200000,"rateLimits":[{"kind":"five_hour",…56},{"kind":"seven_day",…13}],"requests":9}. /costclaw, run a few seconds later against a fresher $.session.usage(), printed ours $0.10 · engine $.session.usage().cost $0.15 · delta -$0.05 followed by the DISCREPANCY line explaining it (ours starts at plugin load and is list-price; neither is an invoice) — the two snapshots differ because the engine's figure kept moving, which is itself the point of printing both. Model breakdown on screen: claude-haiku-4-5-20251001 requests 9 · in 74 · out 931 · cacheR 495564 · cacheW 39294 · $0.10, and 9x Cache writes without TTL metadata were estimated at the standard five-minute rate — turn.step's TurnUsage is the four-field shape, so the ported normalizeUsage takes its legacy branch and says so instead of silently guessing.
  1. The HUD draws, in the required format (pty session ced1cacd, 150 columns):
   ╭────────────────────────────────────────────────────────────────────────────────────────────────╮
   │ COSTCLAW LIVE · efficiency 88/100 · repeated reads 4 (~6330 tok) · cache health 93% ·           │
   │                 tool loop none · context 34% · $0.10                                           │
   │ REPEATED_TARGET [high] $0.0063 — Read hit …/fixtures/probe-hud.txt 5x                          │
   │ SERVED_FROM_CACHE [high] -$0.0053 — 2 Read calls answered by costclaw-live, not by the tool    │
   │ all components assessed · saved so far ~5276 tok ($0.0053) · serve-after 4 · /costclaw …       │
   ╰────────────────────────────────────────────────────────────────────────────────────────────────╯

With a 10 KB file the two served reads kept ~5,276 tokens ($0.0053) out of the context.

  1. /costclaw prints evidence, not assertions — which file, how many times, which request indices, the confidence grade, the threshold that did or did not fire, and decay NOT ASSESSED (n/6 requests) where there is not enough evidence yet.
  1. Result bloat fires on a real result. Session ca13ccc2: {"kind":"bloat","tool":"Read","target":"…/probe-bloat.txt","chars":55317,"tokens":13830}.
  1. AgentLens's confidence grading becomes a fact. Every (tool,target) carries the set of request indices it spanned, taken from turn.step, so requestSpread >= 3 && targetSpread <= 1 → high is measured, not inferred from repeated requestId strings in a JSONL file. Session 99dff7b0, the findings row:
   REPEATED_TARGET high 0.00633 | Read hit …/fixtures/probe-hud.txt 5x
     | 5 Read calls across 5 model requests on the same target — strong signal of
       repeated work without progress. requests=[1,3,5,7,8]
   SERVED_FROM_CACHE high -0.005276 | 2 Read calls answered by costclaw-live, not by the tool
     | #7 …/probe-hud.txt (read 4) ~2638 tok · #8 …/probe-hud.txt (read 5) ~2638 tok

Note requests=[1,3,5,7,8] — five distinct model requests, so the run is high and is allowed a dollar figure. Had all five landed inside one request (a parallel batch) it would have graded low and carried none. The same session reproduces 88/100 headlessly, matching the pty run.

What it does NOT prove

  • Not an invoice reconciliation. The delta against $.session.usage().cost was −$0.05 on $0.15. Our sum covers only requests seen since the plugin loaded and uses a list-price snapshot.
  • ~chars/4 is not a tokenizer. Every saving figure is an estimate, printed with ~, and the dollar figure is an upper bound (the tokens might have been served as a cache read anyway).
  • The cache-serve was exercised on small and 10 KB files, one path, one offset/limit combination, with serveAfter 4. Not exercised: offset/limit variants (they key separately and were never collided), a file deleted between reads, a subagent reading the same file (e.agentId is recorded but the serve path is not agent-scoped), concurrent parallel reads of the same path inside one request.
  • MARATHON_SESSION never fired. The decay branch was only observed in its not fired and NOT ASSESSED states; no session here lost cache. The firing branch is untested.
  • HEAVY_TOOL_LOOPS (≥100 fires) never fired, and no high-confidence consecutive run ever formed — every session here interleaved Bash between Reads, so detectRepeatedRuns saw runs of
  • tool loop stayed none throughout. The active/emerging branches are untested at runtime.
  • $.store persistence across sessions was not verified. /costclaw serve <n> writes it and session.start reads it, but no two-session test was run.
  • No subagent was spawned, so agentId attribution on turn.step is carried but unexercised.
  • **Nothing was measured about whether the intervention makes the model *worse*** — a stale-but- unchanged file is by definition the same bytes, but the model also loses the fresh system-reminders a real Read carries.
  • Known human-surface defect: the /costclaw report is ~45 lines and does not fit a 45-row terminal. In the pty capture the SCORE, FINDINGS and REPEATED sections scrolled off the top; only INTERVENTIONS downward stayed on screen. The HUD band carries the top-five findings so nothing is unreachable, but the report itself wants paging or a $.ui.open pane (which needs ≥110 columns from a command, per §10) before it is a good human artifact.

API assumptions

Status column is from MOD_CAPABILITY_MAP.md §20 as of build 2.1.273.

Event / methodUsed forCapability-map status
on("turn.step") async generator, next(e) stream + stream.resultper-request usage, the request boundaryCONFIRMED [RUN]
on("tool.call") — observe + await next(e)ledger, bloat, cachingCONFIRMED [RUN]
on("tool.call") — answer without next, {result, context}the interventionCONFIRMED [RUN] (result replacement + hidden context); verified again here
on("turn.complete")usage cross-check, evidence flushCONFIRMED [RUN]
on("ui.render", {component:"AbovePrompt", surface:"terminal"}) + $.ui.resolveHUDCONFIRMED [RUN] — the band always draws
on("command.run", {command:"costclaw"}) → {text}the reportCONFIRMED [RUN]
on("session.start")boot, register, timerCONFIRMED [RUN]
$.command.register({immediate:true})/costclawCONFIRMED [RUN]
$.session.usage() → {context, rateLimits, cost}context %, five-hour %, costCONFIRMED [RUN]
$.session.id(), $.session.repo()evidence filename, project keyCONFIRMED [RUN]
$.fs.stat(path) → {kind, size, mtimeMs}the unchanged checkop event, [RUN] here (fs.* listed §2c)
$.fs.write(path, text)evidence JSONLCONFIRMED [RUN]
$.ui.log / $.ui.status / $.ui.invalidatetranscript lines, status row, redrawCONFIRMED [RUN]
$.clock.every(2000, fn)status refresh[DECL] in §20 (experimental) — [RUN] here, ticked for the whole pty session
$.store.get / $.store.setserveAfter across sessions[DECL] in §20 — read/written here, cross-session persistence not verified
plugin.json userConfig → register(on, options)serveAfter default[DECL] — declared and accepted by the validator; a value was never set through /config, so the options path is unexercised

Not used, deliberately: classic.*, prompt.section, prompt.context, skill.prompt (withheld from the user tier on this machine); agent.spawn / tool.check / $.session.compact() (the natural next steps — see below); $.http.fetch (this plugin never leaves the process).


Failure behaviour

  • A hook throws → the engine skips that frame and the chain continues; the debug log names the plugin, the event and the reason. Observed count of hook failures across the 5 session debug logs from these runs (ad705704, f7bb2c37, 8749421a, ca13ccc2, ced1cacd): 0 each (grep -icE 'hooks module (failed|threw)|hook failed|costclaw-live.*(threw|failed)'). Concretely: if turn.step throws, the model request still happens and only the accounting is lost; if the tool.call frame throws before next, the tool runs normally.
  • $.fs.stat fails or the path is not a plain file → statUnchanged returns false and the call goes to the real tool. The cach
Source 1 files
hooks/index.tsx 1009 lines
1// ============================================================================
2// COSTCLAW LIVE — "Claude is wasting tokens RIGHT NOW"
3//
4// A hooks module that does live what C:\Projects\costclaw does post-hoc over
5// ~/.claude/projects/**.jsonl, and then acts on one finding instead of filing
6// it: the Nth identical Read of an unchanged file is answered from this
7// plugin's own cache and never reaches the tool or the bill.
8//
9// PORTED (a Mod cannot import; this is a copy, read-only origins):
10//   - C:\Projects\costclaw\packages\engine\src\pricing.ts
11//       PRICES_PER_MTOK rate card (snapshot 2026-09-14), canonicalModel/lookup/
12//       priceFor, readCounter, normalizeUsage (incl. the cache_creation TTL
13//       split + invalid-counter rejection + conflict warnings),
14//       pricingMultiplier, costForNormalizedUsage,
15//       uncachedInputExposureForNormalizedUsage, cacheHitRate, isPremiumModel.
16//   - C:\Projects\costclaw\packages\engine\src\parser.ts:65-100
17//       commandTarget() (the "cd X && ..." hop stripper, 2026-09-11) and
18//       targetFor() — the (tool, target) grain the thrash rules run on.
19//   - C:\Projects\claude-mods-rnd\archaeology\clones\AgentLens\src\repeated-runs.js
20//       classifyRun() / detectRepeatedRuns() — the high/medium/low confidence
21//       taxonomy by REQUEST SPREAD. Live this stops being a heuristic: the
22//       request boundary is an event (turn.step), not a guess about requestIds.
23//   - C:\Projects\costclaw\packages\engine\src\optimizer.ts thresholds
24//       REDUNDANT_TARGET_MIN_FIRES=10, HEAVY_TOOL_FIRES=100, DECAY_DROP=0.25,
25//       LATE_FLOOR=0.5, and severityForSavings' bands.
26//
27// DELIBERATELY NOT PORTED: costclaw's share gates (HEAVY_TOOL_MIN_SHARE 0.6,
28// REDUNDANT_TARGET_MIN_SHARE 0.25). They exist to suppress false positives that
29// AgentLens solves by grading the evidence instead of raising the bar, and a
30// hook has the request boundary that made the gates necessary. See README.
31//
32// costclaw's rule: a Finding carries a dollar figure only where one is
33// DERIVABLE from measured numbers, and 0/none otherwise. Kept. Every $ printed
34// below comes from measured chars or measured tokens times the dated rate card.
35// ============================================================================
36
37const EVIDENCE_DIR = "C:/Projects/claude-mods-rnd/prototypes/costclaw-live/evidence/";
38
39// ---------------------------------------------------------------------------
40// PORTED from costclaw packages/engine/src/pricing.ts
41// ---------------------------------------------------------------------------
42const PRICING_SNAPSHOT_DATE = "2026-09-14";
43
44// USD per 1,000,000 tokens. Dated rate card; update in lockstep with costclaw.
45const PRICES_PER_MTOK = {
46  "claude-fable-5-1": { input: 10, output: 50, cache_write: 12.5, cache_write_1h: 20, cache_read: 0.25 },
47  "claude-fable-5": { input: 10, output: 50, cache_write: 12.5, cache_write_1h: 20, cache_read: 1 },
48  "claude-mythos-5-1": { input: 10, output: 50, cache_write: 12.5, cache_write_1h: 20, cache_read: 0.25 },
49  "claude-mythos-5": { input: 10, output: 50, cache_write: 12.5, cache_write_1h: 20, cache_read: 1 },
50  "claude-opus-5": { input: 5, output: 25, cache_write: 6.25, cache_write_1h: 10, cache_read: 0.5 },
51  "claude-opus-4-8": { input: 5, output: 25, cache_write: 6.25, cache_write_1h: 10, cache_read: 0.5 },
52  "claude-sonnet-5": { input: 2, output: 10, cache_write: 2.5, cache_write_1h: 4, cache_read: 0.2 },
53  "claude-opus-4-7": { input: 5, output: 25, cache_write: 6.25, cache_write_1h: 10, cache_read: 0.5 },
54  "claude-opus-4-7[1m]": { input: 5, output: 25, cache_write: 6.25, cache_write_1h: 10, cache_read: 0.5 },
55  "claude-opus-4-6": { input: 5, output: 25, cache_write: 6.25, cache_write_1h: 10, cache_read: 0.5 },
56  "claude-opus-4-5": { input: 5, output: 25, cache_write: 6.25, cache_write_1h: 10, cache_read: 0.5 },
57  "claude-opus-4-1": { input: 15, output: 75, cache_write: 18.75, cache_write_1h: 30, cache_read: 1.5 },
58  "claude-opus-4": { input: 15, output: 75, cache_write: 18.75, cache_write_1h: 30, cache_read: 1.5 },
59  "claude-sonnet-4-6": { input: 3, output: 15, cache_write: 3.75, cache_write_1h: 6, cache_read: 0.3 },
60  "claude-sonnet-4-5": { input: 3, output: 15, cache_write: 3.75, cache_write_1h: 6, cache_read: 0.3 },
61  "claude-sonnet-4": { input: 3, output: 15, cache_write: 3.75, cache_write_1h: 6, cache_read: 0.3 },
62  "claude-haiku-4-5": { input: 1, output: 5, cache_write: 1.25, cache_write_1h: 2, cache_read: 0.1 },
63  "claude-haiku-4-5-20251001": { input: 1, output: 5, cache_write: 1.25, cache_write_1h: 2, cache_read: 0.1 },
64  "claude-haiku-3-5": { input: 0.8, output: 4, cache_write: 1, cache_write_1h: 1.6, cache_read: 0.08 },
65};
66const FALLBACK = { input: 3, output: 15, cache_write: 3.75, cache_write_1h: 6, cache_read: 0.3 };
67
68const INVALID_USAGE_WARNING = "Invalid usage counters were ignored.";
69const CACHE_SPLIT_CONFLICT_WARNING = "Conflicting cache-write totals were valued at the conservative one-hour rate.";
70const CACHE_TTL_ASSUMED_WARNING = "Cache writes without TTL metadata were estimated at the standard five-minute rate.";
71const CACHE_TTL_FALLBACK_WARNING = "Cache writes with unusable TTL metadata were conservatively estimated at the one-hour rate.";
72const UNSUPPORTED_PRICING_WARNING = "Unsupported pricing metadata was ignored and standard current rates were used.";
73
74function canonicalModel(model) {
75  if (!model) return null;
76  const noTag = model.replace(/\[[^\]]*\]$/, "");
77  return noTag.replace(/-\d{8}$/, "");
78}
79function lookupRates(model) {
80  if (!model) return null;
81  return PRICES_PER_MTOK[model]
82    || PRICES_PER_MTOK[model.replace(/\[[^\]]*\]$/, "")]
83    || PRICES_PER_MTOK[canonicalModel(model) || ""]
84    || null;
85}
86function priceFor(model) { return lookupRates(model) || FALLBACK; }
87function isKnownModel(model) { return lookupRates(model) !== null; }
88function isPremiumModel(model) {
89  const base = canonicalModel(model);
90  if (!base) return false;
91  return base.indexOf("claude-opus") === 0 || base.indexOf("claude-fable") === 0 || base.indexOf("claude-mythos") === 0;
92}
93
94function readCounter(value, present) {
95  if (!present) return { value: 0, present: false, invalid: false };
96  return typeof value === "number" && Number.isSafeInteger(value) && value >= 0
97    ? { value, present: true, invalid: false }
98    : { value: 0, present: true, invalid: true };
99}
100function own(value, key) { return Object.prototype.hasOwnProperty.call(value, key); }
101
102// normalizeUsage — ported whole. turn.step's TurnUsage is the FOUR-FIELD shape
103// (input/output/cache_read/cache_creation, no ephemeral_5m/1h split), which is
104// exactly the legacy branch below: it values the writes at the 5-minute rate
105// and records CACHE_TTL_ASSUMED_WARNING so the estimate is never silent.
106function normalizeUsage(raw) {
107  const usage = raw && typeof raw === "object" ? raw : {};
108  const warnings = new Set();
109  const input = readCounter(usage.input_tokens, own(usage, "input_tokens"));
110  const output = readCounter(usage.output_tokens, own(usage, "output_tokens"));
111  const cacheRead = readCounter(usage.cache_read_input_tokens, own(usage, "cache_read_input_tokens"));
112  const cacheTop = readCounter(usage.cache_creation_input_tokens, own(usage, "cache_creation_input_tokens"));
113
114  let splitPresent = false;
115  let splitMalformed = false;
116  let cache5m = { value: 0, present: false, invalid: false };
117  let cache1h = { value: 0, present: false, invalid: false };
118  const splitProvided = own(usage, "cache_creation");
119  if (splitProvided) {
120    if (usage.cache_creation && typeof usage.cache_creation === "object" && !Array.isArray(usage.cache_creation)) {
121      const split = usage.cache_creation;
122      cache5m = readCounter(split.ephemeral_5m_input_tokens, own(split, "ephemeral_5m_input_tokens"));
123      cache1h = readCounter(split.ephemeral_1h_input_tokens, own(split, "ephemeral_1h_input_tokens"));
124      splitPresent = cache5m.present || cache1h.present;
125    } else {
126      splitMalformed = true;
127    }
128  }
129
130  let invalid = input.invalid || output.invalid || cacheRead.invalid || cacheTop.invalid || cache5m.invalid || cache1h.invalid || splitMalformed;
131  if (invalid) warnings.add(INVALID_USAGE_WARNING);
132
133  const validSplitSum = cache5m.value + cache1h.value;
134  let creationTotal = 0;
135  let creation5m = 0;
136  let creation1h = 0;
137  let conflicting = false;
138  if (cacheTop.present && !cacheTop.invalid) {
139    creationTotal = cacheTop.value;
140    if (splitPresent && !cache5m.invalid && !cache1h.invalid) {
141      if (validSplitSum === cacheTop.value) {
142        creation5m = cache5m.value;
143        creation1h = cache1h.value;
144      } else {
145        conflicting = true;
146        creation1h = cacheTop.value;
147        warnings.add(CACHE_SPLIT_CONFLICT_WARNING);
148      }
149    } else if (splitProvided) {
150      creation1h = cacheTop.value;
151      warnings.add(CACHE_TTL_FALLBACK_WARNING);
152    } else {
153      creation5m = cacheTop.value;
154      if (cacheTop.value > 0) warnings.add(CACHE_TTL_ASSUMED_WARNING);
155    }
156  } else if (splitPresent) {
157    creationTotal = validSplitSum;
158    creation5m = cache5m.value;
159    creation1h = cache1h.value;
160  }
161
162  let speed = "standard";
163  if (usage.speed === "fast") speed = "fast";
164  else if (usage.speed != null && usage.speed !== "standard") warnings.add(UNSUPPORTED_PRICING_WARNING);
165
166  let inferenceGeo = "global";
167  if (usage.inference_geo === "us") inferenceGeo = "us";
168  else if (usage.inference_geo != null && usage.inference_geo !== "global") warnings.add(UNSUPPORTED_PRICING_WARNING);
169
170  let batch = false;
171  if (usage.batch === true) batch = true;
172  else if (usage.batch != null && usage.batch !== false) warnings.add(UNSUPPORTED_PRICING_WARNING);
173
174  const normalized = {
175    input_tokens: input.value,
176    output_tokens: output.value,
177    cache_creation_input_tokens: creationTotal,
178    cache_read_input_tokens: cacheRead.value,
179    cache_creation_5m_input_tokens: creation5m,
180    cache_creation_1h_input_tokens: creation1h,
181    cache_ttl_known: splitPresent || splitProvided,
182    web_search_requests: 0,
183    speed,
184    inference_geo: inferenceGeo,
185    batch,
186  };
187  const billable = normalized.input_tokens > 0
188    || normalized.output_tokens > 0
189    || normalized.cache_creation_input_tokens > 0
190    || normalized.cache_read_input_tokens > 0;
191  return { usage: normalized, billable, warnings: [...warnings], invalid, conflicting };
192}
193
194function pricingMultiplier(model, usage) {
195  let multiplier = 1;
196  const base = canonicalModel(model);
197  if (usage.speed === "fast" && (base === "claude-opus-5" || base === "claude-opus-4-8")) multiplier *= 2;
198  if (usage.inference_geo === "us") {
199    const geoEligible = base != null && (
200      base.indexOf("claude-fable-5") === 0
201      || base.indexOf("claude-mythos-5") === 0
202      || base === "claude-opus-5" || base === "claude-opus-4-8"
203      || base === "claude-opus-4-7" || base === "claude-opus-4-6"
204      || base === "claude-sonnet-5" || base === "claude-sonnet-4-6"
205    );
206    if (geoEligible) multiplier *= 1.1;
207  }
208  if (usage.batch) multiplier *= 0.5;
209  return multiplier;
210}
211
212function costForNormalizedUsage(model, usage) {
213  const p = priceFor(model);
214  const tokenCost = (
215    usage.input_tokens * p.input
216    + usage.output_tokens * p.output
217    + usage.cache_creation_5m_input_tokens * p.cache_write
218    + usage.cache_creation_1h_input_tokens * p.cache_write_1h
219    + usage.cache_read_input_tokens * p.cache_read
220  ) / 1000000;
221  return tokenCost * pricingMultiplier(model, usage);
222}
223
224// The genuinely recoverable half: raw input that should have been a cache read.
225function uncachedInputExposureForNormalizedUsage(model, usage) {
226  const p = priceFor(model);
227  return usage.input_tokens * (p.input - p.cache_read) * pricingMultiplier(model, usage) / 1000000;
228}
229
230// cacheHitRate — reads over everything that entered the context this request.
231function cacheHitRateOf(inputTok, creationTok, readTok) {
232  const denom = inputTok + creationTok + readTok;
233  if (!Number.isFinite(denom) || denom <= 0) return 0;
234  return readTok / denom;
235}
236
237function roundMoneyHalfUp(value, places) {
238  const n = places === undefined ? 2 : places;
239  if (!Number.isFinite(value)) return 0;
240  const factor = Math.pow(10, n);
241  const scaled = Math.abs(value) * factor;
242  const tolerance = Number.EPSILON * Math.max(1, scaled) * 2;
243  return Math.sign(value) * Math.floor(scaled + 0.5 + tolerance) / factor;
244}
245
246// ---------------------------------------------------------------------------
247// PORTED from costclaw packages/engine/src/parser.ts:65-100
248// The (tool, target) grain feeds the same-target thrash rule, so a shell target
249// must identify WHAT ran: drop a leading "cd <dir> &&" / "Set-Location <dir>;"
250// hop and env assignments, then keep the command plus its first argument.
251// First-word-only collided every "cd repo && ..." call on "cd" (2026-09-11).
252// ---------------------------------------------------------------------------
253function commandTarget(command) {
254  let s = command.trim();
255  const hop = /^(?:cd|Set-Location|pushd)\s+(?:"[^"]*"|'[^']*'|\S+)\s*(?:&&|;)\s*/i;
256  const env = /^[A-Za-z_][A-Za-z0-9_]*=(?:"[^"]*"|'[^']*'|\S*)\s+/;
257  for (;;) {
258    const next2 = s.replace(hop, "").replace(env, "");
259    if (next2 === s) break;
260    s = next2;
261  }
262  const words = s.split(/\s+/).filter(Boolean);
263  if (words.length === 0) return null;
264  return words.slice(0, 2).join(" ").slice(0, 60);
265}
266
267function targetFor(name, input) {
268  const s = (v, n) => (typeof v === "string" && v.length ? v.slice(0, n) : null);
269  if (["Read", "Edit", "Write", "NotebookEdit"].includes(name)) return s(input.file_path, 200);
270  if (["Grep", "Glob"].includes(name)) return s(input.path || input.glob || input.pattern, 160);
271  if (["Bash", "PowerShell"].includes(name)) {
272    return typeof input.command === "string" ? commandTarget(input.command) : null;
273  }
274  if (name.indexOf("Task") === 0) return s(input.subject || input.description, 80);
275  if (["WebFetch", "WebSearch"].includes(name)) return s(input.url || input.query, 60);
276  if (name === "Agent") return s(input.subagent_type || input.description, 60);
277  return null;
278}
279
280// ---------------------------------------------------------------------------
281// PORTED from AgentLens src/repeated-runs.js — the confidence taxonomy.
282// Inputs: tool events in session order, each {name, requestId?, target?}.
283//   high   — same target across >= 3 distinct model requests (repeated work,
284//            no progress). ONLY this grade may carry a dollar figure.
285//   medium — >= 2 distinct requests but the target signal is weak.
286//   low    — the run sits inside ONE model request: a batch, not a stuck loop.
287// "callers should NOT compute savings from low-confidence runs."
288// ---------------------------------------------------------------------------
289const RUN_THRESHOLD = 3;
290
291function classifyRun(run, startIndex, endIndex) {
292  const name = run[0].name;
293  const requestIds = new Set(run.map(x => x.requestId).filter(Boolean));
294  const targets = new Set(run.map(x => x.target).filter(t => t !== null && t !== undefined && t !== ""));
295  const requestSpread = requestIds.size;
296  const targetSpread = targets.size;
297  const allInOneRequest = requestSpread === 1 || (requestSpread === 0 && run.every(x => x.requestId === run[0].requestId));
298
299  let confidence = "low";
300  let evidence;
301  if (requestSpread >= 3 && targetSpread <= 1) {
302    confidence = "high";
303    evidence = `${run.length} ${name} calls across ${requestSpread} model requests on the same target — strong signal of repeated work without progress.`;
304  } else if (requestSpread >= 2) {
305    confidence = "medium";
306    evidence = targetSpread > 1
307      ? `${run.length} ${name} calls across ${requestSpread} requests, touching ${targetSpread} different targets — could be batch follow-ups rather than a loop.`
308      : `${run.length} ${name} calls across ${requestSpread} requests — same tool but unknown target similarity.`;
309  } else {
310    confidence = "low";
311    evidence = allInOneRequest
312      ? `${run.length} ${name} calls inside a single model request — almost certainly a batch, not a stuck loop.`
313      : `${run.length} ${name} calls with no request_id evidence — cannot tell loop from batch.`;
314  }
315  return {
316    name, count: run.length, startIndex, endIndex, confidence, evidence,
317    requestSpread, targetSpread, allInOneRequest: !!allInOneRequest,
318    targets: [...targets].slice(0, 8),
319  };
320}
321
322function detectRepeatedRuns(events, threshold) {
323  const th = threshold === undefined ? RUN_THRESHOLD : threshold;
324  if (!Array.isArray(events) || events.length < th) return [];
325  const signals = [];
326  let runName = null;
327  let runStart = -1;
328  for (let i = 0; i <= events.length; i++) {
329    const cur = i < events.length ? events[i] : null;
330    if (cur && cur.name === runName) continue;
331    if (runName !== null && (i - runStart) >= th) {
332      signals.push(classifyRun(events.slice(runStart, i), runStart, i - 1));
333    }
334    runName = cur ? cur.name : null;
335    runStart = i;
336  }
337  return signals;
338}
339
340// ---------------------------------------------------------------------------
341// PORTED thresholds — costclaw packages/engine/src/optimizer.ts:10-37
342// ---------------------------------------------------------------------------
343const REDUNDANT_TARGET_MIN_FIRES = 10; // 10, not lower: re-reading a file you are
344                                       // actively editing 5-9 times in one session
345                                       // is normal iterative work, not thrash.
346const HEAVY_TOOL_FIRES = 100;
347const DECAY_DROP = 0.25;
348const LATE_FLOOR = 0.5;
349const MARATHON_MIN_REQUESTS = 6;       // lowered from costclaw's MARATHON_MIN_TURNS=40:
350                                       // the batch rule needs 40 turns of history to
351                                       // trust thirds; live we only claim a decay
352                                       // SIGNAL (never a monthly figure) so 6 requests
353                                       // is the smallest window with 2 per third.
354const RESULT_BLOAT_CHARS = 20000;      // a single tool result over 20k chars
355const DEFAULT_SERVE_AFTER = 4;         // the Nth identical Read gets served from cache
356const CHARS_PER_TOKEN = 4;             // tokenizer estimate; every figure using it is "~"
357
358function severityForSavings(usd) {
359  if (usd >= 50) return "critical";
360  if (usd >= 10) return "warn";
361  return "info";
362}
363
364// ---------------------------------------------------------------------------
365// Session state (in-memory for this session; nothing is uploaded anywhere).
366// ---------------------------------------------------------------------------
367const state = {
368  sessionId: "",
369  repo: "",
370  startedAt: Date.now(),
371  serveAfter: DEFAULT_SERVE_AFTER,
372
373  requestIndex: 0,          // increments on every turn.step — THE request boundary
374  requests: [],             // {idx, model, agentId, usage, costUsd, hitRate, exposureUsd}
375  byModel: {},              // model -> {input, output, cacheRead, cacheWrite, costUsd, requests}
376  costUsd: 0,
377  warnings: {},             // normalizeUsage warning -> count
378  unknownModels: {},
379
380  toolEvents: [],           // {name, requestId, target, chars, t, served}
381  targetCount: {},          // "tool|target" -> fires
382  targetRequests: {},       // "tool|target" -> Set of request indices
383  bloat: [],                // {tool, target, chars, requestId}
384
385  readCache: {},            // readKey -> {result, chars, size, mtimeMs, firstRequest, target}
386  mutated: {},              // target -> count of Write/Edit/NotebookEdit
387  served: [],               // {target, requestId, tokens, usd, n}
388  savedTokens: 0,
389  savedUsd: 0,
390
391  lastUsage: null,          // $.session.usage() snapshot
392  usageChecks: [],          // {ours, engine, deltaUsd}
393  log: [],                  // evidence JSONL rows
394  dirty: false,
395};
396
397function primaryModel() {
398  let best = null;
399  let bestTok = -1;
400  for (const m of Object.keys(state.byModel)) {
401    const e = state.byModel[m];
402    const tok = e.input + e.cacheRead + e.cacheWrite + e.output;
403    if (tok > bestTok) { bestTok = tok; best = m; }
404  }
405  return best;
406}
407
408function totals() {
409  let input = 0, output = 0, cacheRead = 0, cacheWrite = 0;
410  for (const m of Object.keys(state.byModel)) {
411    const e = state.byModel[m];
412    input += e.input; output += e.output; cacheRead += e.cacheRead; cacheWrite += e.cacheWrite;
413  }
414  return { input, output, cacheRead, cacheWrite };
415}
416
417function sessionHitRate() {
418  const t = totals();
419  return cacheHitRateOf(t.input, t.cacheWrite, t.cacheRead);
420}
421
422// Early-third vs late-third cache decay (costclaw MARATHON_SESSION, live).
423// Returns {assessed:false} until there are enough requests — costclaw
424// scoring.ts:28 makeCheck discipline: a check with no evidence is NOT assessed,
425// it is not defaulted to zero.
426function cacheDecay() {
427  const n = state.requests.length;
428  if (n < MARATHON_MIN_REQUESTS) return { assessed: false, reason: `${n}/${MARATHON_MIN_REQUESTS} requests` };
429  const third = Math.floor(n / 3);
430  if (third === 0) return { assessed: false, reason: "third is empty" };
431  let eRead = 0, eTot = 0, lRead = 0, lTot = 0, lateExposure = 0;
432  for (let i = 0; i < third; i++) { eRead += state.requests[i].read; eTot += state.requests[i].total; }
433  for (let i = n - third; i < n; i++) { lRead += state.requests[i].read; lTot += state.requests[i].total; lateExposure += state.requests[i].exposureUsd; }
434  if (lTot === 0) return { assessed: false, reason: "late third had no context tokens" };
435  const earlyHit = eTot > 0 ? eRead / eTot : 0;
436  const lateHit = lRead / lTot;
437  const drop = earlyHit - lateHit;
438  const fired = drop >= DECAY_DROP && lateHit < LATE_FLOOR;
439  // Savings, derivable: the share of the late third's uncached-input exposure
440  // attributable to the shortfall against the early hit rate.
441  const expectedReads = lTot * earlyHit;
442  const shortfall = Math.max(0, expectedReads - lRead);
443  const lateMiss = lTot - lRead;
444  const affectedShare = lateMiss > 0 ? Math.min(1, shortfall / lateMiss) : 0;
445  return { assessed: true, fired, earlyHit, lateHit, drop, requests: n, usd: lateExposure * affectedShare };
446}
447
448// Repeated reads: same (tool,target) more than once, graded with classifyRun
449// over that target's own events, so "same target across >=3 model requests"
450// is a FACT (the request index came from turn.step), not an inference.
451function repeatedTargets() {
452  const out = [];
453  for (const key of Object.keys(state.targetCount)) {
454    const fires = state.targetCount[key];
455    if (fires < 2) continue;
456    const sep = key.indexOf("|");
457    const tool = key.slice(0, sep);
458    const target = key.slice(sep + 1);
459    const run = state.toolEvents.filter(x => x.name === tool && x.target === target);
460    const graded = classifyRun(run, 0, run.length - 1);
461    const redundant = fires - 1;
462    const chars = run.reduce((s, x) => s + (x.chars || 0), 0);
463    const perCall = run.length ? Math.round(chars / run.length) : 0;
464    const wastedTokens = Math.ceil((perCall * redundant) / CHARS_PER_TOKEN);
465    out.push({ tool, target, fires, redundant, graded, wastedTokens, thrash: fires >= REDUNDANT_TARGET_MIN_FIRES });
466  }
467  out.sort((a, b) => b.fires - a.fires);
468  return out;
469}
470
471function toolLoopStatus() {
472  const runs = detectRepeatedRuns(state.toolEvents, RUN_THRESHOLD);
473  let heavy = null;
474  const byTool = {};
475  for (const x of state.toolEvents) byTool[x.name] = (byTool[x.name] || 0) + 1;
476  for (const t of Object.keys(byTool)) if (byTool[t] >= HEAVY_TOOL_FIRES) heavy = { tool: t, fires: byTool[t] };
477  const high = runs.filter(r => r.confidence === "high");
478  const medium = runs.filter(r => r.confidence === "medium");
479  let status = "none";
480  if (high.length || heavy) status = "active";
481  else if (medium.length) status = "emerging";
482  return { status, runs, high, medium, heavy };
483}
484
485// Tokens a repeated read would have cost, priced at the session's primary model
486// input rate. Upper bound (the tokens might have landed as a cache read on a
487// later request); stated as such in the README and as "~" everywhere it prints.
488function usdForTokens(tokens) {
489  const p = priceFor(primaryModel());
490  return tokens * p.input / 1000000;
491}
492
493// ---------------------------------------------------------------------------
494// Efficiency score /100 — this plugin's own rubric, deterministic, documented.
495// Start at 100, deduct. A component with no evidence yet is NOT assessed and
496// deducts nothing (costclaw scoring.ts:28 makeCheck discipline); the HUD says
497// which components were skipped rather than pretending they scored full marks.
498//   cache      up to -30  max(0, 0.80 - hitRate) * 100 * 0.5, needs >= 2 requests
499//   repeats    up to -25  3 per redundant high/medium-confidence same-target read
500//   loop       up to -20  active -20, emerging -10
501//   bloat      up to -15  5 per tool result over 20k chars
502//   decay      up to -10  -10 when DECAY_DROP crossed and late third < LATE_FLOOR
503// ---------------------------------------------------------------------------
504function efficiencyScore() {
505  const parts = [];
506  let score = 100;
507
508  if (state.requests.length >= 2) {
509    const hit = sessionHitRate();
510    const pen = Math.min(30, Math.max(0, Math.round((0.80 - hit) * 100 * 0.5)));
511    score -= pen; parts.push({ name: "cache", penalty: pen, assessed: true });
512  } else parts.push({ name: "cache", penalty: 0, assessed: false });
513
514  const reps = repeatedTargets();
515  const graded = reps.filter(r => r.graded.confidence !== "low");
516  if (state.toolEvents.length > 0) {
517    const redundant = graded.reduce((s, r) => s + r.redundant, 0);
518    const pen = Math.min(25, redundant * 3);
519    score -= pen; parts.push({ name: "repeats", penalty: pen, assessed: true });
520  } else parts.push({ name: "repeats", penalty: 0, assessed: false });
521
522  const loop = toolLoopStatus();
523  if (state.toolEvents.length >= RUN_THRESHOLD) {
524    const pen = loop.status === "active" ? 20 : loop.status === "emerging" ? 10 : 0;
525    score -= pen; parts.push({ name: "loop", penalty: pen, assessed: true });
526  } else parts.push({ name: "loop", penalty: 0, assessed: false });
527
528  if (state.toolEvents.length > 0) {
529    const pen = Math.min(15, state.bloat.length * 5);
530    score -= pen; parts.push({ name: "bloat", penalty: pen, assessed: true });
531  } else parts.push({ name: "bloat", penalty: 0, assessed: false });
532
533  const decay = cacheDecay();
534  if (decay.assessed) {
535    const pen = decay.fired ? 10 : 0;
536    score -= pen; parts.push({ name: "decay", penalty: pen, assessed: true });
537  } else parts.push({ name: "decay", penalty: 0, assessed: false });
538
539  return { score: Math.max(0, Math.min(100, Math.round(score))), parts, notAssessed: parts.filter(p => !p.assessed).map(p => p.name) };
540}
541
542// ---------------------------------------------------------------------------
543// Findings. A dollar figure ONLY where it is derivable from measured numbers.
544// costclaw's rule, kept verbatim in spirit: never invented.
545// ---------------------------------------------------------------------------
546function buildFindings() {
547  const out = [];
548
549  for (const r of repeatedTargets()) {
550    if (r.redundant < 1) continue;
551    const derivable = r.graded.confidence === "high";
552    out.push({
553      ruleId: r.thrash ? "REDUNDANT_TARGET_THRASH" : "REPEATED_TARGET",
554      confidence: r.graded.confidence,
555      title: `${r.tool} hit ${r.target} ${r.fires}x`,
556      evidence: `${r.graded.evidence} requests=[${[...(state.targetRequests[r.tool + "|" + r.target] || [])].join(",")}]`,
557      usd: derivable ? usdForTokens(r.wastedTokens) : null,
558      tokens: derivable ? r.wastedTokens : null,
559      severity: derivable ? severityForSavings(usdForTokens(r.wastedTokens)) : "info",
560    });
561  }
562
563  const loop = toolLoopStatus();
564  for (const r of loop.high) {
565    out.push({
566      ruleId: "HEAVY_TOOL_LOOPS", confidence: "high",
567      title: `Consecutive ${r.name} run (${r.count} calls)`,
568      evidence: r.evidence, usd: null, tokens: null, severity: "warn",
569    });
570  }
571  if (loop.heavy) {
572    out.push({
573      ruleId: "HEAVY_TOOL_LOOPS", confidence: "high",
574      title: `${loop.heavy.tool} fired ${loop.heavy.fires}x (>= ${HEAVY_TOOL_FIRES})`,
575      evidence: `costclaw HEAVY_TOOL_FIRES threshold crossed.`, usd: null, tokens: null, severity: "warn",
576    });
577  }
578
579  for (const b of state.bloat) {
580    const tokens = Math.ceil(b.chars / CHARS_PER_TOKEN);
581    out.push({
582      ruleId: "RESULT_BLOAT", confidence: "high",
583      title: `${b.tool} returned ${b.chars} chars from ${b.target}`,
584      evidence: `A single tool result over ${RESULT_BLOAT_CHARS} chars entered the context at request #${b.requestId}; ~${tokens} tokens.`,
585      usd: usdForTokens(tokens), tokens, severity: severityForSavings(usdForTokens(tokens)),
586    });
587  }
588
589  const decay = cacheDecay();
590  if (decay.assessed && decay.fired) {
591    out.push({
592      ruleId: "MARATHON_SESSION", confidence: "high",
593      title: `Cache hit fell ${Math.round(decay.earlyHit * 100)}% -> ${Math.round(decay.lateHit * 100)}%`,
594      evidence: `${decay.requests} model requests; early-third vs late-third drop ${Math.round(decay.drop * 100)}pp crossed DECAY_DROP=${DECAY_DROP * 100}pp with the late third under LATE_FLOOR=${LATE_FLOOR * 100}%.`,
595      usd: decay.usd, tokens: null, severity: severityForSavings(decay.usd),
596    });
597  }
598
599  if (state.served.length) {
600    out.push({
601      ruleId: "SERVED_FROM_CACHE", confidence: "high",
602      title: `${state.served.length} Read call${state.served.length === 1 ? "" : "s"} answered by costclaw-live, not by the tool`,
603      evidence: state.served.map(s => `#${s.requestId} ${s.target} (read ${s.n}) ~${s.tokens} tok`).join(" · "),
604      usd: -state.savedUsd, tokens: -state.savedTokens, severity: "info",
605    });
606  }
607
608  return out;
609}
610
611// ---------------------------------------------------------------------------
612// Rendering helpers
613// ---------------------------------------------------------------------------
614// Two decimals for real money, more for the sub-cent figures a single cached
615// read is worth: rounding those to $0.00 hid the only number the saving has.
616function money(usd) {
617  const a = Math.abs(usd);
618  const places = a >= 0.01 ? 2 : a >= 0.0001 ? 4 : 6;
619  const v = roundMoneyHalfUp(usd, places);
620  return (v < 0 ? "-$" : "$") + Math.abs(v).toFixed(places);
621}
622function pct(x) { return Math.round(x * 100) + "%"; }
623
624function headline() {
625  const s = efficiencyScore();
626  const reps = repeatedTargets().filter(r => r.redundant >= 1);
627  const redundant = reps.reduce((a, r) => a + r.redundant, 0);
628  const redundantTokens = reps.reduce((a, r) => a + (r.graded.confidence === "high" ? r.wastedTokens : 0), 0);
629  const loop = toolLoopStatus();
630  const ctx = state.lastUsage && state.lastUsage.context && state.lastUsage.context.percent !== undefined
631    ? state.lastUsage.context.percent : "?";
632  return `COSTCLAW LIVE · efficiency ${s.score}/100 · repeated reads ${redundant} (~${redundantTokens} tok) · cache health ${pct(sessionHitRate())} · tool loop ${loop.status} · context ${ctx}% · ${money(state.costUsd)}`;
633}
634
635function findingLines() {
636  return buildFindings().map(f => {
637    const dollars = f.usd === null || f.usd === undefined ? "" : ` ${money(f.usd)}`;
638    return `${f.ruleId} [${f.confidence}]${dollars} — ${f.title}`;
639  });
640}
641
642// ---------------------------------------------------------------------------
643// $-using helpers. The loader's static scan allows `$` as the parameter of a
644// TOP-LEVEL function declaration, used only as `$.noun.method(...)`.
645// ---------------------------------------------------------------------------
646function note($, line) { $.ui.log(line); }
647
648async function flushEvidence($) {
649  if (!state.dirty) return;
650  state.dirty = false;
651  const name = state.sessionId ? state.sessionId : "unknown";
652  const body = state.log.map(x => JSON.stringify(x)).join("\n") + "\n";
653  try { await $.fs.write(EVIDENCE_DIR + "live-" + name + ".jsonl", body); } catch (err) { /* evidence is best-effort */ }
654}
655
656async function statUnchanged($, path, entry) {
657  try {
658    const st = await $.fs.stat(path);
659    return st && st.kind === "file" && st.size === entry.size && st.mtimeMs === entry.mtimeMs;
660  } catch (err) {
661    return false;
662  }
663}
664
665async function readServeAfter($) {
666  try {
667    const v = await $.store.get("serveAfter");
668    if (typeof v === "number" && Number.isSafeInteger(v) && v >= 2 && v <= 50) return v;
669  } catch (err) { /* store is optional */ }
670  return DEFAULT_SERVE_AFTER;
671}
672
673async function writeServeAfter($, n) {
674  try { await $.store.set("serveAfter", n); } catch (err) { /* store is optional */ }
675}
676
677async function refreshUsage($) {
678  try {
679    state.lastUsage = await $.session.usage();
680    return state.lastUsage;
681  } catch (err) {
682    return null;
683  }
684}
685
686async function bootSession($) {
687  try { state.sessionId = await $.session.id(); } catch (err) { state.sessionId = "unknown"; }
688  try { const r = await $.session.repo(); state.repo = r && r.name ? r.name : ""; } catch (err) { state.repo = ""; }
689}
690
691function pushLog(kind, data) {
692  state.log.push(Object.assign({ t: Date.now(), kind }, data));
693  if (state.log.length > 4000) state.log.shift();
694  state.dirty = true;
695}
696
697function tick($) { $.ui.status(headline()); }
698
699// ---------------------------------------------------------------------------
700export const register = (on, options) => {
701
702  // ---- 1. turn.step: per-request usage, cost, cache hit rate ----------------
703  // turn.step is THE request boundary. Every tool.call after it belongs to this
704  // index, which is what makes AgentLens's confidence grading a fact.
705  on("turn.step", async function* ($, e, next) {
706    state.requestIndex += 1;
707    const myIndex = state.requestIndex;
708    const stream = next(e);
709    for await (const c of stream) yield c;
710    const r = await stream.result;
711
712    const raw = r && r.usage ? r.usage : null;
713    if (raw) {
714      const model = raw.model || e.model;
715      const norm = normalizeUsage(raw);
716      for (const w of norm.warnings) state.warnings[w] = (state.warnings[w] || 0) + 1;
717      if (!isKnownModel(model)) state.unknownModels[model] = (state.unknownModels[model] || 0) + 1;
718      const u = norm.usage;
719      const costUsd = costForNormalizedUsage(model, u);
720      const exposureUsd = uncachedInputExposureForNormalizedUsage(model, u);
721      const read = u.cache_read_input_tokens;
722      const total = u.input_tokens + u.cache_creation_input_tokens + u.cache_read_input_tokens;
723      const hit = cacheHitRateOf(u.input_tokens, u.cache_creation_input_tokens, read);
724
725      state.costUsd += costUsd;
726      const m = state.byModel[model] || { input: 0, output: 0, cacheRead: 0, cacheWrite: 0, costUsd: 0, requests: 0 };
727      m.input += u.input_tokens; m.output += u.output_tokens;
728      m.cacheRead += u.cache_read_input_tokens; m.cacheWrite += u.cache_creation_input_tokens;
729      m.costUsd += costUsd; m.requests += 1;
730      state.byModel[model] = m;
731
732      state.requests.push({
733        idx: myIndex, model, agentId: e.agentId || null,
734        input: u.input_tokens, output: u.output_tokens,
735        cacheRead: read, cacheWrite: u.cache_creation_input_tokens,
736        read, total, hitRate: hit, costUsd, exposureUsd,
737        premium: isPremiumModel(model),
738      });
739      pushLog("request", {
740        idx: myIndex, model, agentId: e.agentId || null, messageCount: e.messageCount,
741        input: u.input_tokens, output: u.output_tokens, cacheRead: read,
742        cacheWrite: u.cache_creation_input_tokens, hitRate: hit,
743        costUsd, exposureUsd, sessionCostUsd: state.costUsd, warnings: norm.warnings,
744      });
745    }
746    return r;
747  });
748
749  // ---- 2 + 3. tool.call: record, detect, and INTERVENE ----------------------
750  on("tool.call", async ($, e, next) => {
751    const tool = e.tool;
752    const target = targetFor(tool, e);
753    const requestId = state.requestIndex;
754
755    // A Write/Edit invalidates the cache for that path, always, first.
756    if ((tool === "Write" || tool === "Edit" || tool === "NotebookEdit") && target) {
757      state.mutated[target] = (state.mutated[target] || 0) + 1;
758      for (const k of Object.keys(state.readCache)) {
759        if (state.readCache[k].target === target) delete state.readCache[k];
760      }
761      pushLog("invalidate", { tool, target, requestId });
762    }
763
764    // --- THE INTERVENTION -------------------------------------------------
765    // The Nth identical Read of a file nothing wrote to, whose size+mtime are
766    // still what they were at the first read, is answered from this plugin's
767    // cache: no tool run, no tokens, and the model is told in a hidden context
768    // block that the result is the cached one.
769    if (tool === "Read" && target) {
770      const readKey = target + "|" + (e.offset === undefined ? "" : e.offset) + "|" + (e.limit === undefined ? "" : e.limit);
771      const entry = state.readCache[readKey];
772      const priorFires = state.targetCount["Read|" + target] || 0;
773      const nth = priorFires + 1;
774      if (entry && nth >= state.serveAfter && !state.mutated[target]) {
775        const unchanged = await statUnchanged($, target, entry);
776        if (unchanged) {
777          const tokens = Math.ceil(entry.chars / CHARS_PER_TOKEN);
778          const usd = usdForTokens(tokens);
779          state.savedTokens += tokens; state.savedUsd += usd;
780          state.served.push({ target, requestId, tokens, usd, n: nth });
781          state.targetCount["Read|" + target] = nth;
782          const reqs = state.targetRequests["Read|" + target] || new Set();
783          reqs.add(requestId); state.targetRequests["Read|" + target] = reqs;
784          state.toolEvents.push({ name: "Read", requestId, target, chars: 0, t: Date.now(), served: true });
785          pushLog("served", { target, requestId, nth, tokens, usd, savedTokensTotal: state.savedTokens, savedUsdTotal: state.savedUsd, firstReadAtRequest: entry.firstRequest });
786          note($, `⟦costclaw-live⟧ SERVED FROM CACHE  read #${nth} of ${target}  saved ~${tokens} tokens (${money(usd)})  · file unchanged since request #${entry.firstRequest}`);
787          await flushEvidence($);
788          return {
789            result: entry.result,
790            context: [`costclaw-live: served from cache, file unchanged since first read (saved ~${tokens} tokens)`],
791          };
792        }
793      }
794    }
795
796    // --- normal path ------------------------------------------------------
797    const r = await next(e);
798
799    const text = r && typeof r.text === "string" ? r.text : (r && r.result !== undefined ? safeLen(r.result) : "");
800    const chars = typeof text === "string" ? text.length : 0;
801
802    if (target) {
803      const key = tool + "|" + target;
804      state.targetCount[key] = (state.targetCount[key] || 0) + 1;
805      const reqs = state.targetRequests[key] || new Set();
806      reqs.add(requestId);
807      state.targetRequests[key] = reqs;
808    }
809    state.toolEvents.push({ name: tool, requestId, target, chars, t: Date.now(), served: false });
810    if (state.toolEvents.length > 5000) state.toolEvents.shift();
811
812    if (chars > RESULT_BLOAT_CHARS) {
813      state.bloat.push({ tool, target, chars, requestId });
814      pushLog("bloat", { tool, target, chars, requestId, tokens: Math.ceil(chars / CHARS_PER_TOKEN) });
815      note($, `⟦costclaw-live⟧ RESULT BLOAT  ${tool} ${target} returned ${chars} chars (~${Math.ceil(chars / CHARS_PER_TOKEN)} tokens)`);
816    }
817
818    // Cache the first successful Read so the Nth can be served from it.
819    if (tool === "Read" && target && r && !r.isError && r.result !== undefined) {
820      const readKey = target + "|" + (e.offset === undefined ? "" : e.offset) + "|" + (e.limit === undefined ? "" : e.limit);
821      if (!state.readCache[readKey] && !state.mutated[target]) {
822        let size = -1, mtimeMs = -1;
823        try { const st = await $.fs.stat(target); if (st && st.kind === "file") { size = st.size; mtimeMs = st.mtimeMs; } } catch (err) { /* uncacheable */ }
824        if (size >= 0) {
825          state.readCache[readKey] = { result: r.result, chars, size, mtimeMs, firstRequest: requestId, target };
826          pushLog("cached", { target, requestId, chars, size, mtimeMs });
827        }
828      }
829    }
830
831    pushLog("tool", { tool, target, requestId, chars, isError: !!(r && r.isError), deny: r && r.deny ? String(r.deny).slice(0, 120) : null });
832    state.dirty = true;
833    return r;
834  });
835
836  // ---- 6. turn.complete: $.session.usage() cross-check ---------------------
837  on("turn.complete", async ($, e, next) => {
838    const r = await next(e);
839    if (!e.agentId) {
840      const u = await refreshUsage($);
841      if (u) {
842        const engineUsd = u.cost && typeof u.cost.usd === "number" ? u.cost.usd : null;
843        const check = {
844          oursUsd: roundMoneyHalfUp(state.costUsd, 4),
845          engineUsd,
846          deltaUsd: engineUsd === null ? null : roundMoneyHalfUp(state.costUsd - engineUsd, 4),
847          contextPercent: u.context ? u.context.percent : null,
848          contextTokens: u.context ? u.context.tokens : null,
849          contextWindow: u.context ? u.context.window : null,
850          rateLimits: (u.rateLimits || []).map(x => ({ kind: x.kind, percentUsed: Math.round(x.percentUsed) })),
851          requests: state.requests.length,
852        };
853        state.usageChecks.push(check);
854        pushLog("usage-check", check);
855        // The findings snapshot goes in the ledger too: there is no HUD in a
856        // headless session, so this is the only place a `-p` run can read them.
857        pushLog("findings", { score: efficiencyScore().score, headline: headline(), findings: buildFindings() });
858        const five = (u.rateLimits || []).filter(x => x.kind === "five_hour")[0];
859        note($, `⟦costclaw-live⟧ ${headline()}`);
860        note($, `⟦costclaw-live⟧ cost cross-check: ours ${money(state.costUsd)} (${state.requests.length} requests, rate card ${PRICING_SNAPSHOT_DATE}) vs engine ${engineUsd === null ? "n/a" : money(engineUsd)}${check.deltaUsd === null ? "" : ` · delta ${money(check.deltaUsd)}`} · five-hour ${five ? five.percentUsed + "%" : "n/a"}`);
861      }
862    }
863    await flushEvidence($);
864    return r;
865  });
866
867  // ---- 4. HUD above the prompt --------------------------------------------
868  on("ui.render", { component: "AbovePrompt", surface: "terminal" }, async ($, e, next) => {
869    const { Box, Text } = $.ui.resolve(e);
870    const width = Math.max(40, (e.props && e.props.bodyColumns ? e.props.bodyColumns : 100) - 4);
871    const lines = findingLines().slice(0, 5);
872    const s = efficiencyScore();
873    const rows = lines.map((t, i) => (
874      <Text key={"f" + i} color={t.indexOf("SERVED_FROM_CACHE") === 0 ? "green" : "yellow"} wrap="truncate-end">{t.slice(0, width)}</Text>
875    ));
876    const skipped = s.notAssessed.length ? `not assessed: ${s.notAssessed.join(",")}` : "all components assessed";
877    return (
878      <Box flexDirection="column" borderStyle="round" borderColor="cyan" paddingX={1}>
879        <Text bold color="cyan">{headline().slice(0, width)}</Text>
880        {rows}
881        <Text dimColor wrap="truncate-end">{`${skipped} · saved so far ~${state.savedTokens} tok (${money(state.savedUsd)}) · serve-after ${state.serveAfter} · /costclaw for evidence`}</Text>
882      </Box>
883    );
884  });
885
886  // ---- 5. /costclaw --------------------------------------------------------
887  on("command.run", { command: "costclaw" }, async ($, e, next) => {
888    const arg = (e.args || "").trim().toLowerCase();
889    if (arg.indexOf("serve") === 0) {
890      const n = parseInt(arg.split(/\s+/)[1], 10);
891      if (!Number.isSafeInteger(n) || n < 2 || n > 50) return { text: `costclaw-live: serve-after must be an integer 2..50 (currently ${state.serveAfter})` };
892      state.serveAfter = n;
893      await writeServeAfter($, n);
894      $.ui.invalidate("ui.render");
895      return { text: `costclaw-live: serve-after -> ${n} (the ${n}th identical Read of an unchanged file is answered from cache)` };
896    }
897
898    await refreshUsage($);
899    const s = efficiencyScore();
900    const t = totals();
901    const decay = cacheDecay();
902    const loop = toolLoopStatus();
903    const findings = buildFindings();
904
905    const modelRows = Object.keys(state.byModel).map(m => {
906      const x = state.byModel[m];
907      return `  ${m}${isKnownModel(m) ? "" : "  (NO RATE-CARD ENTRY — priced at the fallback, figure is an estimate)"}\n    requests ${x.requests} · in ${x.input} · out ${x.output} · cacheR ${x.cacheRead} · cacheW ${x.cacheWrite} · ${money(x.costUsd)}`;
908    }).join("\n");
909
910    const findingRows = findings.length ? findings.map((f, i) => {
911      const dollars = f.usd === null || f.usd === undefined ? "(no dollar figure is derivable)" : money(f.usd);
912      return `  ${i + 1}. ${f.ruleId} [${f.confidence}/${f.severity}] ${dollars}\n     ${f.title}\n     EVIDENCE: ${f.evidence}`;
913    }).join("\n") : "  none yet";
914
915    const repRows = repeatedTargets().length ? repeatedTargets().map(r =>
916      `  ${r.tool} x${r.fires} ${r.target}\n     requests [${[...(state.targetRequests[r.tool + "|" + r.target] || [])].join(",")}] · confidence ${r.graded.confidence}${r.thrash ? ` · >= REDUNDANT_TARGET_MIN_FIRES(${REDUNDANT_TARGET_MIN_FIRES})` : ""}`
917    ).join("\n") : "  none";
918
919    const servedRows = state.served.length ? state.served.map(x =>
920      `  request #${x.requestId}: read ${x.n} of ${x.target} answered from cache, ~${x.tokens} tokens (${money(x.usd)}) not spent`
921    ).join("\n") : "  none (no file has been read " + state.serveAfter + " times unchanged yet)";
922
923    const u = state.lastUsage;
924    const engineUsd = u && u.cost && typeof u.cost.usd === "number" ? u.cost.usd : null;
925    const deltaLine = engineUsd === null
926      ? "  engine cost unavailable"
927      : `  ours ${money(state.costUsd)} · engine $.session.usage().cost ${money(engineUsd)} · delta ${money(state.costUsd - engineUsd)}` +
928        (Math.abs(state.costUsd - engineUsd) > Math.max(0.01, Math.abs(engineUsd) * 0.15)
929          ? "\n  DISCREPANCY: ours counts only requests seen through turn.step since this plugin loaded and prices them from the "
930            + PRICING_SNAPSHOT_DATE + " list-price snapshot; the engine's figure covers the whole session and its own accounting. Neither is an invoice."
931          : "\n  within 15% — the two agree.");
932
933    const warnRows = Object.keys(state.warnings).length
934      ? Object.keys(state.warnings).map(w => `  ${state.warnings[w]}x ${w}`).join("\n")
935      : "  none";
936
937    const text = [
938      `COSTCLAW LIVE — session ${state.sessionId}${state.repo ? " · repo " + state.repo : ""}`,
939      headline(),
940      "",
941      `SCORE ${s.score}/100 (rubric: 100 minus cache<=30, repeats<=25, loop<=20, bloat<=15, decay<=10)`,
942      `  ${s.parts.map(p => `${p.name} ${p.assessed ? "-" + p.penalty : "not assessed"}`).join(" · ")}`,
943      "",
944      "FINDINGS",
945      findingRows,
946      "",
947      "REPEATED (tool,target)",
948      repRows,
949      "",
950      "INTERVENTIONS (answered by costclaw-live, not by the tool)",
951      servedRows,
952      `  total saved ~${state.savedTokens} tokens (${money(state.savedUsd)}), serve-after = ${state.serveAfter}`,
953      "",
954      "CACHE",
955      `  session hit rate ${pct(sessionHitRate())} over ${state.requests.length} model requests`,
956      decay.assessed
957        ? `  early third ${pct(decay.earlyHit)} -> late third ${pct(decay.lateHit)} (drop ${Math.round(decay.drop * 100)}pp; DECAY_DROP=${DECAY_DROP * 100}pp, LATE_FLOOR=${LATE_FLOOR * 100}%) ${decay.fired ? "FIRED" : "not fired"}`
958        : `  decay NOT ASSESSED (${decay.reason})`,
959      "",
960      "TOOL LOOP",
961      `  status ${loop.status}${loop.heavy ? ` · ${loop.heavy.tool} x${loop.heavy.fires}` : ""}`,
962      loop.runs.length ? loop.runs.map(r => `  [${r.confidence}] ${r.evidence}`).join("\n") : "  no consecutive run of >= " + RUN_THRESHOLD,
963      "",
964      "MODELS (rate card " + PRICING_SNAPSHOT_DATE + ")",
965      modelRows || "  none yet",
966      `  totals: in ${t.input} · out ${t.output} · cacheR ${t.cacheRead} · cacheW ${t.cacheWrite}`,
967      "",
968      "COST CROSS-CHECK",
969      deltaLine,
970      u && u.context ? `  context ${u.context.percent}% (${u.context.tokens} of ${u.context.window})` : "  context unavailable",
971      u && u.rateLimits ? "  " + u.rateLimits.map(x => `${x.kind} ${Math.round(x.percentUsed)}%`).join(" · ") : "  rate limits unavailable",
972      "",
973      "PRICING WARNINGS",
974      warnRows,
975      Object.keys(state.unknownModels).length ? "  UNPRICED MODELS: " + Object.keys(state.unknownModels).join(", ") : "",
976      "",
977      `EVIDENCE: ${EVIDENCE_DIR}live-${state.sessionId}.jsonl (${state.log.length} rows)`,
978    ].join("\n");
979
980    await flushEvidence($);
981    return { text };
982  });
983
984  // ---- session.start -------------------------------------------------------
985  on("session.start", async ($, e, next) => {
986    await bootSession($);
987    state.serveAfter = await readServeAfter($);
988    if (options && typeof options.serveAfter === "number" && Number.isSafeInteger(options.serveAfter) && options.serveAfter >= 2) {
989      state.serveAfter = options.serveAfter;
990    }
991    await $.command.register({
992      name: "costclaw",
993      description: "COSTCLAW LIVE: findings with evidence, cost cross-check, cache decay. /costclaw serve <n> sets the read-cache threshold.",
994      argumentHint: "[serve <n>]",
995      immediate: true,
996    });
997    $.ui.status(headline());
998    $.clock.every(2000, () => tick($));
999    note($, `⟦costclaw-live⟧ armed · serve-after=${state.serveAfter} · rate card ${PRICING_SNAPSHOT_DATE} · evidence -> ${EVIDENCE_DIR}live-${state.sessionId}.jsonl`);
1000    pushLog("start", { sessionId: state.sessionId, repo: state.repo, serveAfter: state.serveAfter, snapshot: PRICING_SNAPSHOT_DATE });
1001    await flushEvidence($);
1002    return next(e);
1003  });
1004};
1005
1006function safeLen(v) {
1007  try { const s = JSON.stringify(v); return typeof s === "string" ? s : ""; } catch (err) { return ""; }
1008}
1009