SLOPSHOPPER

clef-model-router

Picks the Claude model and reasoning effort for each turn with Cloudflare Clef, a fast decision model on Workers AI.

newspinnercommandtoastpromptnetwork
v0.1.5Apache-2.0updated 2026-10-04AbelNavarro/clef-model-router
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · clef-model-router
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /clef ⎿ clef-model-router: Clef router ⎿ clef-model-router: mode auto ⎿ clef-model-router: decider clef-flash · timeout 1500 ms · account not set · token not set ⎿ clef-model-router: today 0 Clef calls · 0 input tokens · ~0 neurons of 9000 budget ⎿ clef-model-router: billing subscription (detected: no API key, provider or gateway is set) · downgrade patience 1 ⎿ clef-model-router: ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts
README

clef-model-router

A Claude Code mod that picks the Claude model and reasoning effort for each turn, using Cloudflare Clef as the decision model.

You keep using Claude Code as usual. When you send a prompt, the mod asks Clef-flash how much capability the request needs. It then runs that turn on the cheapest configuration that should be enough:

Fix the spelling of "recieve" in README.md.                         Clef → Haiku · 94%
Add pagination to /orders following the other list endpoints.       Clef → Sonnet · medium · 81%
Debug why these tests intermittently deadlock only in parallel.     Clef → Opus · high · 72%
Study this subsystem, find why it cascades under partitions, ...    Clef → Opus · xhigh · 88%

Routing is an optimization. If Clef is slow, down, unconfigured or out of free quota, the turn still runs on a deterministic fallback. Claude Code always keeps working.

Status: v0.1, in dogfooding. The mod, the policy and the failure paths are tested against Claude Code 2.1.289's own test host, in live sessions against a mock Workers AI endpoint, and with live Clef-flash calls on a handful of prompts. How well Clef routes real coding work is not measured yet; that is what the local log and /clef feedback are for. The examples above show the output format; see Calibration.

Why

High-capability models and high effort are worth it for difficult debugging, architecture and unfamiliar code. They are wasted on a rename. Nobody switches /model and /effort before every prompt, so this mod does it for you, once per turn, using a decision model built for exactly this kind of typed judgment.

The goal is not "always the cheapest model". It is the least expensive configuration that is sufficiently capable, with every policy decision shown to you.

It is no silver bullet. In a long session most of the cost is the conversation being re-read on every request, whatever model reads it, and moving a warm conversation to another model has a cost of its own. /clear between tasks often saves more than any routing. Routing, prompt caching and cost explains what a router can and cannot save, depending on how you pay.

Requirements

  • Claude Code 2.1.287 or later (mods are on by default from that version). Check with claude --version.
  • A Cloudflare account. The Workers AI free allocation (10,000 neurons/day) is enough for normal personal use. See Cost.

Install

claude plugin marketplace add AbelNavarro/clef-model-router
claude plugin install clef-model-router@clef-model-router

Then give it your Cloudflare credentials. In Claude Code:

/plugin configure clef-model-router@clef-model-router

If the mod is already loaded, run /reload-plugins; otherwise start a new session. You should see ↳ Clef: awaiting prompt at the end of the hint line under the prompt.

To uninstall: claude plugin uninstall clef-model-router@clef-model-router. To stop it without uninstalling, use /clef off (this session) or set Routing enabled to off in /config.

git clone https://github.com/AbelNavarro/clef-model-router
CLOUDFLARE_ACCOUNT_ID=... CLOUDFLARE_API_TOKEN=... claude --plugin-dir ./clef-model-router

Credentials in environment variables are inherited by every command Claude runs, so prefer /plugin configure for regular use.

Get your Cloudflare account ID and API token

These steps follow Cloudflare's Workers AI REST API guide.

1. Create a Cloudflare account (skip if you have one)

  1. Sign up at <https://dash.cloudflare.com/sign-up>.
  2. Verify your email.

No credit card is needed. The free Workers plan includes 10,000 Workers AI neurons per day, roughly 2,000 routed prompts.

2. Open the Workers AI page

  1. Log in at <https://dash.cloudflare.com>.
  2. In the sidebar, go to AI → Workers AI, or use the direct link: <https://dash.cloudflare.com/?to=/:account/ai/workers-ai>.
  3. If you have several accounts, pick the one to use.

3. Get the API token

  1. On the Workers AI page, select Use REST API.
  2. Select Create a Workers AI API Token.
  3. Review the prefilled settings. The template grants Workers AI access only.
  4. Select Create API Token.
  5. Select Copy API Token. Cloudflare shows the token only once; if you lose it, create a new one.

4. Get the Account ID

On the same Use REST API panel, under Get Account ID, copy the Account ID: a 32-character hex string such as 0123456789abcdef0123456789abcdef. It is also on the account home page under Account details → Account ID, and in the dashboard URL right after dash.cloudflare.com/.

5. Give them to the mod

  1. In Claude Code, run:
   /plugin configure clef-model-router@clef-model-router
  1. Paste the Account ID into Cloudflare account ID, and the token into Cloudflare API token. The token field is masked, and the value is kept in your system's secure credential store.
  2. Run /reload-plugins. The line under the prompt should read Clef: awaiting prompt.

6. Check that it works

  • Inside Claude Code: /clef test fix the typo in README shows Clef's probabilities and latency.
  • From a clone of this repository: npm run smoke makes one real call. Export CLOUDFLARE_ACCOUNT_ID and CLOUDFLARE_API_TOKEN first.

If you prefer a custom token

  1. Go to My Profile → API Tokens → Create Token → Create Custom Token.
  2. Add Account → Workers AI → Read and Account → Workers AI → Edit. Cloudflare requires both for tokens made outside the template.
  3. Limit Account Resources to your account.

Good practice

  • Don't share the token. Never paste it into a chat, a file in a repository, or a command line, where process lists and shell history can expose it.
  • If it leaks, delete it under My Profile → API Tokens and create a new one.
  • To keep routing free, stay on the free Workers plan. On Workers Paid, usage beyond the daily allocation is billed. The mod's local budget (9,000 neurons/day) guards against that by default.

What you see

  • Under the prompt, dimmed at the end of the hint line, the current route: ↳ Clef → Sonnet · medium · 87% (terminal; on other surfaces set announce to answer). The percentage is the probability Clef gave the level it picked. When policy changed Clef's pick, the reason follows in brackets, for example (Sonnet deferred) or (unsure). When Clef could not answer: Clef ✕ timeout → Sonnet · medium.
  • /clef shows the full picture: mode, the last decision with Clef's full probability distribution, latency, every policy adjustment and why, today's Clef usage, and the cache state.
Last turn
  route     Opus · high  (claude-opus-5-5, profile hard)
  source    clef
  clef      clef-flash: 72% on hard · clef confidence 52% · score 2.88/4 · follow-up 4% · 410 ms · 580 tokens
    trivial   ····················   1%
    simple    █···················   3%
    standard  ███·················  16%
    hard      ██████████████······  72%  ← Clef
    deep      ██··················   8%
Command
/clefStatus and the last decision
/clef historyThis session's turns: latency, route, confidence, source, and any policy change
/clef stats [days]Totals from the local log: by model, effort, profile, source; latency; fallbacks; overrides; turns held on a warm model and downgrades taken; your feedback
/clef test <prompt>Ask Clef about a prompt without sending it to Claude
/clef profilesWhat each difficulty level runs on here
/clef pin <target>Use one target for the rest of the session (/clef pin opus:high, /clef pin hard, /clef pin :low)
/clef autoUnpin, resume after a /model change, and hand effort back to Clef after /effort
/clef off · /clef onStop or resume routing for this session
`/clef feedback under\ok\over [note]`Rate the last route, for later analysis of whether Clef was right

One turn only: start a prompt with +target. The prefix is removed before Claude sees the prompt.

+opus:max why does this deadlock only on ARM?
+haiku list the files in src/
+off explain this stack trace          (this turn runs exactly as Claude Code would)

How it decides

prompt ──► turn.start ──► Clef-flash: difficulty 0-4 (+ "is this a follow-up?")   one call, ~0.3–0.5 s end to end
                  │
                  ▼
           policy (deterministic): overrides → continuation → confidence → follow-up floor
                                   → availability → context window → downgrade timing → effort clamp
                  │
                  ▼
           turn.step ×N: every main-loop request of the turn sent with that model + effort
  • Clef judges; code decides. Clef answers one semantic question: how demanding is this request, on a five-level rubric. Overrides, failures, unavailable models, context windows, cache economics and effort limits are all deterministic rules. See Architecture.
  • Once per turn. A turn may make many model requests (one after each tool result). The decision is made once and reused for all of them, so Clef's latency is paid once and the model never changes mid-turn. A go-ahead (yes, do it, continue) and background-task notifications reuse the last route without calling Clef.
  • Profiles, not free combinations. Clef picks one of five difficulty levels, and each maps to a model and effort you can change:
LevelDefaultFor
trivialhaikutypo, rename, format, quick lookup
simplesonnet:lowsmall, well-specified change in one place
standardsonnet:mediumordinary feature or bug work
hardopus:hightricky debugging, refactors, unfamiliar code
deepopus:xhighopen-ended investigation and design

Haiku 4.5 takes no effort setting, so the trivial level sends none.

  • Cache-aware, by billing. Each model has its own prompt cache, so moving a warm conversation to a cheaper model writes all of it again. A downgrade is held on the warm model until staying has cost what the switch costs, then taken: a one-off easy question stays put, and a stretch of routine work moves. With an API key, "cost" is dollars. Opus 5.5 and Sonnet 5.5 cost the same to re-read, so the mod mostly stays and lowers the effort. On a subscription it is plan usage, and the mod moves after a few turns to save the stronger model's allowance. Billing is detected, or set with billing. A held turn still gets the effort Clef asked for, which keeps the cache on Opus 5.5, Sonnet 5.5 and Fable 5.1. Upgrades are never held back. See ADR 0001.
  • Low confidence (Clef gives its pick less than 55%): by default the mod takes the more capable of Clef's two likeliest levels. The threshold and policy are configurable, and every change is logged.
  • Latency. Clef-flash's model time is about 40 ms (Cloudflare's figure), but a routed prompt waits for the whole round trip: 340–530 ms in the first live tests. Go-aheads, overrides and pinned sessions skip the call.
  • You stay in control. +target beats /clef pin, which beats Clef. A /model change mid-session pauses routing until /clef auto. An /effort change sets the effort while Clef keeps choosing the model. Subagents keep their own models.

Cost

Routing itself runs on your Cloudflare account. Clef-flash costs $0.09 per million input tokens and has no charged output. One routing call used about 580 input tokens for typical prompts in live tests (the rubric plus your prompt), and up to about 2,000 for a long one, since prompts are cut to 6,000 characters. That works out to roughly 5 neurons per call, or about 2,000 routed prompts a day inside Workers AI's free allocation of 10,000 neurons per day (resets 00:00 UTC). These are estimates derived from Cloudflare's published prices; Clef is not yet in Cloudflare's per-model neuron table.

What happens at the limit depends on your Cloudflare plan (pricing):

  • Workers Free: requests beyond the allocation fail (error 3036). The mod notices, stops calling Clef until 00:00 UTC, and uses the fallback route. You are not billed.
  • Workers Paid: usage beyond the allocation is billed at $0.011 per 1,000 neurons. To keep routing free, the mod has a local daily budget of 9,000 neurons (estimated), after which it stops calling Clef for the day. It counts only this mod's calls, not other Workers AI use on the account.

/clef shows today's calls, tokens and estimated neurons.

The mod also changes what you spend on Claude itself. That is the point, and /clef stats shows where your turns went.

Privacy

For each prompt you type, the mod sends that prompt's text to Cloudflare Workers AI (cut to its first 4,500 and last 1,500 characters if longer than 6,000), along with the fixed rubric questions. Nothing else is sent: no conversation history, file contents, tool output, repository name or metadata. Go-aheads, task notifications, +model or +profile prompts, model pins and /clef off send nothing.

Everything else stays on your machine. The local log keeps a hash and the length of each prompt, not its text, unless you turn on Log prompt text. Cloudflare states that Workers AI does not use your inputs or outputs to train models, and the Clef announcement says Cloudflare does not read, store or train on Clef requests. Details and sources are in docs/PRIVACY.md.

Configuration

Everyday options are plugin options. Set them with /plugin configure or in /config: the Cloudflare account ID and token, the decision model (clef-flash or clef), the profiles, your billing (detected by default), how to show the route, routing on or off, and whether to log prompt text. Tuning knobs go in an optional ~/.claude/clef-model-router.json. Thresholds, policies, timeout, budget, downgrade patience and the rubric are all set there. The full reference, including precedence rules, is in docs/CONFIGURATION.md.

Documentation

Prior art

The earlier routers this project studied, and what it took from each, are listed in Architecture → Prior art: jev-model-router, jev-claude-router, pi-auto-router, clef-router, claude-code-model-router and Morph's router.

License

Apache-2.0. Not affiliated with Anthropic or Cloudflare.

Source 15 files
hooks/register.ts 734 lines
1// clef-model-router: picks the Claude model and effort for each turn with
2// Cloudflare Clef.
3//
4//   prompt.submit  a `+target ` prefix is taken off the prompt and kept for the turn
5//   turn.start     one Clef call (or none), then the policy decides the turn's route
6//   turn.step      every main-loop request of the turn is sent with that route
7//   turn.complete  the outcome goes to the local log
8//   /clef          status, history, stats, test, pin, auto, off, on, feedback
9//
10// The decision is made once per turn, at turn.start, where the person's text
11// is: a turn's later requests (after each tool result) reuse it, so a turn
12// pays Clef's latency once and never changes model halfway. Subagents keep
13// their own model. Every failure path leaves the request as Claude Code made
14// it, or on a deterministic fallback.
15
16import type { EngineInterface, PluginOptions, Register } from "claude-code"
17
18import { clefProvider } from "./lib/clef.ts"
19import { parseAdvancedFile, parseConfig, type Config } from "./lib/config.ts"
20import { cacheTtlMs, detectBilling, modelEnvFrom, type BillingFacts, type EnvValues, type RateLimitWindow } from "./lib/env.ts"
21import {
22  HELP,
23  answerLine,
24  explain,
25  historyReport,
26  profilesReport,
27  statsReport,
28  statusLine,
29  statusReport,
30  targetText,
31  type HistoryRow,
32} from "./lib/format.ts"
33import { blockedReason, estimatedNeurons, normaliseGuard, recordFailure, recordSuccess, type GuardState } from "./lib/guard.ts"
34import { aggregate, logFileName, parseLines, promptHash, turnRecord, type AnsweredUsage, type FeedbackRecord } from "./lib/log.ts"
35import { clampEffort, effortsFor, isEffort, sameModel, type ModelEnv } from "./lib/models.ts"
36import { parseCommand, parsePrefix, trackNativeEffort, turnKind, type ClefCommand } from "./lib/overrides.ts"
37import { decide, describe, needsClef, type CacheState, type HoldState, type RouterMode } from "./lib/policy.ts"
38import { priceOf, stayCost, isOneHour, type Billing } from "./lib/pricing.ts"
39import { DEFAULT_RUBRIC, parseRubric, type Rubric } from "./lib/rubric.ts"
40import type { Decision, ProviderResult, Route, Target } from "./lib/types.ts"
41
42const PLUGIN = "clef-model-router"
43const SESSION_REF = { plugin: "clef-model-router", key: "session" } as const
44const GUARD_KEY = "guard"
45const HISTORY_LIMIT = 50
46const RUN_LIMIT = 32
47/** A routed model that fails this many requests in a row is not used again this session. */
48const FAILURES_BEFORE_UNAVAILABLE = 2
49
50/** What survives a hot reload (in `$.state`) for the rest of the session. */
51type Persisted = {
52  mode: RouterMode
53  pin?: Target
54  pendingOverride?: Target | "off"
55  last?: Route
56  cache?: CacheState
57  /** A downgrade held for the cache, and what staying has cost so far. */
58  hold?: HoldState
59  unavailable: string[]
60  failures: Record<string, number>
61  /** The session model Claude Code reported at the last turn. */
62  baselineModel?: string
63  /** Effort as Claude Code itself would send it, and any /effort the person set. */
64  effortBaseline?: string | number
65  nativeEffort?: string | number
66  /** A model Claude Code fell back to during the last turn, so it is not mistaken for a /model change. */
67  engineFallback?: string
68  history: HistoryRow[]
69  warned: string[]
70}
71
72/** One turn in flight. Module memory only: a reload mid-turn leaves it unrouted. */
73type Run = {
74  decision: Decision
75  prompt: string
76  hash?: string
77  engineModel?: string
78  passthrough: boolean
79  failedRewrite: boolean
80  steps: number
81  answered?: AnsweredUsage
82  /** Cache read and write of the turn's first request: what a switch (or a return) cost. */
83  firstStep?: { cacheReadTokens: number; cacheWriteTokens: number }
84  rateLimits?: RateLimitWindow[]
85}
86
87// Module state. `register` runs again on every reload, which resets these;
88// `load` then restores the session's part from `$.state`.
89let options: PluginOptions = {}
90let state: Persisted = freshState()
91let loaded = false
92let config: Config = parseConfig({}).config
93let problems: string[] = []
94let modelEnv: ModelEnv = modelEnvFrom({})
95let env: EnvValues = {}
96let settingsTtl: unknown
97let apiKeyHelper = false
98let detected: BillingFacts = detectBilling({})
99let billing: Billing = detected.billing
100let ttlMs = 5 * 60_000
101let rubric: Rubric = DEFAULT_RUBRIC
102let logDir: string | undefined
103let apiBase: string | undefined
104let advancedPath: string | undefined
105let logFile: string | undefined
106let logLines: string[] = []
107const runs = new Map<string, Run>()
108
109function freshState(): Persisted {
110  return { mode: "auto", unavailable: [], failures: {}, history: [], warned: [] }
111}
112
113async function save($: EngineInterface): Promise<void> {
114  try {
115    await $.state.set(SESSION_REF, JSON.stringify(state))
116  } catch {
117    // Losing the snapshot only matters on a hot reload; routing goes on.
118  }
119}
120
121async function readEnv($: EngineInterface): Promise<EnvValues> {
122  return {
123    ANTHROPIC_DEFAULT_HAIKU_MODEL: await $.env.get("ANTHROPIC_DEFAULT_HAIKU_MODEL"),
124    ANTHROPIC_DEFAULT_SONNET_MODEL: await $.env.get("ANTHROPIC_DEFAULT_SONNET_MODEL"),
125    ANTHROPIC_DEFAULT_OPUS_MODEL: await $.env.get("ANTHROPIC_DEFAULT_OPUS_MODEL"),
126    ANTHROPIC_DEFAULT_FABLE_MODEL: await $.env.get("ANTHROPIC_DEFAULT_FABLE_MODEL"),
127    CLAUDE_CODE_USE_BEDROCK: await $.env.get("CLAUDE_CODE_USE_BEDROCK"),
128    CLAUDE_CODE_USE_VERTEX: await $.env.get("CLAUDE_CODE_USE_VERTEX"),
129    CLAUDE_CODE_USE_FOUNDRY: await $.env.get("CLAUDE_CODE_USE_FOUNDRY"),
130    CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS: await $.env.get("CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS"),
131    // Only whether it is set: the key itself is never kept.
132    ANTHROPIC_API_KEY: (await $.env.get("ANTHROPIC_API_KEY")) ? "set" : undefined,
133    CLAUDE_CODE_PROMPT_CACHE_TTL: await $.env.get("CLAUDE_CODE_PROMPT_CACHE_TTL"),
134    FORCE_PROMPT_CACHING_5M: await $.env.get("FORCE_PROMPT_CACHING_5M"),
135    ENABLE_PROMPT_CACHING_1H: await $.env.get("ENABLE_PROMPT_CACHING_1H"),
136    // Only whether they are set: they say a gateway bills per token.
137    ANTHROPIC_AUTH_TOKEN: (await $.env.get("ANTHROPIC_AUTH_TOKEN")) ? "set" : undefined,
138    ANTHROPIC_BASE_URL: (await $.env.get("ANTHROPIC_BASE_URL")) ? "set" : undefined,
139  }
140}
141
142/** Reads configuration and restores the session snapshot, once per load. */
143async function load($: EngineInterface): Promise<void> {
144  if (loaded) return
145  loaded = true
146  try {
147    const home = (await $.env.get("HOME")) ?? (await $.env.get("USERPROFILE")) ?? "."
148    const configDir = (await $.env.get("CLAUDE_CONFIG_DIR")) ?? `${home}/.claude`
149    advancedPath = (await $.env.get("CLEF_ROUTER_CONFIG")) ?? `${configDir}/${PLUGIN}.json`
150    const advancedText = await $.fs.read(advancedPath).catch(() => undefined)
151    const advanced = parseAdvancedFile(typeof advancedText === "string" ? advancedText : undefined)
152    const parsed = parseConfig(
153      { ...advanced.values, ...options },
154      { accountId: await $.env.get("CLOUDFLARE_ACCOUNT_ID"), apiToken: await $.env.get("CLOUDFLARE_API_TOKEN") },
155    )
156    config = parsed.config
157    problems = [...advanced.problems, ...parsed.problems]
158    env = await readEnv($)
159    modelEnv = modelEnvFrom(env)
160    const settings = (await $.settings.read().catch(() => ({}))) as Record<string, unknown>
161    settingsTtl = settings.promptCacheTtl
162    apiKeyHelper = typeof settings.apiKeyHelper === "string" && settings.apiKeyHelper !== ""
163    refreshBilling(undefined)
164    if (config.rubricFile) {
165      const text = await $.fs.read(config.rubricFile).catch(() => undefined)
166      const result = typeof text === "string" ? parseRubric(text) : { problems: [`cannot read rubric_file ${config.rubricFile}`] }
167      if ("rubric" in result) rubric = result.rubric
168      else problems.push(...result.problems.map((p) => `${p}; using the built-in rubric`))
169    }
170    const base = await $.env.get("CLEF_ROUTER_API_BASE")
171    if (base && /^https?:\/\//.test(base)) apiBase = base
172    logDir = config.logDir ?? `${configDir}/plugins/data/${PLUGIN}`
173    const snapshot = await $.state.get(SESSION_REF)
174    if (typeof snapshot.value === "string") state = { ...freshState(), ...(JSON.parse(snapshot.value) as Partial<Persisted>) }
175  } catch (error) {
176    problems.push(`setup: ${error instanceof Error ? error.message : String(error)}`)
177  }
178}
179
180/**
181 * How the person pays and the cache TTL that follows from it, refreshed from
182 * the rate-limit windows each turn: they appear after the session's first
183 * response on a subscription, and a window past 100% means usage credits.
184 */
185function refreshBilling(rateLimits: readonly RateLimitWindow[] | undefined): void {
186  detected = detectBilling(env, { apiKeyHelper, ...(rateLimits ? { rateLimits } : {}) })
187  billing = config.billing === "auto" ? detected.billing : config.billing
188  ttlMs = cacheTtlMs(config.cacheTtlMinutes, env, settingsTtl, detected)
189}
190
191async function readGuard($: EngineInterface, now: number): Promise<GuardState> {
192  return normaliseGuard(await $.store.get(GUARD_KEY).catch(() => undefined), now)
193}
194
195async function askClef($: EngineInterface, prompt: string): Promise<ProviderResult> {
196  const now = await $.clock.now()
197  const guard = await readGuard($, now)
198  const blocked = blockedReason(guard, { now, model: config.decisionModel, dailyNeuronBudget: config.dailyNeuronBudget })
199  if (blocked) return { ok: false, failure: blocked }
200  const provider = clefProvider({
201    accountId: config.accountId,
202    apiToken: config.apiToken,
203    model: config.decisionModel,
204    rubric,
205    timeoutMs: config.timeoutMs,
206    maxPromptChars: config.maxPromptChars,
207    ...(apiBase ? { apiBase } : {}),
208    fetch: async (url, init) => {
209      const r = await $.http.fetch(url, init)
210      return { status: r.status, ok: r.ok, text: r.text }
211    },
212    sleep: (ms, signal) => $.clock.sleep(ms, { signal }),
213    now: () => $.clock.now(),
214  })
215  const result = await provider.decide(prompt)
216  const after = result.ok ? recordSuccess(guard, result.recommendation.inputTokens) : recordFailure(guard, result.failure, await $.clock.now())
217  if (after !== guard) await $.store.set(GUARD_KEY, after).catch(() => {})
218  return result
219}
220
221/** One-time notices, so a misconfiguration is said once, not every turn. */
222function warnOnce($: EngineInterface, key: string, text: string): void {
223  if (state.warned.includes(key)) return
224  state.warned.push(key)
225  $.ui.toast(text, { timeoutMs: 8000 })
226}
227
228/** The route line, drawn dim at the end of the hint line under the prompt. */
229let hint: string | undefined
230
231/**
232 * Shows the route without Claude Code's status-line marker (a ⚠ that reads as
233 * a warning): the text joins the prompt's hint line as its tail, prefixed ↳.
234 * The terminal draws that tail; elsewhere `announce: "answer"` shows the route.
235 */
236function showStatus($: EngineInterface, text: string | undefined): void {
237  const next = text === undefined ? undefined : `↳ ${text}`
238  if (next === hint) return
239  hint = next
240  $.ui.invalidate("ui.render")
241}
242
243function announce($: EngineInterface, d: Decision): void {
244  if (config.announce === "status" || config.announce === "both") showStatus($, statusLine(d))
245}
246
247async function appendLog($: EngineInterface, line: string, ts: string): Promise<void> {
248  if (!config.logEnabled || !logDir) return
249  try {
250    const file = `${logDir}/${logFileName(ts, await $.session.id())}`
251    if (file !== logFile) {
252      logFile = file
253      const existing = await $.fs.read(file).catch(() => "")
254      logLines = typeof existing === "string" && existing !== "" ? existing.trimEnd().split("\n") : []
255    }
256    logLines.push(line)
257    await $.fs.write(file, logLines.join("\n") + "\n")
258  } catch {
259    // A log that cannot be written must not cost the turn anything.
260  }
261}
262
263/** Notices a /model change made between turns: the person taking over the model. */
264async function noticeNativeModel($: EngineInterface): Promise<void> {
265  const sessionModel = await $.session.model().catch(() => undefined)
266  if (sessionModel && state.baselineModel && !sameModel(sessionModel, state.baselineModel)) {
267    // A fallback Claude Code made itself (a safety classifier moving the
268    // session) is not one.
269    const engineMoved = state.engineFallback !== undefined && sameModel(sessionModel, state.engineFallback)
270    if (!engineMoved && state.mode === "auto" && config.pauseOnNativeChange) {
271      state.mode = "paused-native"
272      $.ui.toast(`Clef paused: you switched to ${sessionModel}. /clef auto resumes routing.`, { timeoutMs: 8000 })
273    }
274  }
275  delete state.engineFallback
276  if (sessionModel) state.baselineModel = sessionModel
277}
278
279async function routeTurn($: EngineInterface, turnId: string, text: string): Promise<void> {
280  const kind = turnKind(text)
281  const override = state.pendingOverride
282  delete state.pendingOverride
283  await noticeNativeModel($)
284
285  const base = { turnId, kind, config, modelEnv, session: state, ...(override ? { override } : {}) }
286  const result = needsClef(base) ? await askClef($, text) : undefined
287  const usage = await $.session.usage().catch(() => undefined)
288  const rateLimits = usage?.rateLimits?.map((w) => ({ kind: w.kind, percentUsed: w.percentUsed }))
289  refreshBilling(rateLimits)
290  const now = await $.clock.now()
291  const decision = decide({
292    ...base,
293    ...(result ? { result } : {}),
294    ...(usage?.context?.tokens ? { contextTokens: usage.context.tokens } : {}),
295    now,
296    cacheTtlMs: ttlMs,
297    billing,
298  })
299
300  const kindOfFailure = decision.failure?.kind
301  if (kindOfFailure === "not-configured")
302    warnOnce($, "not-configured", `Clef router: set your Cloudflare account ID and API token with /plugin configure ${PLUGIN}. Using the ${config.fallbackLevel} profile meanwhile.`)
303  else if (kindOfFailure === "auth")
304    warnOnce($, "auth", `Clef router: Cloudflare rejected the API token (${decision.failure?.message}). Falling back until it is fixed.`)
305  else if (kindOfFailure === "quota")
306    warnOnce($, `quota-${new Date(now).toISOString().slice(0, 10)}`, "Clef router: Workers AI's free daily allocation is used up; falling back until 00:00 UTC.")
307
308  if (decision.final) state.last = decision.final
309  // A held turn carries the stretch on (its cost is added when it completes);
310  // any other turn ends it.
311  const deferral = decision.deferral
312  if (deferral?.held && decision.final) {
313    state.hold = { model: decision.final.model, wanted: deferral.wanted, spent: deferral.spent, turns: deferral.turns }
314  } else delete state.hold
315  const run: Run = { decision, prompt: text, passthrough: !decision.final, failedRewrite: false, steps: 0 }
316  if (rateLimits && rateLimits.length > 0) run.rateLimits = rateLimits
317  const hash = await promptHash(text).catch(() => undefined)
318  if (hash) run.hash = hash
319  runs.set(turnId, run)
320  while (runs.size > RUN_LIMIT) runs.delete(runs.keys().next().value!)
321  state.history.push({ decision, prompt: text.slice(0, 200) })
322  while (state.history.length > HISTORY_LIMIT) state.history.shift()
323  announce($, decision)
324  await save($)
325}
326
327/** Watches Claude Code's own model and effort at a main-loop step, before any rewrite. */
328function noticeStep($: EngineInterface, run: Run, model: string, effort: string | number | undefined, index: number): void {
329  if (index === 0) {
330    run.engineModel = model
331    const tracked = trackNativeEffort({ baseline: state.effortBaseline, native: state.nativeEffort }, effort)
332    state.effortBaseline = tracked.baseline
333    if (tracked.native === undefined) delete state.nativeEffort
334    else state.nativeEffort = tracked.native
335    if (!config.pauseOnNativeChange) return
336    if (tracked.change === "set" && state.mode === "auto")
337      $.ui.toast(`Clef: using your effort ${String(tracked.native)}; Clef still picks the model. /clef auto hands effort back.`, { timeoutMs: 8000 })
338    // The person's /effort beats Clef's, not an explicit +target or pin.
339    const d = run.decision
340    if (state.nativeEffort !== undefined && d.final && (d.source === "clef" || d.source === "continuation" || d.source === "fallback")) {
341      const wanted = typeof state.nativeEffort === "string" && isEffort(state.nativeEffort) ? state.nativeEffort : undefined
342      const effortNow = wanted ? clampEffort(d.final.model, wanted) : undefined
343      if (effortNow !== d.final.effort) {
344        const final = { ...d.final }
345        if (effortNow) final.effort = effortNow
346        else delete final.effort
347        run.decision = {
348          ...d,
349          final,
350          adjustments: [...d.adjustments, { rule: "pinned-effort", from: d.final.effort ?? "default", to: effortNow ?? "default", reason: "your /effort" }],
351        }
352        state.last = final
353        announce($, run.decision)
354      }
355    }
356  } else if (run.engineModel && !sameModel(model, run.engineModel)) {
357    // Claude Code moved the turn to a fallback model (an error, or a safety
358    // classifier). That is never overridden.
359    run.passthrough = true
360    state.engineFallback = model
361  }
362}
363
364/** Records what a step's response says: failures of a routed model, usage, cache. */
365async function afterStep(
366  $: EngineInterface,
367  run: Run,
368  sent: { model: string; effort?: unknown },
369  rewritten: boolean,
370  result: { stopReason: string | null; usage: { model: string; input_tokens: number; output_tokens: number; cache_read_input_tokens: number; cache_creation_input_tokens: number } | null },
371  aborted: boolean,
372): Promise<void> {
373  const now = await $.clock.now()
374  if (rewritten) {
375    if (result.stopReason === null && !aborted) {
376      // The routed request got no response: stop routing this turn, and stop
377      // using the model after repeated failures.
378      run.failedRewrite = true
379      const n = (state.failures[sent.model] ?? 0) + 1
380      state.failures[sent.model] = n
381      if (n >= FAILURES_BEFORE_UNAVAILABLE && !state.unavailable.includes(sent.model)) {
382        state.unavailable.push(sent.model)
383        $.ui.toast(`Clef router: ${sent.model} failed ${n} times; not routing to it again this session.`, { timeoutMs: 8000 })
384      }
385    } else if (result.stopReason !== null) {
386      state.failures[sent.model] = 0
387    }
388  }
389  const u = result.usage
390  if (u) {
391    const a = run.answered ?? { model: u.model, inputTokens: 0, outputTokens: 0, cacheReadTokens: 0, cacheWriteTokens: 0 }
392    run.answered = {
393      model: u.model,
394      inputTokens: a.inputTokens + u.input_tokens,
395      outputTokens: a.outputTokens + u.output_tokens,
396      cacheReadTokens: a.cacheReadTokens + u.cache_read_input_tokens,
397      cacheWriteTokens: a.cacheWriteTokens + u.cache_creation_input_tokens,
398    }
399    if (!run.firstStep) run.firstStep = { cacheReadTokens: u.cache_read_input_tokens, cacheWriteTokens: u.cache_creation_input_tokens }
400    const cache: CacheState = {
401      model: sent.model,
402      at: now,
403      promptTokens: u.input_tokens + u.cache_read_input_tokens + u.cache_creation_input_tokens,
404      caching: u.cache_read_input_tokens + u.cache_creation_input_tokens > 0,
405    }
406    if (typeof sent.effort === "string" && isEffort(sent.effort)) cache.effort = sent.effort
407    state.cache = cache
408  }
409  await save($)
410}
411
412/** Adds what a held turn cost to its stretch, from the usage the API reported. */
413function accrueHold(run: Run): void {
414  const d = run.decision
415  const hold = state.hold
416  if (!d.deferral?.held || !d.final || !run.answered || !hold || !sameModel(hold.model, run.answered.model)) return
417  const held = priceOf(d.final.model)
418  const wanted = priceOf(d.deferral.wanted.model)
419  if (!held || !wanted) return
420  hold.spent = d.deferral.spent + stayCost(d.deferral.billing, held, wanted, run.answered, isOneHour(ttlMs))
421}
422
423async function completeTurn($: EngineInterface, run: Run, durationMs: number, reason: string): Promise<void> {
424  accrueHold(run)
425  const ts = new Date(await $.clock.now()).toISOString()
426  const record = turnRecord({
427    decision: run.decision,
428    session: await $.session.id(),
429    ts,
430    promptText: run.prompt,
431    ...(run.hash ? { hash: run.hash } : {}),
432    logPrompts: config.logPrompts,
433    ...(run.answered ? { answered: run.answered } : {}),
434    ...(run.firstStep ? { firstStep: run.firstStep } : {}),
435    ...(run.rateLimits ? { rateLimits: run.rateLimits } : {}),
436    steps: run.steps,
437    durationMs,
438    endReason: reason,
439  })
440  await appendLog($, JSON.stringify(record), ts)
441  const row = state.history.find((h) => h.decision.turnId === run.decision.turnId)
442  if (row) {
443    row.decision = run.decision
444    if (run.answered) row.answeredModel = run.answered.model
445  }
446  await save($)
447}
448
449async function registerCommand($: EngineInterface): Promise<void> {
450  try {
451    await $.command.register({
452      name: "clef",
453      description: "Clef router: status, history, stats, test, pin, auto, off, on, feedback",
454      argumentHint: "[status|history|stats|profiles|test|pin|auto|off|on|feedback|help]",
455      immediate: true,
456    })
457  } catch {
458    // Without the command the router still routes.
459  }
460  if (!config.enabled) showStatus($, undefined)
461  else if (!config.accountId || !config.apiToken) showStatus($, "Clef: not configured")
462  else {
463    // Always replace what an earlier load showed (a stale "not configured"
464    // survives a reload otherwise): the last route if there is one.
465    const last = state.history.at(-1)?.decision
466    showStatus($, (last && statusLine(last)) ?? "Clef: awaiting prompt")
467  }
468}
469
470async function onClear($: EngineInterface): Promise<void> {
471  // /clear starts a new conversation: nothing is cached and nothing continues.
472  delete state.last
473  delete state.cache
474  delete state.hold
475  delete state.pendingOverride
476  state.history = []
477  runs.clear()
478  logFile = undefined
479  if (config.enabled && config.accountId && config.apiToken) showStatus($, "Clef: awaiting prompt")
480  await save($)
481}
482
483async function setPendingOverride($: EngineInterface, override: Target | "off"): Promise<void> {
484  await load($)
485  state.pendingOverride = override
486  await save($)
487}
488
489async function statusText($: EngineInterface, now: number): Promise<string> {
490  const guard = await readGuard($, now)
491  const blocked = blockedReason(guard, { now, model: config.decisionModel, dailyNeuronBudget: config.dailyNeuronBudget })
492  const last = state.history.at(-1)?.decision
493  return statusReport({
494    config,
495    configProblems: problems,
496    modelEnv,
497    mode: state.mode,
498    ...(state.pin ? { pin: state.pin } : {}),
499    ...(last ? { last } : {}),
500    guard: {
501      calls: guard.calls,
502      inputTokens: guard.inputTokens,
503      neurons: estimatedNeurons(guard, config.decisionModel),
504      ...(blocked ? { blocked: blocked.message } : {}),
505    },
506    ...(state.cache
507      ? {
508          cache: {
509            model: state.cache.model,
510            promptTokens: state.cache.promptTokens,
511            ageSeconds: Math.round((now - state.cache.at) / 1000),
512            ttlSeconds: Math.round(ttlMs / 1000),
513          },
514        }
515      : {}),
516    billing: { billing, detected, configured: config.billing, patience: config.downgradePatience },
517    ...(state.hold ? { hold: state.hold } : {}),
518    unavailable: state.unavailable,
519    logDir,
520    ...(advancedPath ? { advancedPath } : {}),
521  })
522}
523
524async function statsText($: EngineInterface, now: number, days: number): Promise<string> {
525  if (!logDir) return "No log directory."
526  const since = new Date(now - (days - 1) * 86_400_000).toISOString().slice(0, 10)
527  const entries = await $.fs.list(logDir).catch(() => [])
528  const files = entries.filter((f) => /^routing-\d{4}-\d{2}-\d{2}-/.test(f.name) && f.name.slice(8, 18) >= since)
529  const records = []
530  for (const f of files) {
531    const text = await $.fs.read(`${logDir}/${f.name}`).catch(() => "")
532    if (typeof text === "string") records.push(...parseLines(text))
533  }
534  return statsReport(aggregate(records), days, files.length)
535}
536
537async function testText($: EngineInterface, now: number, prompt: string): Promise<string> {
538  const result = await askClef($, prompt)
539  const d = decide({
540    turnId: "test",
541    kind: turnKind(prompt),
542    config,
543    modelEnv,
544    session: { ...state, mode: "auto" },
545    result,
546    now,
547    cacheTtlMs: ttlMs,
548    billing,
549  })
550  return [`Clef on: ${prompt.slice(0, 80)}`, ...explain(d), "", "(Not sent to Claude. Counts toward today's Clef usage.)"].join("\n")
551}
552
553async function feedbackText($: EngineInterface, now: number, cmd: Extract<ClefCommand, { kind: "feedback" }>): Promise<string> {
554  const last = state.history.at(-1)
555  const record: FeedbackRecord = {
556    v: 1,
557    type: "feedback",
558    ts: new Date(now).toISOString(),
559    session: await $.session.id(),
560    verdict: cmd.verdict,
561    ...(last ? { turn: last.decision.turnId } : {}),
562    ...(cmd.note ? { note: cmd.note } : {}),
563  }
564  await appendLog($, JSON.stringify(record), record.ts)
565  const route = last?.decision.final ? describe(last.decision.final) : "the last turn"
566  const verdict = cmd.verdict === "ok" ? "about right" : cmd.verdict === "under" ? "not capable enough" : "more than needed"
567  return `Noted: ${route} was ${verdict}.`
568}
569
570async function runCommand($: EngineInterface, args: string): Promise<string> {
571  await load($)
572  const cmd = parseCommand(args)
573  const now = await $.clock.now()
574  switch (cmd.kind) {
575    case "help":
576      return HELP
577    case "error":
578      return cmd.message
579    case "status":
580      return statusText($, now)
581    case "history":
582      return historyReport(state.history)
583    case "profiles":
584      return profilesReport(config, modelEnv, state.unavailable)
585    case "stats":
586      return statsText($, now, cmd.days)
587    case "test":
588      return testText($, now, cmd.prompt)
589    case "feedback":
590      return feedbackText($, now, cmd)
591    case "auto":
592      state.mode = "auto"
593      delete state.pin
594      delete state.nativeEffort
595      delete state.effortBaseline
596      await save($)
597      showStatus($, config.enabled ? `Clef auto · ${config.decisionModel}` : undefined)
598      return config.enabled ? "Routing is automatic again." : "Routing is disabled in the plugin config (enabled = false)."
599    case "on":
600      state.mode = "auto"
601      await save($)
602      showStatus($, `Clef on · ${config.decisionModel}`)
603      return state.pin ? `Routing on, still pinned to ${targetText(state.pin)} (/clef auto to unpin).` : "Routing on."
604    case "off":
605      state.mode = "off"
606      await save($)
607      showStatus($, "Clef off")
608      return "Routing off for this session: Claude Code's own model and effort apply. /clef on resumes."
609    case "pin":
610      state.pin = cmd.target
611      state.mode = "auto"
612      await save($)
613      showStatus($, `Pinned → ${targetText(cmd.target)}`)
614      return `Pinned to ${targetText(cmd.target)} for this session. /clef auto unpins.`
615  }
616}
617
618/** The request a step is sent with: the turn's route, effort fitted to the model. */
619function routed<E extends { model: string; effort?: unknown }>(e: E, route: Route): E {
620  const request = { ...e, model: route.model } as E & { effort?: unknown }
621  if (effortsFor(route.model) === null) delete request.effort
622  else if (route.effort) request.effort = route.effort
623  else if (typeof e.effort === "string" && isEffort(e.effort)) {
624    const clamped = clampEffort(route.model, e.effort)
625    if (clamped) request.effort = clamped
626  }
627  return request
628}
629
630export const register: Register = (on, pluginOptions) => {
631  options = pluginOptions
632  state = freshState()
633  hint = undefined
634  apiBase = undefined
635  loaded = false
636  logFile = undefined
637  logLines = []
638  runs.clear()
639
640  on("ui.render", { component: "PromptHint" }, async ($, e, next) => {
641    if (hint === undefined) return next(e)
642    const tail = e.props.tail ? `${e.props.tail} · ${hint}` : `  ${hint}`
643    return next({ ...e, props: { ...e.props, tail } })
644  })
645
646  on("session.start", async ($, e, next) => {
647    await load($)
648    await registerCommand($)
649    return next(e)
650  })
651
652  // A compaction replaces the conversation, so the next request writes a new
653  // cache whatever the model: a downgrade then costs nothing to take.
654  on("session.compact", async ($, e, next) => {
655    const result = await next(e)
656    if (e.agentId === undefined && e.trigger !== "precompute" && result.messages !== undefined) {
657      try {
658        await load($)
659        delete state.cache
660        delete state.hold
661        await save($)
662      } catch {
663        // Bookkeeping only.
664      }
665    }
666    return result
667  })
668
669  on("session.end", async ($, e, next) => {
670    if (e.reason === "clear") await onClear($)
671    return next(e)
672  })
673
674  on("prompt.submit", async ($, e, next) => {
675    // Only a prompt that starts a turn of its own; one typed into a running
676    // turn joins that turn, which keeps its route.
677    if (e.turnId !== undefined) return next(e)
678    const parsed = parsePrefix(e.text)
679    if (parsed.override === undefined) return next(e)
680    await setPendingOverride($, parsed.override)
681    return next({ ...e, text: parsed.text })
682  })
683
684  on("turn.start", async ($, e, next) => {
685    try {
686      await load($)
687      await routeTurn($, e.turnId, e.text)
688    } catch {
689      // Unrouted: Claude Code's own model and effort apply.
690    }
691    return next(e)
692  })
693
694  on("turn.step", async function* ($, e, next) {
695    // Subagents run on the model their definition or Claude Code gives them.
696    const run = e.agentId === undefined ? runs.get(e.turnId) : undefined
697    if (!run) return yield* next(e)
698    run.steps++
699    try {
700      noticeStep($, run, e.model, e.effort, e.index)
701    } catch {
702      run.passthrough = true
703    }
704    const route = run.decision.final
705    const rewrite = route !== undefined && !run.passthrough && !run.failedRewrite
706    const request = rewrite ? routed(e, route) : e
707    const result = yield* next(request)
708    try {
709      await afterStep($, run, request, rewrite, result, next.signal?.aborted === true)
710    } catch {
711      // Bookkeeping only.
712    }
713    return result
714  })
715
716  on("turn.complete", async ($, e, next) => {
717    const result = await next(e)
718    const run = e.agentId === undefined ? runs.get(e.turnId) : undefined
719    if (!run) return result
720    try {
721      await completeTurn($, run, e.durationMs, e.reason)
722    } catch {
723      // Logging only.
724    }
725    if (config.announce === "answer" || config.announce === "both") {
726      const line = answerLine(run.decision, run.decision.recommendation?.latencyMs)
727      if (line) return { ...result, text: line }
728    }
729    return result
730  })
731
732  on("command.run", { command: "clef" }, async ($, e) => ({ text: await runCommand($, e.args) }))
733}
734
hooks/lib/clef.ts 252 lines
1// Cloudflare Clef on Workers AI: the request, the response, and every way the
2// exchange can fail, turned into a normalised Recommendation or a classified
3// ProviderFailure. Nothing here throws.
4//
5// API (developers.cloudflare.com/workers-ai/models/clef-flash, Oct 2026):
6//   POST https://api.cloudflare.com/client/v4/accounts/{account}/ai/run/@cf/cloudflare/{model}
7//   Authorization: Bearer {token}
8//   { "model": "clef-flash", "state": ..., "questions": { id: {type, instructions, criteria} } }
9// → { "result": { "model", "answers": { id: answer }, "usage": { input_tokens, output_tokens } },
10//     "success": true, "errors": [], "messages": [] }
11
12import { buildQuestions, Q_DIFFICULTY, Q_FOLLOW_UP, type QuestionStyle, type Rubric } from "./rubric.ts"
13import { redact } from "./redact.ts"
14import { LEVELS, type DecisionProvider, type Level, type ProviderFailure, type ProviderResult, type Recommendation } from "./types.ts"
15
16export const CLEF_MODELS = ["clef-flash", "clef"] as const
17export type ClefModel = (typeof CLEF_MODELS)[number]
18
19export const DEFAULT_API_BASE = "https://api.cloudflare.com/client/v4"
20
21/** The slice of `$.http.fetch` (or the global fetch, in scripts) this needs. */
22export type HttpLike = (
23  url: string,
24  init: { method: string; headers: Record<string, string>; body: string },
25) => Promise<{ status: number; ok: boolean; text: string }>
26
27export type ClefOptions = {
28  accountId: string | undefined
29  apiToken: string | undefined
30  model: ClefModel
31  rubric: Rubric
32  style?: QuestionStyle
33  timeoutMs: number
34  maxPromptChars: number
35  apiBase?: string
36  fetch: HttpLike
37  /** Resolves after `ms` (rejects if `signal` aborts); the timeout races the request against it. */
38  sleep: (ms: number, signal: AbortSignal) => Promise<void>
39  now: () => number | Promise<number>
40}
41
42export function endpoint(accountId: string, model: ClefModel, apiBase = DEFAULT_API_BASE): string {
43  return `${apiBase.replace(/\/$/, "")}/accounts/${encodeURIComponent(accountId)}/ai/run/@cf/cloudflare/${model}`
44}
45
46/**
47 * Keeps the head and the tail of a long prompt. The request's intent is
48 * usually stated at one end; the middle of a long paste rarely changes how
49 * hard the task is, and every character sent is billed and leaves the machine.
50 */
51export function truncatePrompt(text: string, maxChars: number): string {
52  if (text.length <= maxChars) return text
53  const head = Math.floor(maxChars * 0.75)
54  const tail = maxChars - head
55  const omitted = text.length - head - tail
56  return `${text.slice(0, head)}\n[... ${omitted} characters omitted ...]\n${text.slice(text.length - tail)}`
57}
58
59export function buildRequestBody(prompt: string, opts: Pick<ClefOptions, "model" | "rubric" | "style" | "maxPromptChars">): string {
60  return JSON.stringify({
61    model: opts.model,
62    state: truncatePrompt(prompt, opts.maxPromptChars),
63    questions: buildQuestions(opts.rubric, opts.style ?? "score"),
64  })
65}
66
67type Envelope = {
68  success?: unknown
69  result?: unknown
70  errors?: unknown
71}
72
73function errorsOf(envelope: Envelope): { code?: number; message: string }[] {
74  if (!Array.isArray(envelope.errors)) return []
75  return envelope.errors
76    .filter((e): e is Record<string, unknown> => typeof e === "object" && e !== null)
77    .map((e) => ({
78      code: typeof e.code === "number" ? e.code : undefined,
79      message: typeof e.message === "string" ? e.message : "",
80    }))
81}
82
83/** Maps an HTTP status and Cloudflare error envelope to a failure kind. */
84export function classifyHttpFailure(
85  status: number,
86  bodyText: string,
87  latencyMs: number,
88  secrets: readonly (string | undefined)[] = [],
89): ProviderFailure {
90  let envelope: Envelope = {}
91  try {
92    envelope = JSON.parse(bodyText) as Envelope
93  } catch {
94    // Not JSON (a proxy's HTML page, say); the status alone decides.
95  }
96  const errors = errorsOf(envelope)
97  const codes = errors.map((e) => e.code)
98  const text = errors.map((e) => (e.code === undefined ? e.message : `${e.code}: ${e.message}`)).join("; ")
99  const message = redact(text || `HTTP ${status}`, secrets).slice(0, 200)
100  const quotaText = /daily free allocation|neurons/i.test(text)
101
102  if (codes.includes(3036) || (status === 429 && quotaText)) return { kind: "quota", message, status, latencyMs }
103  if (status === 429) return { kind: "rate-limited", message, status, latencyMs }
104  if (status === 401 || status === 403 || codes.includes(10000)) return { kind: "auth", message, status, latencyMs }
105  if (status === 408 || codes.includes(3007)) return { kind: "timeout", message, status, latencyMs }
106  if (status >= 500) return { kind: "server", message, status, latencyMs }
107  if (status >= 400) return { kind: "bad-request", message, status, latencyMs }
108  return { kind: "malformed", message, status, latencyMs }
109}
110
111function num(value: unknown): number | undefined {
112  return typeof value === "number" && Number.isFinite(value) ? value : undefined
113}
114
115/**
116 * Reads a difficulty answer's per-level probabilities. A score answer keys
117 * them by level index ("0".."4", per the schema), a choice answer by option
118 * id (our level names). A 1-based index set is accepted too, defensively.
119 */
120export function levelProbabilities(answer: Record<string, unknown>, rubric: Rubric): Record<Level, number> | undefined {
121  const raw = answer.probabilities
122  if (typeof raw !== "object" || raw === null) return undefined
123  const entries = Object.entries(raw as Record<string, unknown>)
124  const out = Object.fromEntries(LEVELS.map((l) => [l, 0])) as Record<Level, number>
125  const numericKeys = entries.every(([k]) => /^\d+$/.test(k))
126  const base = numericKeys ? Math.min(...entries.map(([k]) => Number(k))) : 0
127  let matched = 0
128  for (const [key, value] of entries) {
129    const p = num(value)
130    if (p === undefined || p < 0 || p > 1.0001) return undefined
131    let index: number
132    if (numericKeys) index = Number(key) - (base === 1 && entries.length === LEVELS.length ? 1 : 0)
133    else if ((LEVELS as readonly string[]).includes(key)) index = LEVELS.indexOf(key as Level)
134    else index = rubric.levels.indexOf(key)
135    const level = LEVELS[index]
136    if (level === undefined) return undefined
137    out[level] += p
138    matched++
139  }
140  if (matched === 0) return undefined
141  const total = LEVELS.reduce((s, l) => s + out[l], 0)
142  if (total < 0.98 || total > 1.02) return undefined
143  return out
144}
145
146/** The most probable level; a tie goes to the more capable one. */
147export function topLevel(probabilities: Record<Level, number>): Level {
148  let best: Level = LEVELS[0]
149  for (const level of LEVELS) if (probabilities[level] >= probabilities[best]) best = level
150  return best
151}
152
153/**
154 * Turns a Workers AI success envelope into a Recommendation, or says why it
155 * cannot. Exported for tests and the calibration script.
156 */
157export function parseResponse(
158  bodyText: string,
159  opts: { provider: string; rubric: Rubric; latencyMs: number },
160): ProviderResult {
161  const fail = (message: string): ProviderResult => ({
162    ok: false,
163    failure: { kind: "malformed", message, latencyMs: opts.latencyMs },
164  })
165  let envelope: Envelope
166  try {
167    envelope = JSON.parse(bodyText) as Envelope
168  } catch {
169    return fail("response is not JSON")
170  }
171  if (envelope.success === false) return { ok: false, failure: classifyHttpFailure(200, bodyText, opts.latencyMs) }
172  const result = envelope.result as Record<string, unknown> | undefined
173  if (typeof result !== "object" || result === null) return fail("response has no result")
174  const answers = result.answers as Record<string, unknown> | undefined
175  if (typeof answers !== "object" || answers === null) return fail("result has no answers")
176  const difficulty = answers[Q_DIFFICULTY] as Record<string, unknown> | undefined
177  if (typeof difficulty !== "object" || difficulty === null) return fail(`no answer for "${Q_DIFFICULTY}"`)
178
179  const probabilities = levelProbabilities(difficulty, opts.rubric)
180  if (probabilities === undefined) return fail("difficulty answer has no usable probabilities")
181  const level = topLevel(probabilities)
182  const confidence = num(difficulty.confidence)
183  const score = num(difficulty.score)
184
185  const followUp = answers[Q_FOLLOW_UP] as Record<string, unknown> | undefined
186  const contextDependent = followUp && typeof followUp === "object" ? num(followUp.noul) : undefined
187
188  const usage = result.usage as Record<string, unknown> | undefined
189  const recommendation: Recommendation = {
190    provider: opts.provider,
191    level,
192    confidence: probabilities[level],
193    probabilities,
194    latencyMs: opts.latencyMs,
195  }
196  if (confidence !== undefined && confidence >= 0 && confidence <= 1) recommendation.providerConfidence = confidence
197  if (score !== undefined) recommendation.score = score
198  if (contextDependent !== undefined && contextDependent >= 0 && contextDependent <= 1) recommendation.contextDependent = contextDependent
199  const inputTokens = usage ? num(usage.input_tokens) : undefined
200  if (inputTokens !== undefined) recommendation.inputTokens = inputTokens
201  return { ok: true, recommendation }
202}
203
204const TIMED_OUT: unique symbol = Symbol("timeout")
205
206/** The Clef provider. One HTTP request per call, no retries: a retry would
207 * only add latency in the interactive path, and the fallback is cheap. */
208export function clefProvider(opts: ClefOptions): DecisionProvider {
209  const secrets = [opts.apiToken, opts.accountId]
210  return {
211    name: opts.model,
212    async decide(prompt: string): Promise<ProviderResult> {
213      if (!opts.accountId || !opts.apiToken) {
214        return {
215          ok: false,
216          failure: { kind: "not-configured", message: "Cloudflare account ID or API token not set", latencyMs: 0 },
217        }
218      }
219      const started = await opts.now()
220      const elapsed = async () => Math.round((await opts.now()) - started)
221      let response: { status: number; ok: boolean; text: string } | typeof TIMED_OUT
222      const timer = new AbortController()
223      try {
224        const request = opts.fetch(endpoint(opts.accountId, opts.model, opts.apiBase), {
225          method: "POST",
226          headers: { Authorization: `Bearer ${opts.apiToken}`, "Content-Type": "application/json" },
227          body: buildRequestBody(prompt, opts),
228        })
229        // $.http.fetch takes no abort signal: on a timeout the request is left
230        // to finish on its own, and its answer is ignored.
231        request.catch(() => {})
232        const deadline = opts.sleep(opts.timeoutMs, timer.signal).then(
233          (): typeof TIMED_OUT => TIMED_OUT,
234          (): typeof TIMED_OUT => TIMED_OUT,
235        )
236        response = await Promise.race([request, deadline])
237      } catch (error) {
238        const message = redact(error instanceof Error ? error.message : String(error), secrets).slice(0, 200)
239        return { ok: false, failure: { kind: "network", message, latencyMs: await elapsed() } }
240      } finally {
241        timer.abort()
242      }
243      const latencyMs = await elapsed()
244      if (response === TIMED_OUT) {
245        return { ok: false, failure: { kind: "timeout", message: `no answer within ${opts.timeoutMs} ms`, latencyMs } }
246      }
247      if (!response.ok) return { ok: false, failure: classifyHttpFailure(response.status, response.text, latencyMs, secrets) }
248      return parseResponse(response.text, { provider: opts.model, rubric: opts.rubric, latencyMs })
249    },
250  }
251}
252
hooks/lib/config.ts 240 lines
1// Reads the plugin's `userConfig` values (set with /plugin configure or
2// /config) into a validated Config. Bad values fall back to defaults and are
3// reported, never thrown: a typo in one field must not stop Claude Code.
4
5import { isEffort } from "./models.ts"
6import { CLEF_MODELS, type ClefModel } from "./clef.ts"
7import { BILLINGS, type Billing } from "./pricing.ts"
8import { EFFORTS, LEVELS, type Effort, type Level, type ProfileSpec } from "./types.ts"
9
10export const LOW_CONFIDENCE_POLICIES = ["upper-of-top-two", "bump", "hold", "fallback", "obey"] as const
11export type LowConfidencePolicy = (typeof LOW_CONFIDENCE_POLICIES)[number]
12
13export const ANNOUNCE_MODES = ["status", "answer", "both", "off"] as const
14export type AnnounceMode = (typeof ANNOUNCE_MODES)[number]
15
16export const BILLING_OPTIONS = ["auto", ...BILLINGS] as const
17export type BillingOption = "auto" | Billing
18
19export type Config = {
20  enabled: boolean
21  accountId?: string
22  apiToken?: string
23  decisionModel: ClefModel
24  timeoutMs: number
25  profiles: Record<Level, ProfileSpec>
26  confidenceThreshold: number
27  lowConfidencePolicy: LowConfidencePolicy
28  fallbackLevel: Level
29  /** Honour /model (pause) and /effort (effort only) changes made mid-session. */
30  pauseOnNativeChange: boolean
31  /** P(follow-up) at or above which a prompt never routes below the last route. */
32  followUpThreshold: number
33  /** How the person pays for Claude; `auto` detects it. Decides what a held downgrade costs. */
34  billing: BillingOption
35  /**
36   * A downgrade off a warm cache is taken once staying has cost this many
37   * times what the switch costs; 0 = take every downgrade at once.
38   */
39  downgradePatience: number
40  /** Prompt-cache TTL in minutes; 0 = work it out from the environment. */
41  cacheTtlMinutes: number
42  maxEffort?: Effort
43  dailyNeuronBudget: number
44  maxPromptChars: number
45  announce: AnnounceMode
46  logEnabled: boolean
47  logPrompts: boolean
48  logDir?: string
49  rubricFile?: string
50}
51
52export const DEFAULT_PROFILES: Record<Level, string> = {
53  trivial: "haiku",
54  simple: "sonnet:low",
55  standard: "sonnet:medium",
56  hard: "opus:high",
57  deep: "opus:xhigh",
58}
59
60export const DEFAULTS = {
61  decisionModel: "clef-flash" as ClefModel,
62  timeoutMs: 1500,
63  confidenceThreshold: 0.55,
64  lowConfidencePolicy: "upper-of-top-two" as LowConfidencePolicy,
65  fallbackLevel: "standard" as Level,
66  followUpThreshold: 0.6,
67  billing: "auto" as BillingOption,
68  downgradePatience: 1,
69  cacheTtlMinutes: 0,
70  dailyNeuronBudget: 9_000,
71  maxPromptChars: 6_000,
72  announce: "status" as AnnounceMode,
73}
74
75/** "opus:high" → { model: "opus", effort: "high" }; "haiku" → { model: "haiku" }. */
76export function parseProfile(level: Level, text: string): ProfileSpec | string {
77  const trimmed = text.trim()
78  if (trimmed === "") return `profile_${level} is empty`
79  const split = splitTarget(trimmed)
80  if (split.model === "") return `profile_${level} "${text}" names no model`
81  if (split.badEffort) return `profile_${level} "${text}": effort must be one of ${EFFORTS.join(", ")}`
82  return split.effort ? { level, model: split.model, effort: split.effort } : { level, model: split.model }
83}
84
85/**
86 * Splits "model:effort". A suffix that is not an effort name stays part of
87 * the model, since provider IDs carry colons ("...-v1:0" on Bedrock); a
88 * suffix that looks like a word but is no effort is reported.
89 */
90export function splitTarget(text: string): { model: string; effort?: Effort; badEffort?: boolean } {
91  const colon = text.lastIndexOf(":")
92  if (colon === -1) return { model: text.trim() }
93  const model = text.slice(0, colon).trim()
94  const suffix = text.slice(colon + 1).trim().toLowerCase()
95  if (suffix === "" || suffix === "default") return { model }
96  if (isEffort(suffix)) return { model, effort: suffix }
97  if (/^\d+$/.test(suffix)) return { model: text.trim() }
98  return { model, badEffort: true }
99}
100
101type Options = Readonly<Record<string, unknown>>
102
103function str(options: Options, key: string): string | undefined {
104  const v = options[key]
105  return typeof v === "string" && v.trim() !== "" ? v.trim() : undefined
106}
107
108function numIn(options: Options, key: string, min: number, max: number, fallback: number, problems: string[]): number {
109  const v = options[key]
110  if (v === undefined || v === "") return fallback
111  const n = typeof v === "number" ? v : Number(v)
112  if (!Number.isFinite(n) || n < min || n > max) {
113    problems.push(`${key} must be a number from ${min} to ${max}; using ${fallback}`)
114    return fallback
115  }
116  return n
117}
118
119function oneOf<T extends string>(options: Options, key: string, allowed: readonly T[], fallback: T, problems: string[]): T {
120  const v = str(options, key)
121  if (v === undefined) return fallback
122  if ((allowed as readonly string[]).includes(v)) return v as T
123  problems.push(`${key} must be one of ${allowed.join(", ")}; using ${fallback}`)
124  return fallback
125}
126
127function bool(options: Options, key: string, fallback: boolean): boolean {
128  const v = options[key]
129  if (typeof v === "boolean") return v
130  if (v === "true") return true
131  if (v === "false") return false
132  return fallback
133}
134
135/**
136 * `downgrade_patience`, or what the old `cache_hold_min_tokens` meant by 0
137 * (never hold). Its other values have no equivalent: holding now depends on
138 * what staying has cost, not on the context's size.
139 */
140function downgradePatience(options: Options, problems: string[]): number {
141  if (options.downgrade_patience !== undefined) return numIn(options, "downgrade_patience", 0, 100, DEFAULTS.downgradePatience, problems)
142  const old = options.cache_hold_min_tokens
143  if (old === undefined) return DEFAULTS.downgradePatience
144  if (Number(old) === 0) {
145    problems.push("cache_hold_min_tokens is replaced by downgrade_patience; reading 0 as downgrade_patience 0")
146    return 0
147  }
148  problems.push(`cache_hold_min_tokens is replaced by downgrade_patience and ignored; using ${DEFAULTS.downgradePatience}`)
149  return DEFAULTS.downgradePatience
150}
151
152/**
153 * Reads the optional advanced-settings file (`clef-model-router.json`): a
154 * JSON object with the same keys as the plugin options. Plugin options win
155 * over it; it wins over the defaults.
156 */
157export function parseAdvancedFile(text: string | undefined): { values: Options; problems: string[] } {
158  if (text === undefined) return { values: {}, problems: [] }
159  try {
160    const value = JSON.parse(text) as unknown
161    if (typeof value !== "object" || value === null || Array.isArray(value)) {
162      return { values: {}, problems: ["clef-model-router.json must hold a JSON object; ignoring it"] }
163    }
164    const values = { ...(value as Record<string, unknown>) }
165    // Credentials belong in the plugin's secure storage, not in a plain file.
166    const problems: string[] = []
167    if ("cloudflare_api_token" in values) {
168      delete values.cloudflare_api_token
169      problems.push("clef-model-router.json: cloudflare_api_token is ignored there; set it with /plugin configure")
170    }
171    return { values, problems }
172  } catch {
173    return { values: {}, problems: ["clef-model-router.json is not valid JSON; ignoring it"] }
174  }
175}
176
177/**
178 * Builds the Config from plugin options plus environment fallbacks for the
179 * credentials (CLOUDFLARE_ACCOUNT_ID / CLOUDFLARE_API_TOKEN), which the
180 * caller reads and passes in.
181 */
182export function parseConfig(
183  options: Options,
184  envFallback: { accountId?: string; apiToken?: string } = {},
185): { config: Config; problems: string[] } {
186  const problems: string[] = []
187  const profiles = {} as Record<Level, ProfileSpec>
188  // `profiles` is the five of them in one line, lowest first; a `profile_<level>`
189  // key (in the advanced file) names one.
190  const list = str(options, "profiles")?.split(",").map((s) => s.trim())
191  if (list && list.length !== LEVELS.length) {
192    problems.push(`profiles must list ${LEVELS.length} entries (${LEVELS.join(", ")}), got ${list.length}; using the defaults`)
193  }
194  for (const [i, level] of LEVELS.entries()) {
195    const fromList = list && list.length === LEVELS.length ? list[i] : undefined
196    const raw = str(options, `profile_${level}`) ?? fromList ?? DEFAULT_PROFILES[level]
197    const parsed = parseProfile(level, raw)
198    if (typeof parsed === "string") {
199      problems.push(`${parsed}; using "${DEFAULT_PROFILES[level]}"`)
200      profiles[level] = parseProfile(level, DEFAULT_PROFILES[level]) as ProfileSpec
201    } else profiles[level] = parsed
202  }
203  const maxEffortRaw = str(options, "max_effort")
204  let maxEffort: Effort | undefined
205  if (maxEffortRaw !== undefined && maxEffortRaw !== "none") {
206    if (isEffort(maxEffortRaw)) maxEffort = maxEffortRaw
207    else problems.push(`max_effort must be one of ${EFFORTS.join(", ")} or none; ignoring it`)
208  }
209
210  const config: Config = {
211    enabled: bool(options, "enabled", true),
212    decisionModel: oneOf(options, "decision_model", CLEF_MODELS, DEFAULTS.decisionModel, problems),
213    timeoutMs: numIn(options, "timeout_ms", 100, 10_000, DEFAULTS.timeoutMs, problems),
214    profiles,
215    confidenceThreshold: numIn(options, "confidence_threshold", 0, 1, DEFAULTS.confidenceThreshold, problems),
216    lowConfidencePolicy: oneOf(options, "low_confidence_policy", LOW_CONFIDENCE_POLICIES, DEFAULTS.lowConfidencePolicy, problems),
217    fallbackLevel: oneOf(options, "fallback_profile", LEVELS, DEFAULTS.fallbackLevel, problems),
218    followUpThreshold: numIn(options, "follow_up_threshold", 0, 1, DEFAULTS.followUpThreshold, problems),
219    billing: oneOf(options, "billing", BILLING_OPTIONS, DEFAULTS.billing, problems),
220    downgradePatience: downgradePatience(options, problems),
221    cacheTtlMinutes: numIn(options, "cache_ttl_minutes", 0, 1440, DEFAULTS.cacheTtlMinutes, problems),
222    dailyNeuronBudget: numIn(options, "daily_neuron_budget", 0, 1_000_000_000, DEFAULTS.dailyNeuronBudget, problems),
223    maxPromptChars: numIn(options, "max_prompt_chars", 200, 200_000, DEFAULTS.maxPromptChars, problems),
224    announce: oneOf(options, "announce", ANNOUNCE_MODES, DEFAULTS.announce, problems),
225    pauseOnNativeChange: bool(options, "pause_on_native_change", true),
226    logEnabled: bool(options, "log_enabled", true),
227    logPrompts: bool(options, "log_prompts", false),
228  }
229  if (maxEffort) config.maxEffort = maxEffort
230  const accountId = str(options, "cloudflare_account_id") ?? envFallback.accountId
231  const apiToken = str(options, "cloudflare_api_token") ?? envFallback.apiToken
232  if (accountId) config.accountId = accountId
233  if (apiToken) config.apiToken = apiToken
234  const logDir = str(options, "log_dir")
235  if (logDir) config.logDir = logDir
236  const rubricFile = str(options, "rubric_file")
237  if (rubricFile) config.rubricFile = rubricFile
238  return { config, problems }
239}
240
hooks/lib/env.ts 115 lines
1// Turns the environment variables Claude Code itself honours into the facts
2// the policy needs. Pure: the hooks module reads the variables (by literal
3// name, as the engine requires) and passes them here.
4
5import type { ModelEnv } from "./models.ts"
6import type { Billing } from "./pricing.ts"
7
8export type EnvValues = {
9  ANTHROPIC_DEFAULT_HAIKU_MODEL?: string
10  ANTHROPIC_DEFAULT_SONNET_MODEL?: string
11  ANTHROPIC_DEFAULT_OPUS_MODEL?: string
12  ANTHROPIC_DEFAULT_FABLE_MODEL?: string
13  CLAUDE_CODE_USE_BEDROCK?: string
14  CLAUDE_CODE_USE_VERTEX?: string
15  CLAUDE_CODE_USE_FOUNDRY?: string
16  CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS?: string
17  ANTHROPIC_API_KEY?: string
18  CLAUDE_CODE_PROMPT_CACHE_TTL?: string
19  FORCE_PROMPT_CACHING_5M?: string
20  ENABLE_PROMPT_CACHING_1H?: string
21  ANTHROPIC_AUTH_TOKEN?: string
22  ANTHROPIC_BASE_URL?: string
23}
24
25const truthy = (v: string | undefined) => v !== undefined && v !== "" && v !== "0" && v.toLowerCase() !== "false"
26
27export function modelEnvFrom(env: EnvValues): ModelEnv {
28  const defaults: ModelEnv["defaults"] = {}
29  if (env.ANTHROPIC_DEFAULT_HAIKU_MODEL) defaults.haiku = env.ANTHROPIC_DEFAULT_HAIKU_MODEL
30  if (env.ANTHROPIC_DEFAULT_SONNET_MODEL) defaults.sonnet = env.ANTHROPIC_DEFAULT_SONNET_MODEL
31  if (env.ANTHROPIC_DEFAULT_OPUS_MODEL) defaults.opus = env.ANTHROPIC_DEFAULT_OPUS_MODEL
32  if (env.ANTHROPIC_DEFAULT_FABLE_MODEL) defaults.fable = env.ANTHROPIC_DEFAULT_FABLE_MODEL
33  return {
34    defaults,
35    thirdParty: truthy(env.CLAUDE_CODE_USE_BEDROCK) || truthy(env.CLAUDE_CODE_USE_VERTEX) || truthy(env.CLAUDE_CODE_USE_FOUNDRY),
36    betasDisabled: truthy(env.CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS),
37  }
38}
39
40const FIVE_MIN = 5 * 60_000
41const ONE_HOUR = 60 * 60_000
42
43function ttlValue(v: string | undefined): number | undefined {
44  if (v === "5m") return FIVE_MIN
45  if (v === "1h") return ONE_HOUR
46  return undefined
47}
48
49/** One rate-limit window as `$.session.usage()` reports it. */
50export type RateLimitWindow = { kind: string; percentUsed: number }
51
52/** What the session shows about how the person pays for Claude. */
53export type BillingFacts = {
54  billing: Billing
55  /** Why, in a few words, for /clef. */
56  why: string
57  /** A subscription past its included usage, drawing usage credits billed per token. */
58  overage: boolean
59}
60
61/** The windows a Claude subscription reports; an API key or cloud provider reports none. */
62const PLAN_WINDOWS = ["five_hour", "seven_day"]
63
64/**
65 * How the person pays, from the strongest evidence available:
66 *
67 * 1. The plan's rate-limit windows (`five_hour`, `seven_day`), which Claude
68 *    Code reports only on a subscription. One at 100% or more means the plan's
69 *    included usage is spent and further requests are usage credits, billed
70 *    per token like the API.
71 * 2. A gateway's `spend_limit` window: billed per token.
72 * 3. The environment: a cloud provider, an API key (or `apiKeyHelper`), or a
73 *    gateway (`ANTHROPIC_AUTH_TOKEN`, `ANTHROPIC_BASE_URL`) bill per token.
74 *    Nothing set means a claude.ai login, which is a subscription.
75 *
76 * The windows are empty until the session's first response, so the first
77 * turn is judged on the environment alone.
78 */
79export function detectBilling(
80  env: EnvValues,
81  opts: { apiKeyHelper?: boolean; rateLimits?: readonly RateLimitWindow[] } = {},
82): BillingFacts {
83  const windows = opts.rateLimits ?? []
84  const plan = windows.filter((w) => PLAN_WINDOWS.includes(w.kind))
85  if (plan.length > 0) {
86    if (plan.some((w) => w.percentUsed >= 100)) return { billing: "api", why: "plan limit reached: usage credits bill per token", overage: true }
87    return { billing: "subscription", why: "the plan's rate limits are reported", overage: false }
88  }
89  if (windows.some((w) => w.kind === "spend_limit")) return { billing: "api", why: "a gateway spend limit is reported", overage: false }
90  if (modelEnvFrom(env).thirdParty) return { billing: "api", why: "a cloud provider bills per token", overage: false }
91  if (env.ANTHROPIC_API_KEY || opts.apiKeyHelper) return { billing: "api", why: "an API key is set", overage: false }
92  if (env.ANTHROPIC_AUTH_TOKEN || env.ANTHROPIC_BASE_URL) return { billing: "api", why: "a gateway is set (ANTHROPIC_BASE_URL)", overage: false }
93  return { billing: "subscription", why: "no API key, provider or gateway is set", overage: false }
94}
95
96/**
97 * The main conversation's prompt-cache TTL, resolved in the order Claude
98 * Code's prompt-caching docs give: FORCE_PROMPT_CACHING_5M, the TTL variable,
99 * the `promptCacheTtl` setting, ENABLE_PROMPT_CACHING_1H, then the default:
100 * one hour on a subscription within its included usage, five minutes
101 * otherwise (an API key, a cloud provider, usage credits). The default follows
102 * what was detected, not the `billing` option: the option says what the
103 * person wants optimised, the TTL is what Claude Code actually requests.
104 */
105export function cacheTtlMs(configMinutes: number, env: EnvValues, settingsTtl?: unknown, detected: BillingFacts = detectBilling(env)): number {
106  if (configMinutes > 0) return configMinutes * 60_000
107  if (truthy(env.FORCE_PROMPT_CACHING_5M)) return FIVE_MIN
108  const fromEnv = ttlValue(env.CLAUDE_CODE_PROMPT_CACHE_TTL)
109  if (fromEnv) return fromEnv
110  const fromSettings = typeof settingsTtl === "string" ? ttlValue(settingsTtl) : undefined
111  if (fromSettings) return fromSettings
112  if (truthy(env.ENABLE_PROMPT_CACHING_1H)) return ONE_HOUR
113  return detected.billing === "subscription" ? ONE_HOUR : FIVE_MIN
114}
115
hooks/lib/format.ts 242 lines
1// Everything the router shows: the one-line status, and the text of each
2// /clef subcommand. Plain text, so it reads the same on every surface.
3
4import type { BillingOption, Config } from "./config.ts"
5import type { BillingFacts } from "./env.ts"
6import { displayName, resolveModel, type ModelEnv } from "./models.ts"
7import { describe, kTokens, pct, type HoldState, type RouterMode } from "./policy.ts"
8import { dollars, type Billing } from "./pricing.ts"
9import { presence } from "./redact.ts"
10import type { Stats } from "./log.ts"
11import { LEVELS, type Decision, type Level, type Target } from "./types.ts"
12
13const RULE_LABEL: Record<string, string> = {
14  "low-confidence": "unsure",
15  "context-dependent": "follow-up",
16  unavailable: "unavailable",
17  "context-window": "window",
18  "cache-hold": "held for cache",
19  "effort-clamp": "effort clamped",
20  "effort-cap": "effort capped",
21  "pinned-effort": "pinned effort",
22}
23
24/**
25 * A held downgrade names what it held back, so a route that differs from
26 * Clef's pick says so: "(Sonnet deferred)", or "(low effort deferred)".
27 */
28function holdLabel(d: Decision): string {
29  const w = d.deferral?.wanted
30  if (!w || !d.final) return RULE_LABEL["cache-hold"]!
31  return w.model === d.final.model ? `${w.effort ?? "lower"} effort deferred` : `${displayName(w.model)} deferred`
32}
33
34/** The one line under the prompt, e.g. "Clef → Sonnet · medium · 87%". */
35export function statusLine(d: Decision): string | undefined {
36  const route = d.final ? describe(d.final) : undefined
37  const notes = d.adjustments.map((a) => (a.rule === "cache-hold" ? holdLabel(d) : (RULE_LABEL[a.rule] ?? a.rule)))
38  const tail = notes.length > 0 ? ` (${[...new Set(notes)].join(", ")})` : ""
39  switch (d.source) {
40    case "clef": {
41      const conf = d.recommendation ? ` · ${pct(d.recommendation.confidence)}` : ""
42      return `Clef → ${route}${conf}${tail}`
43    }
44    case "fallback":
45      return `Clef ✕ ${d.failure?.kind ?? "no answer"} → ${route}${tail}`
46    case "continuation":
47      return route ? `Clef ↻ ${route}${tail}` : undefined
48    case "override":
49      return route ? `+ ${route}${tail}` : `+ ${d.note ?? "override"}`
50    case "pin":
51      return route ? `Pinned → ${route}${tail}` : `Pinned: ${d.note ?? ""}`
52    case "native":
53      return "Clef paused (you chose /model)"
54    case "disabled":
55      return d.note?.startsWith("+off") ? "Clef skipped this turn" : "Clef off"
56  }
57}
58
59/** A line under the answer, when `announce` asks for one. */
60export function answerLine(d: Decision, latencyMs?: number): string | undefined {
61  const line = statusLine(d)
62  if (!line) return undefined
63  return latencyMs !== undefined && d.source === "clef" ? `${line} · ${latencyMs} ms` : line
64}
65
66function bar(p: number, width = 20): string {
67  const filled = Math.round(p * width)
68  return "█".repeat(filled) + "·".repeat(width - filled)
69}
70
71export function distribution(probabilities: Record<Level, number>, mark?: Level, final?: Level): string[] {
72  return LEVELS.map((level) => {
73    const tags = [level === mark ? "← Clef" : "", level === final && final !== mark ? "← routed" : ""].filter(Boolean).join(" ")
74    return `    ${level.padEnd(9)} ${bar(probabilities[level])} ${pct(probabilities[level]).padStart(4)}  ${tags}`.trimEnd()
75  })
76}
77
78export function explain(d: Decision): string[] {
79  const lines: string[] = []
80  lines.push(`  route     ${d.final ? `${describe(d.final)}  (${d.final.model}${d.final.level ? `, profile ${d.final.level}` : ""})` : "untouched (Claude Code's own model and effort)"}`)
81  lines.push(`  source    ${d.source}${d.note ? ` — ${d.note}` : ""}`)
82  const rec = d.recommendation
83  if (rec) {
84    const extras = [
85      `${pct(rec.confidence)} on ${rec.level}`,
86      rec.providerConfidence !== undefined ? `clef confidence ${pct(rec.providerConfidence)}` : "",
87      rec.score !== undefined ? `score ${rec.score.toFixed(2)}/4` : "",
88      rec.contextDependent !== undefined ? `follow-up ${pct(rec.contextDependent)}` : "",
89      `${rec.latencyMs} ms`,
90      rec.inputTokens !== undefined ? `${rec.inputTokens} tokens` : "",
91    ].filter(Boolean)
92    lines.push(`  clef      ${rec.provider}: ${extras.join(" · ")}`)
93    lines.push(...distribution(rec.probabilities, rec.level, d.final?.level))
94  }
95  if (d.failure) lines.push(`  failure   ${d.failure.kind}: ${d.failure.message}${d.failure.latencyMs ? ` (${d.failure.latencyMs} ms)` : ""}`)
96  for (const a of d.adjustments) lines.push(`  policy    ${a.rule}: ${a.from} → ${a.to} — ${a.reason}`)
97  const f = d.deferral
98  if (f) {
99    const what = f.wanted.model === f.from ? `${f.wanted.effort ?? "lower"} effort` : displayName(f.wanted.model)
100    lines.push(
101      f.held
102        ? `  downgrade ${what} deferred (held turn ${f.turns}): staying has cost ${dollars(f.spent)} so far, a switch costs ${dollars(f.cost)} now (${f.billing}, list prices)`
103        : `  downgrade ${what} taken after ${f.turns} held turn${f.turns === 1 ? "" : "s"}: staying had cost ${dollars(f.spent)}, the switch ${dollars(f.cost)} (${f.billing}, list prices)`,
104    )
105  }
106  return lines
107}
108
109export type StatusArgs = {
110  config: Config
111  configProblems: readonly string[]
112  modelEnv: ModelEnv
113  mode: RouterMode
114  pin?: Target
115  last?: Decision
116  guard: { calls: number; inputTokens: number; neurons: number; blocked?: string }
117  cache?: { model: string; promptTokens: number; ageSeconds: number; ttlSeconds: number }
118  billing: { billing: Billing; detected: BillingFacts; configured: BillingOption; patience: number }
119  hold?: HoldState
120  unavailable: readonly string[]
121  logDir: string | undefined
122  advancedPath?: string
123}
124
125export function targetText(t: Target): string {
126  return [t.level ?? t.model ?? "", t.effort ? `:${t.effort}` : ""].join("")
127}
128
129export function statusReport(a: StatusArgs): string {
130  const c = a.config
131  const lines = ["Clef router"]
132  const mode =
133    !c.enabled ? "disabled in plugin config" : a.mode === "auto" ? (a.pin ? `pinned to ${targetText(a.pin)} (/clef auto to unpin)` : "auto") : a.mode === "off" ? "off for this session (/clef on)" : "paused: you changed /model (/clef auto to resume)"
134  lines.push(`  mode      ${mode}`)
135  lines.push(`  decider   ${c.decisionModel} · timeout ${c.timeoutMs} ms · account ${presence(c.accountId)} · token ${presence(c.apiToken)}`)
136  const budget = c.dailyNeuronBudget > 0 ? ` of ${c.dailyNeuronBudget} budget` : ""
137  lines.push(`  today     ${a.guard.calls} Clef calls · ${kTokens(a.guard.inputTokens)} input tokens · ~${Math.round(a.guard.neurons)} neurons${budget}${a.guard.blocked ? ` · ${a.guard.blocked}` : ""}`)
138  const b = a.billing
139  const source = b.configured === "auto" ? `detected: ${b.detected.why}` : `set in config; detected ${b.detected.billing} (${b.detected.why})`
140  const patience = b.patience === 0 ? "downgrades taken at once" : `downgrade patience ${b.patience}`
141  lines.push(`  billing   ${b.billing} (${source}) · ${patience}`)
142  if (a.cache) {
143    const warm = a.cache.ageSeconds < a.cache.ttlSeconds
144    lines.push(`  cache     ${displayName(a.cache.model)} · ${kTokens(a.cache.promptTokens)} context · ${warm ? `warm (${a.cache.ageSeconds}s of ${a.cache.ttlSeconds}s)` : "cold"}`)
145  }
146  if (a.hold) {
147    const what = a.hold.wanted.model === a.hold.model ? `${a.hold.wanted.effort ?? "lower"} effort` : displayName(a.hold.wanted.model)
148    lines.push(`  deferred  ${what} · ${a.hold.turns} held turn${a.hold.turns === 1 ? "" : "s"} on ${displayName(a.hold.model)} · staying has cost ${dollars(a.hold.spent)} (list prices)`)
149  }
150  if (a.unavailable.length > 0) lines.push(`  unusable  ${a.unavailable.join(", ")}`)
151  for (const p of a.configProblems) lines.push(`  config!   ${p}`)
152  if (a.last) {
153    lines.push("", "Last turn")
154    lines.push(...explain(a.last))
155  } else {
156    lines.push("", "No turn routed yet this session.")
157  }
158  if (a.advancedPath) lines.push("", `Advanced settings: ${a.advancedPath} (optional)`)
159  lines.push(`Log: ${c.logEnabled ? (a.logDir ?? "(unavailable)") : "off"}${c.logEnabled && c.logPrompts ? " (with prompt text)" : ""}`)
160  lines.push("Commands: /clef history · stats · profiles · test <prompt> · pin <target> · auto · off · on · feedback under|ok|over")
161  return lines.join("\n")
162}
163
164export function profilesReport(config: Config, env: ModelEnv, unavailable: readonly string[]): string {
165  const lines = ["Profiles, lowest first. Change them with /config (Profiles) or /plugin configure."]
166  for (const level of LEVELS) {
167    const spec = config.profiles[level]
168    const id = resolveModel(spec.model, env)
169    const state = id === undefined ? "cannot resolve here (set a full model ID)" : unavailable.includes(id) ? `${id} (failed this session)` : id
170    lines.push(`  ${level.padEnd(9)} ${`${spec.model}${spec.effort ? `:${spec.effort}` : ""}`.padEnd(16)} → ${state}`)
171  }
172  lines.push(`  fallback  ${config.fallbackLevel} · low confidence (< ${pct(config.confidenceThreshold)}): ${config.lowConfidencePolicy}`)
173  return lines.join("\n")
174}
175
176export type HistoryRow = { decision: Decision; prompt: string; answeredModel?: string }
177
178export function historyReport(rows: readonly HistoryRow[]): string {
179  if (rows.length === 0) return "No turns yet this session."
180  const lines = ["Turns this session, newest first"]
181  lines.push("     ms  route                  conf  source        prompt")
182  for (const row of [...rows].reverse()) {
183    const d = row.decision
184    const ms = d.recommendation?.latencyMs ?? d.failure?.latencyMs
185    const conf = d.recommendation ? pct(d.recommendation.confidence) : "—"
186    const route = d.final ? describe(d.final) : "untouched"
187    const flag = d.adjustments.length > 0 ? "*" : " "
188    const prompt = row.prompt.replace(/\s+/g, " ").slice(0, 48)
189    lines.push(`  ${String(ms ?? "—").padStart(5)}  ${(route + flag).padEnd(22)} ${conf.padStart(4)}  ${d.source.padEnd(12)}  ${prompt}`)
190    for (const a of d.adjustments) lines.push(`         ${a.rule}: ${a.from} → ${a.to}`)
191    if (d.failure) lines.push(`         ${d.failure.kind}: ${d.failure.message}`)
192    if (row.answeredModel && d.final && !row.answeredModel.startsWith(d.final.model)) lines.push(`         answered by ${row.answeredModel}`)
193  }
194  lines.push("  * policy changed Clef's recommendation")
195  return lines.join("\n")
196}
197
198function table(title: string, map: Record<string, number>, total: number): string[] {
199  const entries = Object.entries(map).sort((a, b) => b[1] - a[1])
200  if (entries.length === 0) return []
201  return [`  ${title}`, ...entries.map(([k, v]) => `    ${k.padEnd(24)} ${String(v).padStart(5)}  ${pct(total ? v / total : 0).padStart(4)}`)]
202}
203
204export function statsReport(s: Stats, days: number, files: number): string {
205  if (s.turns === 0 && Object.keys(s.feedback).length === 0) return `No routing log entries in the last ${days} day(s).`
206  const lines = [`Routing over the last ${days} day(s): ${s.turns} turns in ${files} log file(s)`]
207  lines.push(...table("by model", s.byModel, s.turns))
208  lines.push(...table("by effort", s.byEffort, s.turns))
209  lines.push(...table("by profile", s.byLevel, s.turns))
210  lines.push(...table("by source", s.bySource, s.turns))
211  lines.push("  clef")
212  lines.push(`    calls ${s.clefCalls} · ${kTokens(s.clefInputTokens)} input tokens`)
213  if (s.latency) lines.push(`    latency mean ${s.latency.mean} ms · p50 ${s.latency.p50} ms · p95 ${s.latency.p95} ms`)
214  if (s.meanConfidence !== undefined) lines.push(`    mean probability of Clef's pick ${pct(s.meanConfidence)}`)
215  lines.push(`    recommendation changed by policy: ${s.recommendationChanged} · manual overrides: ${s.overrides}`)
216  lines.push(`    turns held on a warm model: ${s.cacheHolds} · downgrades taken after a hold: ${s.downgradesTaken}`)
217  const failures = Object.entries(s.failures)
218  if (failures.length > 0) lines.push(`    fallbacks: ${failures.map(([k, v]) => `${k} ${v}`).join(", ")}`)
219  const fb = Object.entries(s.feedback)
220  if (fb.length > 0) lines.push(`  your feedback: ${fb.map(([k, v]) => `${k} ${v}`).join(", ")}`)
221  return lines.join("\n")
222}
223
224export const HELP = [
225  "Clef router — picks the Claude model and effort for each turn.",
226  "",
227  "  /clef                  status and the last decision, with Clef's probabilities",
228  "  /clef history          this session's turns",
229  "  /clef stats [days]     totals from the local log (default 7 days)",
230  "  /clef profiles         what each difficulty level runs on",
231  "  /clef test <prompt>    ask Clef about a prompt without sending it to Claude",
232  "  /clef pin <target>     use one target for the rest of the session",
233  "                         target: trivial|simple|standard|hard|deep, haiku|sonnet|opus|fable,",
234  "                         a model ID, with optional :effort; or :effort alone",
235  "  /clef auto             unpin, resume after /model, and hand effort back after /effort",
236  "  /clef off | on         stop or resume routing for this session",
237  "  /clef feedback under|ok|over [note]   rate the last route, for later analysis",
238  "",
239  "One turn only: start a prompt with +target, e.g. `+opus:max why does this deadlock?`,",
240  "or `+off ...` to leave that turn to Claude Code.",
241].join("\n")
242
hooks/lib/guard.ts 131 lines
1// Deterministic gates in front of the Clef call: a local daily budget that
2// keeps usage inside Workers AI's free allocation, a pause after the
3// allocation is reported exhausted, and a circuit breaker so a dead endpoint
4// costs one timeout, not one per prompt.
5//
6// The state is plain JSON kept in `$.store`, shared by every session on the
7// machine (best effort: two sessions writing at once may lose a count).
8
9import type { FailureKind, ProviderFailure } from "./types.ts"
10
11export type GuardState = {
12  /** UTC day (YYYY-MM-DD) the counters belong to. */
13  day: string
14  calls: number
15  inputTokens: number
16  /** Clef said the day's free allocation is used up; no calls until `day` changes. */
17  quotaExhausted?: boolean
18  consecutiveFailures: number
19  /** Epoch ms until which no call is made. */
20  pausedUntil?: number
21  pauseReason?: FailureKind
22}
23
24/**
25 * Neurons per million input tokens, derived from Cloudflare's published
26 * prices ($0.09/M for clef-flash, $0.24/M for clef) at $0.011 per 1,000
27 * neurons. Clef is not yet in the per-model neuron table; this is an
28 * estimate, and the Workers AI dashboard is the authority.
29 */
30export const NEURONS_PER_M_INPUT: Record<string, number> = {
31  "clef-flash": (0.09 / 0.011) * 1000,
32  clef: (0.24 / 0.011) * 1000,
33}
34
35export function utcDay(epochMs: number): string {
36  return new Date(epochMs).toISOString().slice(0, 10)
37}
38
39export function freshGuard(epochMs: number): GuardState {
40  return { day: utcDay(epochMs), calls: 0, inputTokens: 0, consecutiveFailures: 0 }
41}
42
43/** Rolls the counters over at 00:00 UTC, when Workers AI's allocation resets. */
44export function normaliseGuard(state: unknown, epochMs: number): GuardState {
45  const today = utcDay(epochMs)
46  if (typeof state !== "object" || state === null) return freshGuard(epochMs)
47  const s = state as Partial<GuardState>
48  if (s.day !== today) {
49    const next = freshGuard(epochMs)
50    if (typeof s.pausedUntil === "number" && s.pausedUntil > epochMs && s.pauseReason !== "quota") {
51      next.pausedUntil = s.pausedUntil
52      next.pauseReason = s.pauseReason
53    }
54    return next
55  }
56  return {
57    day: today,
58    calls: typeof s.calls === "number" ? s.calls : 0,
59    inputTokens: typeof s.inputTokens === "number" ? s.inputTokens : 0,
60    consecutiveFailures: typeof s.consecutiveFailures === "number" ? s.consecutiveFailures : 0,
61    ...(s.quotaExhausted ? { quotaExhausted: true } : {}),
62    ...(typeof s.pausedUntil === "number" ? { pausedUntil: s.pausedUntil } : {}),
63    ...(s.pauseReason ? { pauseReason: s.pauseReason } : {}),
64  }
65}
66
67export function estimatedNeurons(state: GuardState, model: string): number {
68  return (state.inputTokens / 1_000_000) * (NEURONS_PER_M_INPUT[model] ?? NEURONS_PER_M_INPUT.clef!)
69}
70
71/** Why no call should be made now, or undefined to go ahead. */
72export function blockedReason(
73  state: GuardState,
74  opts: { now: number; model: string; dailyNeuronBudget: number },
75): ProviderFailure | undefined {
76  if (state.quotaExhausted) {
77    return { kind: "quota", message: "Workers AI daily free allocation used up; resets 00:00 UTC", latencyMs: 0 }
78  }
79  if (opts.dailyNeuronBudget > 0 && estimatedNeurons(state, opts.model) >= opts.dailyNeuronBudget) {
80    return {
81      kind: "budget",
82      message: `local daily budget of ${opts.dailyNeuronBudget} neurons reached; resets 00:00 UTC`,
83      latencyMs: 0,
84    }
85  }
86  if (state.pausedUntil !== undefined && state.pausedUntil > opts.now) {
87    const seconds = Math.ceil((state.pausedUntil - opts.now) / 1000)
88    return {
89      kind: "circuit-open",
90      message: `paused ${seconds}s after ${state.pauseReason ?? "repeated failures"}`,
91      latencyMs: 0,
92    }
93  }
94  return undefined
95}
96
97/** How long to stop calling after a failure, in ms; 0 for none. */
98export const PAUSES = {
99  /** After this many failures in a row, stop calling for a while. */
100  breakerThreshold: 3,
101  breakerMs: 5 * 60_000,
102  rateLimitedMs: 60_000,
103  /** Auth and request errors need the user to fix configuration. */
104  configMs: 30 * 60_000,
105}
106
107export function recordSuccess(state: GuardState, inputTokens: number | undefined): GuardState {
108  const next: GuardState = {
109    ...state,
110    calls: state.calls + 1,
111    inputTokens: state.inputTokens + (inputTokens ?? 0),
112    consecutiveFailures: 0,
113  }
114  delete next.pausedUntil
115  delete next.pauseReason
116  return next
117}
118
119export function recordFailure(state: GuardState, failure: ProviderFailure, now: number): GuardState {
120  // Failures that never reached Cloudflare count toward nothing.
121  if (failure.kind === "not-configured" || failure.kind === "budget" || failure.kind === "circuit-open") return state
122  const next: GuardState = { ...state, calls: state.calls + 1, consecutiveFailures: state.consecutiveFailures + 1 }
123  if (failure.kind === "quota") return { ...next, quotaExhausted: true }
124  let pause = 0
125  if (failure.kind === "rate-limited") pause = PAUSES.rateLimitedMs
126  else if (failure.kind === "auth" || failure.kind === "bad-request") pause = PAUSES.configMs
127  else if (next.consecutiveFailures >= PAUSES.breakerThreshold) pause = PAUSES.breakerMs
128  if (pause > 0) return { ...next, pausedUntil: now + pause, pauseReason: failure.kind }
129  return next
130}
131
hooks/lib/log.ts 248 lines
1// The local routing log: one JSON object per line, one file per UTC day and
2// session, under the log directory. Nothing is sent anywhere. Prompt text is
3// left out unless `log_prompts` is on; a short SHA-256 prefix lets repeated
4// prompts be recognised without storing them.
5
6import type { Decision, Deferral, Effort, Level, Source } from "./types.ts"
7
8export const LOG_VERSION = 1
9
10export type AnsweredUsage = {
11  model: string
12  inputTokens: number
13  outputTokens: number
14  cacheReadTokens: number
15  cacheWriteTokens: number
16}
17
18export type TurnRecord = {
19  v: number
20  type: "turn"
21  ts: string
22  session: string
23  turn: string
24  kind: Decision["kind"]
25  source: Source
26  promptHash?: string
27  promptChars: number
28  prompt?: string
29  provider?: string
30  recommendation?: {
31    level: Level
32    /** Probability of `level`: what the threshold compares. */
33    confidence: number
34    /** Clef's own `confidence` field (entropy-like; not a probability). */
35    clefConfidence?: number
36    probabilities: Record<Level, number>
37    score?: number
38    followUp?: number
39  }
40  latencyMs?: number
41  clefInputTokens?: number
42  proposed?: { level?: Level; model: string; effort?: Effort }
43  final?: { level?: Level; model: string; effort?: Effort }
44  adjustments: { rule: string; from: string; to: string; reason: string }[]
45  failure?: { kind: string; message: string; status?: number }
46  note?: string
47  /** What the API said answered, summed over the turn's main-loop requests. */
48  answered?: AnsweredUsage
49  /** Cache read and write of the turn's first request: what a model switch, or a return, cost. */
50  firstStep?: { cacheReadTokens: number; cacheWriteTokens: number }
51  /** How the person pays, as the policy saw it. */
52  billing?: "subscription" | "api"
53  /** A downgrade weighed against a warm cache: held, or taken after a held stretch. */
54  deferral?: Deferral
55  /** The plan's rate-limit windows at the start of the turn (subscriptions only). */
56  rateLimits?: { kind: string; percentUsed: number }[]
57  steps?: number
58  durationMs?: number
59  endReason?: string
60}
61
62export type FeedbackRecord = {
63  v: number
64  type: "feedback"
65  ts: string
66  session: string
67  turn?: string
68  verdict: "under" | "ok" | "over"
69  note?: string
70}
71
72export type LogRecord = TurnRecord | FeedbackRecord
73
74export async function promptHash(text: string): Promise<string> {
75  const bytes = new TextEncoder().encode(text)
76  const digest = await crypto.subtle.digest("SHA-256", bytes)
77  return Array.from(new Uint8Array(digest).slice(0, 8), (b) => b.toString(16).padStart(2, "0")).join("")
78}
79
80export function turnRecord(args: {
81  decision: Decision
82  session: string
83  ts: string
84  promptText: string
85  hash?: string
86  logPrompts: boolean
87  answered?: AnsweredUsage
88  firstStep?: { cacheReadTokens: number; cacheWriteTokens: number }
89  rateLimits?: { kind: string; percentUsed: number }[]
90  steps?: number
91  durationMs?: number
92  endReason?: string
93}): TurnRecord {
94  const { decision: d } = args
95  const record: TurnRecord = {
96    v: LOG_VERSION,
97    type: "turn",
98    ts: args.ts,
99    session: args.session,
100    turn: d.turnId,
101    kind: d.kind,
102    source: d.source,
103    promptChars: args.promptText.length,
104    adjustments: d.adjustments.map((a) => ({ ...a })),
105  }
106  if (args.hash) record.promptHash = args.hash
107  if (args.logPrompts) record.prompt = args.promptText
108  const rec = d.recommendation
109  if (rec) {
110    record.provider = rec.provider
111    record.recommendation = { level: rec.level, confidence: rec.confidence, probabilities: { ...rec.probabilities } }
112    if (rec.providerConfidence !== undefined) record.recommendation.clefConfidence = rec.providerConfidence
113    if (rec.score !== undefined) record.recommendation.score = rec.score
114    if (rec.contextDependent !== undefined) record.recommendation.followUp = rec.contextDependent
115    record.latencyMs = rec.latencyMs
116    if (rec.inputTokens !== undefined) record.clefInputTokens = rec.inputTokens
117  }
118  if (d.failure) {
119    record.failure = { kind: d.failure.kind, message: d.failure.message }
120    if (d.failure.status !== undefined) record.failure.status = d.failure.status
121    if (d.failure.latencyMs > 0) record.latencyMs = d.failure.latencyMs
122  }
123  if (d.proposed) record.proposed = { ...d.proposed }
124  if (d.final) record.final = { ...d.final }
125  if (d.note) record.note = d.note
126  if (args.answered) record.answered = { ...args.answered }
127  if (args.firstStep) record.firstStep = { ...args.firstStep }
128  if (d.billing) record.billing = d.billing
129  if (d.deferral) record.deferral = { ...d.deferral, wanted: { ...d.deferral.wanted } }
130  if (args.rateLimits && args.rateLimits.length > 0) record.rateLimits = args.rateLimits.map((w) => ({ ...w }))
131  if (args.steps !== undefined) record.steps = args.steps
132  if (args.durationMs !== undefined) record.durationMs = args.durationMs
133  if (args.endReason) record.endReason = args.endReason
134  return record
135}
136
137export function logFileName(ts: string, session: string): string {
138  const safe = session.replace(/[^A-Za-z0-9_-]/g, "").slice(0, 12) || "session"
139  return `routing-${ts.slice(0, 10)}-${safe}.jsonl`
140}
141
142export function parseLines(text: string): LogRecord[] {
143  const out: LogRecord[] = []
144  for (const line of text.split("\n")) {
145    if (line.trim() === "") continue
146    try {
147      const value = JSON.parse(line) as LogRecord
148      if (value && typeof value === "object" && (value.type === "turn" || value.type === "feedback")) out.push(value)
149    } catch {
150      // A torn line from a crash mid-write; skip it.
151    }
152  }
153  return out
154}
155
156export type Stats = {
157  turns: number
158  bySource: Record<string, number>
159  byModel: Record<string, number>
160  byEffort: Record<string, number>
161  byLevel: Record<string, number>
162  clefCalls: number
163  latency: { mean: number; p50: number; p95: number } | undefined
164  meanConfidence: number | undefined
165  failures: Record<string, number>
166  overrides: number
167  /** Turns held on a warm model instead of the cheaper one Clef's level asked for. */
168  cacheHolds: number
169  /** Held stretches that ended in the downgrade being taken. */
170  downgradesTaken: number
171  adjusted: number
172  /** Turns where Clef's raw level differs from the final route's level. */
173  recommendationChanged: number
174  feedback: Record<string, number>
175  clefInputTokens: number
176}
177
178function quantile(sorted: number[], q: number): number {
179  if (sorted.length === 0) return 0
180  const i = Math.min(sorted.length - 1, Math.max(0, Math.ceil(q * sorted.length) - 1))
181  return sorted[i]!
182}
183
184function bump(map: Record<string, number>, key: string) {
185  map[key] = (map[key] ?? 0) + 1
186}
187
188export function aggregate(records: readonly LogRecord[]): Stats {
189  const stats: Stats = {
190    turns: 0,
191    bySource: {},
192    byModel: {},
193    byEffort: {},
194    byLevel: {},
195    clefCalls: 0,
196    latency: undefined,
197    meanConfidence: undefined,
198    failures: {},
199    overrides: 0,
200    cacheHolds: 0,
201    downgradesTaken: 0,
202    adjusted: 0,
203    recommendationChanged: 0,
204    feedback: {},
205    clefInputTokens: 0,
206  }
207  const latencies: number[] = []
208  let confidenceSum = 0
209  let confidenceN = 0
210  for (const r of records) {
211    if (r.type === "feedback") {
212      bump(stats.feedback, r.verdict)
213      continue
214    }
215    stats.turns++
216    bump(stats.bySource, r.source)
217    const model = r.answered?.model ?? r.final?.model ?? "session default"
218    bump(stats.byModel, model.replace(/-\d{8}$/, ""))
219    bump(stats.byEffort, r.final ? (r.final.effort ?? "default") : "untouched")
220    if (r.final?.level) bump(stats.byLevel, r.final.level)
221    if (r.recommendation || r.failure) {
222      if (r.failure?.kind !== "not-configured" && r.failure?.kind !== "budget" && r.failure?.kind !== "circuit-open") stats.clefCalls++
223    }
224    if (r.recommendation) {
225      if (r.latencyMs !== undefined) latencies.push(r.latencyMs)
226      confidenceSum += r.recommendation.confidence
227      confidenceN++
228      if (r.final?.level !== r.recommendation.level || r.final?.model !== r.proposed?.model) stats.recommendationChanged++
229    }
230    if (r.failure) bump(stats.failures, r.failure.kind)
231    if (r.source === "override" || r.source === "pin") stats.overrides++
232    if (r.adjustments.some((a) => a.rule === "cache-hold")) stats.cacheHolds++
233    if (r.deferral && !r.deferral.held) stats.downgradesTaken++
234    if (r.adjustments.length > 0) stats.adjusted++
235    stats.clefInputTokens += r.clefInputTokens ?? 0
236  }
237  if (latencies.length > 0) {
238    const sorted = [...latencies].sort((a, b) => a - b)
239    stats.latency = {
240      mean: Math.round(sorted.reduce((s, x) => s + x, 0) / sorted.length),
241      p50: quantile(sorted, 0.5),
242      p95: quantile(sorted, 0.95),
243    }
244  }
245  if (confidenceN > 0) stats.meanConfidence = confidenceSum / confidenceN
246  return stats
247}
248
hooks/lib/models.ts 134 lines
1// What this router knows about Claude models: how an alias resolves, which
2// effort levels a model takes, how big its window is, and how it ranks.
3//
4// Source: Claude Code model configuration docs (Claude Code 2.1.289,
5// October 2026). `turn.step` does not resolve aliases (a request for
6// "haiku" fails with unrecognized_model), so every route is resolved here to
7// a full ID before it is sent.
8
9import { EFFORTS, type Effort } from "./types.ts"
10
11export type Family = "haiku" | "sonnet" | "opus" | "fable"
12
13/** What each alias resolves to on the Anthropic API, as Claude Code does. */
14export const FIRST_PARTY_ALIASES: Record<Family, string> = {
15  haiku: "claude-haiku-4-5",
16  sonnet: "claude-sonnet-5-5",
17  opus: "claude-opus-5-5",
18  fable: "claude-fable-5-1",
19}
20
21/**
22 * The environment the resolution depends on: Claude Code's own
23 * ANTHROPIC_DEFAULT_*_MODEL pins, and whether a third-party provider is in use
24 * (where first-party IDs do not exist).
25 */
26export type ModelEnv = {
27  defaults: Partial<Record<Family, string>>
28  thirdParty: boolean
29  /** Prompt-cache beta features off (CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS). */
30  betasDisabled: boolean
31}
32
33export const FIRST_PARTY_ENV: ModelEnv = { defaults: {}, thirdParty: false, betasDisabled: false }
34
35export function isAlias(model: string): model is Family {
36  return model === "haiku" || model === "sonnet" || model === "opus" || model === "fable"
37}
38
39/** Resolves an alias or ID to a concrete model ID, or undefined if it cannot be. */
40export function resolveModel(model: string, env: ModelEnv): string | undefined {
41  const m = model.trim()
42  if (m === "") return undefined
43  if (!isAlias(m)) return m
44  const pinned = env.defaults[m]
45  if (pinned) return pinned
46  // On Bedrock, Vertex or Foundry, a first-party ID would fail.
47  return env.thirdParty ? undefined : FIRST_PARTY_ALIASES[m]
48}
49
50export function familyOf(modelId: string): Family | undefined {
51  const id = modelId.toLowerCase()
52  if (id.includes("haiku")) return "haiku"
53  if (id.includes("sonnet")) return "sonnet"
54  if (id.includes("fable")) return "fable"
55  if (id.includes("opus")) return "opus"
56  return undefined
57}
58
59const RANK: Record<Family, number> = { haiku: 0, sonnet: 1, opus: 2, fable: 3 }
60
61/** Capability/price rank for comparing two models; undefined if unknown. */
62export function rankOf(modelId: string): number | undefined {
63  const f = familyOf(modelId)
64  return f === undefined ? undefined : RANK[f]
65}
66
67/** Display name for the status line: "Sonnet", "Opus", or the raw ID. */
68export function displayName(modelId: string): string {
69  const f = familyOf(modelId)
70  return f === undefined ? modelId : f[0]!.toUpperCase() + f.slice(1)
71}
72
73/**
74 * The effort levels a model accepts; null when it takes no effort at all,
75 * undefined when the model is unknown (send what was asked; Claude Code
76 * clamps it).
77 */
78export function effortsFor(modelId: string): readonly Effort[] | null | undefined {
79  const id = modelId.toLowerCase()
80  if (id.includes("haiku")) return null
81  if (/(opus|sonnet)-4-6/.test(id)) return ["low", "medium", "high", "max"]
82  if (/fable|opus-5|sonnet-5|opus-4-[78]/.test(id)) return EFFORTS
83  return undefined
84}
85
86/**
87 * The highest supported level at or below the one asked, which is what
88 * Claude Code itself does; undefined when the model takes no effort.
89 */
90export function clampEffort(modelId: string, effort: Effort | undefined): Effort | undefined {
91  if (effort === undefined) return undefined
92  const supported = effortsFor(modelId)
93  if (supported === null) return undefined
94  if (supported === undefined) return effort
95  for (let i = EFFORTS.indexOf(effort); i >= 0; i--) {
96    const level = EFFORTS[i]!
97    if (supported.includes(level)) return level
98  }
99  return supported[0]
100}
101
102export function capEffort(effort: Effort | undefined, cap: Effort | undefined): Effort | undefined {
103  if (effort === undefined || cap === undefined) return effort
104  return EFFORTS.indexOf(effort) > EFFORTS.indexOf(cap) ? cap : effort
105}
106
107/** Context window in tokens where it is known to be smaller than 1M. */
108export function windowOf(modelId: string): number | undefined {
109  const id = modelId.toLowerCase()
110  if (id.includes("haiku")) return 200_000
111  return undefined
112}
113
114/**
115 * Whether changing effort between requests keeps the prompt cache. Per the
116 * Claude Code prompt-caching docs, it does on Opus 5.5, Sonnet 5.5 and
117 * Fable 5.1 with an API key or subscription, and not on Bedrock, Vertex, a
118 * gateway, or with experimental betas disabled. Elsewhere it is a full miss.
119 */
120export function effortChangeKeepsCache(modelId: string, env: ModelEnv): boolean {
121  if (env.thirdParty || env.betasDisabled) return false
122  return /claude-(opus|sonnet)-5-5|claude-fable-5-1/.test(modelId.toLowerCase())
123}
124
125/** Compares IDs ignoring a date suffix and a [1m] marker. */
126export function sameModel(a: string, b: string): boolean {
127  const norm = (id: string) => id.toLowerCase().replace(/\[1m\]$/, "").replace(/-\d{8}$/, "")
128  return norm(a) === norm(b)
129}
130
131export function isEffort(value: string): value is Effort {
132  return (EFFORTS as readonly string[]).includes(value)
133}
134
hooks/lib/overrides.ts 183 lines
1// Explicit user intent, parsed deterministically: the one-turn `+target`
2// prompt prefix, `/clef` command targets, and turns that only continue the
3// previous one. No model is involved in any of this.
4
5import { splitTarget } from "./config.ts"
6import { isAlias, isEffort } from "./models.ts"
7import { LEVELS, type Level, type Target, type TurnKind } from "./types.ts"
8
9/**
10 * Parses a target: a profile (`hard`), an alias or model ID with an optional
11 * effort (`opus`, `opus:max`, `claude-sonnet-5-5:low`), or an effort alone
12 * (`:high`). Undefined when the text is none of those.
13 */
14export function parseTarget(text: string): Target | undefined {
15  const t = text.trim()
16  if (t === "") return undefined
17  if (t.startsWith(":")) {
18    const effort = t.slice(1).toLowerCase()
19    return isEffort(effort) ? { effort } : undefined
20  }
21  const { model, effort, badEffort } = splitTarget(t)
22  if (badEffort || model === "") return undefined
23  const lower = model.toLowerCase()
24  const target: Target = {}
25  if ((LEVELS as readonly string[]).includes(lower)) target.level = lower as Level
26  else if (isAlias(lower)) target.model = lower
27  else if (/^claude-[a-z0-9.-]+(\[1m\])?$/i.test(model) || /anthropic\./i.test(model)) target.model = model
28  else return undefined
29  if (effort) target.effort = effort
30  return target
31}
32
33export type PrefixResult = { text: string; override?: Target | "off" }
34
35const PREFIX = /^\+(\S+)(?:\s+|$)/
36
37/**
38 * A prompt that starts with `+target ` routes that one turn to the target and
39 * reaches Claude without the prefix; `+off ` runs the turn as Claude Code
40 * would. Anything else (`+1`, `+x`) is left untouched.
41 */
42export function parsePrefix(text: string): PrefixResult {
43  const match = PREFIX.exec(text)
44  if (!match) return { text }
45  const token = match[1]!
46  const rest = text.slice(match[0].length)
47  if (token.toLowerCase() === "off" || token.toLowerCase() === "noroute") return { text: rest, override: "off" }
48  const target = parseTarget(token)
49  if (!target) return { text }
50  return { text: rest, override: target }
51}
52
53/** Words a go-ahead is made of ("yes, do it", "ok go ahead", "lgtm, ship it"). */
54const GO_AHEAD_WORDS = new Set([
55  "y", "ya", "yes", "yep", "yeah", "yup", "ok", "okay", "k", "sure", "alright", "fine", "cool", "great", "perfect",
56  "go", "ahead", "for", "it", "do", "that", "this", "continue", "proceed", "carry", "on", "keep", "going", "next",
57  "lgtm", "looks", "sounds", "good", "ship", "please", "approved", "approve", "confirm", "confirmed", "thanks",
58])
59/** A go-ahead says yes to something; "it", "on" or "good" alone do not. */
60const GO_AHEAD_ANCHORS = new Set([
61  "y", "ya", "yes", "yep", "yeah", "yup", "ok", "okay", "k", "sure", "alright", "go", "do", "continue", "proceed",
62  "carry", "keep", "next", "lgtm", "ship", "approved", "approve", "confirm", "confirmed", "sounds", "looks",
63])
64
65/** Classifies a turn's text before anything is asked of Clef. */
66export function turnKind(text: string): TurnKind {
67  const t = text.trim()
68  if (t === "") return "empty"
69  if (t.startsWith("<task-notification>")) return "notification"
70  if (t.length <= 40) {
71    const words = t.toLowerCase().replace(/[.,!;:'"]+/g, " ").split(/\s+/).filter(Boolean)
72    if (words.length > 0 && words.length <= 6 && words.every((w) => GO_AHEAD_WORDS.has(w)) && words.some((w) => GO_AHEAD_ANCHORS.has(w)))
73      return "go-ahead"
74  }
75  return "prompt"
76}
77
78export type ClefCommand =
79  | { kind: "status" }
80  | { kind: "history" }
81  | { kind: "stats"; days: number }
82  | { kind: "profiles" }
83  | { kind: "test"; prompt: string }
84  | { kind: "auto" }
85  | { kind: "on" }
86  | { kind: "off" }
87  | { kind: "pin"; target: Target }
88  | { kind: "feedback"; verdict: "under" | "ok" | "over"; note?: string }
89  | { kind: "help" }
90  | { kind: "error"; message: string }
91
92const FEEDBACK: Record<string, "under" | "ok" | "over"> = {
93  under: "under",
94  underpowered: "under",
95  weak: "under",
96  "too-weak": "under",
97  ok: "ok",
98  right: "ok",
99  good: "ok",
100  over: "over",
101  overpowered: "over",
102  strong: "over",
103  "too-strong": "over",
104}
105
106export function parseCommand(args: string): ClefCommand {
107  const trimmed = args.trim()
108  const space = trimmed.search(/\s/)
109  const head = (space === -1 ? trimmed : trimmed.slice(0, space)).toLowerCase()
110  const rest = space === -1 ? "" : trimmed.slice(space + 1).trim()
111  switch (head) {
112    case "":
113    case "status":
114      return { kind: "status" }
115    case "history":
116    case "log":
117      return { kind: "history" }
118    case "stats": {
119      const days = rest === "" ? 7 : Number(rest)
120      return Number.isInteger(days) && days > 0 && days <= 366
121        ? { kind: "stats", days }
122        : { kind: "error", message: "usage: /clef stats [days]" }
123    }
124    case "profiles":
125      return { kind: "profiles" }
126    case "test":
127      return rest === "" ? { kind: "error", message: "usage: /clef test <prompt>" } : { kind: "test", prompt: rest }
128    case "auto":
129    case "unpin":
130      return { kind: "auto" }
131    case "on":
132      return { kind: "on" }
133    case "off":
134      return { kind: "off" }
135    case "pin": {
136      const target = parseTarget(rest)
137      return target
138        ? { kind: "pin", target }
139        : { kind: "error", message: `usage: /clef pin <${LEVELS.join("|")}|haiku|sonnet|opus|fable|model-id>[:effort] or /clef pin :<effort>` }
140    }
141    case "feedback": {
142      const space2 = rest.search(/\s/)
143      const word = (space2 === -1 ? rest : rest.slice(0, space2)).toLowerCase()
144      const verdict = FEEDBACK[word]
145      if (!verdict) return { kind: "error", message: "usage: /clef feedback under|ok|over [note]" }
146      const note = space2 === -1 ? "" : rest.slice(space2 + 1).trim()
147      return note ? { kind: "feedback", verdict, note } : { kind: "feedback", verdict }
148    }
149    case "help":
150      return { kind: "help" }
151    default: {
152      // `/clef opus:high` as shorthand for `/clef pin opus:high`.
153      const target = parseTarget(trimmed)
154      return target ? { kind: "pin", target } : { kind: "error", message: `unknown subcommand "${head}"; try /clef help` }
155    }
156  }
157}
158
159export type NativeEffort = {
160  /** The effort Claude Code sent when the router started watching. */
161  baseline?: string | number
162  /** An effort the person set with /effort since, in force until /clef auto. */
163  native?: string | number
164}
165
166/**
167 * Tracks the effort Claude Code itself would send, seen at each turn's first
168 * request. A change from the baseline is the person's /effort (or a skill's
169 * `effort` for one turn): it is honoured, and dropped again when the engine's
170 * effort returns to the baseline, so a one-turn skill does not stick.
171 */
172export function trackNativeEffort(prev: NativeEffort, engine: string | number | undefined): NativeEffort & { change?: "set" | "cleared" } {
173  const keep = (): NativeEffort => ({
174    ...(prev.baseline !== undefined ? { baseline: prev.baseline } : {}),
175    ...(prev.native !== undefined ? { native: prev.native } : {}),
176  })
177  if (engine === undefined) return keep()
178  if (prev.baseline === undefined) return { baseline: engine }
179  if (engine === prev.native) return keep()
180  if (engine === prev.baseline) return prev.native === undefined ? keep() : { baseline: prev.baseline, change: "cleared" }
181  return { baseline: prev.baseline, native: engine, change: "set" }
182}
183
hooks/lib/policy.ts 438 lines
1// The routing policy: a pure function from what is known at the start of a
2// turn to the route its requests will use, with every change it made on the
3// way recorded. Clef makes the semantic judgment (how hard is this?); this
4// file makes the deterministic ones (what is allowed, what is safe, what is
5// worth the cache), in this order:
6//
7//   1. Routing off (config, +off; /clef off and a mid-session /model change, unless +model/+profile)
8//   2. An explicit choice: a +target prefix this turn, then a /clef pin
9//   3. A continuation (go-ahead, task notification, empty) reuses the last route
10//   4. Clef's recommendation, or the fallback when Clef did not answer
11//   5. Low confidence → the configured confidence policy
12//   6. A follow-up never routes below the route it follows
13//   7. A profile whose model is unavailable → the nearest available one, upward first
14//   8. A context too big for the model's window → the nearest profile that fits
15//   9. A downgrade off a warm prompt cache → held until staying has cost what
16//      switching costs, in the unit the person's billing makes scarce
17//      (docs/adr/0001-downgrade-timing-by-billing-mode.md)
18//  10. Effort is capped and clamped to what the model takes
19
20import type { Config } from "./config.ts"
21import {
22  capEffort,
23  clampEffort,
24  displayName,
25  effortChangeKeepsCache,
26  effortsFor,
27  rankOf,
28  resolveModel,
29  sameModel,
30  windowOf,
31  type ModelEnv,
32} from "./models.ts"
33import { dollars, isOneHour, priceOf, switchCost, type Billing } from "./pricing.ts"
34import {
35  EFFORTS,
36  LEVELS,
37  type Adjustment,
38  type Decision,
39  type Effort,
40  type Level,
41  type ProviderResult,
42  type Route,
43  type Source,
44  type Target,
45  type TurnKind,
46} from "./types.ts"
47
48export type RouterMode = "auto" | "off" | "paused-native"
49
50/** The last request the main loop sent, which is what the prompt cache holds. */
51export type CacheState = {
52  model: string
53  effort?: Effort
54  /** Epoch ms the response finished. */
55  at: number
56  /** Input + cache read + cache write tokens of that request. */
57  promptTokens: number
58  /**
59   * Whether the API reported any cache read or write for it. False where
60   * nothing is cached (a gateway that strips cache markers): there is then no
61   * cache to keep. Absent in state saved by earlier versions: assumed true.
62   */
63  caching?: boolean
64}
65
66/** A stretch of held downgrades: what staying on the warm model has cost so far. */
67export type HoldState = {
68  /** The model held, whose cache is warm. */
69  model: string
70  /** What Clef's level asked for on the last held turn; a continuation weighs it again. */
71  wanted: Route
72  /** List-price dollars staying has cost over the stretch's completed turns. */
73  spent: number
74  turns: number
75}
76
77export type PolicySession = {
78  mode: RouterMode
79  pin?: Target
80  last?: Route
81  cache?: CacheState
82  hold?: HoldState
83  /** Models that failed when routed to this session. */
84  unavailable: readonly string[]
85}
86
87export type PolicyInput = {
88  turnId: string
89  kind: TurnKind
90  config: Config
91  modelEnv: ModelEnv
92  session: PolicySession
93  /** A one-turn +target prefix; "off" leaves this turn alone. */
94  override?: Target | "off"
95  /** Clef's answer; absent when Clef was not asked. */
96  result?: ProviderResult
97  /** Tokens in the conversation now, when known. */
98  contextTokens?: number
99  now: number
100  cacheTtlMs: number
101  /** How the person pays: what a held downgrade costs them. */
102  billing: Billing
103}
104
105/** Headroom kept under a model's window for the reply and tool results. */
106export const WINDOW_HEADROOM = 20_000
107
108const idx = (level: Level) => LEVELS.indexOf(level)
109const atIdx = (i: number): Level => LEVELS[Math.max(0, Math.min(LEVELS.length - 1, i))]!
110const higher = (a: Level, b: Level): Level => (idx(a) >= idx(b) ? a : b)
111
112export function describe(route: Route | undefined): string {
113  if (!route) return "session default"
114  return route.effort ? `${displayName(route.model)} · ${route.effort}` : displayName(route.model)
115}
116
117/**
118 * Whether the policy should ask Clef at all for this turn. Explicit choices,
119 * continuations and a disabled router make no network call.
120 */
121export function needsClef(input: Omit<PolicyInput, "result" | "now" | "cacheTtlMs" | "contextTokens" | "billing">): boolean {
122  const { config, session, override, kind } = input
123  if (!config.enabled || override === "off") return false
124  if (override && (override.level || override.model)) return false
125  if (session.mode !== "auto") return false
126  if (session.pin && (session.pin.level || session.pin.model)) return false
127  if (kind !== "prompt" && session.last) return false
128  if ((kind === "notification" || kind === "empty") && !session.last) return false
129  return true
130}
131
132function isUnavailable(model: string, session: PolicySession): boolean {
133  return session.unavailable.some((u) => sameModel(u, model))
134}
135
136/** The route for a profile, or the nearest usable profile (upward first). */
137function routeForLevel(
138  level: Level,
139  input: PolicyInput,
140  adjustments: Adjustment[],
141): Route | undefined {
142  const order = [idx(level), ...LEVELS.map((_, i) => i).filter((i) => i > idx(level)), ...LEVELS.map((_, i) => i).filter((i) => i < idx(level)).reverse()]
143  for (const i of order) {
144    const lv = atIdx(i)
145    const spec = input.config.profiles[lv]
146    const model = resolveModel(spec.model, input.modelEnv)
147    if (model === undefined || isUnavailable(model, input.session)) continue
148    const route: Route = { level: lv, model }
149    if (spec.effort) route.effort = spec.effort
150    if (lv !== level) {
151      adjustments.push({
152        rule: "unavailable",
153        from: level,
154        to: lv,
155        reason: `profile ${level} (${input.config.profiles[level].model}) has no usable model here`,
156      })
157    }
158    return route
159  }
160  return undefined
161}
162
163function routeForTarget(target: Target, input: PolicyInput, adjustments: Adjustment[]): Route | undefined {
164  if (target.level) {
165    const route = routeForLevel(target.level, input, adjustments)
166    if (route && target.effort) route.effort = target.effort
167    return route
168  }
169  if (target.model) {
170    const model = resolveModel(target.model, input.modelEnv)
171    if (model === undefined) return undefined
172    return target.effort ? { model, effort: target.effort } : { model }
173  }
174  return undefined
175}
176
177/**
178 * The effort a turn held on its warm model runs at. Where an effort change
179 * keeps the cache, it still comes down to what Clef's level asked for, and a
180 * level whose model takes no effort (Haiku) gets the held model's lowest: the
181 * turn is held for the cache, not for more thinking. Elsewhere the cached
182 * effort stays, since changing it would cost the cache the hold is keeping.
183 */
184function heldEffort(route: Route, cache: CacheState, env: ModelEnv): Effort | undefined {
185  if (!effortChangeKeepsCache(cache.model, env)) return cache.effort
186  if (route.effort) return route.effort
187  if (effortsFor(route.model) === null) return effortsFor(cache.model)?.[0] ?? "low"
188  return undefined
189}
190
191/** The more capable of the two most probable levels. */
192function upperOfTopTwo(probabilities: Record<Level, number>, top: Level): Level {
193  let second: Level | undefined
194  for (const level of LEVELS) {
195    if (level === top) continue
196    if (second === undefined || probabilities[level] >= probabilities[second]) second = level
197  }
198  return second === undefined ? top : higher(top, second)
199}
200
201export function decide(input: PolicyInput): Decision {
202  const { config, session, turnId, kind } = input
203  const adjustments: Adjustment[] = []
204  const base = { turnId, kind, adjustments }
205  const disabled = (source: Source, note: string): Decision => ({ ...base, source, note })
206
207  // 1. Off. A one-turn +model or +profile still applies while the session is
208  // off or paused: it is the most explicit request there is.
209  if (!config.enabled) return disabled("disabled", "routing disabled in plugin config")
210  if (input.override === "off") return disabled("disabled", "+off: this turn runs as Claude Code would")
211  const explicitTurn = input.override !== undefined && (input.override.level !== undefined || input.override.model !== undefined)
212  if (session.mode === "off" && !explicitTurn) return disabled("disabled", "routing off for this session (/clef on)")
213  if (session.mode === "paused-native" && !explicitTurn)
214    return disabled("native", "paused: you changed /model (/clef auto to resume)")
215
216  let source: Source
217  let route: Route | undefined
218  let proposed: Route | undefined
219  let note: string | undefined
220  const decision: Decision = { ...base, source: "clef" }
221
222  // 2. Explicit choice.
223  const explicit = input.override ?? (session.mode === "auto" ? session.pin : undefined)
224  const explicitSource: Source = input.override ? "override" : "pin"
225  if (explicit && (explicit.level || explicit.model)) {
226    source = explicitSource
227    route = routeForTarget(explicit, input, adjustments)
228    if (!route) return disabled(explicitSource, `cannot resolve ${explicit.model ?? explicit.level} here`)
229    proposed = { ...route }
230  } else if (kind !== "prompt" && session.last) {
231    // 3. Continuation. One that continues a held turn weighs the deferred
232    // downgrade again (step 9), so a run of go-aheads cannot outlast it.
233    source = "continuation"
234    const hold = session.hold && sameModel(session.last.model, session.hold.model) ? session.hold : undefined
235    route = { ...(hold ? hold.wanted : session.last) }
236    proposed = { ...route }
237    note = kind === "go-ahead" ? "go-ahead continues the last route" : `${kind} continues the last route`
238  } else if ((kind === "notification" || kind === "empty") && !session.last) {
239    return disabled("continuation", "nothing to continue yet")
240  } else {
241    // 4. Clef, or the fallback.
242    const result = input.result
243    let level: Level
244    if (result?.ok) {
245      source = "clef"
246      const rec = result.recommendation
247      decision.recommendation = rec
248      level = rec.level
249      proposed = routeForLevel(level, input, [])
250
251      // 5. Low confidence.
252      if (rec.confidence < config.confidenceThreshold) {
253        let adjusted: Level = level
254        switch (config.lowConfidencePolicy) {
255          case "upper-of-top-two":
256            adjusted = upperOfTopTwo(rec.probabilities, level)
257            break
258          case "bump":
259            adjusted = atIdx(idx(level) + 1)
260            break
261          case "hold":
262            adjusted = session.last?.level ?? config.fallbackLevel
263            break
264          case "fallback":
265            adjusted = config.fallbackLevel
266            break
267          case "obey":
268            break
269        }
270        if (adjusted !== level) {
271          adjustments.push({
272            rule: "low-confidence",
273            from: level,
274            to: adjusted,
275            reason: `confidence ${pct(rec.confidence)} < ${pct(config.confidenceThreshold)} (${config.lowConfidencePolicy})`,
276          })
277          level = adjusted
278        }
279      }
280
281      // 6. Follow-up floor.
282      const lastLevel = session.last?.level
283      if (
284        rec.contextDependent !== undefined &&
285        rec.contextDependent >= config.followUpThreshold &&
286        lastLevel !== undefined &&
287        idx(lastLevel) > idx(level)
288      ) {
289        adjustments.push({
290          rule: "context-dependent",
291          from: level,
292          to: lastLevel,
293          reason: `follow-up (${pct(rec.contextDependent)}) to a ${lastLevel} turn`,
294        })
295        level = lastLevel
296      }
297    } else {
298      source = "fallback"
299      if (result && !result.ok) decision.failure = result.failure
300      const lastLevel = session.last?.level
301      level = lastLevel ? higher(lastLevel, config.fallbackLevel) : config.fallbackLevel
302      note = `${result && !result.ok ? result.failure.kind : "no answer"}: using ${level}`
303    }
304
305    // 7. Availability.
306    route = routeForLevel(level, input, adjustments)
307    if (!route) return { ...decision, ...disabled(source, "no configured profile resolves to a usable model here") }
308    if (source === "fallback") proposed = { ...route }
309  }
310
311  // 8. Context window.
312  const tokens = input.contextTokens ?? session.cache?.promptTokens
313  if (tokens !== undefined) {
314    const fits = (model: string) => {
315      const w = windowOf(model)
316      return w === undefined || tokens + WINDOW_HEADROOM <= w
317    }
318    if (!fits(route.model)) {
319      const start = route.level ? idx(route.level) : 0
320      let moved: Route | undefined
321      for (let i = start + 1; i < LEVELS.length && !moved; i++) {
322        const candidate = routeForLevel(atIdx(i), input, [])
323        if (candidate && fits(candidate.model)) moved = candidate
324      }
325      if (moved) {
326        adjustments.push({
327          rule: "context-window",
328          from: describe(route),
329          to: describe(moved),
330          reason: `${kTokens(tokens)} context does not fit ${displayName(route.model)}'s window`,
331        })
332        route = moved
333      }
334    }
335  }
336
337  // 9. A downgrade off a warm cache, for routes the router chose itself. Each
338  // model has its own prompt cache, so moving a conversation writes all of it
339  // again. That pays off over a stretch of cheaper turns, not over one: the
340  // downgrade is held while what staying has cost so far is less than what the
341  // switch costs now (times downgrade_patience), then taken. Short dips stay
342  // put; long stretches move. What "cost" means depends on the billing.
343  const cache = session.cache
344  const warm = cache !== undefined && input.now - cache.at < input.cacheTtlMs
345  if ((source === "clef" || source === "continuation") && cache && warm && cache.caching !== false && !isUnavailable(cache.model, session)) {
346    const fromRank = rankOf(cache.model)
347    const toRank = rankOf(route.model)
348    const modelDown = fromRank !== undefined && toRank !== undefined && toRank < fromRank
349    const effortDown =
350      !modelDown &&
351      sameModel(route.model, cache.model) &&
352      route.effort !== undefined &&
353      cache.effort !== undefined &&
354      EFFORTS.indexOf(route.effort) < EFFORTS.indexOf(cache.effort) &&
355      !effortChangeKeepsCache(cache.model, input.modelEnv)
356    // An effort change that breaks the cache rewrites it on the same model.
357    const toPrice = priceOf(modelDown ? route.model : cache.model)
358    if ((modelDown || effortDown) && toPrice) {
359      const tokens = input.contextTokens ?? cache.promptTokens
360      const cost = switchCost(toPrice, tokens, isOneHour(input.cacheTtlMs))
361      const prior = session.hold && sameModel(session.hold.model, cache.model) ? session.hold : undefined
362      const spent = prior?.spent ?? 0
363      const turns = prior?.turns ?? 0
364      const wanted: Route = { ...route }
365      const billing = input.billing
366      if (spent >= cost * config.downgradePatience) {
367        if (prior) decision.deferral = { wanted, from: cache.model, billing, spent, cost, turns, held: false }
368      } else {
369        const held: Route = { model: cache.model }
370        if (route.level) held.level = route.level
371        const effort = modelDown ? heldEffort(route, cache, input.modelEnv) : cache.effort
372        if (effort) held.effort = effort
373        const sofar = `staying has cost ${dollars(spent)} so far (${billing}, list prices)`
374        adjustments.push({
375          rule: "cache-hold",
376          from: describe(route),
377          to: describe(held),
378          reason: modelDown
379            ? `${kTokens(tokens)} context is cached on ${displayName(cache.model)}; moving to ${displayName(route.model)} writes it again (${dollars(cost)}); ${sofar}`
380            : `changing effort on ${displayName(cache.model)} here rewrites its ${kTokens(tokens)} cached context (${dollars(cost)}); ${sofar}`,
381        })
382        decision.deferral = { wanted, from: cache.model, billing, spent, cost, turns: turns + 1, held: true }
383        route = held
384      }
385    }
386  }
387
388  // An effort-only pin applies to whatever model was chosen.
389  if (session.pin?.effort && !session.pin.level && !session.pin.model && !(input.override && input.override.effort)) {
390    if (route.effort !== session.pin.effort) {
391      adjustments.push({ rule: "pinned-effort", from: route.effort ?? "default", to: session.pin.effort, reason: "/clef pin" })
392      route = { ...route, effort: session.pin.effort }
393    }
394  }
395  if (input.override && !input.override.level && !input.override.model && input.override.effort) {
396    adjustments.push({ rule: "pinned-effort", from: route.effort ?? "default", to: input.override.effort, reason: "+ prefix" })
397    route = { ...route, effort: input.override.effort }
398  }
399
400  // 10. Effort cap and clamp.
401  if (route.effort) {
402    const capped = capEffort(route.effort, config.maxEffort)
403    if (capped !== route.effort) {
404      adjustments.push({ rule: "effort-cap", from: route.effort, to: capped ?? "none", reason: `max_effort is ${config.maxEffort}` })
405    }
406    const clamped = clampEffort(route.model, capped)
407    if (clamped !== capped) {
408      adjustments.push({
409        rule: "effort-clamp",
410        from: capped ?? "none",
411        to: clamped ?? "none",
412        reason: `${displayName(route.model)} ${clamped ? `tops out at ${clamped}` : "takes no effort setting"}`,
413      })
414    }
415    route = { ...route }
416    if (clamped) route.effort = clamped
417    else delete route.effort
418  }
419
420  const out: Decision = { ...decision, source, final: route, billing: input.billing }
421  if (proposed) out.proposed = proposed
422  if (note) out.note = note
423  return out
424}
425
426export function pct(p: number): string {
427  return `${Math.round(p * 100)}%`
428}
429
430export function kTokens(n: number): string {
431  return n >= 1000 ? `${Math.round(n / 1000)}k` : String(n)
432}
433
434/** Whether policy changed what Clef (or the override) proposed. */
435export function wasAdjusted(decision: Decision): boolean {
436  return decision.adjustments.length > 0
437}
438
hooks/lib/pricing.ts 115 lines
1// What a turn costs, and what a model switch costs, in list-price dollars.
2//
3// These figures are a yardstick for comparing staying on a model with leaving
4// it, not a bill: cloud providers price differently (regional endpoints cost
5// 10% more), negotiated rates differ, and a subscription does not bill per
6// token at all. What matters to the policy is the ratio between the two sides,
7// which holds wherever the price list is proportional to Anthropic's. See
8// docs/COSTS.md and docs/adr/0001-downgrade-timing-by-billing-mode.md.
9//
10// Source: Anthropic API pricing (platform.claude.com/docs/en/about-claude/pricing),
11// October 2026. Prices are dollars per million tokens.
12
13import { familyOf, type Family } from "./models.ts"
14
15/** How the person pays for Claude, which decides what a held turn costs them. */
16export const BILLINGS = ["subscription", "api"] as const
17export type Billing = (typeof BILLINGS)[number]
18
19export type Price = {
20  input: number
21  write5m: number
22  write1h: number
23  /** Cache hits and refreshes. */
24  read: number
25  output: number
26}
27
28/** Cache writes are 1.25× input (5 minutes) or 2× (1 hour); the read multiplier varies by model. */
29const price = (input: number, read: number, output: number): Price => ({ input, write5m: input * 1.25, write1h: input * 2, read, output })
30
31// Most specific first: "opus-5-5" before "opus-5", "opus-4-5" before "opus-4".
32const TABLE: readonly [RegExp, Price][] = [
33  [/(fable|mythos)-5-1/, price(10, 0.25, 50)],
34  [/(fable|mythos)-5/, price(10, 1, 50)],
35  [/opus-5-5/, price(4, 0.2, 20)],
36  [/opus-(5|4-[5-8])/, price(5, 0.5, 25)],
37  [/opus-4/, price(15, 1.5, 75)],
38  [/sonnet-5/, price(2, 0.2, 10)],
39  [/sonnet-4/, price(3, 0.3, 15)],
40  [/haiku-4-5/, price(1, 0.1, 5)],
41  [/haiku-3-5/, price(0.8, 0.08, 4)],
42]
43
44/** An unrecognised ID of a known family is priced as that family's current model. */
45const BY_FAMILY: Record<Family, Price> = {
46  haiku: price(1, 0.1, 5),
47  sonnet: price(2, 0.2, 10),
48  opus: price(4, 0.2, 20),
49  fable: price(10, 0.25, 50),
50}
51
52export function priceOf(modelId: string): Price | undefined {
53  const id = modelId.toLowerCase()
54  for (const [pattern, p] of TABLE) if (pattern.test(id)) return p
55  const family = familyOf(id)
56  return family === undefined ? undefined : BY_FAMILY[family]
57}
58
59/** The token counts the API reports for a turn, summed over its requests. */
60export type TurnTokens = {
61  inputTokens: number
62  outputTokens: number
63  cacheReadTokens: number
64  cacheWriteTokens: number
65}
66
67/** A cache written with a TTL over five minutes is billed at the one-hour rate. */
68export function isOneHour(ttlMs: number): boolean {
69  return ttlMs > 5 * 60_000
70}
71
72/** List-price dollars for a turn's tokens on a model. */
73export function turnCost(p: Price, t: TurnTokens, oneHour: boolean): number {
74  const write = oneHour ? p.write1h : p.write5m
75  return (t.inputTokens * p.input + t.cacheWriteTokens * write + t.cacheReadTokens * p.read + t.outputTokens * p.output) / 1e6
76}
77
78/**
79 * What a switch costs: the whole context written to the new model's cache
80 * (each model has its own), at that model's write price. The way back is not
81 * counted: a later upgrade is a turn that needs the stronger model, and is
82 * never held for price.
83 */
84export function switchCost(target: Price, contextTokens: number, oneHour: boolean): number {
85  return (contextTokens * (oneHour ? target.write1h : target.write5m)) / 1e6
86}
87
88/**
89 * What one held turn cost the person, beyond what the cheaper route would
90 * have cost, in the unit their billing makes scarce:
91 *
92 * - `api`: dollars. The turn's tokens priced on the held model, minus the same
93 *   tokens on the wanted model as if its cache were warm. On models whose
94 *   cache reads cost the same (Opus 5.5 and Sonnet 5.5), only writes and
95 *   output differ, so staying costs little.
96 * - `subscription`: plan usage. All of the held turn counts, since it is drawn
97 *   from the stronger model's allowance, which the plan meters separately and
98 *   which a turn on the cheaper model would not touch.
99 *
100 * On `api`, keeping a higher effort on the same model (where an effort change
101 * breaks the cache) costs nothing by this measure, so it stays held while the
102 * cache is warm: no price list says what a lower effort saves.
103 */
104export function stayCost(billing: Billing, held: Price, wanted: Price, t: TurnTokens, oneHour: boolean): number {
105  const onHeld = turnCost(held, t, oneHour)
106  if (billing === "subscription") return onHeld
107  return Math.max(0, onHeld - turnCost(wanted, t, oneHour))
108}
109
110export function dollars(n: number): string {
111  if (n === 0) return "$0"
112  if (n < 0.01) return "<$0.01"
113  return `$${n < 10 ? n.toFixed(2) : n.toFixed(0)}`
114}
115
hooks/lib/rubric.ts 108 lines
1// The questions Clef is asked. This is the calibration surface: change the
2// wording here (or point `rubric_file` at a JSON file with the same shape)
3// and nothing else needs to change. `npm run calibrate` shows the effect.
4//
5// Clef answers every question in one forward pass, so the second question
6// costs a few extra input tokens and no extra latency.
7
8import { LEVELS } from "./types.ts"
9
10export type Rubric = {
11  /** What Clef is asked to rate. The prompt itself is sent as the `state`. */
12  instructions: string
13  /** One description per level, lowest first; exactly five. */
14  levels: readonly string[]
15  /** The follow-up question; empty string to not ask it. */
16  followUp: string
17}
18
19export const DEFAULT_RUBRIC: Rubric = {
20  instructions:
21    "A developer sent this message to Claude Code, an AI coding agent working inside their software " +
22    "repository with tools to read, search, edit and run code. Rate how much model capability and " +
23    "reasoning effort the agent needs to do this well. Judge the work the message asks for, not the " +
24    "length of the message: a short request can be hard, and a long paste can still be a trivial task.",
25  levels: [
26    "Trivial: mechanical and obvious, no judgment. Fix a typo, rename one symbol, reformat, run a known " +
27      "command, answer a quick factual or yes/no question, find where something is defined.",
28    "Simple: a small, well-specified change or question in one place. A one-line fix with a clear cause, " +
29      "add a log line or one simple test, explain a short function, a small config edit.",
30    "Moderate: ordinary feature or bug work with a clear goal, a few files, following existing patterns. " +
31      "Add an endpoint or pagination, write tests for a module, fix a reproducible bug, a routine refactor.",
32    "Hard: tricky, ambiguous or multi-step work where a wrong answer is costly. Debug intermittent, " +
33      "concurrency or non-obvious failures, significant refactors, unfamiliar code, performance, " +
34      "security-sensitive changes, choosing between designs.",
35    "Very hard: open-ended investigation or design across a whole subsystem. Root-cause cascading or " +
36      "distributed failures, architecture or migration plans, weigh alternatives then implement and " +
37      "verify, long autonomous work.",
38  ],
39  followUp:
40    "Is this message a short follow-up whose actual task is defined by earlier conversation rather than " +
41    "by the message itself, such as 'yes do it', 'try again', 'go with option 2', 'that didn't work', " +
42    "or 'continue'?",
43}
44
45export type ClefQuestion =
46  | { type: "score"; instructions: string; criteria: string[] }
47  | { type: "choice"; instructions: string; criteria: Record<string, string> }
48  | { type: "noul"; instructions: string; criteria?: { true: string; false: string } }
49
50export type QuestionStyle = "score" | "choice"
51
52/** Question IDs, shared by the request builder and the response parser. */
53export const Q_DIFFICULTY = "difficulty"
54export const Q_FOLLOW_UP = "follow_up"
55
56/**
57 * Builds Clef's question map. `score` (the default) treats the levels as an
58 * ordered rubric, which is what they are; `choice` is kept so calibration can
59 * compare the two on the same corpus.
60 */
61export function buildQuestions(rubric: Rubric, style: QuestionStyle = "score"): Record<string, ClefQuestion> {
62  const questions: Record<string, ClefQuestion> = {}
63  questions[Q_DIFFICULTY] =
64    style === "score"
65      ? { type: "score", instructions: rubric.instructions, criteria: [...rubric.levels] }
66      : {
67          type: "choice",
68          instructions: rubric.instructions,
69          criteria: Object.fromEntries(LEVELS.map((level, i) => [level, rubric.levels[i]!])),
70        }
71  if (rubric.followUp.trim() !== "") {
72    questions[Q_FOLLOW_UP] = {
73      type: "noul",
74      instructions: rubric.followUp,
75      criteria: {
76        true: "The message only makes sense together with earlier conversation.",
77        false: "The message states a self-contained request.",
78      },
79    }
80  }
81  return questions
82}
83
84/** Validates a rubric loaded from a file; returns the problems found. */
85export function rubricProblems(value: unknown): string[] {
86  const problems: string[] = []
87  if (typeof value !== "object" || value === null) return ["rubric must be a JSON object"]
88  const r = value as Record<string, unknown>
89  if (typeof r.instructions !== "string" || r.instructions.trim() === "") problems.push("`instructions` must be a non-empty string")
90  if (!Array.isArray(r.levels) || r.levels.length !== LEVELS.length || !r.levels.every((l) => typeof l === "string" && l.trim() !== ""))
91    problems.push(`\`levels\` must be ${LEVELS.length} non-empty strings, lowest first`)
92  if (r.followUp !== undefined && typeof r.followUp !== "string") problems.push("`followUp` must be a string when present")
93  return problems
94}
95
96export function parseRubric(text: string): { rubric: Rubric } | { problems: string[] } {
97  let value: unknown
98  try {
99    value = JSON.parse(text)
100  } catch {
101    return { problems: ["rubric file is not valid JSON"] }
102  }
103  const problems = rubricProblems(value)
104  if (problems.length > 0) return { problems }
105  const r = value as { instructions: string; levels: string[]; followUp?: string }
106  return { rubric: { instructions: r.instructions, levels: r.levels, followUp: r.followUp ?? DEFAULT_RUBRIC.followUp } }
107}
108