Picks the Claude model and reasoning effort for each turn with Cloudflare Clef, a fast decision model on Workers AI.

A Claude Code mod that picks the Claude model and reasoning effort for each turn, using Cloudflare Clef as the decision model.
You keep using Claude Code as usual. When you send a prompt, the mod asks Clef-flash how much capability the request needs. It then runs that turn on the cheapest configuration that should be enough:
Fix the spelling of "recieve" in README.md. Clef → Haiku · 94%
Add pagination to /orders following the other list endpoints. Clef → Sonnet · medium · 81%
Debug why these tests intermittently deadlock only in parallel. Clef → Opus · high · 72%
Study this subsystem, find why it cascades under partitions, ... Clef → Opus · xhigh · 88%
Routing is an optimization. If Clef is slow, down, unconfigured or out of free quota, the turn still runs on a deterministic fallback. Claude Code always keeps working.
Status: v0.1, in dogfooding. The mod, the policy and the failure paths are tested against Claude Code 2.1.289's own test host, in live sessions against a mock Workers AI endpoint, and with live Clef-flash calls on a handful of prompts. How well Clef routes real coding work is not measured yet; that is what the local log and
/clef feedbackare for. The examples above show the output format; see Calibration.
High-capability models and high effort are worth it for difficult debugging, architecture and unfamiliar code. They are wasted on a rename. Nobody switches /model and /effort before every prompt, so this mod does it for you, once per turn, using a decision model built for exactly this kind of typed judgment.
The goal is not "always the cheapest model". It is the least expensive configuration that is sufficiently capable, with every policy decision shown to you.
It is no silver bullet. In a long session most of the cost is the conversation being re-read on every request, whatever model reads it, and moving a warm conversation to another model has a cost of its own. /clear between tasks often saves more than any routing. Routing, prompt caching and cost explains what a router can and cannot save, depending on how you pay.
claude --version.claude plugin marketplace add AbelNavarro/clef-model-router
claude plugin install clef-model-router@clef-model-router
Then give it your Cloudflare credentials. In Claude Code:
/plugin configure clef-model-router@clef-model-router
settings.json.If the mod is already loaded, run /reload-plugins; otherwise start a new session. You should see ↳ Clef: awaiting prompt at the end of the hint line under the prompt.
To uninstall: claude plugin uninstall clef-model-router@clef-model-router. To stop it without uninstalling, use /clef off (this session) or set Routing enabled to off in /config.
git clone https://github.com/AbelNavarro/clef-model-router
CLOUDFLARE_ACCOUNT_ID=... CLOUDFLARE_API_TOKEN=... claude --plugin-dir ./clef-model-router
Credentials in environment variables are inherited by every command Claude runs, so prefer /plugin configure for regular use.
These steps follow Cloudflare's Workers AI REST API guide.
1. Create a Cloudflare account (skip if you have one)
No credit card is needed. The free Workers plan includes 10,000 Workers AI neurons per day, roughly 2,000 routed prompts.
2. Open the Workers AI page
3. Get the API token
4. Get the Account ID
On the same Use REST API panel, under Get Account ID, copy the Account ID: a 32-character hex string such as 0123456789abcdef0123456789abcdef. It is also on the account home page under Account details → Account ID, and in the dashboard URL right after dash.cloudflare.com/.
5. Give them to the mod
/plugin configure clef-model-router@clef-model-router
/reload-plugins. The line under the prompt should read Clef: awaiting prompt.6. Check that it works
/clef test fix the typo in README shows Clef's probabilities and latency.npm run smoke makes one real call. Export CLOUDFLARE_ACCOUNT_ID and CLOUDFLARE_API_TOKEN first.If you prefer a custom token
Good practice
↳ Clef → Sonnet · medium · 87% (terminal; on other surfaces set announce to answer). The percentage is the probability Clef gave the level it picked. When policy changed Clef's pick, the reason follows in brackets, for example (Sonnet deferred) or (unsure). When Clef could not answer: Clef ✕ timeout → Sonnet · medium./clef shows the full picture: mode, the last decision with Clef's full probability distribution, latency, every policy adjustment and why, today's Clef usage, and the cache state.Last turn
route Opus · high (claude-opus-5-5, profile hard)
source clef
clef clef-flash: 72% on hard · clef confidence 52% · score 2.88/4 · follow-up 4% · 410 ms · 580 tokens
trivial ···················· 1%
simple █··················· 3%
standard ███················· 16%
hard ██████████████······ 72% ← Clef
deep ██·················· 8%
| Command | |||
|---|---|---|---|
/clef | Status and the last decision | ||
/clef history | This session's turns: latency, route, confidence, source, and any policy change | ||
/clef stats [days] | Totals from the local log: by model, effort, profile, source; latency; fallbacks; overrides; turns held on a warm model and downgrades taken; your feedback | ||
/clef test <prompt> | Ask Clef about a prompt without sending it to Claude | ||
/clef profiles | What each difficulty level runs on here | ||
/clef pin <target> | Use one target for the rest of the session (/clef pin opus:high, /clef pin hard, /clef pin :low) | ||
/clef auto | Unpin, resume after a /model change, and hand effort back to Clef after /effort | ||
/clef off · /clef on | Stop or resume routing for this session | ||
| `/clef feedback under\ | ok\ | over [note]` | Rate the last route, for later analysis of whether Clef was right |
One turn only: start a prompt with +target. The prefix is removed before Claude sees the prompt.
+opus:max why does this deadlock only on ARM?
+haiku list the files in src/
+off explain this stack trace (this turn runs exactly as Claude Code would)
prompt ──► turn.start ──► Clef-flash: difficulty 0-4 (+ "is this a follow-up?") one call, ~0.3–0.5 s end to end
│
▼
policy (deterministic): overrides → continuation → confidence → follow-up floor
→ availability → context window → downgrade timing → effort clamp
│
▼
turn.step ×N: every main-loop request of the turn sent with that model + effort
yes, do it, continue) and background-task notifications reuse the last route without calling Clef.| Level | Default | For |
|---|---|---|
| trivial | haiku | typo, rename, format, quick lookup |
| simple | sonnet:low | small, well-specified change in one place |
| standard | sonnet:medium | ordinary feature or bug work |
| hard | opus:high | tricky debugging, refactors, unfamiliar code |
| deep | opus:xhigh | open-ended investigation and design |
Haiku 4.5 takes no effort setting, so the trivial level sends none.
billing. A held turn still gets the effort Clef asked for, which keeps the cache on Opus 5.5, Sonnet 5.5 and Fable 5.1. Upgrades are never held back. See ADR 0001.+target beats /clef pin, which beats Clef. A /model change mid-session pauses routing until /clef auto. An /effort change sets the effort while Clef keeps choosing the model. Subagents keep their own models.Routing itself runs on your Cloudflare account. Clef-flash costs $0.09 per million input tokens and has no charged output. One routing call used about 580 input tokens for typical prompts in live tests (the rubric plus your prompt), and up to about 2,000 for a long one, since prompts are cut to 6,000 characters. That works out to roughly 5 neurons per call, or about 2,000 routed prompts a day inside Workers AI's free allocation of 10,000 neurons per day (resets 00:00 UTC). These are estimates derived from Cloudflare's published prices; Clef is not yet in Cloudflare's per-model neuron table.
What happens at the limit depends on your Cloudflare plan (pricing):
/clef shows today's calls, tokens and estimated neurons.
The mod also changes what you spend on Claude itself. That is the point, and /clef stats shows where your turns went.
For each prompt you type, the mod sends that prompt's text to Cloudflare Workers AI (cut to its first 4,500 and last 1,500 characters if longer than 6,000), along with the fixed rubric questions. Nothing else is sent: no conversation history, file contents, tool output, repository name or metadata. Go-aheads, task notifications, +model or +profile prompts, model pins and /clef off send nothing.
Everything else stays on your machine. The local log keeps a hash and the length of each prompt, not its text, unless you turn on Log prompt text. Cloudflare states that Workers AI does not use your inputs or outputs to train models, and the Clef announcement says Cloudflare does not read, store or train on Clef requests. Details and sources are in docs/PRIVACY.md.
Everyday options are plugin options. Set them with /plugin configure or in /config: the Cloudflare account ID and token, the decision model (clef-flash or clef), the profiles, your billing (detected by default), how to show the route, routing on or off, and whether to log prompt text. Tuning knobs go in an optional ~/.claude/clef-model-router.json. Thresholds, policies, timeout, budget, downgrade patience and the rubric are all set there. The full reference, including precedence rules, is in docs/CONFIGURATION.md.
npm run calibrate, judging whether Clef was rightThe earlier routers this project studied, and what it took from each, are listed in Architecture → Prior art: jev-model-router, jev-claude-router, pi-auto-router, clef-router, claude-code-model-router and Morph's router.
Apache-2.0. Not affiliated with Anthropic or Cloudflare.
hooks/register.ts 734 lines1// clef-model-router: picks the Claude model and effort for each turn with
2// Cloudflare Clef.
3//
4// prompt.submit a `+target ` prefix is taken off the prompt and kept for the turn
5// turn.start one Clef call (or none), then the policy decides the turn's route
6// turn.step every main-loop request of the turn is sent with that route
7// turn.complete the outcome goes to the local log
8// /clef status, history, stats, test, pin, auto, off, on, feedback
9//
10// The decision is made once per turn, at turn.start, where the person's text
11// is: a turn's later requests (after each tool result) reuse it, so a turn
12// pays Clef's latency once and never changes model halfway. Subagents keep
13// their own model. Every failure path leaves the request as Claude Code made
14// it, or on a deterministic fallback.
15
16import type { EngineInterface, PluginOptions, Register } from "claude-code"
17
18import { clefProvider } from "./lib/clef.ts"
19import { parseAdvancedFile, parseConfig, type Config } from "./lib/config.ts"
20import { cacheTtlMs, detectBilling, modelEnvFrom, type BillingFacts, type EnvValues, type RateLimitWindow } from "./lib/env.ts"
21import {
22 HELP,
23 answerLine,
24 explain,
25 historyReport,
26 profilesReport,
27 statsReport,
28 statusLine,
29 statusReport,
30 targetText,
31 type HistoryRow,
32} from "./lib/format.ts"
33import { blockedReason, estimatedNeurons, normaliseGuard, recordFailure, recordSuccess, type GuardState } from "./lib/guard.ts"
34import { aggregate, logFileName, parseLines, promptHash, turnRecord, type AnsweredUsage, type FeedbackRecord } from "./lib/log.ts"
35import { clampEffort, effortsFor, isEffort, sameModel, type ModelEnv } from "./lib/models.ts"
36import { parseCommand, parsePrefix, trackNativeEffort, turnKind, type ClefCommand } from "./lib/overrides.ts"
37import { decide, describe, needsClef, type CacheState, type HoldState, type RouterMode } from "./lib/policy.ts"
38import { priceOf, stayCost, isOneHour, type Billing } from "./lib/pricing.ts"
39import { DEFAULT_RUBRIC, parseRubric, type Rubric } from "./lib/rubric.ts"
40import type { Decision, ProviderResult, Route, Target } from "./lib/types.ts"
41
42const PLUGIN = "clef-model-router"
43const SESSION_REF = { plugin: "clef-model-router", key: "session" } as const
44const GUARD_KEY = "guard"
45const HISTORY_LIMIT = 50
46const RUN_LIMIT = 32
47/** A routed model that fails this many requests in a row is not used again this session. */
48const FAILURES_BEFORE_UNAVAILABLE = 2
49
50/** What survives a hot reload (in `$.state`) for the rest of the session. */
51type Persisted = {
52 mode: RouterMode
53 pin?: Target
54 pendingOverride?: Target | "off"
55 last?: Route
56 cache?: CacheState
57 /** A downgrade held for the cache, and what staying has cost so far. */
58 hold?: HoldState
59 unavailable: string[]
60 failures: Record<string, number>
61 /** The session model Claude Code reported at the last turn. */
62 baselineModel?: string
63 /** Effort as Claude Code itself would send it, and any /effort the person set. */
64 effortBaseline?: string | number
65 nativeEffort?: string | number
66 /** A model Claude Code fell back to during the last turn, so it is not mistaken for a /model change. */
67 engineFallback?: string
68 history: HistoryRow[]
69 warned: string[]
70}
71
72/** One turn in flight. Module memory only: a reload mid-turn leaves it unrouted. */
73type Run = {
74 decision: Decision
75 prompt: string
76 hash?: string
77 engineModel?: string
78 passthrough: boolean
79 failedRewrite: boolean
80 steps: number
81 answered?: AnsweredUsage
82 /** Cache read and write of the turn's first request: what a switch (or a return) cost. */
83 firstStep?: { cacheReadTokens: number; cacheWriteTokens: number }
84 rateLimits?: RateLimitWindow[]
85}
86
87// Module state. `register` runs again on every reload, which resets these;
88// `load` then restores the session's part from `$.state`.
89let options: PluginOptions = {}
90let state: Persisted = freshState()
91let loaded = false
92let config: Config = parseConfig({}).config
93let problems: string[] = []
94let modelEnv: ModelEnv = modelEnvFrom({})
95let env: EnvValues = {}
96let settingsTtl: unknown
97let apiKeyHelper = false
98let detected: BillingFacts = detectBilling({})
99let billing: Billing = detected.billing
100let ttlMs = 5 * 60_000
101let rubric: Rubric = DEFAULT_RUBRIC
102let logDir: string | undefined
103let apiBase: string | undefined
104let advancedPath: string | undefined
105let logFile: string | undefined
106let logLines: string[] = []
107const runs = new Map<string, Run>()
108
109function freshState(): Persisted {
110 return { mode: "auto", unavailable: [], failures: {}, history: [], warned: [] }
111}
112
113async function save($: EngineInterface): Promise<void> {
114 try {
115 await $.state.set(SESSION_REF, JSON.stringify(state))
116 } catch {
117 // Losing the snapshot only matters on a hot reload; routing goes on.
118 }
119}
120
121async function readEnv($: EngineInterface): Promise<EnvValues> {
122 return {
123 ANTHROPIC_DEFAULT_HAIKU_MODEL: await $.env.get("ANTHROPIC_DEFAULT_HAIKU_MODEL"),
124 ANTHROPIC_DEFAULT_SONNET_MODEL: await $.env.get("ANTHROPIC_DEFAULT_SONNET_MODEL"),
125 ANTHROPIC_DEFAULT_OPUS_MODEL: await $.env.get("ANTHROPIC_DEFAULT_OPUS_MODEL"),
126 ANTHROPIC_DEFAULT_FABLE_MODEL: await $.env.get("ANTHROPIC_DEFAULT_FABLE_MODEL"),
127 CLAUDE_CODE_USE_BEDROCK: await $.env.get("CLAUDE_CODE_USE_BEDROCK"),
128 CLAUDE_CODE_USE_VERTEX: await $.env.get("CLAUDE_CODE_USE_VERTEX"),
129 CLAUDE_CODE_USE_FOUNDRY: await $.env.get("CLAUDE_CODE_USE_FOUNDRY"),
130 CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS: await $.env.get("CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS"),
131 // Only whether it is set: the key itself is never kept.
132 ANTHROPIC_API_KEY: (await $.env.get("ANTHROPIC_API_KEY")) ? "set" : undefined,
133 CLAUDE_CODE_PROMPT_CACHE_TTL: await $.env.get("CLAUDE_CODE_PROMPT_CACHE_TTL"),
134 FORCE_PROMPT_CACHING_5M: await $.env.get("FORCE_PROMPT_CACHING_5M"),
135 ENABLE_PROMPT_CACHING_1H: await $.env.get("ENABLE_PROMPT_CACHING_1H"),
136 // Only whether they are set: they say a gateway bills per token.
137 ANTHROPIC_AUTH_TOKEN: (await $.env.get("ANTHROPIC_AUTH_TOKEN")) ? "set" : undefined,
138 ANTHROPIC_BASE_URL: (await $.env.get("ANTHROPIC_BASE_URL")) ? "set" : undefined,
139 }
140}
141
142/** Reads configuration and restores the session snapshot, once per load. */
143async function load($: EngineInterface): Promise<void> {
144 if (loaded) return
145 loaded = true
146 try {
147 const home = (await $.env.get("HOME")) ?? (await $.env.get("USERPROFILE")) ?? "."
148 const configDir = (await $.env.get("CLAUDE_CONFIG_DIR")) ?? `${home}/.claude`
149 advancedPath = (await $.env.get("CLEF_ROUTER_CONFIG")) ?? `${configDir}/${PLUGIN}.json`
150 const advancedText = await $.fs.read(advancedPath).catch(() => undefined)
151 const advanced = parseAdvancedFile(typeof advancedText === "string" ? advancedText : undefined)
152 const parsed = parseConfig(
153 { ...advanced.values, ...options },
154 { accountId: await $.env.get("CLOUDFLARE_ACCOUNT_ID"), apiToken: await $.env.get("CLOUDFLARE_API_TOKEN") },
155 )
156 config = parsed.config
157 problems = [...advanced.problems, ...parsed.problems]
158 env = await readEnv($)
159 modelEnv = modelEnvFrom(env)
160 const settings = (await $.settings.read().catch(() => ({}))) as Record<string, unknown>
161 settingsTtl = settings.promptCacheTtl
162 apiKeyHelper = typeof settings.apiKeyHelper === "string" && settings.apiKeyHelper !== ""
163 refreshBilling(undefined)
164 if (config.rubricFile) {
165 const text = await $.fs.read(config.rubricFile).catch(() => undefined)
166 const result = typeof text === "string" ? parseRubric(text) : { problems: [`cannot read rubric_file ${config.rubricFile}`] }
167 if ("rubric" in result) rubric = result.rubric
168 else problems.push(...result.problems.map((p) => `${p}; using the built-in rubric`))
169 }
170 const base = await $.env.get("CLEF_ROUTER_API_BASE")
171 if (base && /^https?:\/\//.test(base)) apiBase = base
172 logDir = config.logDir ?? `${configDir}/plugins/data/${PLUGIN}`
173 const snapshot = await $.state.get(SESSION_REF)
174 if (typeof snapshot.value === "string") state = { ...freshState(), ...(JSON.parse(snapshot.value) as Partial<Persisted>) }
175 } catch (error) {
176 problems.push(`setup: ${error instanceof Error ? error.message : String(error)}`)
177 }
178}
179
180/**
181 * How the person pays and the cache TTL that follows from it, refreshed from
182 * the rate-limit windows each turn: they appear after the session's first
183 * response on a subscription, and a window past 100% means usage credits.
184 */
185function refreshBilling(rateLimits: readonly RateLimitWindow[] | undefined): void {
186 detected = detectBilling(env, { apiKeyHelper, ...(rateLimits ? { rateLimits } : {}) })
187 billing = config.billing === "auto" ? detected.billing : config.billing
188 ttlMs = cacheTtlMs(config.cacheTtlMinutes, env, settingsTtl, detected)
189}
190
191async function readGuard($: EngineInterface, now: number): Promise<GuardState> {
192 return normaliseGuard(await $.store.get(GUARD_KEY).catch(() => undefined), now)
193}
194
195async function askClef($: EngineInterface, prompt: string): Promise<ProviderResult> {
196 const now = await $.clock.now()
197 const guard = await readGuard($, now)
198 const blocked = blockedReason(guard, { now, model: config.decisionModel, dailyNeuronBudget: config.dailyNeuronBudget })
199 if (blocked) return { ok: false, failure: blocked }
200 const provider = clefProvider({
201 accountId: config.accountId,
202 apiToken: config.apiToken,
203 model: config.decisionModel,
204 rubric,
205 timeoutMs: config.timeoutMs,
206 maxPromptChars: config.maxPromptChars,
207 ...(apiBase ? { apiBase } : {}),
208 fetch: async (url, init) => {
209 const r = await $.http.fetch(url, init)
210 return { status: r.status, ok: r.ok, text: r.text }
211 },
212 sleep: (ms, signal) => $.clock.sleep(ms, { signal }),
213 now: () => $.clock.now(),
214 })
215 const result = await provider.decide(prompt)
216 const after = result.ok ? recordSuccess(guard, result.recommendation.inputTokens) : recordFailure(guard, result.failure, await $.clock.now())
217 if (after !== guard) await $.store.set(GUARD_KEY, after).catch(() => {})
218 return result
219}
220
221/** One-time notices, so a misconfiguration is said once, not every turn. */
222function warnOnce($: EngineInterface, key: string, text: string): void {
223 if (state.warned.includes(key)) return
224 state.warned.push(key)
225 $.ui.toast(text, { timeoutMs: 8000 })
226}
227
228/** The route line, drawn dim at the end of the hint line under the prompt. */
229let hint: string | undefined
230
231/**
232 * Shows the route without Claude Code's status-line marker (a ⚠ that reads as
233 * a warning): the text joins the prompt's hint line as its tail, prefixed ↳.
234 * The terminal draws that tail; elsewhere `announce: "answer"` shows the route.
235 */
236function showStatus($: EngineInterface, text: string | undefined): void {
237 const next = text === undefined ? undefined : `↳ ${text}`
238 if (next === hint) return
239 hint = next
240 $.ui.invalidate("ui.render")
241}
242
243function announce($: EngineInterface, d: Decision): void {
244 if (config.announce === "status" || config.announce === "both") showStatus($, statusLine(d))
245}
246
247async function appendLog($: EngineInterface, line: string, ts: string): Promise<void> {
248 if (!config.logEnabled || !logDir) return
249 try {
250 const file = `${logDir}/${logFileName(ts, await $.session.id())}`
251 if (file !== logFile) {
252 logFile = file
253 const existing = await $.fs.read(file).catch(() => "")
254 logLines = typeof existing === "string" && existing !== "" ? existing.trimEnd().split("\n") : []
255 }
256 logLines.push(line)
257 await $.fs.write(file, logLines.join("\n") + "\n")
258 } catch {
259 // A log that cannot be written must not cost the turn anything.
260 }
261}
262
263/** Notices a /model change made between turns: the person taking over the model. */
264async function noticeNativeModel($: EngineInterface): Promise<void> {
265 const sessionModel = await $.session.model().catch(() => undefined)
266 if (sessionModel && state.baselineModel && !sameModel(sessionModel, state.baselineModel)) {
267 // A fallback Claude Code made itself (a safety classifier moving the
268 // session) is not one.
269 const engineMoved = state.engineFallback !== undefined && sameModel(sessionModel, state.engineFallback)
270 if (!engineMoved && state.mode === "auto" && config.pauseOnNativeChange) {
271 state.mode = "paused-native"
272 $.ui.toast(`Clef paused: you switched to ${sessionModel}. /clef auto resumes routing.`, { timeoutMs: 8000 })
273 }
274 }
275 delete state.engineFallback
276 if (sessionModel) state.baselineModel = sessionModel
277}
278
279async function routeTurn($: EngineInterface, turnId: string, text: string): Promise<void> {
280 const kind = turnKind(text)
281 const override = state.pendingOverride
282 delete state.pendingOverride
283 await noticeNativeModel($)
284
285 const base = { turnId, kind, config, modelEnv, session: state, ...(override ? { override } : {}) }
286 const result = needsClef(base) ? await askClef($, text) : undefined
287 const usage = await $.session.usage().catch(() => undefined)
288 const rateLimits = usage?.rateLimits?.map((w) => ({ kind: w.kind, percentUsed: w.percentUsed }))
289 refreshBilling(rateLimits)
290 const now = await $.clock.now()
291 const decision = decide({
292 ...base,
293 ...(result ? { result } : {}),
294 ...(usage?.context?.tokens ? { contextTokens: usage.context.tokens } : {}),
295 now,
296 cacheTtlMs: ttlMs,
297 billing,
298 })
299
300 const kindOfFailure = decision.failure?.kind
301 if (kindOfFailure === "not-configured")
302 warnOnce($, "not-configured", `Clef router: set your Cloudflare account ID and API token with /plugin configure ${PLUGIN}. Using the ${config.fallbackLevel} profile meanwhile.`)
303 else if (kindOfFailure === "auth")
304 warnOnce($, "auth", `Clef router: Cloudflare rejected the API token (${decision.failure?.message}). Falling back until it is fixed.`)
305 else if (kindOfFailure === "quota")
306 warnOnce($, `quota-${new Date(now).toISOString().slice(0, 10)}`, "Clef router: Workers AI's free daily allocation is used up; falling back until 00:00 UTC.")
307
308 if (decision.final) state.last = decision.final
309 // A held turn carries the stretch on (its cost is added when it completes);
310 // any other turn ends it.
311 const deferral = decision.deferral
312 if (deferral?.held && decision.final) {
313 state.hold = { model: decision.final.model, wanted: deferral.wanted, spent: deferral.spent, turns: deferral.turns }
314 } else delete state.hold
315 const run: Run = { decision, prompt: text, passthrough: !decision.final, failedRewrite: false, steps: 0 }
316 if (rateLimits && rateLimits.length > 0) run.rateLimits = rateLimits
317 const hash = await promptHash(text).catch(() => undefined)
318 if (hash) run.hash = hash
319 runs.set(turnId, run)
320 while (runs.size > RUN_LIMIT) runs.delete(runs.keys().next().value!)
321 state.history.push({ decision, prompt: text.slice(0, 200) })
322 while (state.history.length > HISTORY_LIMIT) state.history.shift()
323 announce($, decision)
324 await save($)
325}
326
327/** Watches Claude Code's own model and effort at a main-loop step, before any rewrite. */
328function noticeStep($: EngineInterface, run: Run, model: string, effort: string | number | undefined, index: number): void {
329 if (index === 0) {
330 run.engineModel = model
331 const tracked = trackNativeEffort({ baseline: state.effortBaseline, native: state.nativeEffort }, effort)
332 state.effortBaseline = tracked.baseline
333 if (tracked.native === undefined) delete state.nativeEffort
334 else state.nativeEffort = tracked.native
335 if (!config.pauseOnNativeChange) return
336 if (tracked.change === "set" && state.mode === "auto")
337 $.ui.toast(`Clef: using your effort ${String(tracked.native)}; Clef still picks the model. /clef auto hands effort back.`, { timeoutMs: 8000 })
338 // The person's /effort beats Clef's, not an explicit +target or pin.
339 const d = run.decision
340 if (state.nativeEffort !== undefined && d.final && (d.source === "clef" || d.source === "continuation" || d.source === "fallback")) {
341 const wanted = typeof state.nativeEffort === "string" && isEffort(state.nativeEffort) ? state.nativeEffort : undefined
342 const effortNow = wanted ? clampEffort(d.final.model, wanted) : undefined
343 if (effortNow !== d.final.effort) {
344 const final = { ...d.final }
345 if (effortNow) final.effort = effortNow
346 else delete final.effort
347 run.decision = {
348 ...d,
349 final,
350 adjustments: [...d.adjustments, { rule: "pinned-effort", from: d.final.effort ?? "default", to: effortNow ?? "default", reason: "your /effort" }],
351 }
352 state.last = final
353 announce($, run.decision)
354 }
355 }
356 } else if (run.engineModel && !sameModel(model, run.engineModel)) {
357 // Claude Code moved the turn to a fallback model (an error, or a safety
358 // classifier). That is never overridden.
359 run.passthrough = true
360 state.engineFallback = model
361 }
362}
363
364/** Records what a step's response says: failures of a routed model, usage, cache. */
365async function afterStep(
366 $: EngineInterface,
367 run: Run,
368 sent: { model: string; effort?: unknown },
369 rewritten: boolean,
370 result: { stopReason: string | null; usage: { model: string; input_tokens: number; output_tokens: number; cache_read_input_tokens: number; cache_creation_input_tokens: number } | null },
371 aborted: boolean,
372): Promise<void> {
373 const now = await $.clock.now()
374 if (rewritten) {
375 if (result.stopReason === null && !aborted) {
376 // The routed request got no response: stop routing this turn, and stop
377 // using the model after repeated failures.
378 run.failedRewrite = true
379 const n = (state.failures[sent.model] ?? 0) + 1
380 state.failures[sent.model] = n
381 if (n >= FAILURES_BEFORE_UNAVAILABLE && !state.unavailable.includes(sent.model)) {
382 state.unavailable.push(sent.model)
383 $.ui.toast(`Clef router: ${sent.model} failed ${n} times; not routing to it again this session.`, { timeoutMs: 8000 })
384 }
385 } else if (result.stopReason !== null) {
386 state.failures[sent.model] = 0
387 }
388 }
389 const u = result.usage
390 if (u) {
391 const a = run.answered ?? { model: u.model, inputTokens: 0, outputTokens: 0, cacheReadTokens: 0, cacheWriteTokens: 0 }
392 run.answered = {
393 model: u.model,
394 inputTokens: a.inputTokens + u.input_tokens,
395 outputTokens: a.outputTokens + u.output_tokens,
396 cacheReadTokens: a.cacheReadTokens + u.cache_read_input_tokens,
397 cacheWriteTokens: a.cacheWriteTokens + u.cache_creation_input_tokens,
398 }
399 if (!run.firstStep) run.firstStep = { cacheReadTokens: u.cache_read_input_tokens, cacheWriteTokens: u.cache_creation_input_tokens }
400 const cache: CacheState = {
401 model: sent.model,
402 at: now,
403 promptTokens: u.input_tokens + u.cache_read_input_tokens + u.cache_creation_input_tokens,
404 caching: u.cache_read_input_tokens + u.cache_creation_input_tokens > 0,
405 }
406 if (typeof sent.effort === "string" && isEffort(sent.effort)) cache.effort = sent.effort
407 state.cache = cache
408 }
409 await save($)
410}
411
412/** Adds what a held turn cost to its stretch, from the usage the API reported. */
413function accrueHold(run: Run): void {
414 const d = run.decision
415 const hold = state.hold
416 if (!d.deferral?.held || !d.final || !run.answered || !hold || !sameModel(hold.model, run.answered.model)) return
417 const held = priceOf(d.final.model)
418 const wanted = priceOf(d.deferral.wanted.model)
419 if (!held || !wanted) return
420 hold.spent = d.deferral.spent + stayCost(d.deferral.billing, held, wanted, run.answered, isOneHour(ttlMs))
421}
422
423async function completeTurn($: EngineInterface, run: Run, durationMs: number, reason: string): Promise<void> {
424 accrueHold(run)
425 const ts = new Date(await $.clock.now()).toISOString()
426 const record = turnRecord({
427 decision: run.decision,
428 session: await $.session.id(),
429 ts,
430 promptText: run.prompt,
431 ...(run.hash ? { hash: run.hash } : {}),
432 logPrompts: config.logPrompts,
433 ...(run.answered ? { answered: run.answered } : {}),
434 ...(run.firstStep ? { firstStep: run.firstStep } : {}),
435 ...(run.rateLimits ? { rateLimits: run.rateLimits } : {}),
436 steps: run.steps,
437 durationMs,
438 endReason: reason,
439 })
440 await appendLog($, JSON.stringify(record), ts)
441 const row = state.history.find((h) => h.decision.turnId === run.decision.turnId)
442 if (row) {
443 row.decision = run.decision
444 if (run.answered) row.answeredModel = run.answered.model
445 }
446 await save($)
447}
448
449async function registerCommand($: EngineInterface): Promise<void> {
450 try {
451 await $.command.register({
452 name: "clef",
453 description: "Clef router: status, history, stats, test, pin, auto, off, on, feedback",
454 argumentHint: "[status|history|stats|profiles|test|pin|auto|off|on|feedback|help]",
455 immediate: true,
456 })
457 } catch {
458 // Without the command the router still routes.
459 }
460 if (!config.enabled) showStatus($, undefined)
461 else if (!config.accountId || !config.apiToken) showStatus($, "Clef: not configured")
462 else {
463 // Always replace what an earlier load showed (a stale "not configured"
464 // survives a reload otherwise): the last route if there is one.
465 const last = state.history.at(-1)?.decision
466 showStatus($, (last && statusLine(last)) ?? "Clef: awaiting prompt")
467 }
468}
469
470async function onClear($: EngineInterface): Promise<void> {
471 // /clear starts a new conversation: nothing is cached and nothing continues.
472 delete state.last
473 delete state.cache
474 delete state.hold
475 delete state.pendingOverride
476 state.history = []
477 runs.clear()
478 logFile = undefined
479 if (config.enabled && config.accountId && config.apiToken) showStatus($, "Clef: awaiting prompt")
480 await save($)
481}
482
483async function setPendingOverride($: EngineInterface, override: Target | "off"): Promise<void> {
484 await load($)
485 state.pendingOverride = override
486 await save($)
487}
488
489async function statusText($: EngineInterface, now: number): Promise<string> {
490 const guard = await readGuard($, now)
491 const blocked = blockedReason(guard, { now, model: config.decisionModel, dailyNeuronBudget: config.dailyNeuronBudget })
492 const last = state.history.at(-1)?.decision
493 return statusReport({
494 config,
495 configProblems: problems,
496 modelEnv,
497 mode: state.mode,
498 ...(state.pin ? { pin: state.pin } : {}),
499 ...(last ? { last } : {}),
500 guard: {
501 calls: guard.calls,
502 inputTokens: guard.inputTokens,
503 neurons: estimatedNeurons(guard, config.decisionModel),
504 ...(blocked ? { blocked: blocked.message } : {}),
505 },
506 ...(state.cache
507 ? {
508 cache: {
509 model: state.cache.model,
510 promptTokens: state.cache.promptTokens,
511 ageSeconds: Math.round((now - state.cache.at) / 1000),
512 ttlSeconds: Math.round(ttlMs / 1000),
513 },
514 }
515 : {}),
516 billing: { billing, detected, configured: config.billing, patience: config.downgradePatience },
517 ...(state.hold ? { hold: state.hold } : {}),
518 unavailable: state.unavailable,
519 logDir,
520 ...(advancedPath ? { advancedPath } : {}),
521 })
522}
523
524async function statsText($: EngineInterface, now: number, days: number): Promise<string> {
525 if (!logDir) return "No log directory."
526 const since = new Date(now - (days - 1) * 86_400_000).toISOString().slice(0, 10)
527 const entries = await $.fs.list(logDir).catch(() => [])
528 const files = entries.filter((f) => /^routing-\d{4}-\d{2}-\d{2}-/.test(f.name) && f.name.slice(8, 18) >= since)
529 const records = []
530 for (const f of files) {
531 const text = await $.fs.read(`${logDir}/${f.name}`).catch(() => "")
532 if (typeof text === "string") records.push(...parseLines(text))
533 }
534 return statsReport(aggregate(records), days, files.length)
535}
536
537async function testText($: EngineInterface, now: number, prompt: string): Promise<string> {
538 const result = await askClef($, prompt)
539 const d = decide({
540 turnId: "test",
541 kind: turnKind(prompt),
542 config,
543 modelEnv,
544 session: { ...state, mode: "auto" },
545 result,
546 now,
547 cacheTtlMs: ttlMs,
548 billing,
549 })
550 return [`Clef on: ${prompt.slice(0, 80)}`, ...explain(d), "", "(Not sent to Claude. Counts toward today's Clef usage.)"].join("\n")
551}
552
553async function feedbackText($: EngineInterface, now: number, cmd: Extract<ClefCommand, { kind: "feedback" }>): Promise<string> {
554 const last = state.history.at(-1)
555 const record: FeedbackRecord = {
556 v: 1,
557 type: "feedback",
558 ts: new Date(now).toISOString(),
559 session: await $.session.id(),
560 verdict: cmd.verdict,
561 ...(last ? { turn: last.decision.turnId } : {}),
562 ...(cmd.note ? { note: cmd.note } : {}),
563 }
564 await appendLog($, JSON.stringify(record), record.ts)
565 const route = last?.decision.final ? describe(last.decision.final) : "the last turn"
566 const verdict = cmd.verdict === "ok" ? "about right" : cmd.verdict === "under" ? "not capable enough" : "more than needed"
567 return `Noted: ${route} was ${verdict}.`
568}
569
570async function runCommand($: EngineInterface, args: string): Promise<string> {
571 await load($)
572 const cmd = parseCommand(args)
573 const now = await $.clock.now()
574 switch (cmd.kind) {
575 case "help":
576 return HELP
577 case "error":
578 return cmd.message
579 case "status":
580 return statusText($, now)
581 case "history":
582 return historyReport(state.history)
583 case "profiles":
584 return profilesReport(config, modelEnv, state.unavailable)
585 case "stats":
586 return statsText($, now, cmd.days)
587 case "test":
588 return testText($, now, cmd.prompt)
589 case "feedback":
590 return feedbackText($, now, cmd)
591 case "auto":
592 state.mode = "auto"
593 delete state.pin
594 delete state.nativeEffort
595 delete state.effortBaseline
596 await save($)
597 showStatus($, config.enabled ? `Clef auto · ${config.decisionModel}` : undefined)
598 return config.enabled ? "Routing is automatic again." : "Routing is disabled in the plugin config (enabled = false)."
599 case "on":
600 state.mode = "auto"
601 await save($)
602 showStatus($, `Clef on · ${config.decisionModel}`)
603 return state.pin ? `Routing on, still pinned to ${targetText(state.pin)} (/clef auto to unpin).` : "Routing on."
604 case "off":
605 state.mode = "off"
606 await save($)
607 showStatus($, "Clef off")
608 return "Routing off for this session: Claude Code's own model and effort apply. /clef on resumes."
609 case "pin":
610 state.pin = cmd.target
611 state.mode = "auto"
612 await save($)
613 showStatus($, `Pinned → ${targetText(cmd.target)}`)
614 return `Pinned to ${targetText(cmd.target)} for this session. /clef auto unpins.`
615 }
616}
617
618/** The request a step is sent with: the turn's route, effort fitted to the model. */
619function routed<E extends { model: string; effort?: unknown }>(e: E, route: Route): E {
620 const request = { ...e, model: route.model } as E & { effort?: unknown }
621 if (effortsFor(route.model) === null) delete request.effort
622 else if (route.effort) request.effort = route.effort
623 else if (typeof e.effort === "string" && isEffort(e.effort)) {
624 const clamped = clampEffort(route.model, e.effort)
625 if (clamped) request.effort = clamped
626 }
627 return request
628}
629
630export const register: Register = (on, pluginOptions) => {
631 options = pluginOptions
632 state = freshState()
633 hint = undefined
634 apiBase = undefined
635 loaded = false
636 logFile = undefined
637 logLines = []
638 runs.clear()
639
640 on("ui.render", { component: "PromptHint" }, async ($, e, next) => {
641 if (hint === undefined) return next(e)
642 const tail = e.props.tail ? `${e.props.tail} · ${hint}` : ` ${hint}`
643 return next({ ...e, props: { ...e.props, tail } })
644 })
645
646 on("session.start", async ($, e, next) => {
647 await load($)
648 await registerCommand($)
649 return next(e)
650 })
651
652 // A compaction replaces the conversation, so the next request writes a new
653 // cache whatever the model: a downgrade then costs nothing to take.
654 on("session.compact", async ($, e, next) => {
655 const result = await next(e)
656 if (e.agentId === undefined && e.trigger !== "precompute" && result.messages !== undefined) {
657 try {
658 await load($)
659 delete state.cache
660 delete state.hold
661 await save($)
662 } catch {
663 // Bookkeeping only.
664 }
665 }
666 return result
667 })
668
669 on("session.end", async ($, e, next) => {
670 if (e.reason === "clear") await onClear($)
671 return next(e)
672 })
673
674 on("prompt.submit", async ($, e, next) => {
675 // Only a prompt that starts a turn of its own; one typed into a running
676 // turn joins that turn, which keeps its route.
677 if (e.turnId !== undefined) return next(e)
678 const parsed = parsePrefix(e.text)
679 if (parsed.override === undefined) return next(e)
680 await setPendingOverride($, parsed.override)
681 return next({ ...e, text: parsed.text })
682 })
683
684 on("turn.start", async ($, e, next) => {
685 try {
686 await load($)
687 await routeTurn($, e.turnId, e.text)
688 } catch {
689 // Unrouted: Claude Code's own model and effort apply.
690 }
691 return next(e)
692 })
693
694 on("turn.step", async function* ($, e, next) {
695 // Subagents run on the model their definition or Claude Code gives them.
696 const run = e.agentId === undefined ? runs.get(e.turnId) : undefined
697 if (!run) return yield* next(e)
698 run.steps++
699 try {
700 noticeStep($, run, e.model, e.effort, e.index)
701 } catch {
702 run.passthrough = true
703 }
704 const route = run.decision.final
705 const rewrite = route !== undefined && !run.passthrough && !run.failedRewrite
706 const request = rewrite ? routed(e, route) : e
707 const result = yield* next(request)
708 try {
709 await afterStep($, run, request, rewrite, result, next.signal?.aborted === true)
710 } catch {
711 // Bookkeeping only.
712 }
713 return result
714 })
715
716 on("turn.complete", async ($, e, next) => {
717 const result = await next(e)
718 const run = e.agentId === undefined ? runs.get(e.turnId) : undefined
719 if (!run) return result
720 try {
721 await completeTurn($, run, e.durationMs, e.reason)
722 } catch {
723 // Logging only.
724 }
725 if (config.announce === "answer" || config.announce === "both") {
726 const line = answerLine(run.decision, run.decision.recommendation?.latencyMs)
727 if (line) return { ...result, text: line }
728 }
729 return result
730 })
731
732 on("command.run", { command: "clef" }, async ($, e) => ({ text: await runCommand($, e.args) }))
733}
734hooks/lib/clef.ts 252 lines1// Cloudflare Clef on Workers AI: the request, the response, and every way the
2// exchange can fail, turned into a normalised Recommendation or a classified
3// ProviderFailure. Nothing here throws.
4//
5// API (developers.cloudflare.com/workers-ai/models/clef-flash, Oct 2026):
6// POST https://api.cloudflare.com/client/v4/accounts/{account}/ai/run/@cf/cloudflare/{model}
7// Authorization: Bearer {token}
8// { "model": "clef-flash", "state": ..., "questions": { id: {type, instructions, criteria} } }
9// → { "result": { "model", "answers": { id: answer }, "usage": { input_tokens, output_tokens } },
10// "success": true, "errors": [], "messages": [] }
11
12import { buildQuestions, Q_DIFFICULTY, Q_FOLLOW_UP, type QuestionStyle, type Rubric } from "./rubric.ts"
13import { redact } from "./redact.ts"
14import { LEVELS, type DecisionProvider, type Level, type ProviderFailure, type ProviderResult, type Recommendation } from "./types.ts"
15
16export const CLEF_MODELS = ["clef-flash", "clef"] as const
17export type ClefModel = (typeof CLEF_MODELS)[number]
18
19export const DEFAULT_API_BASE = "https://api.cloudflare.com/client/v4"
20
21/** The slice of `$.http.fetch` (or the global fetch, in scripts) this needs. */
22export type HttpLike = (
23 url: string,
24 init: { method: string; headers: Record<string, string>; body: string },
25) => Promise<{ status: number; ok: boolean; text: string }>
26
27export type ClefOptions = {
28 accountId: string | undefined
29 apiToken: string | undefined
30 model: ClefModel
31 rubric: Rubric
32 style?: QuestionStyle
33 timeoutMs: number
34 maxPromptChars: number
35 apiBase?: string
36 fetch: HttpLike
37 /** Resolves after `ms` (rejects if `signal` aborts); the timeout races the request against it. */
38 sleep: (ms: number, signal: AbortSignal) => Promise<void>
39 now: () => number | Promise<number>
40}
41
42export function endpoint(accountId: string, model: ClefModel, apiBase = DEFAULT_API_BASE): string {
43 return `${apiBase.replace(/\/$/, "")}/accounts/${encodeURIComponent(accountId)}/ai/run/@cf/cloudflare/${model}`
44}
45
46/**
47 * Keeps the head and the tail of a long prompt. The request's intent is
48 * usually stated at one end; the middle of a long paste rarely changes how
49 * hard the task is, and every character sent is billed and leaves the machine.
50 */
51export function truncatePrompt(text: string, maxChars: number): string {
52 if (text.length <= maxChars) return text
53 const head = Math.floor(maxChars * 0.75)
54 const tail = maxChars - head
55 const omitted = text.length - head - tail
56 return `${text.slice(0, head)}\n[... ${omitted} characters omitted ...]\n${text.slice(text.length - tail)}`
57}
58
59export function buildRequestBody(prompt: string, opts: Pick<ClefOptions, "model" | "rubric" | "style" | "maxPromptChars">): string {
60 return JSON.stringify({
61 model: opts.model,
62 state: truncatePrompt(prompt, opts.maxPromptChars),
63 questions: buildQuestions(opts.rubric, opts.style ?? "score"),
64 })
65}
66
67type Envelope = {
68 success?: unknown
69 result?: unknown
70 errors?: unknown
71}
72
73function errorsOf(envelope: Envelope): { code?: number; message: string }[] {
74 if (!Array.isArray(envelope.errors)) return []
75 return envelope.errors
76 .filter((e): e is Record<string, unknown> => typeof e === "object" && e !== null)
77 .map((e) => ({
78 code: typeof e.code === "number" ? e.code : undefined,
79 message: typeof e.message === "string" ? e.message : "",
80 }))
81}
82
83/** Maps an HTTP status and Cloudflare error envelope to a failure kind. */
84export function classifyHttpFailure(
85 status: number,
86 bodyText: string,
87 latencyMs: number,
88 secrets: readonly (string | undefined)[] = [],
89): ProviderFailure {
90 let envelope: Envelope = {}
91 try {
92 envelope = JSON.parse(bodyText) as Envelope
93 } catch {
94 // Not JSON (a proxy's HTML page, say); the status alone decides.
95 }
96 const errors = errorsOf(envelope)
97 const codes = errors.map((e) => e.code)
98 const text = errors.map((e) => (e.code === undefined ? e.message : `${e.code}: ${e.message}`)).join("; ")
99 const message = redact(text || `HTTP ${status}`, secrets).slice(0, 200)
100 const quotaText = /daily free allocation|neurons/i.test(text)
101
102 if (codes.includes(3036) || (status === 429 && quotaText)) return { kind: "quota", message, status, latencyMs }
103 if (status === 429) return { kind: "rate-limited", message, status, latencyMs }
104 if (status === 401 || status === 403 || codes.includes(10000)) return { kind: "auth", message, status, latencyMs }
105 if (status === 408 || codes.includes(3007)) return { kind: "timeout", message, status, latencyMs }
106 if (status >= 500) return { kind: "server", message, status, latencyMs }
107 if (status >= 400) return { kind: "bad-request", message, status, latencyMs }
108 return { kind: "malformed", message, status, latencyMs }
109}
110
111function num(value: unknown): number | undefined {
112 return typeof value === "number" && Number.isFinite(value) ? value : undefined
113}
114
115/**
116 * Reads a difficulty answer's per-level probabilities. A score answer keys
117 * them by level index ("0".."4", per the schema), a choice answer by option
118 * id (our level names). A 1-based index set is accepted too, defensively.
119 */
120export function levelProbabilities(answer: Record<string, unknown>, rubric: Rubric): Record<Level, number> | undefined {
121 const raw = answer.probabilities
122 if (typeof raw !== "object" || raw === null) return undefined
123 const entries = Object.entries(raw as Record<string, unknown>)
124 const out = Object.fromEntries(LEVELS.map((l) => [l, 0])) as Record<Level, number>
125 const numericKeys = entries.every(([k]) => /^\d+$/.test(k))
126 const base = numericKeys ? Math.min(...entries.map(([k]) => Number(k))) : 0
127 let matched = 0
128 for (const [key, value] of entries) {
129 const p = num(value)
130 if (p === undefined || p < 0 || p > 1.0001) return undefined
131 let index: number
132 if (numericKeys) index = Number(key) - (base === 1 && entries.length === LEVELS.length ? 1 : 0)
133 else if ((LEVELS as readonly string[]).includes(key)) index = LEVELS.indexOf(key as Level)
134 else index = rubric.levels.indexOf(key)
135 const level = LEVELS[index]
136 if (level === undefined) return undefined
137 out[level] += p
138 matched++
139 }
140 if (matched === 0) return undefined
141 const total = LEVELS.reduce((s, l) => s + out[l], 0)
142 if (total < 0.98 || total > 1.02) return undefined
143 return out
144}
145
146/** The most probable level; a tie goes to the more capable one. */
147export function topLevel(probabilities: Record<Level, number>): Level {
148 let best: Level = LEVELS[0]
149 for (const level of LEVELS) if (probabilities[level] >= probabilities[best]) best = level
150 return best
151}
152
153/**
154 * Turns a Workers AI success envelope into a Recommendation, or says why it
155 * cannot. Exported for tests and the calibration script.
156 */
157export function parseResponse(
158 bodyText: string,
159 opts: { provider: string; rubric: Rubric; latencyMs: number },
160): ProviderResult {
161 const fail = (message: string): ProviderResult => ({
162 ok: false,
163 failure: { kind: "malformed", message, latencyMs: opts.latencyMs },
164 })
165 let envelope: Envelope
166 try {
167 envelope = JSON.parse(bodyText) as Envelope
168 } catch {
169 return fail("response is not JSON")
170 }
171 if (envelope.success === false) return { ok: false, failure: classifyHttpFailure(200, bodyText, opts.latencyMs) }
172 const result = envelope.result as Record<string, unknown> | undefined
173 if (typeof result !== "object" || result === null) return fail("response has no result")
174 const answers = result.answers as Record<string, unknown> | undefined
175 if (typeof answers !== "object" || answers === null) return fail("result has no answers")
176 const difficulty = answers[Q_DIFFICULTY] as Record<string, unknown> | undefined
177 if (typeof difficulty !== "object" || difficulty === null) return fail(`no answer for "${Q_DIFFICULTY}"`)
178
179 const probabilities = levelProbabilities(difficulty, opts.rubric)
180 if (probabilities === undefined) return fail("difficulty answer has no usable probabilities")
181 const level = topLevel(probabilities)
182 const confidence = num(difficulty.confidence)
183 const score = num(difficulty.score)
184
185 const followUp = answers[Q_FOLLOW_UP] as Record<string, unknown> | undefined
186 const contextDependent = followUp && typeof followUp === "object" ? num(followUp.noul) : undefined
187
188 const usage = result.usage as Record<string, unknown> | undefined
189 const recommendation: Recommendation = {
190 provider: opts.provider,
191 level,
192 confidence: probabilities[level],
193 probabilities,
194 latencyMs: opts.latencyMs,
195 }
196 if (confidence !== undefined && confidence >= 0 && confidence <= 1) recommendation.providerConfidence = confidence
197 if (score !== undefined) recommendation.score = score
198 if (contextDependent !== undefined && contextDependent >= 0 && contextDependent <= 1) recommendation.contextDependent = contextDependent
199 const inputTokens = usage ? num(usage.input_tokens) : undefined
200 if (inputTokens !== undefined) recommendation.inputTokens = inputTokens
201 return { ok: true, recommendation }
202}
203
204const TIMED_OUT: unique symbol = Symbol("timeout")
205
206/** The Clef provider. One HTTP request per call, no retries: a retry would
207 * only add latency in the interactive path, and the fallback is cheap. */
208export function clefProvider(opts: ClefOptions): DecisionProvider {
209 const secrets = [opts.apiToken, opts.accountId]
210 return {
211 name: opts.model,
212 async decide(prompt: string): Promise<ProviderResult> {
213 if (!opts.accountId || !opts.apiToken) {
214 return {
215 ok: false,
216 failure: { kind: "not-configured", message: "Cloudflare account ID or API token not set", latencyMs: 0 },
217 }
218 }
219 const started = await opts.now()
220 const elapsed = async () => Math.round((await opts.now()) - started)
221 let response: { status: number; ok: boolean; text: string } | typeof TIMED_OUT
222 const timer = new AbortController()
223 try {
224 const request = opts.fetch(endpoint(opts.accountId, opts.model, opts.apiBase), {
225 method: "POST",
226 headers: { Authorization: `Bearer ${opts.apiToken}`, "Content-Type": "application/json" },
227 body: buildRequestBody(prompt, opts),
228 })
229 // $.http.fetch takes no abort signal: on a timeout the request is left
230 // to finish on its own, and its answer is ignored.
231 request.catch(() => {})
232 const deadline = opts.sleep(opts.timeoutMs, timer.signal).then(
233 (): typeof TIMED_OUT => TIMED_OUT,
234 (): typeof TIMED_OUT => TIMED_OUT,
235 )
236 response = await Promise.race([request, deadline])
237 } catch (error) {
238 const message = redact(error instanceof Error ? error.message : String(error), secrets).slice(0, 200)
239 return { ok: false, failure: { kind: "network", message, latencyMs: await elapsed() } }
240 } finally {
241 timer.abort()
242 }
243 const latencyMs = await elapsed()
244 if (response === TIMED_OUT) {
245 return { ok: false, failure: { kind: "timeout", message: `no answer within ${opts.timeoutMs} ms`, latencyMs } }
246 }
247 if (!response.ok) return { ok: false, failure: classifyHttpFailure(response.status, response.text, latencyMs, secrets) }
248 return parseResponse(response.text, { provider: opts.model, rubric: opts.rubric, latencyMs })
249 },
250 }
251}
252hooks/lib/config.ts 240 lines1// Reads the plugin's `userConfig` values (set with /plugin configure or
2// /config) into a validated Config. Bad values fall back to defaults and are
3// reported, never thrown: a typo in one field must not stop Claude Code.
4
5import { isEffort } from "./models.ts"
6import { CLEF_MODELS, type ClefModel } from "./clef.ts"
7import { BILLINGS, type Billing } from "./pricing.ts"
8import { EFFORTS, LEVELS, type Effort, type Level, type ProfileSpec } from "./types.ts"
9
10export const LOW_CONFIDENCE_POLICIES = ["upper-of-top-two", "bump", "hold", "fallback", "obey"] as const
11export type LowConfidencePolicy = (typeof LOW_CONFIDENCE_POLICIES)[number]
12
13export const ANNOUNCE_MODES = ["status", "answer", "both", "off"] as const
14export type AnnounceMode = (typeof ANNOUNCE_MODES)[number]
15
16export const BILLING_OPTIONS = ["auto", ...BILLINGS] as const
17export type BillingOption = "auto" | Billing
18
19export type Config = {
20 enabled: boolean
21 accountId?: string
22 apiToken?: string
23 decisionModel: ClefModel
24 timeoutMs: number
25 profiles: Record<Level, ProfileSpec>
26 confidenceThreshold: number
27 lowConfidencePolicy: LowConfidencePolicy
28 fallbackLevel: Level
29 /** Honour /model (pause) and /effort (effort only) changes made mid-session. */
30 pauseOnNativeChange: boolean
31 /** P(follow-up) at or above which a prompt never routes below the last route. */
32 followUpThreshold: number
33 /** How the person pays for Claude; `auto` detects it. Decides what a held downgrade costs. */
34 billing: BillingOption
35 /**
36 * A downgrade off a warm cache is taken once staying has cost this many
37 * times what the switch costs; 0 = take every downgrade at once.
38 */
39 downgradePatience: number
40 /** Prompt-cache TTL in minutes; 0 = work it out from the environment. */
41 cacheTtlMinutes: number
42 maxEffort?: Effort
43 dailyNeuronBudget: number
44 maxPromptChars: number
45 announce: AnnounceMode
46 logEnabled: boolean
47 logPrompts: boolean
48 logDir?: string
49 rubricFile?: string
50}
51
52export const DEFAULT_PROFILES: Record<Level, string> = {
53 trivial: "haiku",
54 simple: "sonnet:low",
55 standard: "sonnet:medium",
56 hard: "opus:high",
57 deep: "opus:xhigh",
58}
59
60export const DEFAULTS = {
61 decisionModel: "clef-flash" as ClefModel,
62 timeoutMs: 1500,
63 confidenceThreshold: 0.55,
64 lowConfidencePolicy: "upper-of-top-two" as LowConfidencePolicy,
65 fallbackLevel: "standard" as Level,
66 followUpThreshold: 0.6,
67 billing: "auto" as BillingOption,
68 downgradePatience: 1,
69 cacheTtlMinutes: 0,
70 dailyNeuronBudget: 9_000,
71 maxPromptChars: 6_000,
72 announce: "status" as AnnounceMode,
73}
74
75/** "opus:high" → { model: "opus", effort: "high" }; "haiku" → { model: "haiku" }. */
76export function parseProfile(level: Level, text: string): ProfileSpec | string {
77 const trimmed = text.trim()
78 if (trimmed === "") return `profile_${level} is empty`
79 const split = splitTarget(trimmed)
80 if (split.model === "") return `profile_${level} "${text}" names no model`
81 if (split.badEffort) return `profile_${level} "${text}": effort must be one of ${EFFORTS.join(", ")}`
82 return split.effort ? { level, model: split.model, effort: split.effort } : { level, model: split.model }
83}
84
85/**
86 * Splits "model:effort". A suffix that is not an effort name stays part of
87 * the model, since provider IDs carry colons ("...-v1:0" on Bedrock); a
88 * suffix that looks like a word but is no effort is reported.
89 */
90export function splitTarget(text: string): { model: string; effort?: Effort; badEffort?: boolean } {
91 const colon = text.lastIndexOf(":")
92 if (colon === -1) return { model: text.trim() }
93 const model = text.slice(0, colon).trim()
94 const suffix = text.slice(colon + 1).trim().toLowerCase()
95 if (suffix === "" || suffix === "default") return { model }
96 if (isEffort(suffix)) return { model, effort: suffix }
97 if (/^\d+$/.test(suffix)) return { model: text.trim() }
98 return { model, badEffort: true }
99}
100
101type Options = Readonly<Record<string, unknown>>
102
103function str(options: Options, key: string): string | undefined {
104 const v = options[key]
105 return typeof v === "string" && v.trim() !== "" ? v.trim() : undefined
106}
107
108function numIn(options: Options, key: string, min: number, max: number, fallback: number, problems: string[]): number {
109 const v = options[key]
110 if (v === undefined || v === "") return fallback
111 const n = typeof v === "number" ? v : Number(v)
112 if (!Number.isFinite(n) || n < min || n > max) {
113 problems.push(`${key} must be a number from ${min} to ${max}; using ${fallback}`)
114 return fallback
115 }
116 return n
117}
118
119function oneOf<T extends string>(options: Options, key: string, allowed: readonly T[], fallback: T, problems: string[]): T {
120 const v = str(options, key)
121 if (v === undefined) return fallback
122 if ((allowed as readonly string[]).includes(v)) return v as T
123 problems.push(`${key} must be one of ${allowed.join(", ")}; using ${fallback}`)
124 return fallback
125}
126
127function bool(options: Options, key: string, fallback: boolean): boolean {
128 const v = options[key]
129 if (typeof v === "boolean") return v
130 if (v === "true") return true
131 if (v === "false") return false
132 return fallback
133}
134
135/**
136 * `downgrade_patience`, or what the old `cache_hold_min_tokens` meant by 0
137 * (never hold). Its other values have no equivalent: holding now depends on
138 * what staying has cost, not on the context's size.
139 */
140function downgradePatience(options: Options, problems: string[]): number {
141 if (options.downgrade_patience !== undefined) return numIn(options, "downgrade_patience", 0, 100, DEFAULTS.downgradePatience, problems)
142 const old = options.cache_hold_min_tokens
143 if (old === undefined) return DEFAULTS.downgradePatience
144 if (Number(old) === 0) {
145 problems.push("cache_hold_min_tokens is replaced by downgrade_patience; reading 0 as downgrade_patience 0")
146 return 0
147 }
148 problems.push(`cache_hold_min_tokens is replaced by downgrade_patience and ignored; using ${DEFAULTS.downgradePatience}`)
149 return DEFAULTS.downgradePatience
150}
151
152/**
153 * Reads the optional advanced-settings file (`clef-model-router.json`): a
154 * JSON object with the same keys as the plugin options. Plugin options win
155 * over it; it wins over the defaults.
156 */
157export function parseAdvancedFile(text: string | undefined): { values: Options; problems: string[] } {
158 if (text === undefined) return { values: {}, problems: [] }
159 try {
160 const value = JSON.parse(text) as unknown
161 if (typeof value !== "object" || value === null || Array.isArray(value)) {
162 return { values: {}, problems: ["clef-model-router.json must hold a JSON object; ignoring it"] }
163 }
164 const values = { ...(value as Record<string, unknown>) }
165 // Credentials belong in the plugin's secure storage, not in a plain file.
166 const problems: string[] = []
167 if ("cloudflare_api_token" in values) {
168 delete values.cloudflare_api_token
169 problems.push("clef-model-router.json: cloudflare_api_token is ignored there; set it with /plugin configure")
170 }
171 return { values, problems }
172 } catch {
173 return { values: {}, problems: ["clef-model-router.json is not valid JSON; ignoring it"] }
174 }
175}
176
177/**
178 * Builds the Config from plugin options plus environment fallbacks for the
179 * credentials (CLOUDFLARE_ACCOUNT_ID / CLOUDFLARE_API_TOKEN), which the
180 * caller reads and passes in.
181 */
182export function parseConfig(
183 options: Options,
184 envFallback: { accountId?: string; apiToken?: string } = {},
185): { config: Config; problems: string[] } {
186 const problems: string[] = []
187 const profiles = {} as Record<Level, ProfileSpec>
188 // `profiles` is the five of them in one line, lowest first; a `profile_<level>`
189 // key (in the advanced file) names one.
190 const list = str(options, "profiles")?.split(",").map((s) => s.trim())
191 if (list && list.length !== LEVELS.length) {
192 problems.push(`profiles must list ${LEVELS.length} entries (${LEVELS.join(", ")}), got ${list.length}; using the defaults`)
193 }
194 for (const [i, level] of LEVELS.entries()) {
195 const fromList = list && list.length === LEVELS.length ? list[i] : undefined
196 const raw = str(options, `profile_${level}`) ?? fromList ?? DEFAULT_PROFILES[level]
197 const parsed = parseProfile(level, raw)
198 if (typeof parsed === "string") {
199 problems.push(`${parsed}; using "${DEFAULT_PROFILES[level]}"`)
200 profiles[level] = parseProfile(level, DEFAULT_PROFILES[level]) as ProfileSpec
201 } else profiles[level] = parsed
202 }
203 const maxEffortRaw = str(options, "max_effort")
204 let maxEffort: Effort | undefined
205 if (maxEffortRaw !== undefined && maxEffortRaw !== "none") {
206 if (isEffort(maxEffortRaw)) maxEffort = maxEffortRaw
207 else problems.push(`max_effort must be one of ${EFFORTS.join(", ")} or none; ignoring it`)
208 }
209
210 const config: Config = {
211 enabled: bool(options, "enabled", true),
212 decisionModel: oneOf(options, "decision_model", CLEF_MODELS, DEFAULTS.decisionModel, problems),
213 timeoutMs: numIn(options, "timeout_ms", 100, 10_000, DEFAULTS.timeoutMs, problems),
214 profiles,
215 confidenceThreshold: numIn(options, "confidence_threshold", 0, 1, DEFAULTS.confidenceThreshold, problems),
216 lowConfidencePolicy: oneOf(options, "low_confidence_policy", LOW_CONFIDENCE_POLICIES, DEFAULTS.lowConfidencePolicy, problems),
217 fallbackLevel: oneOf(options, "fallback_profile", LEVELS, DEFAULTS.fallbackLevel, problems),
218 followUpThreshold: numIn(options, "follow_up_threshold", 0, 1, DEFAULTS.followUpThreshold, problems),
219 billing: oneOf(options, "billing", BILLING_OPTIONS, DEFAULTS.billing, problems),
220 downgradePatience: downgradePatience(options, problems),
221 cacheTtlMinutes: numIn(options, "cache_ttl_minutes", 0, 1440, DEFAULTS.cacheTtlMinutes, problems),
222 dailyNeuronBudget: numIn(options, "daily_neuron_budget", 0, 1_000_000_000, DEFAULTS.dailyNeuronBudget, problems),
223 maxPromptChars: numIn(options, "max_prompt_chars", 200, 200_000, DEFAULTS.maxPromptChars, problems),
224 announce: oneOf(options, "announce", ANNOUNCE_MODES, DEFAULTS.announce, problems),
225 pauseOnNativeChange: bool(options, "pause_on_native_change", true),
226 logEnabled: bool(options, "log_enabled", true),
227 logPrompts: bool(options, "log_prompts", false),
228 }
229 if (maxEffort) config.maxEffort = maxEffort
230 const accountId = str(options, "cloudflare_account_id") ?? envFallback.accountId
231 const apiToken = str(options, "cloudflare_api_token") ?? envFallback.apiToken
232 if (accountId) config.accountId = accountId
233 if (apiToken) config.apiToken = apiToken
234 const logDir = str(options, "log_dir")
235 if (logDir) config.logDir = logDir
236 const rubricFile = str(options, "rubric_file")
237 if (rubricFile) config.rubricFile = rubricFile
238 return { config, problems }
239}
240hooks/lib/env.ts 115 lines1// Turns the environment variables Claude Code itself honours into the facts
2// the policy needs. Pure: the hooks module reads the variables (by literal
3// name, as the engine requires) and passes them here.
4
5import type { ModelEnv } from "./models.ts"
6import type { Billing } from "./pricing.ts"
7
8export type EnvValues = {
9 ANTHROPIC_DEFAULT_HAIKU_MODEL?: string
10 ANTHROPIC_DEFAULT_SONNET_MODEL?: string
11 ANTHROPIC_DEFAULT_OPUS_MODEL?: string
12 ANTHROPIC_DEFAULT_FABLE_MODEL?: string
13 CLAUDE_CODE_USE_BEDROCK?: string
14 CLAUDE_CODE_USE_VERTEX?: string
15 CLAUDE_CODE_USE_FOUNDRY?: string
16 CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS?: string
17 ANTHROPIC_API_KEY?: string
18 CLAUDE_CODE_PROMPT_CACHE_TTL?: string
19 FORCE_PROMPT_CACHING_5M?: string
20 ENABLE_PROMPT_CACHING_1H?: string
21 ANTHROPIC_AUTH_TOKEN?: string
22 ANTHROPIC_BASE_URL?: string
23}
24
25const truthy = (v: string | undefined) => v !== undefined && v !== "" && v !== "0" && v.toLowerCase() !== "false"
26
27export function modelEnvFrom(env: EnvValues): ModelEnv {
28 const defaults: ModelEnv["defaults"] = {}
29 if (env.ANTHROPIC_DEFAULT_HAIKU_MODEL) defaults.haiku = env.ANTHROPIC_DEFAULT_HAIKU_MODEL
30 if (env.ANTHROPIC_DEFAULT_SONNET_MODEL) defaults.sonnet = env.ANTHROPIC_DEFAULT_SONNET_MODEL
31 if (env.ANTHROPIC_DEFAULT_OPUS_MODEL) defaults.opus = env.ANTHROPIC_DEFAULT_OPUS_MODEL
32 if (env.ANTHROPIC_DEFAULT_FABLE_MODEL) defaults.fable = env.ANTHROPIC_DEFAULT_FABLE_MODEL
33 return {
34 defaults,
35 thirdParty: truthy(env.CLAUDE_CODE_USE_BEDROCK) || truthy(env.CLAUDE_CODE_USE_VERTEX) || truthy(env.CLAUDE_CODE_USE_FOUNDRY),
36 betasDisabled: truthy(env.CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS),
37 }
38}
39
40const FIVE_MIN = 5 * 60_000
41const ONE_HOUR = 60 * 60_000
42
43function ttlValue(v: string | undefined): number | undefined {
44 if (v === "5m") return FIVE_MIN
45 if (v === "1h") return ONE_HOUR
46 return undefined
47}
48
49/** One rate-limit window as `$.session.usage()` reports it. */
50export type RateLimitWindow = { kind: string; percentUsed: number }
51
52/** What the session shows about how the person pays for Claude. */
53export type BillingFacts = {
54 billing: Billing
55 /** Why, in a few words, for /clef. */
56 why: string
57 /** A subscription past its included usage, drawing usage credits billed per token. */
58 overage: boolean
59}
60
61/** The windows a Claude subscription reports; an API key or cloud provider reports none. */
62const PLAN_WINDOWS = ["five_hour", "seven_day"]
63
64/**
65 * How the person pays, from the strongest evidence available:
66 *
67 * 1. The plan's rate-limit windows (`five_hour`, `seven_day`), which Claude
68 * Code reports only on a subscription. One at 100% or more means the plan's
69 * included usage is spent and further requests are usage credits, billed
70 * per token like the API.
71 * 2. A gateway's `spend_limit` window: billed per token.
72 * 3. The environment: a cloud provider, an API key (or `apiKeyHelper`), or a
73 * gateway (`ANTHROPIC_AUTH_TOKEN`, `ANTHROPIC_BASE_URL`) bill per token.
74 * Nothing set means a claude.ai login, which is a subscription.
75 *
76 * The windows are empty until the session's first response, so the first
77 * turn is judged on the environment alone.
78 */
79export function detectBilling(
80 env: EnvValues,
81 opts: { apiKeyHelper?: boolean; rateLimits?: readonly RateLimitWindow[] } = {},
82): BillingFacts {
83 const windows = opts.rateLimits ?? []
84 const plan = windows.filter((w) => PLAN_WINDOWS.includes(w.kind))
85 if (plan.length > 0) {
86 if (plan.some((w) => w.percentUsed >= 100)) return { billing: "api", why: "plan limit reached: usage credits bill per token", overage: true }
87 return { billing: "subscription", why: "the plan's rate limits are reported", overage: false }
88 }
89 if (windows.some((w) => w.kind === "spend_limit")) return { billing: "api", why: "a gateway spend limit is reported", overage: false }
90 if (modelEnvFrom(env).thirdParty) return { billing: "api", why: "a cloud provider bills per token", overage: false }
91 if (env.ANTHROPIC_API_KEY || opts.apiKeyHelper) return { billing: "api", why: "an API key is set", overage: false }
92 if (env.ANTHROPIC_AUTH_TOKEN || env.ANTHROPIC_BASE_URL) return { billing: "api", why: "a gateway is set (ANTHROPIC_BASE_URL)", overage: false }
93 return { billing: "subscription", why: "no API key, provider or gateway is set", overage: false }
94}
95
96/**
97 * The main conversation's prompt-cache TTL, resolved in the order Claude
98 * Code's prompt-caching docs give: FORCE_PROMPT_CACHING_5M, the TTL variable,
99 * the `promptCacheTtl` setting, ENABLE_PROMPT_CACHING_1H, then the default:
100 * one hour on a subscription within its included usage, five minutes
101 * otherwise (an API key, a cloud provider, usage credits). The default follows
102 * what was detected, not the `billing` option: the option says what the
103 * person wants optimised, the TTL is what Claude Code actually requests.
104 */
105export function cacheTtlMs(configMinutes: number, env: EnvValues, settingsTtl?: unknown, detected: BillingFacts = detectBilling(env)): number {
106 if (configMinutes > 0) return configMinutes * 60_000
107 if (truthy(env.FORCE_PROMPT_CACHING_5M)) return FIVE_MIN
108 const fromEnv = ttlValue(env.CLAUDE_CODE_PROMPT_CACHE_TTL)
109 if (fromEnv) return fromEnv
110 const fromSettings = typeof settingsTtl === "string" ? ttlValue(settingsTtl) : undefined
111 if (fromSettings) return fromSettings
112 if (truthy(env.ENABLE_PROMPT_CACHING_1H)) return ONE_HOUR
113 return detected.billing === "subscription" ? ONE_HOUR : FIVE_MIN
114}
115hooks/lib/format.ts 242 lines1// Everything the router shows: the one-line status, and the text of each
2// /clef subcommand. Plain text, so it reads the same on every surface.
3
4import type { BillingOption, Config } from "./config.ts"
5import type { BillingFacts } from "./env.ts"
6import { displayName, resolveModel, type ModelEnv } from "./models.ts"
7import { describe, kTokens, pct, type HoldState, type RouterMode } from "./policy.ts"
8import { dollars, type Billing } from "./pricing.ts"
9import { presence } from "./redact.ts"
10import type { Stats } from "./log.ts"
11import { LEVELS, type Decision, type Level, type Target } from "./types.ts"
12
13const RULE_LABEL: Record<string, string> = {
14 "low-confidence": "unsure",
15 "context-dependent": "follow-up",
16 unavailable: "unavailable",
17 "context-window": "window",
18 "cache-hold": "held for cache",
19 "effort-clamp": "effort clamped",
20 "effort-cap": "effort capped",
21 "pinned-effort": "pinned effort",
22}
23
24/**
25 * A held downgrade names what it held back, so a route that differs from
26 * Clef's pick says so: "(Sonnet deferred)", or "(low effort deferred)".
27 */
28function holdLabel(d: Decision): string {
29 const w = d.deferral?.wanted
30 if (!w || !d.final) return RULE_LABEL["cache-hold"]!
31 return w.model === d.final.model ? `${w.effort ?? "lower"} effort deferred` : `${displayName(w.model)} deferred`
32}
33
34/** The one line under the prompt, e.g. "Clef → Sonnet · medium · 87%". */
35export function statusLine(d: Decision): string | undefined {
36 const route = d.final ? describe(d.final) : undefined
37 const notes = d.adjustments.map((a) => (a.rule === "cache-hold" ? holdLabel(d) : (RULE_LABEL[a.rule] ?? a.rule)))
38 const tail = notes.length > 0 ? ` (${[...new Set(notes)].join(", ")})` : ""
39 switch (d.source) {
40 case "clef": {
41 const conf = d.recommendation ? ` · ${pct(d.recommendation.confidence)}` : ""
42 return `Clef → ${route}${conf}${tail}`
43 }
44 case "fallback":
45 return `Clef ✕ ${d.failure?.kind ?? "no answer"} → ${route}${tail}`
46 case "continuation":
47 return route ? `Clef ↻ ${route}${tail}` : undefined
48 case "override":
49 return route ? `+ ${route}${tail}` : `+ ${d.note ?? "override"}`
50 case "pin":
51 return route ? `Pinned → ${route}${tail}` : `Pinned: ${d.note ?? ""}`
52 case "native":
53 return "Clef paused (you chose /model)"
54 case "disabled":
55 return d.note?.startsWith("+off") ? "Clef skipped this turn" : "Clef off"
56 }
57}
58
59/** A line under the answer, when `announce` asks for one. */
60export function answerLine(d: Decision, latencyMs?: number): string | undefined {
61 const line = statusLine(d)
62 if (!line) return undefined
63 return latencyMs !== undefined && d.source === "clef" ? `${line} · ${latencyMs} ms` : line
64}
65
66function bar(p: number, width = 20): string {
67 const filled = Math.round(p * width)
68 return "█".repeat(filled) + "·".repeat(width - filled)
69}
70
71export function distribution(probabilities: Record<Level, number>, mark?: Level, final?: Level): string[] {
72 return LEVELS.map((level) => {
73 const tags = [level === mark ? "← Clef" : "", level === final && final !== mark ? "← routed" : ""].filter(Boolean).join(" ")
74 return ` ${level.padEnd(9)} ${bar(probabilities[level])} ${pct(probabilities[level]).padStart(4)} ${tags}`.trimEnd()
75 })
76}
77
78export function explain(d: Decision): string[] {
79 const lines: string[] = []
80 lines.push(` route ${d.final ? `${describe(d.final)} (${d.final.model}${d.final.level ? `, profile ${d.final.level}` : ""})` : "untouched (Claude Code's own model and effort)"}`)
81 lines.push(` source ${d.source}${d.note ? ` — ${d.note}` : ""}`)
82 const rec = d.recommendation
83 if (rec) {
84 const extras = [
85 `${pct(rec.confidence)} on ${rec.level}`,
86 rec.providerConfidence !== undefined ? `clef confidence ${pct(rec.providerConfidence)}` : "",
87 rec.score !== undefined ? `score ${rec.score.toFixed(2)}/4` : "",
88 rec.contextDependent !== undefined ? `follow-up ${pct(rec.contextDependent)}` : "",
89 `${rec.latencyMs} ms`,
90 rec.inputTokens !== undefined ? `${rec.inputTokens} tokens` : "",
91 ].filter(Boolean)
92 lines.push(` clef ${rec.provider}: ${extras.join(" · ")}`)
93 lines.push(...distribution(rec.probabilities, rec.level, d.final?.level))
94 }
95 if (d.failure) lines.push(` failure ${d.failure.kind}: ${d.failure.message}${d.failure.latencyMs ? ` (${d.failure.latencyMs} ms)` : ""}`)
96 for (const a of d.adjustments) lines.push(` policy ${a.rule}: ${a.from} → ${a.to} — ${a.reason}`)
97 const f = d.deferral
98 if (f) {
99 const what = f.wanted.model === f.from ? `${f.wanted.effort ?? "lower"} effort` : displayName(f.wanted.model)
100 lines.push(
101 f.held
102 ? ` downgrade ${what} deferred (held turn ${f.turns}): staying has cost ${dollars(f.spent)} so far, a switch costs ${dollars(f.cost)} now (${f.billing}, list prices)`
103 : ` downgrade ${what} taken after ${f.turns} held turn${f.turns === 1 ? "" : "s"}: staying had cost ${dollars(f.spent)}, the switch ${dollars(f.cost)} (${f.billing}, list prices)`,
104 )
105 }
106 return lines
107}
108
109export type StatusArgs = {
110 config: Config
111 configProblems: readonly string[]
112 modelEnv: ModelEnv
113 mode: RouterMode
114 pin?: Target
115 last?: Decision
116 guard: { calls: number; inputTokens: number; neurons: number; blocked?: string }
117 cache?: { model: string; promptTokens: number; ageSeconds: number; ttlSeconds: number }
118 billing: { billing: Billing; detected: BillingFacts; configured: BillingOption; patience: number }
119 hold?: HoldState
120 unavailable: readonly string[]
121 logDir: string | undefined
122 advancedPath?: string
123}
124
125export function targetText(t: Target): string {
126 return [t.level ?? t.model ?? "", t.effort ? `:${t.effort}` : ""].join("")
127}
128
129export function statusReport(a: StatusArgs): string {
130 const c = a.config
131 const lines = ["Clef router"]
132 const mode =
133 !c.enabled ? "disabled in plugin config" : a.mode === "auto" ? (a.pin ? `pinned to ${targetText(a.pin)} (/clef auto to unpin)` : "auto") : a.mode === "off" ? "off for this session (/clef on)" : "paused: you changed /model (/clef auto to resume)"
134 lines.push(` mode ${mode}`)
135 lines.push(` decider ${c.decisionModel} · timeout ${c.timeoutMs} ms · account ${presence(c.accountId)} · token ${presence(c.apiToken)}`)
136 const budget = c.dailyNeuronBudget > 0 ? ` of ${c.dailyNeuronBudget} budget` : ""
137 lines.push(` today ${a.guard.calls} Clef calls · ${kTokens(a.guard.inputTokens)} input tokens · ~${Math.round(a.guard.neurons)} neurons${budget}${a.guard.blocked ? ` · ${a.guard.blocked}` : ""}`)
138 const b = a.billing
139 const source = b.configured === "auto" ? `detected: ${b.detected.why}` : `set in config; detected ${b.detected.billing} (${b.detected.why})`
140 const patience = b.patience === 0 ? "downgrades taken at once" : `downgrade patience ${b.patience}`
141 lines.push(` billing ${b.billing} (${source}) · ${patience}`)
142 if (a.cache) {
143 const warm = a.cache.ageSeconds < a.cache.ttlSeconds
144 lines.push(` cache ${displayName(a.cache.model)} · ${kTokens(a.cache.promptTokens)} context · ${warm ? `warm (${a.cache.ageSeconds}s of ${a.cache.ttlSeconds}s)` : "cold"}`)
145 }
146 if (a.hold) {
147 const what = a.hold.wanted.model === a.hold.model ? `${a.hold.wanted.effort ?? "lower"} effort` : displayName(a.hold.wanted.model)
148 lines.push(` deferred ${what} · ${a.hold.turns} held turn${a.hold.turns === 1 ? "" : "s"} on ${displayName(a.hold.model)} · staying has cost ${dollars(a.hold.spent)} (list prices)`)
149 }
150 if (a.unavailable.length > 0) lines.push(` unusable ${a.unavailable.join(", ")}`)
151 for (const p of a.configProblems) lines.push(` config! ${p}`)
152 if (a.last) {
153 lines.push("", "Last turn")
154 lines.push(...explain(a.last))
155 } else {
156 lines.push("", "No turn routed yet this session.")
157 }
158 if (a.advancedPath) lines.push("", `Advanced settings: ${a.advancedPath} (optional)`)
159 lines.push(`Log: ${c.logEnabled ? (a.logDir ?? "(unavailable)") : "off"}${c.logEnabled && c.logPrompts ? " (with prompt text)" : ""}`)
160 lines.push("Commands: /clef history · stats · profiles · test <prompt> · pin <target> · auto · off · on · feedback under|ok|over")
161 return lines.join("\n")
162}
163
164export function profilesReport(config: Config, env: ModelEnv, unavailable: readonly string[]): string {
165 const lines = ["Profiles, lowest first. Change them with /config (Profiles) or /plugin configure."]
166 for (const level of LEVELS) {
167 const spec = config.profiles[level]
168 const id = resolveModel(spec.model, env)
169 const state = id === undefined ? "cannot resolve here (set a full model ID)" : unavailable.includes(id) ? `${id} (failed this session)` : id
170 lines.push(` ${level.padEnd(9)} ${`${spec.model}${spec.effort ? `:${spec.effort}` : ""}`.padEnd(16)} → ${state}`)
171 }
172 lines.push(` fallback ${config.fallbackLevel} · low confidence (< ${pct(config.confidenceThreshold)}): ${config.lowConfidencePolicy}`)
173 return lines.join("\n")
174}
175
176export type HistoryRow = { decision: Decision; prompt: string; answeredModel?: string }
177
178export function historyReport(rows: readonly HistoryRow[]): string {
179 if (rows.length === 0) return "No turns yet this session."
180 const lines = ["Turns this session, newest first"]
181 lines.push(" ms route conf source prompt")
182 for (const row of [...rows].reverse()) {
183 const d = row.decision
184 const ms = d.recommendation?.latencyMs ?? d.failure?.latencyMs
185 const conf = d.recommendation ? pct(d.recommendation.confidence) : "—"
186 const route = d.final ? describe(d.final) : "untouched"
187 const flag = d.adjustments.length > 0 ? "*" : " "
188 const prompt = row.prompt.replace(/\s+/g, " ").slice(0, 48)
189 lines.push(` ${String(ms ?? "—").padStart(5)} ${(route + flag).padEnd(22)} ${conf.padStart(4)} ${d.source.padEnd(12)} ${prompt}`)
190 for (const a of d.adjustments) lines.push(` ${a.rule}: ${a.from} → ${a.to}`)
191 if (d.failure) lines.push(` ${d.failure.kind}: ${d.failure.message}`)
192 if (row.answeredModel && d.final && !row.answeredModel.startsWith(d.final.model)) lines.push(` answered by ${row.answeredModel}`)
193 }
194 lines.push(" * policy changed Clef's recommendation")
195 return lines.join("\n")
196}
197
198function table(title: string, map: Record<string, number>, total: number): string[] {
199 const entries = Object.entries(map).sort((a, b) => b[1] - a[1])
200 if (entries.length === 0) return []
201 return [` ${title}`, ...entries.map(([k, v]) => ` ${k.padEnd(24)} ${String(v).padStart(5)} ${pct(total ? v / total : 0).padStart(4)}`)]
202}
203
204export function statsReport(s: Stats, days: number, files: number): string {
205 if (s.turns === 0 && Object.keys(s.feedback).length === 0) return `No routing log entries in the last ${days} day(s).`
206 const lines = [`Routing over the last ${days} day(s): ${s.turns} turns in ${files} log file(s)`]
207 lines.push(...table("by model", s.byModel, s.turns))
208 lines.push(...table("by effort", s.byEffort, s.turns))
209 lines.push(...table("by profile", s.byLevel, s.turns))
210 lines.push(...table("by source", s.bySource, s.turns))
211 lines.push(" clef")
212 lines.push(` calls ${s.clefCalls} · ${kTokens(s.clefInputTokens)} input tokens`)
213 if (s.latency) lines.push(` latency mean ${s.latency.mean} ms · p50 ${s.latency.p50} ms · p95 ${s.latency.p95} ms`)
214 if (s.meanConfidence !== undefined) lines.push(` mean probability of Clef's pick ${pct(s.meanConfidence)}`)
215 lines.push(` recommendation changed by policy: ${s.recommendationChanged} · manual overrides: ${s.overrides}`)
216 lines.push(` turns held on a warm model: ${s.cacheHolds} · downgrades taken after a hold: ${s.downgradesTaken}`)
217 const failures = Object.entries(s.failures)
218 if (failures.length > 0) lines.push(` fallbacks: ${failures.map(([k, v]) => `${k} ${v}`).join(", ")}`)
219 const fb = Object.entries(s.feedback)
220 if (fb.length > 0) lines.push(` your feedback: ${fb.map(([k, v]) => `${k} ${v}`).join(", ")}`)
221 return lines.join("\n")
222}
223
224export const HELP = [
225 "Clef router — picks the Claude model and effort for each turn.",
226 "",
227 " /clef status and the last decision, with Clef's probabilities",
228 " /clef history this session's turns",
229 " /clef stats [days] totals from the local log (default 7 days)",
230 " /clef profiles what each difficulty level runs on",
231 " /clef test <prompt> ask Clef about a prompt without sending it to Claude",
232 " /clef pin <target> use one target for the rest of the session",
233 " target: trivial|simple|standard|hard|deep, haiku|sonnet|opus|fable,",
234 " a model ID, with optional :effort; or :effort alone",
235 " /clef auto unpin, resume after /model, and hand effort back after /effort",
236 " /clef off | on stop or resume routing for this session",
237 " /clef feedback under|ok|over [note] rate the last route, for later analysis",
238 "",
239 "One turn only: start a prompt with +target, e.g. `+opus:max why does this deadlock?`,",
240 "or `+off ...` to leave that turn to Claude Code.",
241].join("\n")
242hooks/lib/guard.ts 131 lines1// Deterministic gates in front of the Clef call: a local daily budget that
2// keeps usage inside Workers AI's free allocation, a pause after the
3// allocation is reported exhausted, and a circuit breaker so a dead endpoint
4// costs one timeout, not one per prompt.
5//
6// The state is plain JSON kept in `$.store`, shared by every session on the
7// machine (best effort: two sessions writing at once may lose a count).
8
9import type { FailureKind, ProviderFailure } from "./types.ts"
10
11export type GuardState = {
12 /** UTC day (YYYY-MM-DD) the counters belong to. */
13 day: string
14 calls: number
15 inputTokens: number
16 /** Clef said the day's free allocation is used up; no calls until `day` changes. */
17 quotaExhausted?: boolean
18 consecutiveFailures: number
19 /** Epoch ms until which no call is made. */
20 pausedUntil?: number
21 pauseReason?: FailureKind
22}
23
24/**
25 * Neurons per million input tokens, derived from Cloudflare's published
26 * prices ($0.09/M for clef-flash, $0.24/M for clef) at $0.011 per 1,000
27 * neurons. Clef is not yet in the per-model neuron table; this is an
28 * estimate, and the Workers AI dashboard is the authority.
29 */
30export const NEURONS_PER_M_INPUT: Record<string, number> = {
31 "clef-flash": (0.09 / 0.011) * 1000,
32 clef: (0.24 / 0.011) * 1000,
33}
34
35export function utcDay(epochMs: number): string {
36 return new Date(epochMs).toISOString().slice(0, 10)
37}
38
39export function freshGuard(epochMs: number): GuardState {
40 return { day: utcDay(epochMs), calls: 0, inputTokens: 0, consecutiveFailures: 0 }
41}
42
43/** Rolls the counters over at 00:00 UTC, when Workers AI's allocation resets. */
44export function normaliseGuard(state: unknown, epochMs: number): GuardState {
45 const today = utcDay(epochMs)
46 if (typeof state !== "object" || state === null) return freshGuard(epochMs)
47 const s = state as Partial<GuardState>
48 if (s.day !== today) {
49 const next = freshGuard(epochMs)
50 if (typeof s.pausedUntil === "number" && s.pausedUntil > epochMs && s.pauseReason !== "quota") {
51 next.pausedUntil = s.pausedUntil
52 next.pauseReason = s.pauseReason
53 }
54 return next
55 }
56 return {
57 day: today,
58 calls: typeof s.calls === "number" ? s.calls : 0,
59 inputTokens: typeof s.inputTokens === "number" ? s.inputTokens : 0,
60 consecutiveFailures: typeof s.consecutiveFailures === "number" ? s.consecutiveFailures : 0,
61 ...(s.quotaExhausted ? { quotaExhausted: true } : {}),
62 ...(typeof s.pausedUntil === "number" ? { pausedUntil: s.pausedUntil } : {}),
63 ...(s.pauseReason ? { pauseReason: s.pauseReason } : {}),
64 }
65}
66
67export function estimatedNeurons(state: GuardState, model: string): number {
68 return (state.inputTokens / 1_000_000) * (NEURONS_PER_M_INPUT[model] ?? NEURONS_PER_M_INPUT.clef!)
69}
70
71/** Why no call should be made now, or undefined to go ahead. */
72export function blockedReason(
73 state: GuardState,
74 opts: { now: number; model: string; dailyNeuronBudget: number },
75): ProviderFailure | undefined {
76 if (state.quotaExhausted) {
77 return { kind: "quota", message: "Workers AI daily free allocation used up; resets 00:00 UTC", latencyMs: 0 }
78 }
79 if (opts.dailyNeuronBudget > 0 && estimatedNeurons(state, opts.model) >= opts.dailyNeuronBudget) {
80 return {
81 kind: "budget",
82 message: `local daily budget of ${opts.dailyNeuronBudget} neurons reached; resets 00:00 UTC`,
83 latencyMs: 0,
84 }
85 }
86 if (state.pausedUntil !== undefined && state.pausedUntil > opts.now) {
87 const seconds = Math.ceil((state.pausedUntil - opts.now) / 1000)
88 return {
89 kind: "circuit-open",
90 message: `paused ${seconds}s after ${state.pauseReason ?? "repeated failures"}`,
91 latencyMs: 0,
92 }
93 }
94 return undefined
95}
96
97/** How long to stop calling after a failure, in ms; 0 for none. */
98export const PAUSES = {
99 /** After this many failures in a row, stop calling for a while. */
100 breakerThreshold: 3,
101 breakerMs: 5 * 60_000,
102 rateLimitedMs: 60_000,
103 /** Auth and request errors need the user to fix configuration. */
104 configMs: 30 * 60_000,
105}
106
107export function recordSuccess(state: GuardState, inputTokens: number | undefined): GuardState {
108 const next: GuardState = {
109 ...state,
110 calls: state.calls + 1,
111 inputTokens: state.inputTokens + (inputTokens ?? 0),
112 consecutiveFailures: 0,
113 }
114 delete next.pausedUntil
115 delete next.pauseReason
116 return next
117}
118
119export function recordFailure(state: GuardState, failure: ProviderFailure, now: number): GuardState {
120 // Failures that never reached Cloudflare count toward nothing.
121 if (failure.kind === "not-configured" || failure.kind === "budget" || failure.kind === "circuit-open") return state
122 const next: GuardState = { ...state, calls: state.calls + 1, consecutiveFailures: state.consecutiveFailures + 1 }
123 if (failure.kind === "quota") return { ...next, quotaExhausted: true }
124 let pause = 0
125 if (failure.kind === "rate-limited") pause = PAUSES.rateLimitedMs
126 else if (failure.kind === "auth" || failure.kind === "bad-request") pause = PAUSES.configMs
127 else if (next.consecutiveFailures >= PAUSES.breakerThreshold) pause = PAUSES.breakerMs
128 if (pause > 0) return { ...next, pausedUntil: now + pause, pauseReason: failure.kind }
129 return next
130}
131hooks/lib/log.ts 248 lines1// The local routing log: one JSON object per line, one file per UTC day and
2// session, under the log directory. Nothing is sent anywhere. Prompt text is
3// left out unless `log_prompts` is on; a short SHA-256 prefix lets repeated
4// prompts be recognised without storing them.
5
6import type { Decision, Deferral, Effort, Level, Source } from "./types.ts"
7
8export const LOG_VERSION = 1
9
10export type AnsweredUsage = {
11 model: string
12 inputTokens: number
13 outputTokens: number
14 cacheReadTokens: number
15 cacheWriteTokens: number
16}
17
18export type TurnRecord = {
19 v: number
20 type: "turn"
21 ts: string
22 session: string
23 turn: string
24 kind: Decision["kind"]
25 source: Source
26 promptHash?: string
27 promptChars: number
28 prompt?: string
29 provider?: string
30 recommendation?: {
31 level: Level
32 /** Probability of `level`: what the threshold compares. */
33 confidence: number
34 /** Clef's own `confidence` field (entropy-like; not a probability). */
35 clefConfidence?: number
36 probabilities: Record<Level, number>
37 score?: number
38 followUp?: number
39 }
40 latencyMs?: number
41 clefInputTokens?: number
42 proposed?: { level?: Level; model: string; effort?: Effort }
43 final?: { level?: Level; model: string; effort?: Effort }
44 adjustments: { rule: string; from: string; to: string; reason: string }[]
45 failure?: { kind: string; message: string; status?: number }
46 note?: string
47 /** What the API said answered, summed over the turn's main-loop requests. */
48 answered?: AnsweredUsage
49 /** Cache read and write of the turn's first request: what a model switch, or a return, cost. */
50 firstStep?: { cacheReadTokens: number; cacheWriteTokens: number }
51 /** How the person pays, as the policy saw it. */
52 billing?: "subscription" | "api"
53 /** A downgrade weighed against a warm cache: held, or taken after a held stretch. */
54 deferral?: Deferral
55 /** The plan's rate-limit windows at the start of the turn (subscriptions only). */
56 rateLimits?: { kind: string; percentUsed: number }[]
57 steps?: number
58 durationMs?: number
59 endReason?: string
60}
61
62export type FeedbackRecord = {
63 v: number
64 type: "feedback"
65 ts: string
66 session: string
67 turn?: string
68 verdict: "under" | "ok" | "over"
69 note?: string
70}
71
72export type LogRecord = TurnRecord | FeedbackRecord
73
74export async function promptHash(text: string): Promise<string> {
75 const bytes = new TextEncoder().encode(text)
76 const digest = await crypto.subtle.digest("SHA-256", bytes)
77 return Array.from(new Uint8Array(digest).slice(0, 8), (b) => b.toString(16).padStart(2, "0")).join("")
78}
79
80export function turnRecord(args: {
81 decision: Decision
82 session: string
83 ts: string
84 promptText: string
85 hash?: string
86 logPrompts: boolean
87 answered?: AnsweredUsage
88 firstStep?: { cacheReadTokens: number; cacheWriteTokens: number }
89 rateLimits?: { kind: string; percentUsed: number }[]
90 steps?: number
91 durationMs?: number
92 endReason?: string
93}): TurnRecord {
94 const { decision: d } = args
95 const record: TurnRecord = {
96 v: LOG_VERSION,
97 type: "turn",
98 ts: args.ts,
99 session: args.session,
100 turn: d.turnId,
101 kind: d.kind,
102 source: d.source,
103 promptChars: args.promptText.length,
104 adjustments: d.adjustments.map((a) => ({ ...a })),
105 }
106 if (args.hash) record.promptHash = args.hash
107 if (args.logPrompts) record.prompt = args.promptText
108 const rec = d.recommendation
109 if (rec) {
110 record.provider = rec.provider
111 record.recommendation = { level: rec.level, confidence: rec.confidence, probabilities: { ...rec.probabilities } }
112 if (rec.providerConfidence !== undefined) record.recommendation.clefConfidence = rec.providerConfidence
113 if (rec.score !== undefined) record.recommendation.score = rec.score
114 if (rec.contextDependent !== undefined) record.recommendation.followUp = rec.contextDependent
115 record.latencyMs = rec.latencyMs
116 if (rec.inputTokens !== undefined) record.clefInputTokens = rec.inputTokens
117 }
118 if (d.failure) {
119 record.failure = { kind: d.failure.kind, message: d.failure.message }
120 if (d.failure.status !== undefined) record.failure.status = d.failure.status
121 if (d.failure.latencyMs > 0) record.latencyMs = d.failure.latencyMs
122 }
123 if (d.proposed) record.proposed = { ...d.proposed }
124 if (d.final) record.final = { ...d.final }
125 if (d.note) record.note = d.note
126 if (args.answered) record.answered = { ...args.answered }
127 if (args.firstStep) record.firstStep = { ...args.firstStep }
128 if (d.billing) record.billing = d.billing
129 if (d.deferral) record.deferral = { ...d.deferral, wanted: { ...d.deferral.wanted } }
130 if (args.rateLimits && args.rateLimits.length > 0) record.rateLimits = args.rateLimits.map((w) => ({ ...w }))
131 if (args.steps !== undefined) record.steps = args.steps
132 if (args.durationMs !== undefined) record.durationMs = args.durationMs
133 if (args.endReason) record.endReason = args.endReason
134 return record
135}
136
137export function logFileName(ts: string, session: string): string {
138 const safe = session.replace(/[^A-Za-z0-9_-]/g, "").slice(0, 12) || "session"
139 return `routing-${ts.slice(0, 10)}-${safe}.jsonl`
140}
141
142export function parseLines(text: string): LogRecord[] {
143 const out: LogRecord[] = []
144 for (const line of text.split("\n")) {
145 if (line.trim() === "") continue
146 try {
147 const value = JSON.parse(line) as LogRecord
148 if (value && typeof value === "object" && (value.type === "turn" || value.type === "feedback")) out.push(value)
149 } catch {
150 // A torn line from a crash mid-write; skip it.
151 }
152 }
153 return out
154}
155
156export type Stats = {
157 turns: number
158 bySource: Record<string, number>
159 byModel: Record<string, number>
160 byEffort: Record<string, number>
161 byLevel: Record<string, number>
162 clefCalls: number
163 latency: { mean: number; p50: number; p95: number } | undefined
164 meanConfidence: number | undefined
165 failures: Record<string, number>
166 overrides: number
167 /** Turns held on a warm model instead of the cheaper one Clef's level asked for. */
168 cacheHolds: number
169 /** Held stretches that ended in the downgrade being taken. */
170 downgradesTaken: number
171 adjusted: number
172 /** Turns where Clef's raw level differs from the final route's level. */
173 recommendationChanged: number
174 feedback: Record<string, number>
175 clefInputTokens: number
176}
177
178function quantile(sorted: number[], q: number): number {
179 if (sorted.length === 0) return 0
180 const i = Math.min(sorted.length - 1, Math.max(0, Math.ceil(q * sorted.length) - 1))
181 return sorted[i]!
182}
183
184function bump(map: Record<string, number>, key: string) {
185 map[key] = (map[key] ?? 0) + 1
186}
187
188export function aggregate(records: readonly LogRecord[]): Stats {
189 const stats: Stats = {
190 turns: 0,
191 bySource: {},
192 byModel: {},
193 byEffort: {},
194 byLevel: {},
195 clefCalls: 0,
196 latency: undefined,
197 meanConfidence: undefined,
198 failures: {},
199 overrides: 0,
200 cacheHolds: 0,
201 downgradesTaken: 0,
202 adjusted: 0,
203 recommendationChanged: 0,
204 feedback: {},
205 clefInputTokens: 0,
206 }
207 const latencies: number[] = []
208 let confidenceSum = 0
209 let confidenceN = 0
210 for (const r of records) {
211 if (r.type === "feedback") {
212 bump(stats.feedback, r.verdict)
213 continue
214 }
215 stats.turns++
216 bump(stats.bySource, r.source)
217 const model = r.answered?.model ?? r.final?.model ?? "session default"
218 bump(stats.byModel, model.replace(/-\d{8}$/, ""))
219 bump(stats.byEffort, r.final ? (r.final.effort ?? "default") : "untouched")
220 if (r.final?.level) bump(stats.byLevel, r.final.level)
221 if (r.recommendation || r.failure) {
222 if (r.failure?.kind !== "not-configured" && r.failure?.kind !== "budget" && r.failure?.kind !== "circuit-open") stats.clefCalls++
223 }
224 if (r.recommendation) {
225 if (r.latencyMs !== undefined) latencies.push(r.latencyMs)
226 confidenceSum += r.recommendation.confidence
227 confidenceN++
228 if (r.final?.level !== r.recommendation.level || r.final?.model !== r.proposed?.model) stats.recommendationChanged++
229 }
230 if (r.failure) bump(stats.failures, r.failure.kind)
231 if (r.source === "override" || r.source === "pin") stats.overrides++
232 if (r.adjustments.some((a) => a.rule === "cache-hold")) stats.cacheHolds++
233 if (r.deferral && !r.deferral.held) stats.downgradesTaken++
234 if (r.adjustments.length > 0) stats.adjusted++
235 stats.clefInputTokens += r.clefInputTokens ?? 0
236 }
237 if (latencies.length > 0) {
238 const sorted = [...latencies].sort((a, b) => a - b)
239 stats.latency = {
240 mean: Math.round(sorted.reduce((s, x) => s + x, 0) / sorted.length),
241 p50: quantile(sorted, 0.5),
242 p95: quantile(sorted, 0.95),
243 }
244 }
245 if (confidenceN > 0) stats.meanConfidence = confidenceSum / confidenceN
246 return stats
247}
248hooks/lib/models.ts 134 lines1// What this router knows about Claude models: how an alias resolves, which
2// effort levels a model takes, how big its window is, and how it ranks.
3//
4// Source: Claude Code model configuration docs (Claude Code 2.1.289,
5// October 2026). `turn.step` does not resolve aliases (a request for
6// "haiku" fails with unrecognized_model), so every route is resolved here to
7// a full ID before it is sent.
8
9import { EFFORTS, type Effort } from "./types.ts"
10
11export type Family = "haiku" | "sonnet" | "opus" | "fable"
12
13/** What each alias resolves to on the Anthropic API, as Claude Code does. */
14export const FIRST_PARTY_ALIASES: Record<Family, string> = {
15 haiku: "claude-haiku-4-5",
16 sonnet: "claude-sonnet-5-5",
17 opus: "claude-opus-5-5",
18 fable: "claude-fable-5-1",
19}
20
21/**
22 * The environment the resolution depends on: Claude Code's own
23 * ANTHROPIC_DEFAULT_*_MODEL pins, and whether a third-party provider is in use
24 * (where first-party IDs do not exist).
25 */
26export type ModelEnv = {
27 defaults: Partial<Record<Family, string>>
28 thirdParty: boolean
29 /** Prompt-cache beta features off (CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS). */
30 betasDisabled: boolean
31}
32
33export const FIRST_PARTY_ENV: ModelEnv = { defaults: {}, thirdParty: false, betasDisabled: false }
34
35export function isAlias(model: string): model is Family {
36 return model === "haiku" || model === "sonnet" || model === "opus" || model === "fable"
37}
38
39/** Resolves an alias or ID to a concrete model ID, or undefined if it cannot be. */
40export function resolveModel(model: string, env: ModelEnv): string | undefined {
41 const m = model.trim()
42 if (m === "") return undefined
43 if (!isAlias(m)) return m
44 const pinned = env.defaults[m]
45 if (pinned) return pinned
46 // On Bedrock, Vertex or Foundry, a first-party ID would fail.
47 return env.thirdParty ? undefined : FIRST_PARTY_ALIASES[m]
48}
49
50export function familyOf(modelId: string): Family | undefined {
51 const id = modelId.toLowerCase()
52 if (id.includes("haiku")) return "haiku"
53 if (id.includes("sonnet")) return "sonnet"
54 if (id.includes("fable")) return "fable"
55 if (id.includes("opus")) return "opus"
56 return undefined
57}
58
59const RANK: Record<Family, number> = { haiku: 0, sonnet: 1, opus: 2, fable: 3 }
60
61/** Capability/price rank for comparing two models; undefined if unknown. */
62export function rankOf(modelId: string): number | undefined {
63 const f = familyOf(modelId)
64 return f === undefined ? undefined : RANK[f]
65}
66
67/** Display name for the status line: "Sonnet", "Opus", or the raw ID. */
68export function displayName(modelId: string): string {
69 const f = familyOf(modelId)
70 return f === undefined ? modelId : f[0]!.toUpperCase() + f.slice(1)
71}
72
73/**
74 * The effort levels a model accepts; null when it takes no effort at all,
75 * undefined when the model is unknown (send what was asked; Claude Code
76 * clamps it).
77 */
78export function effortsFor(modelId: string): readonly Effort[] | null | undefined {
79 const id = modelId.toLowerCase()
80 if (id.includes("haiku")) return null
81 if (/(opus|sonnet)-4-6/.test(id)) return ["low", "medium", "high", "max"]
82 if (/fable|opus-5|sonnet-5|opus-4-[78]/.test(id)) return EFFORTS
83 return undefined
84}
85
86/**
87 * The highest supported level at or below the one asked, which is what
88 * Claude Code itself does; undefined when the model takes no effort.
89 */
90export function clampEffort(modelId: string, effort: Effort | undefined): Effort | undefined {
91 if (effort === undefined) return undefined
92 const supported = effortsFor(modelId)
93 if (supported === null) return undefined
94 if (supported === undefined) return effort
95 for (let i = EFFORTS.indexOf(effort); i >= 0; i--) {
96 const level = EFFORTS[i]!
97 if (supported.includes(level)) return level
98 }
99 return supported[0]
100}
101
102export function capEffort(effort: Effort | undefined, cap: Effort | undefined): Effort | undefined {
103 if (effort === undefined || cap === undefined) return effort
104 return EFFORTS.indexOf(effort) > EFFORTS.indexOf(cap) ? cap : effort
105}
106
107/** Context window in tokens where it is known to be smaller than 1M. */
108export function windowOf(modelId: string): number | undefined {
109 const id = modelId.toLowerCase()
110 if (id.includes("haiku")) return 200_000
111 return undefined
112}
113
114/**
115 * Whether changing effort between requests keeps the prompt cache. Per the
116 * Claude Code prompt-caching docs, it does on Opus 5.5, Sonnet 5.5 and
117 * Fable 5.1 with an API key or subscription, and not on Bedrock, Vertex, a
118 * gateway, or with experimental betas disabled. Elsewhere it is a full miss.
119 */
120export function effortChangeKeepsCache(modelId: string, env: ModelEnv): boolean {
121 if (env.thirdParty || env.betasDisabled) return false
122 return /claude-(opus|sonnet)-5-5|claude-fable-5-1/.test(modelId.toLowerCase())
123}
124
125/** Compares IDs ignoring a date suffix and a [1m] marker. */
126export function sameModel(a: string, b: string): boolean {
127 const norm = (id: string) => id.toLowerCase().replace(/\[1m\]$/, "").replace(/-\d{8}$/, "")
128 return norm(a) === norm(b)
129}
130
131export function isEffort(value: string): value is Effort {
132 return (EFFORTS as readonly string[]).includes(value)
133}
134hooks/lib/overrides.ts 183 lines1// Explicit user intent, parsed deterministically: the one-turn `+target`
2// prompt prefix, `/clef` command targets, and turns that only continue the
3// previous one. No model is involved in any of this.
4
5import { splitTarget } from "./config.ts"
6import { isAlias, isEffort } from "./models.ts"
7import { LEVELS, type Level, type Target, type TurnKind } from "./types.ts"
8
9/**
10 * Parses a target: a profile (`hard`), an alias or model ID with an optional
11 * effort (`opus`, `opus:max`, `claude-sonnet-5-5:low`), or an effort alone
12 * (`:high`). Undefined when the text is none of those.
13 */
14export function parseTarget(text: string): Target | undefined {
15 const t = text.trim()
16 if (t === "") return undefined
17 if (t.startsWith(":")) {
18 const effort = t.slice(1).toLowerCase()
19 return isEffort(effort) ? { effort } : undefined
20 }
21 const { model, effort, badEffort } = splitTarget(t)
22 if (badEffort || model === "") return undefined
23 const lower = model.toLowerCase()
24 const target: Target = {}
25 if ((LEVELS as readonly string[]).includes(lower)) target.level = lower as Level
26 else if (isAlias(lower)) target.model = lower
27 else if (/^claude-[a-z0-9.-]+(\[1m\])?$/i.test(model) || /anthropic\./i.test(model)) target.model = model
28 else return undefined
29 if (effort) target.effort = effort
30 return target
31}
32
33export type PrefixResult = { text: string; override?: Target | "off" }
34
35const PREFIX = /^\+(\S+)(?:\s+|$)/
36
37/**
38 * A prompt that starts with `+target ` routes that one turn to the target and
39 * reaches Claude without the prefix; `+off ` runs the turn as Claude Code
40 * would. Anything else (`+1`, `+x`) is left untouched.
41 */
42export function parsePrefix(text: string): PrefixResult {
43 const match = PREFIX.exec(text)
44 if (!match) return { text }
45 const token = match[1]!
46 const rest = text.slice(match[0].length)
47 if (token.toLowerCase() === "off" || token.toLowerCase() === "noroute") return { text: rest, override: "off" }
48 const target = parseTarget(token)
49 if (!target) return { text }
50 return { text: rest, override: target }
51}
52
53/** Words a go-ahead is made of ("yes, do it", "ok go ahead", "lgtm, ship it"). */
54const GO_AHEAD_WORDS = new Set([
55 "y", "ya", "yes", "yep", "yeah", "yup", "ok", "okay", "k", "sure", "alright", "fine", "cool", "great", "perfect",
56 "go", "ahead", "for", "it", "do", "that", "this", "continue", "proceed", "carry", "on", "keep", "going", "next",
57 "lgtm", "looks", "sounds", "good", "ship", "please", "approved", "approve", "confirm", "confirmed", "thanks",
58])
59/** A go-ahead says yes to something; "it", "on" or "good" alone do not. */
60const GO_AHEAD_ANCHORS = new Set([
61 "y", "ya", "yes", "yep", "yeah", "yup", "ok", "okay", "k", "sure", "alright", "go", "do", "continue", "proceed",
62 "carry", "keep", "next", "lgtm", "ship", "approved", "approve", "confirm", "confirmed", "sounds", "looks",
63])
64
65/** Classifies a turn's text before anything is asked of Clef. */
66export function turnKind(text: string): TurnKind {
67 const t = text.trim()
68 if (t === "") return "empty"
69 if (t.startsWith("<task-notification>")) return "notification"
70 if (t.length <= 40) {
71 const words = t.toLowerCase().replace(/[.,!;:'"]+/g, " ").split(/\s+/).filter(Boolean)
72 if (words.length > 0 && words.length <= 6 && words.every((w) => GO_AHEAD_WORDS.has(w)) && words.some((w) => GO_AHEAD_ANCHORS.has(w)))
73 return "go-ahead"
74 }
75 return "prompt"
76}
77
78export type ClefCommand =
79 | { kind: "status" }
80 | { kind: "history" }
81 | { kind: "stats"; days: number }
82 | { kind: "profiles" }
83 | { kind: "test"; prompt: string }
84 | { kind: "auto" }
85 | { kind: "on" }
86 | { kind: "off" }
87 | { kind: "pin"; target: Target }
88 | { kind: "feedback"; verdict: "under" | "ok" | "over"; note?: string }
89 | { kind: "help" }
90 | { kind: "error"; message: string }
91
92const FEEDBACK: Record<string, "under" | "ok" | "over"> = {
93 under: "under",
94 underpowered: "under",
95 weak: "under",
96 "too-weak": "under",
97 ok: "ok",
98 right: "ok",
99 good: "ok",
100 over: "over",
101 overpowered: "over",
102 strong: "over",
103 "too-strong": "over",
104}
105
106export function parseCommand(args: string): ClefCommand {
107 const trimmed = args.trim()
108 const space = trimmed.search(/\s/)
109 const head = (space === -1 ? trimmed : trimmed.slice(0, space)).toLowerCase()
110 const rest = space === -1 ? "" : trimmed.slice(space + 1).trim()
111 switch (head) {
112 case "":
113 case "status":
114 return { kind: "status" }
115 case "history":
116 case "log":
117 return { kind: "history" }
118 case "stats": {
119 const days = rest === "" ? 7 : Number(rest)
120 return Number.isInteger(days) && days > 0 && days <= 366
121 ? { kind: "stats", days }
122 : { kind: "error", message: "usage: /clef stats [days]" }
123 }
124 case "profiles":
125 return { kind: "profiles" }
126 case "test":
127 return rest === "" ? { kind: "error", message: "usage: /clef test <prompt>" } : { kind: "test", prompt: rest }
128 case "auto":
129 case "unpin":
130 return { kind: "auto" }
131 case "on":
132 return { kind: "on" }
133 case "off":
134 return { kind: "off" }
135 case "pin": {
136 const target = parseTarget(rest)
137 return target
138 ? { kind: "pin", target }
139 : { kind: "error", message: `usage: /clef pin <${LEVELS.join("|")}|haiku|sonnet|opus|fable|model-id>[:effort] or /clef pin :<effort>` }
140 }
141 case "feedback": {
142 const space2 = rest.search(/\s/)
143 const word = (space2 === -1 ? rest : rest.slice(0, space2)).toLowerCase()
144 const verdict = FEEDBACK[word]
145 if (!verdict) return { kind: "error", message: "usage: /clef feedback under|ok|over [note]" }
146 const note = space2 === -1 ? "" : rest.slice(space2 + 1).trim()
147 return note ? { kind: "feedback", verdict, note } : { kind: "feedback", verdict }
148 }
149 case "help":
150 return { kind: "help" }
151 default: {
152 // `/clef opus:high` as shorthand for `/clef pin opus:high`.
153 const target = parseTarget(trimmed)
154 return target ? { kind: "pin", target } : { kind: "error", message: `unknown subcommand "${head}"; try /clef help` }
155 }
156 }
157}
158
159export type NativeEffort = {
160 /** The effort Claude Code sent when the router started watching. */
161 baseline?: string | number
162 /** An effort the person set with /effort since, in force until /clef auto. */
163 native?: string | number
164}
165
166/**
167 * Tracks the effort Claude Code itself would send, seen at each turn's first
168 * request. A change from the baseline is the person's /effort (or a skill's
169 * `effort` for one turn): it is honoured, and dropped again when the engine's
170 * effort returns to the baseline, so a one-turn skill does not stick.
171 */
172export function trackNativeEffort(prev: NativeEffort, engine: string | number | undefined): NativeEffort & { change?: "set" | "cleared" } {
173 const keep = (): NativeEffort => ({
174 ...(prev.baseline !== undefined ? { baseline: prev.baseline } : {}),
175 ...(prev.native !== undefined ? { native: prev.native } : {}),
176 })
177 if (engine === undefined) return keep()
178 if (prev.baseline === undefined) return { baseline: engine }
179 if (engine === prev.native) return keep()
180 if (engine === prev.baseline) return prev.native === undefined ? keep() : { baseline: prev.baseline, change: "cleared" }
181 return { baseline: prev.baseline, native: engine, change: "set" }
182}
183hooks/lib/policy.ts 438 lines1// The routing policy: a pure function from what is known at the start of a
2// turn to the route its requests will use, with every change it made on the
3// way recorded. Clef makes the semantic judgment (how hard is this?); this
4// file makes the deterministic ones (what is allowed, what is safe, what is
5// worth the cache), in this order:
6//
7// 1. Routing off (config, +off; /clef off and a mid-session /model change, unless +model/+profile)
8// 2. An explicit choice: a +target prefix this turn, then a /clef pin
9// 3. A continuation (go-ahead, task notification, empty) reuses the last route
10// 4. Clef's recommendation, or the fallback when Clef did not answer
11// 5. Low confidence → the configured confidence policy
12// 6. A follow-up never routes below the route it follows
13// 7. A profile whose model is unavailable → the nearest available one, upward first
14// 8. A context too big for the model's window → the nearest profile that fits
15// 9. A downgrade off a warm prompt cache → held until staying has cost what
16// switching costs, in the unit the person's billing makes scarce
17// (docs/adr/0001-downgrade-timing-by-billing-mode.md)
18// 10. Effort is capped and clamped to what the model takes
19
20import type { Config } from "./config.ts"
21import {
22 capEffort,
23 clampEffort,
24 displayName,
25 effortChangeKeepsCache,
26 effortsFor,
27 rankOf,
28 resolveModel,
29 sameModel,
30 windowOf,
31 type ModelEnv,
32} from "./models.ts"
33import { dollars, isOneHour, priceOf, switchCost, type Billing } from "./pricing.ts"
34import {
35 EFFORTS,
36 LEVELS,
37 type Adjustment,
38 type Decision,
39 type Effort,
40 type Level,
41 type ProviderResult,
42 type Route,
43 type Source,
44 type Target,
45 type TurnKind,
46} from "./types.ts"
47
48export type RouterMode = "auto" | "off" | "paused-native"
49
50/** The last request the main loop sent, which is what the prompt cache holds. */
51export type CacheState = {
52 model: string
53 effort?: Effort
54 /** Epoch ms the response finished. */
55 at: number
56 /** Input + cache read + cache write tokens of that request. */
57 promptTokens: number
58 /**
59 * Whether the API reported any cache read or write for it. False where
60 * nothing is cached (a gateway that strips cache markers): there is then no
61 * cache to keep. Absent in state saved by earlier versions: assumed true.
62 */
63 caching?: boolean
64}
65
66/** A stretch of held downgrades: what staying on the warm model has cost so far. */
67export type HoldState = {
68 /** The model held, whose cache is warm. */
69 model: string
70 /** What Clef's level asked for on the last held turn; a continuation weighs it again. */
71 wanted: Route
72 /** List-price dollars staying has cost over the stretch's completed turns. */
73 spent: number
74 turns: number
75}
76
77export type PolicySession = {
78 mode: RouterMode
79 pin?: Target
80 last?: Route
81 cache?: CacheState
82 hold?: HoldState
83 /** Models that failed when routed to this session. */
84 unavailable: readonly string[]
85}
86
87export type PolicyInput = {
88 turnId: string
89 kind: TurnKind
90 config: Config
91 modelEnv: ModelEnv
92 session: PolicySession
93 /** A one-turn +target prefix; "off" leaves this turn alone. */
94 override?: Target | "off"
95 /** Clef's answer; absent when Clef was not asked. */
96 result?: ProviderResult
97 /** Tokens in the conversation now, when known. */
98 contextTokens?: number
99 now: number
100 cacheTtlMs: number
101 /** How the person pays: what a held downgrade costs them. */
102 billing: Billing
103}
104
105/** Headroom kept under a model's window for the reply and tool results. */
106export const WINDOW_HEADROOM = 20_000
107
108const idx = (level: Level) => LEVELS.indexOf(level)
109const atIdx = (i: number): Level => LEVELS[Math.max(0, Math.min(LEVELS.length - 1, i))]!
110const higher = (a: Level, b: Level): Level => (idx(a) >= idx(b) ? a : b)
111
112export function describe(route: Route | undefined): string {
113 if (!route) return "session default"
114 return route.effort ? `${displayName(route.model)} · ${route.effort}` : displayName(route.model)
115}
116
117/**
118 * Whether the policy should ask Clef at all for this turn. Explicit choices,
119 * continuations and a disabled router make no network call.
120 */
121export function needsClef(input: Omit<PolicyInput, "result" | "now" | "cacheTtlMs" | "contextTokens" | "billing">): boolean {
122 const { config, session, override, kind } = input
123 if (!config.enabled || override === "off") return false
124 if (override && (override.level || override.model)) return false
125 if (session.mode !== "auto") return false
126 if (session.pin && (session.pin.level || session.pin.model)) return false
127 if (kind !== "prompt" && session.last) return false
128 if ((kind === "notification" || kind === "empty") && !session.last) return false
129 return true
130}
131
132function isUnavailable(model: string, session: PolicySession): boolean {
133 return session.unavailable.some((u) => sameModel(u, model))
134}
135
136/** The route for a profile, or the nearest usable profile (upward first). */
137function routeForLevel(
138 level: Level,
139 input: PolicyInput,
140 adjustments: Adjustment[],
141): Route | undefined {
142 const order = [idx(level), ...LEVELS.map((_, i) => i).filter((i) => i > idx(level)), ...LEVELS.map((_, i) => i).filter((i) => i < idx(level)).reverse()]
143 for (const i of order) {
144 const lv = atIdx(i)
145 const spec = input.config.profiles[lv]
146 const model = resolveModel(spec.model, input.modelEnv)
147 if (model === undefined || isUnavailable(model, input.session)) continue
148 const route: Route = { level: lv, model }
149 if (spec.effort) route.effort = spec.effort
150 if (lv !== level) {
151 adjustments.push({
152 rule: "unavailable",
153 from: level,
154 to: lv,
155 reason: `profile ${level} (${input.config.profiles[level].model}) has no usable model here`,
156 })
157 }
158 return route
159 }
160 return undefined
161}
162
163function routeForTarget(target: Target, input: PolicyInput, adjustments: Adjustment[]): Route | undefined {
164 if (target.level) {
165 const route = routeForLevel(target.level, input, adjustments)
166 if (route && target.effort) route.effort = target.effort
167 return route
168 }
169 if (target.model) {
170 const model = resolveModel(target.model, input.modelEnv)
171 if (model === undefined) return undefined
172 return target.effort ? { model, effort: target.effort } : { model }
173 }
174 return undefined
175}
176
177/**
178 * The effort a turn held on its warm model runs at. Where an effort change
179 * keeps the cache, it still comes down to what Clef's level asked for, and a
180 * level whose model takes no effort (Haiku) gets the held model's lowest: the
181 * turn is held for the cache, not for more thinking. Elsewhere the cached
182 * effort stays, since changing it would cost the cache the hold is keeping.
183 */
184function heldEffort(route: Route, cache: CacheState, env: ModelEnv): Effort | undefined {
185 if (!effortChangeKeepsCache(cache.model, env)) return cache.effort
186 if (route.effort) return route.effort
187 if (effortsFor(route.model) === null) return effortsFor(cache.model)?.[0] ?? "low"
188 return undefined
189}
190
191/** The more capable of the two most probable levels. */
192function upperOfTopTwo(probabilities: Record<Level, number>, top: Level): Level {
193 let second: Level | undefined
194 for (const level of LEVELS) {
195 if (level === top) continue
196 if (second === undefined || probabilities[level] >= probabilities[second]) second = level
197 }
198 return second === undefined ? top : higher(top, second)
199}
200
201export function decide(input: PolicyInput): Decision {
202 const { config, session, turnId, kind } = input
203 const adjustments: Adjustment[] = []
204 const base = { turnId, kind, adjustments }
205 const disabled = (source: Source, note: string): Decision => ({ ...base, source, note })
206
207 // 1. Off. A one-turn +model or +profile still applies while the session is
208 // off or paused: it is the most explicit request there is.
209 if (!config.enabled) return disabled("disabled", "routing disabled in plugin config")
210 if (input.override === "off") return disabled("disabled", "+off: this turn runs as Claude Code would")
211 const explicitTurn = input.override !== undefined && (input.override.level !== undefined || input.override.model !== undefined)
212 if (session.mode === "off" && !explicitTurn) return disabled("disabled", "routing off for this session (/clef on)")
213 if (session.mode === "paused-native" && !explicitTurn)
214 return disabled("native", "paused: you changed /model (/clef auto to resume)")
215
216 let source: Source
217 let route: Route | undefined
218 let proposed: Route | undefined
219 let note: string | undefined
220 const decision: Decision = { ...base, source: "clef" }
221
222 // 2. Explicit choice.
223 const explicit = input.override ?? (session.mode === "auto" ? session.pin : undefined)
224 const explicitSource: Source = input.override ? "override" : "pin"
225 if (explicit && (explicit.level || explicit.model)) {
226 source = explicitSource
227 route = routeForTarget(explicit, input, adjustments)
228 if (!route) return disabled(explicitSource, `cannot resolve ${explicit.model ?? explicit.level} here`)
229 proposed = { ...route }
230 } else if (kind !== "prompt" && session.last) {
231 // 3. Continuation. One that continues a held turn weighs the deferred
232 // downgrade again (step 9), so a run of go-aheads cannot outlast it.
233 source = "continuation"
234 const hold = session.hold && sameModel(session.last.model, session.hold.model) ? session.hold : undefined
235 route = { ...(hold ? hold.wanted : session.last) }
236 proposed = { ...route }
237 note = kind === "go-ahead" ? "go-ahead continues the last route" : `${kind} continues the last route`
238 } else if ((kind === "notification" || kind === "empty") && !session.last) {
239 return disabled("continuation", "nothing to continue yet")
240 } else {
241 // 4. Clef, or the fallback.
242 const result = input.result
243 let level: Level
244 if (result?.ok) {
245 source = "clef"
246 const rec = result.recommendation
247 decision.recommendation = rec
248 level = rec.level
249 proposed = routeForLevel(level, input, [])
250
251 // 5. Low confidence.
252 if (rec.confidence < config.confidenceThreshold) {
253 let adjusted: Level = level
254 switch (config.lowConfidencePolicy) {
255 case "upper-of-top-two":
256 adjusted = upperOfTopTwo(rec.probabilities, level)
257 break
258 case "bump":
259 adjusted = atIdx(idx(level) + 1)
260 break
261 case "hold":
262 adjusted = session.last?.level ?? config.fallbackLevel
263 break
264 case "fallback":
265 adjusted = config.fallbackLevel
266 break
267 case "obey":
268 break
269 }
270 if (adjusted !== level) {
271 adjustments.push({
272 rule: "low-confidence",
273 from: level,
274 to: adjusted,
275 reason: `confidence ${pct(rec.confidence)} < ${pct(config.confidenceThreshold)} (${config.lowConfidencePolicy})`,
276 })
277 level = adjusted
278 }
279 }
280
281 // 6. Follow-up floor.
282 const lastLevel = session.last?.level
283 if (
284 rec.contextDependent !== undefined &&
285 rec.contextDependent >= config.followUpThreshold &&
286 lastLevel !== undefined &&
287 idx(lastLevel) > idx(level)
288 ) {
289 adjustments.push({
290 rule: "context-dependent",
291 from: level,
292 to: lastLevel,
293 reason: `follow-up (${pct(rec.contextDependent)}) to a ${lastLevel} turn`,
294 })
295 level = lastLevel
296 }
297 } else {
298 source = "fallback"
299 if (result && !result.ok) decision.failure = result.failure
300 const lastLevel = session.last?.level
301 level = lastLevel ? higher(lastLevel, config.fallbackLevel) : config.fallbackLevel
302 note = `${result && !result.ok ? result.failure.kind : "no answer"}: using ${level}`
303 }
304
305 // 7. Availability.
306 route = routeForLevel(level, input, adjustments)
307 if (!route) return { ...decision, ...disabled(source, "no configured profile resolves to a usable model here") }
308 if (source === "fallback") proposed = { ...route }
309 }
310
311 // 8. Context window.
312 const tokens = input.contextTokens ?? session.cache?.promptTokens
313 if (tokens !== undefined) {
314 const fits = (model: string) => {
315 const w = windowOf(model)
316 return w === undefined || tokens + WINDOW_HEADROOM <= w
317 }
318 if (!fits(route.model)) {
319 const start = route.level ? idx(route.level) : 0
320 let moved: Route | undefined
321 for (let i = start + 1; i < LEVELS.length && !moved; i++) {
322 const candidate = routeForLevel(atIdx(i), input, [])
323 if (candidate && fits(candidate.model)) moved = candidate
324 }
325 if (moved) {
326 adjustments.push({
327 rule: "context-window",
328 from: describe(route),
329 to: describe(moved),
330 reason: `${kTokens(tokens)} context does not fit ${displayName(route.model)}'s window`,
331 })
332 route = moved
333 }
334 }
335 }
336
337 // 9. A downgrade off a warm cache, for routes the router chose itself. Each
338 // model has its own prompt cache, so moving a conversation writes all of it
339 // again. That pays off over a stretch of cheaper turns, not over one: the
340 // downgrade is held while what staying has cost so far is less than what the
341 // switch costs now (times downgrade_patience), then taken. Short dips stay
342 // put; long stretches move. What "cost" means depends on the billing.
343 const cache = session.cache
344 const warm = cache !== undefined && input.now - cache.at < input.cacheTtlMs
345 if ((source === "clef" || source === "continuation") && cache && warm && cache.caching !== false && !isUnavailable(cache.model, session)) {
346 const fromRank = rankOf(cache.model)
347 const toRank = rankOf(route.model)
348 const modelDown = fromRank !== undefined && toRank !== undefined && toRank < fromRank
349 const effortDown =
350 !modelDown &&
351 sameModel(route.model, cache.model) &&
352 route.effort !== undefined &&
353 cache.effort !== undefined &&
354 EFFORTS.indexOf(route.effort) < EFFORTS.indexOf(cache.effort) &&
355 !effortChangeKeepsCache(cache.model, input.modelEnv)
356 // An effort change that breaks the cache rewrites it on the same model.
357 const toPrice = priceOf(modelDown ? route.model : cache.model)
358 if ((modelDown || effortDown) && toPrice) {
359 const tokens = input.contextTokens ?? cache.promptTokens
360 const cost = switchCost(toPrice, tokens, isOneHour(input.cacheTtlMs))
361 const prior = session.hold && sameModel(session.hold.model, cache.model) ? session.hold : undefined
362 const spent = prior?.spent ?? 0
363 const turns = prior?.turns ?? 0
364 const wanted: Route = { ...route }
365 const billing = input.billing
366 if (spent >= cost * config.downgradePatience) {
367 if (prior) decision.deferral = { wanted, from: cache.model, billing, spent, cost, turns, held: false }
368 } else {
369 const held: Route = { model: cache.model }
370 if (route.level) held.level = route.level
371 const effort = modelDown ? heldEffort(route, cache, input.modelEnv) : cache.effort
372 if (effort) held.effort = effort
373 const sofar = `staying has cost ${dollars(spent)} so far (${billing}, list prices)`
374 adjustments.push({
375 rule: "cache-hold",
376 from: describe(route),
377 to: describe(held),
378 reason: modelDown
379 ? `${kTokens(tokens)} context is cached on ${displayName(cache.model)}; moving to ${displayName(route.model)} writes it again (${dollars(cost)}); ${sofar}`
380 : `changing effort on ${displayName(cache.model)} here rewrites its ${kTokens(tokens)} cached context (${dollars(cost)}); ${sofar}`,
381 })
382 decision.deferral = { wanted, from: cache.model, billing, spent, cost, turns: turns + 1, held: true }
383 route = held
384 }
385 }
386 }
387
388 // An effort-only pin applies to whatever model was chosen.
389 if (session.pin?.effort && !session.pin.level && !session.pin.model && !(input.override && input.override.effort)) {
390 if (route.effort !== session.pin.effort) {
391 adjustments.push({ rule: "pinned-effort", from: route.effort ?? "default", to: session.pin.effort, reason: "/clef pin" })
392 route = { ...route, effort: session.pin.effort }
393 }
394 }
395 if (input.override && !input.override.level && !input.override.model && input.override.effort) {
396 adjustments.push({ rule: "pinned-effort", from: route.effort ?? "default", to: input.override.effort, reason: "+ prefix" })
397 route = { ...route, effort: input.override.effort }
398 }
399
400 // 10. Effort cap and clamp.
401 if (route.effort) {
402 const capped = capEffort(route.effort, config.maxEffort)
403 if (capped !== route.effort) {
404 adjustments.push({ rule: "effort-cap", from: route.effort, to: capped ?? "none", reason: `max_effort is ${config.maxEffort}` })
405 }
406 const clamped = clampEffort(route.model, capped)
407 if (clamped !== capped) {
408 adjustments.push({
409 rule: "effort-clamp",
410 from: capped ?? "none",
411 to: clamped ?? "none",
412 reason: `${displayName(route.model)} ${clamped ? `tops out at ${clamped}` : "takes no effort setting"}`,
413 })
414 }
415 route = { ...route }
416 if (clamped) route.effort = clamped
417 else delete route.effort
418 }
419
420 const out: Decision = { ...decision, source, final: route, billing: input.billing }
421 if (proposed) out.proposed = proposed
422 if (note) out.note = note
423 return out
424}
425
426export function pct(p: number): string {
427 return `${Math.round(p * 100)}%`
428}
429
430export function kTokens(n: number): string {
431 return n >= 1000 ? `${Math.round(n / 1000)}k` : String(n)
432}
433
434/** Whether policy changed what Clef (or the override) proposed. */
435export function wasAdjusted(decision: Decision): boolean {
436 return decision.adjustments.length > 0
437}
438hooks/lib/pricing.ts 115 lines1// What a turn costs, and what a model switch costs, in list-price dollars.
2//
3// These figures are a yardstick for comparing staying on a model with leaving
4// it, not a bill: cloud providers price differently (regional endpoints cost
5// 10% more), negotiated rates differ, and a subscription does not bill per
6// token at all. What matters to the policy is the ratio between the two sides,
7// which holds wherever the price list is proportional to Anthropic's. See
8// docs/COSTS.md and docs/adr/0001-downgrade-timing-by-billing-mode.md.
9//
10// Source: Anthropic API pricing (platform.claude.com/docs/en/about-claude/pricing),
11// October 2026. Prices are dollars per million tokens.
12
13import { familyOf, type Family } from "./models.ts"
14
15/** How the person pays for Claude, which decides what a held turn costs them. */
16export const BILLINGS = ["subscription", "api"] as const
17export type Billing = (typeof BILLINGS)[number]
18
19export type Price = {
20 input: number
21 write5m: number
22 write1h: number
23 /** Cache hits and refreshes. */
24 read: number
25 output: number
26}
27
28/** Cache writes are 1.25× input (5 minutes) or 2× (1 hour); the read multiplier varies by model. */
29const price = (input: number, read: number, output: number): Price => ({ input, write5m: input * 1.25, write1h: input * 2, read, output })
30
31// Most specific first: "opus-5-5" before "opus-5", "opus-4-5" before "opus-4".
32const TABLE: readonly [RegExp, Price][] = [
33 [/(fable|mythos)-5-1/, price(10, 0.25, 50)],
34 [/(fable|mythos)-5/, price(10, 1, 50)],
35 [/opus-5-5/, price(4, 0.2, 20)],
36 [/opus-(5|4-[5-8])/, price(5, 0.5, 25)],
37 [/opus-4/, price(15, 1.5, 75)],
38 [/sonnet-5/, price(2, 0.2, 10)],
39 [/sonnet-4/, price(3, 0.3, 15)],
40 [/haiku-4-5/, price(1, 0.1, 5)],
41 [/haiku-3-5/, price(0.8, 0.08, 4)],
42]
43
44/** An unrecognised ID of a known family is priced as that family's current model. */
45const BY_FAMILY: Record<Family, Price> = {
46 haiku: price(1, 0.1, 5),
47 sonnet: price(2, 0.2, 10),
48 opus: price(4, 0.2, 20),
49 fable: price(10, 0.25, 50),
50}
51
52export function priceOf(modelId: string): Price | undefined {
53 const id = modelId.toLowerCase()
54 for (const [pattern, p] of TABLE) if (pattern.test(id)) return p
55 const family = familyOf(id)
56 return family === undefined ? undefined : BY_FAMILY[family]
57}
58
59/** The token counts the API reports for a turn, summed over its requests. */
60export type TurnTokens = {
61 inputTokens: number
62 outputTokens: number
63 cacheReadTokens: number
64 cacheWriteTokens: number
65}
66
67/** A cache written with a TTL over five minutes is billed at the one-hour rate. */
68export function isOneHour(ttlMs: number): boolean {
69 return ttlMs > 5 * 60_000
70}
71
72/** List-price dollars for a turn's tokens on a model. */
73export function turnCost(p: Price, t: TurnTokens, oneHour: boolean): number {
74 const write = oneHour ? p.write1h : p.write5m
75 return (t.inputTokens * p.input + t.cacheWriteTokens * write + t.cacheReadTokens * p.read + t.outputTokens * p.output) / 1e6
76}
77
78/**
79 * What a switch costs: the whole context written to the new model's cache
80 * (each model has its own), at that model's write price. The way back is not
81 * counted: a later upgrade is a turn that needs the stronger model, and is
82 * never held for price.
83 */
84export function switchCost(target: Price, contextTokens: number, oneHour: boolean): number {
85 return (contextTokens * (oneHour ? target.write1h : target.write5m)) / 1e6
86}
87
88/**
89 * What one held turn cost the person, beyond what the cheaper route would
90 * have cost, in the unit their billing makes scarce:
91 *
92 * - `api`: dollars. The turn's tokens priced on the held model, minus the same
93 * tokens on the wanted model as if its cache were warm. On models whose
94 * cache reads cost the same (Opus 5.5 and Sonnet 5.5), only writes and
95 * output differ, so staying costs little.
96 * - `subscription`: plan usage. All of the held turn counts, since it is drawn
97 * from the stronger model's allowance, which the plan meters separately and
98 * which a turn on the cheaper model would not touch.
99 *
100 * On `api`, keeping a higher effort on the same model (where an effort change
101 * breaks the cache) costs nothing by this measure, so it stays held while the
102 * cache is warm: no price list says what a lower effort saves.
103 */
104export function stayCost(billing: Billing, held: Price, wanted: Price, t: TurnTokens, oneHour: boolean): number {
105 const onHeld = turnCost(held, t, oneHour)
106 if (billing === "subscription") return onHeld
107 return Math.max(0, onHeld - turnCost(wanted, t, oneHour))
108}
109
110export function dollars(n: number): string {
111 if (n === 0) return "$0"
112 if (n < 0.01) return "<$0.01"
113 return `$${n < 10 ? n.toFixed(2) : n.toFixed(0)}`
114}
115hooks/lib/rubric.ts 108 lines1// The questions Clef is asked. This is the calibration surface: change the
2// wording here (or point `rubric_file` at a JSON file with the same shape)
3// and nothing else needs to change. `npm run calibrate` shows the effect.
4//
5// Clef answers every question in one forward pass, so the second question
6// costs a few extra input tokens and no extra latency.
7
8import { LEVELS } from "./types.ts"
9
10export type Rubric = {
11 /** What Clef is asked to rate. The prompt itself is sent as the `state`. */
12 instructions: string
13 /** One description per level, lowest first; exactly five. */
14 levels: readonly string[]
15 /** The follow-up question; empty string to not ask it. */
16 followUp: string
17}
18
19export const DEFAULT_RUBRIC: Rubric = {
20 instructions:
21 "A developer sent this message to Claude Code, an AI coding agent working inside their software " +
22 "repository with tools to read, search, edit and run code. Rate how much model capability and " +
23 "reasoning effort the agent needs to do this well. Judge the work the message asks for, not the " +
24 "length of the message: a short request can be hard, and a long paste can still be a trivial task.",
25 levels: [
26 "Trivial: mechanical and obvious, no judgment. Fix a typo, rename one symbol, reformat, run a known " +
27 "command, answer a quick factual or yes/no question, find where something is defined.",
28 "Simple: a small, well-specified change or question in one place. A one-line fix with a clear cause, " +
29 "add a log line or one simple test, explain a short function, a small config edit.",
30 "Moderate: ordinary feature or bug work with a clear goal, a few files, following existing patterns. " +
31 "Add an endpoint or pagination, write tests for a module, fix a reproducible bug, a routine refactor.",
32 "Hard: tricky, ambiguous or multi-step work where a wrong answer is costly. Debug intermittent, " +
33 "concurrency or non-obvious failures, significant refactors, unfamiliar code, performance, " +
34 "security-sensitive changes, choosing between designs.",
35 "Very hard: open-ended investigation or design across a whole subsystem. Root-cause cascading or " +
36 "distributed failures, architecture or migration plans, weigh alternatives then implement and " +
37 "verify, long autonomous work.",
38 ],
39 followUp:
40 "Is this message a short follow-up whose actual task is defined by earlier conversation rather than " +
41 "by the message itself, such as 'yes do it', 'try again', 'go with option 2', 'that didn't work', " +
42 "or 'continue'?",
43}
44
45export type ClefQuestion =
46 | { type: "score"; instructions: string; criteria: string[] }
47 | { type: "choice"; instructions: string; criteria: Record<string, string> }
48 | { type: "noul"; instructions: string; criteria?: { true: string; false: string } }
49
50export type QuestionStyle = "score" | "choice"
51
52/** Question IDs, shared by the request builder and the response parser. */
53export const Q_DIFFICULTY = "difficulty"
54export const Q_FOLLOW_UP = "follow_up"
55
56/**
57 * Builds Clef's question map. `score` (the default) treats the levels as an
58 * ordered rubric, which is what they are; `choice` is kept so calibration can
59 * compare the two on the same corpus.
60 */
61export function buildQuestions(rubric: Rubric, style: QuestionStyle = "score"): Record<string, ClefQuestion> {
62 const questions: Record<string, ClefQuestion> = {}
63 questions[Q_DIFFICULTY] =
64 style === "score"
65 ? { type: "score", instructions: rubric.instructions, criteria: [...rubric.levels] }
66 : {
67 type: "choice",
68 instructions: rubric.instructions,
69 criteria: Object.fromEntries(LEVELS.map((level, i) => [level, rubric.levels[i]!])),
70 }
71 if (rubric.followUp.trim() !== "") {
72 questions[Q_FOLLOW_UP] = {
73 type: "noul",
74 instructions: rubric.followUp,
75 criteria: {
76 true: "The message only makes sense together with earlier conversation.",
77 false: "The message states a self-contained request.",
78 },
79 }
80 }
81 return questions
82}
83
84/** Validates a rubric loaded from a file; returns the problems found. */
85export function rubricProblems(value: unknown): string[] {
86 const problems: string[] = []
87 if (typeof value !== "object" || value === null) return ["rubric must be a JSON object"]
88 const r = value as Record<string, unknown>
89 if (typeof r.instructions !== "string" || r.instructions.trim() === "") problems.push("`instructions` must be a non-empty string")
90 if (!Array.isArray(r.levels) || r.levels.length !== LEVELS.length || !r.levels.every((l) => typeof l === "string" && l.trim() !== ""))
91 problems.push(`\`levels\` must be ${LEVELS.length} non-empty strings, lowest first`)
92 if (r.followUp !== undefined && typeof r.followUp !== "string") problems.push("`followUp` must be a string when present")
93 return problems
94}
95
96export function parseRubric(text: string): { rubric: Rubric } | { problems: string[] } {
97 let value: unknown
98 try {
99 value = JSON.parse(text)
100 } catch {
101 return { problems: ["rubric file is not valid JSON"] }
102 }
103 const problems = rubricProblems(value)
104 if (problems.length > 0) return { problems }
105 const r = value as { instructions: string; levels: string[]; followUp?: string }
106 return { rubric: { instructions: r.instructions, levels: r.levels, followUp: r.followUp ?? DEFAULT_RUBRIC.followUp } }
107}
108