Picks the Claude model and reasoning effort for each turn with Cloudflare Clef, a fast decision model, on Workers AI or fully local (llama-server, Ollama).

A Claude Code mod that picks the Claude model and reasoning effort for each turn, using Cloudflare Clef as the decision model.
You keep using Claude Code as usual. When you send a prompt, the mod asks Clef-flash how much capability the request needs. It then runs that turn on the cheapest configuration that should be enough:
Fix the spelling of "recieve" in README.md. Clef → Haiku · 94%
Add pagination to /orders following the other list endpoints. Clef → Sonnet · medium · 81%
Debug why these tests intermittently deadlock only in parallel. Clef → Opus · high · 72%
Study this subsystem, find why it cascades under partitions, ... Clef → Opus · xhigh · 88%
Routing is an optimization. If Clef is slow, down, unconfigured or out of free quota, the turn still runs on a deterministic fallback. Claude Code always keeps working.
Status: v0.1, in dogfooding. The mod, the policy and the failure paths are tested against Claude Code 2.1.289's own test host, in live sessions against a mock Workers AI endpoint, and with live Clef-flash calls on a handful of prompts. How well Clef routes real coding work is not measured yet; that is what the local log and
/clef feedbackare for. The examples above show the output format; see Calibration.
High-capability models and high effort are worth it for difficult debugging, architecture and unfamiliar code. They are wasted on a rename. Nobody switches /model and /effort before every prompt, so this mod does it for you, once per turn, using a decision model built for exactly this kind of typed judgment.
The goal is not "always the cheapest model". It is the least expensive configuration that is sufficiently capable, with every policy decision shown to you.
claude --version.claude plugin marketplace add dwain-barnes/clef-model-router
claude plugin install clef-model-router@clef-model-router
This is a fork of AbelNavarro/clef-claude-router that adds the local backend. Install from upstream instead if you only want the Cloudflare backend.
Then give it your Cloudflare credentials. In Claude Code:
/plugin configure clef-model-router@clef-model-router
settings.json.If the mod is already loaded, run /reload-plugins; otherwise start a new session. You should see ↳ Clef: awaiting prompt at the end of the hint line under the prompt.
To uninstall: claude plugin uninstall clef-model-router@clef-model-router. To stop it without uninstalling, use /clef off (this session) or set Routing enabled to off in /config.
git clone https://github.com/dwain-barnes/clef-model-router
CLOUDFLARE_ACCOUNT_ID=... CLOUDFLARE_API_TOKEN=... claude --plugin-dir ./clef-model-router
Credentials in environment variables are inherited by every command Claude runs, so prefer /plugin configure for regular use.
No Cloudflare account, no network. The mod talks to any server that answers POST /v1/systemone on localhost.
llama.cpp (reference): build b11379 or later (Clef support merged 2026-10-03). Download a release for your platform from <https://github.com/ggml-org/llama.cpp/releases>, then:
llama-server -hf ggml-org/Clef-Flash-GGUF:Q4_K_M --alias clef-flash --host 127.0.0.1 --port 8080 -ngl 99 -c 8192 -np 2 -b 8192 -ub 8192 --no-webui
The first run downloads the 6.5 GB Clef-Flash model. -b 8192 -ub 8192 matter: Clef scores every question in one sequence and needs a batch larger than the default.
Ollama 0.35.1 or later: ollama pull clef (the 27B model, 18 GB; Ollama does not yet load the Clef-Flash GGUF from Hugging Face), then use endpoint http://127.0.0.1:11434/v1/systemone and model clef.
Then in Claude Code:
/plugin configure clef-model-router@clef-model-router
Set Backend to local, Local endpoint to your server (the llama-server default is already filled in) and Local model name to the alias or tag. The hint line shows ↳ Clef: awaiting prompt; the first routed turn shows Clef → … as usual, and /clef reports decider local · clef-flash at http://127.0.0.1:8080/v1/systemone.
Speed: on a GPU a decision takes well under a second. On CPU alone (a 24-thread desktop measured 3.5 s for a 300-token prompt) set max_prompt_chars to about 2000 in the advanced file so routing stays under a few seconds, or accept the delay. If the server is down, the mod says so once and routes to the fallback profile until it is back.
Check your server from the command line with npm run smoke -- --local.
These steps follow Cloudflare's Workers AI REST API guide.
1. Create a Cloudflare account (skip if you have one)
No credit card is needed. The free Workers plan includes 10,000 Workers AI neurons per day, roughly 2,000 routed prompts.
2. Open the Workers AI page
3. Get the API token
4. Get the Account ID
On the same Use REST API panel, under Get Account ID, copy the Account ID: a 32-character hex string such as 0123456789abcdef0123456789abcdef. It is also on the account home page under Account details → Account ID, and in the dashboard URL right after dash.cloudflare.com/.
5. Give them to the mod
/plugin configure clef-model-router@clef-model-router
/reload-plugins. The line under the prompt should read Clef: awaiting prompt.6. Check that it works
/clef test fix the typo in README shows Clef's probabilities and latency.npm run smoke makes one real call. Export CLOUDFLARE_ACCOUNT_ID and CLOUDFLARE_API_TOKEN first.If you prefer a custom token
Good practice
↳ Clef → Sonnet · medium · 87% (terminal; on other surfaces set announce to answer). The percentage is the probability Clef gave the level it picked. When policy changed Clef's pick, the reason follows in brackets, for example (held for cache) or (unsure). When Clef could not answer: Clef ✕ timeout → Sonnet · medium./clef shows the full picture: mode, the last decision with Clef's full probability distribution, latency, every policy adjustment and why, today's Clef usage, and the cache state.Last turn
route Opus · high (claude-opus-5-5, profile hard)
source clef
clef clef-flash: 72% on hard · clef confidence 52% · score 2.88/4 · follow-up 4% · 410 ms · 580 tokens
trivial ···················· 1%
simple █··················· 3%
standard ███················· 16%
hard ██████████████······ 72% ← Clef
deep ██·················· 8%
| Command | |||
|---|---|---|---|
/clef | Status and the last decision | ||
/clef history | This session's turns: latency, route, confidence, source, and any policy change | ||
/clef stats [days] | Totals from the local log: by model, effort, profile, source; latency; fallbacks; overrides; cache holds; your feedback | ||
/clef test <prompt> | Ask Clef about a prompt without sending it to Claude | ||
/clef profiles | What each difficulty level runs on here | ||
/clef pin <target> | Use one target for the rest of the session (/clef pin opus:high, /clef pin hard, /clef pin :low) | ||
/clef auto | Unpin, resume after a /model change, and hand effort back to Clef after /effort | ||
/clef off · /clef on | Stop or resume routing for this session | ||
| `/clef feedback under\ | ok\ | over [note]` | Rate the last route, for later analysis of whether Clef was right |
One turn only: start a prompt with +target. The prefix is removed before Claude sees the prompt.
+opus:max why does this deadlock only on ARM?
+haiku list the files in src/
+off explain this stack trace (this turn runs exactly as Claude Code would)
prompt ──► turn.start ──► Clef-flash: difficulty 0-4 (+ "is this a follow-up?") one call, ~0.3–0.5 s end to end
│
▼
policy (deterministic): overrides → continuation → confidence → follow-up floor
→ availability → context window → cache hold → effort clamp
│
▼
turn.step ×N: every main-loop request of the turn sent with that model + effort
yes, do it, continue) and background-task notifications reuse the last route without calling Clef.| Level | Default | For |
|---|---|---|
| trivial | haiku | typo, rename, format, quick lookup |
| simple | sonnet:low | small, well-specified change in one place |
| standard | sonnet:medium | ordinary feature or bug work |
| hard | opus:high | tricky debugging, refactors, unfamiliar code |
| deep | opus:xhigh | open-ended investigation and design |
Haiku 4.5 takes no effort setting, so the trivial level sends none.
+target beats /clef pin, which beats Clef. A /model change mid-session pauses routing until /clef auto. An /effort change sets the effort while Clef keeps choosing the model. Subagents keep their own models.Routing itself runs on your Cloudflare account. Clef-flash costs $0.09 per million input tokens and has no charged output. One routing call used about 580 input tokens for typical prompts in live tests (the rubric plus your prompt), and up to about 2,000 for a long one, since prompts are cut to 6,000 characters. That works out to roughly 5 neurons per call, or about 2,000 routed prompts a day inside Workers AI's free allocation of 10,000 neurons per day (resets 00:00 UTC). These are estimates derived from Cloudflare's published prices; Clef is not yet in Cloudflare's per-model neuron table.
What happens at the limit depends on your Cloudflare plan (pricing):
/clef shows today's calls, tokens and estimated neurons.
The mod also changes what you spend on Claude itself. That is the point, and /clef stats shows where your turns went.
For each prompt you type, the mod sends that prompt's text to Cloudflare Workers AI (cut to its first 4,500 and last 1,500 characters if longer than 6,000), along with the fixed rubric questions. Nothing else is sent: no conversation history, file contents, tool output, repository name or metadata. Go-aheads, task notifications, +model or +profile prompts, model pins and /clef off send nothing.
Everything else stays on your machine. The local log keeps a hash and the length of each prompt, not its text, unless you turn on Log prompt text. Cloudflare states that Workers AI does not use your inputs or outputs to train models, and the Clef announcement says Cloudflare does not read, store or train on Clef requests. Details and sources are in docs/PRIVACY.md.
Everyday options are plugin options. Set them with /plugin configure or in /config: the Cloudflare account ID and token, the decision model (clef-flash or clef), the profiles, how to show the route, routing on or off, and whether to log prompt text. Tuning knobs go in an optional ~/.claude/clef-model-router.json. Thresholds, policies, timeout, budget, cache hold and the rubric are all set there. The full reference, including precedence rules, is in docs/CONFIGURATION.md.
npm run calibrate, judging whether Clef was rightThe earlier routers this project studied, and what it took from each, are listed in Architecture → Prior art: jev-model-router, jev-claude-router, pi-auto-router, clef-router, claude-code-model-router and Morph's router.
Apache-2.0. Not affiliated with Anthropic or Cloudflare.
hooks/register.ts 697 lines1// clef-model-router: picks the Claude model and effort for each turn with
2// Cloudflare Clef.
3//
4// prompt.submit a `+target ` prefix is taken off the prompt and kept for the turn
5// turn.start one Clef call (or none), then the policy decides the turn's route
6// turn.step every main-loop request of the turn is sent with that route
7// turn.complete the outcome goes to the local log
8// /clef status, history, stats, test, pin, auto, off, on, feedback
9//
10// The decision is made once per turn, at turn.start, where the person's text
11// is: a turn's later requests (after each tool result) reuse it, so a turn
12// pays Clef's latency once and never changes model halfway. Subagents keep
13// their own model. Every failure path leaves the request as Claude Code made
14// it, or on a deterministic fallback.
15
16import type { EngineInterface, PluginOptions, Register } from "claude-code"
17
18import { clefProvider } from "./lib/clef.ts"
19import { parseAdvancedFile, parseConfig, type Config } from "./lib/config.ts"
20import { localProvider } from "./lib/local.ts"
21import { cacheTtlMs, modelEnvFrom, type EnvValues } from "./lib/env.ts"
22import {
23 HELP,
24 answerLine,
25 explain,
26 historyReport,
27 profilesReport,
28 statsReport,
29 statusLine,
30 statusReport,
31 targetText,
32 type HistoryRow,
33} from "./lib/format.ts"
34import { blockedReason, estimatedNeurons, normaliseGuard, recordFailure, recordSuccess, type GuardState } from "./lib/guard.ts"
35import { aggregate, logFileName, parseLines, promptHash, turnRecord, type AnsweredUsage, type FeedbackRecord } from "./lib/log.ts"
36import { clampEffort, effortsFor, isEffort, sameModel, type ModelEnv } from "./lib/models.ts"
37import { parseCommand, parsePrefix, trackNativeEffort, turnKind, type ClefCommand } from "./lib/overrides.ts"
38import { decide, describe, needsClef, type CacheState, type RouterMode } from "./lib/policy.ts"
39import { DEFAULT_RUBRIC, parseRubric, type Rubric } from "./lib/rubric.ts"
40import type { Decision, ProviderResult, Route, Target } from "./lib/types.ts"
41
42const PLUGIN = "clef-model-router"
43const SESSION_REF = { plugin: "clef-model-router", key: "session" } as const
44const HISTORY_LIMIT = 50
45const RUN_LIMIT = 32
46/** A routed model that fails this many requests in a row is not used again this session. */
47const FAILURES_BEFORE_UNAVAILABLE = 2
48
49/** What survives a hot reload (in `$.state`) for the rest of the session. */
50type Persisted = {
51 mode: RouterMode
52 pin?: Target
53 pendingOverride?: Target | "off"
54 last?: Route
55 cache?: CacheState
56 unavailable: string[]
57 failures: Record<string, number>
58 /** The session model Claude Code reported at the last turn. */
59 baselineModel?: string
60 /** Effort as Claude Code itself would send it, and any /effort the person set. */
61 effortBaseline?: string | number
62 nativeEffort?: string | number
63 /** A model Claude Code fell back to during the last turn, so it is not mistaken for a /model change. */
64 engineFallback?: string
65 history: HistoryRow[]
66 warned: string[]
67}
68
69/** One turn in flight. Module memory only: a reload mid-turn leaves it unrouted. */
70type Run = {
71 decision: Decision
72 prompt: string
73 hash?: string
74 engineModel?: string
75 passthrough: boolean
76 failedRewrite: boolean
77 steps: number
78 answered?: AnsweredUsage
79}
80
81// Module state. `register` runs again on every reload, which resets these;
82// `load` then restores the session's part from `$.state`.
83let options: PluginOptions = {}
84let state: Persisted = freshState()
85let loaded = false
86let config: Config = parseConfig({}).config
87let problems: string[] = []
88let modelEnv: ModelEnv = modelEnvFrom({})
89let ttlMs = 5 * 60_000
90let rubric: Rubric = DEFAULT_RUBRIC
91let logDir: string | undefined
92let apiBase: string | undefined
93let advancedPath: string | undefined
94let logFile: string | undefined
95let logLines: string[] = []
96const runs = new Map<string, Run>()
97
98function freshState(): Persisted {
99 return { mode: "auto", unavailable: [], failures: {}, history: [], warned: [] }
100}
101
102async function save($: EngineInterface): Promise<void> {
103 try {
104 await $.state.set(SESSION_REF, JSON.stringify(state))
105 } catch {
106 // Losing the snapshot only matters on a hot reload; routing goes on.
107 }
108}
109
110async function readEnv($: EngineInterface): Promise<EnvValues> {
111 return {
112 ANTHROPIC_DEFAULT_HAIKU_MODEL: await $.env.get("ANTHROPIC_DEFAULT_HAIKU_MODEL"),
113 ANTHROPIC_DEFAULT_SONNET_MODEL: await $.env.get("ANTHROPIC_DEFAULT_SONNET_MODEL"),
114 ANTHROPIC_DEFAULT_OPUS_MODEL: await $.env.get("ANTHROPIC_DEFAULT_OPUS_MODEL"),
115 ANTHROPIC_DEFAULT_FABLE_MODEL: await $.env.get("ANTHROPIC_DEFAULT_FABLE_MODEL"),
116 CLAUDE_CODE_USE_BEDROCK: await $.env.get("CLAUDE_CODE_USE_BEDROCK"),
117 CLAUDE_CODE_USE_VERTEX: await $.env.get("CLAUDE_CODE_USE_VERTEX"),
118 CLAUDE_CODE_USE_FOUNDRY: await $.env.get("CLAUDE_CODE_USE_FOUNDRY"),
119 CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS: await $.env.get("CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS"),
120 // Only whether it is set: the key itself is never kept.
121 ANTHROPIC_API_KEY: (await $.env.get("ANTHROPIC_API_KEY")) ? "set" : undefined,
122 CLAUDE_CODE_PROMPT_CACHE_TTL: await $.env.get("CLAUDE_CODE_PROMPT_CACHE_TTL"),
123 FORCE_PROMPT_CACHING_5M: await $.env.get("FORCE_PROMPT_CACHING_5M"),
124 ENABLE_PROMPT_CACHING_1H: await $.env.get("ENABLE_PROMPT_CACHING_1H"),
125 }
126}
127
128/** Reads configuration and restores the session snapshot, once per load. */
129async function load($: EngineInterface): Promise<void> {
130 if (loaded) return
131 loaded = true
132 try {
133 const home = (await $.env.get("HOME")) ?? (await $.env.get("USERPROFILE")) ?? "."
134 const configDir = (await $.env.get("CLAUDE_CONFIG_DIR")) ?? `${home}/.claude`
135 advancedPath = (await $.env.get("CLEF_ROUTER_CONFIG")) ?? `${configDir}/${PLUGIN}.json`
136 const advancedText = await $.fs.read(advancedPath).catch(() => undefined)
137 const advanced = parseAdvancedFile(typeof advancedText === "string" ? advancedText : undefined)
138 const parsed = parseConfig(
139 { ...advanced.values, ...options },
140 { accountId: await $.env.get("CLOUDFLARE_ACCOUNT_ID"), apiToken: await $.env.get("CLOUDFLARE_API_TOKEN") },
141 )
142 config = parsed.config
143 problems = [...advanced.problems, ...parsed.problems]
144 const env = await readEnv($)
145 modelEnv = modelEnvFrom(env)
146 const settings = (await $.settings.read().catch(() => ({}))) as Record<string, unknown>
147 ttlMs = cacheTtlMs(config.cacheTtlMinutes, env, settings.promptCacheTtl)
148 if (config.rubricFile) {
149 const text = await $.fs.read(config.rubricFile).catch(() => undefined)
150 const result = typeof text === "string" ? parseRubric(text) : { problems: [`cannot read rubric_file ${config.rubricFile}`] }
151 if ("rubric" in result) rubric = result.rubric
152 else problems.push(...result.problems.map((p) => `${p}; using the built-in rubric`))
153 }
154 const base = await $.env.get("CLEF_ROUTER_API_BASE")
155 if (base && /^https?:\/\//.test(base)) apiBase = base
156 logDir = config.logDir ?? `${configDir}/plugins/data/${PLUGIN}`
157 const snapshot = await $.state.get(SESSION_REF)
158 if (typeof snapshot.value === "string") state = { ...freshState(), ...(JSON.parse(snapshot.value) as Partial<Persisted>) }
159 } catch (error) {
160 problems.push(`setup: ${error instanceof Error ? error.message : String(error)}`)
161 }
162}
163
164/** Counters per backend: the Cloudflare budget must not count local calls. */
165function guardKey(): string {
166 return config.backend === "local" ? "guard-local" : "guard"
167}
168
169/** A local server has no quota or bill; only the circuit breaker applies. */
170function neuronBudget(): number {
171 return config.backend === "local" ? 0 : config.dailyNeuronBudget
172}
173
174async function readGuard($: EngineInterface, now: number): Promise<GuardState> {
175 return normaliseGuard(await $.store.get(guardKey()).catch(() => undefined), now)
176}
177
178async function askClef($: EngineInterface, prompt: string): Promise<ProviderResult> {
179 const now = await $.clock.now()
180 const guard = await readGuard($, now)
181 const blocked = blockedReason(guard, { now, model: config.decisionModel, dailyNeuronBudget: neuronBudget() })
182 if (blocked) return { ok: false, failure: blocked }
183 const http = async (url: string, init: { method: string; headers: Record<string, string>; body: string }) => {
184 const r = await $.http.fetch(url, init)
185 return { status: r.status, ok: r.ok, text: r.text }
186 }
187 const sleep = (ms: number, signal: AbortSignal) => $.clock.sleep(ms, { signal })
188 const clock = () => $.clock.now()
189 const provider =
190 config.backend === "local"
191 ? localProvider({
192 endpoint: config.localEndpoint,
193 model: config.localModel,
194 rubric,
195 timeoutMs: config.localTimeoutMs,
196 maxPromptChars: config.maxPromptChars,
197 fetch: http,
198 sleep,
199 now: clock,
200 })
201 : clefProvider({
202 accountId: config.accountId,
203 apiToken: config.apiToken,
204 model: config.decisionModel,
205 rubric,
206 timeoutMs: config.timeoutMs,
207 maxPromptChars: config.maxPromptChars,
208 ...(apiBase ? { apiBase } : {}),
209 fetch: http,
210 sleep,
211 now: clock,
212 })
213 const result = await provider.decide(prompt)
214 const after = result.ok ? recordSuccess(guard, result.recommendation.inputTokens) : recordFailure(guard, result.failure, await $.clock.now())
215 if (after !== guard) await $.store.set(guardKey(), after).catch(() => {})
216 return result
217}
218
219/** One-time notices, so a misconfiguration is said once, not every turn. */
220function warnOnce($: EngineInterface, key: string, text: string): void {
221 if (state.warned.includes(key)) return
222 state.warned.push(key)
223 $.ui.toast(text, { timeoutMs: 8000 })
224}
225
226/** Whether the configured backend can be called at all. */
227function configured(): boolean {
228 return config.backend === "local" || (config.accountId !== undefined && config.apiToken !== undefined)
229}
230
231/** What the status line calls the decider. */
232function deciderName(): string {
233 return config.backend === "local" ? `local ${config.localModel}` : config.decisionModel
234}
235
236/** The route line, drawn dim at the end of the hint line under the prompt. */
237let hint: string | undefined
238
239/**
240 * Shows the route without Claude Code's status-line marker (a ⚠ that reads as
241 * a warning): the text joins the prompt's hint line as its tail, prefixed ↳.
242 * The terminal draws that tail; elsewhere `announce: "answer"` shows the route.
243 */
244function showStatus($: EngineInterface, text: string | undefined): void {
245 const next = text === undefined ? undefined : `↳ ${text}`
246 if (next === hint) return
247 hint = next
248 $.ui.invalidate("ui.render")
249}
250
251function announce($: EngineInterface, d: Decision): void {
252 if (config.announce === "status" || config.announce === "both") showStatus($, statusLine(d))
253}
254
255async function appendLog($: EngineInterface, line: string, ts: string): Promise<void> {
256 if (!config.logEnabled || !logDir) return
257 try {
258 const file = `${logDir}/${logFileName(ts, await $.session.id())}`
259 if (file !== logFile) {
260 logFile = file
261 const existing = await $.fs.read(file).catch(() => "")
262 logLines = typeof existing === "string" && existing !== "" ? existing.trimEnd().split("\n") : []
263 }
264 logLines.push(line)
265 await $.fs.write(file, logLines.join("\n") + "\n")
266 } catch {
267 // A log that cannot be written must not cost the turn anything.
268 }
269}
270
271/** Notices a /model change made between turns: the person taking over the model. */
272async function noticeNativeModel($: EngineInterface): Promise<void> {
273 const sessionModel = await $.session.model().catch(() => undefined)
274 if (sessionModel && state.baselineModel && !sameModel(sessionModel, state.baselineModel)) {
275 // A fallback Claude Code made itself (a safety classifier moving the
276 // session) is not one.
277 const engineMoved = state.engineFallback !== undefined && sameModel(sessionModel, state.engineFallback)
278 if (!engineMoved && state.mode === "auto" && config.pauseOnNativeChange) {
279 state.mode = "paused-native"
280 $.ui.toast(`Clef paused: you switched to ${sessionModel}. /clef auto resumes routing.`, { timeoutMs: 8000 })
281 }
282 }
283 delete state.engineFallback
284 if (sessionModel) state.baselineModel = sessionModel
285}
286
287async function routeTurn($: EngineInterface, turnId: string, text: string): Promise<void> {
288 const kind = turnKind(text)
289 const override = state.pendingOverride
290 delete state.pendingOverride
291 await noticeNativeModel($)
292
293 const base = { turnId, kind, config, modelEnv, session: state, ...(override ? { override } : {}) }
294 const result = needsClef(base) ? await askClef($, text) : undefined
295 const usage = await $.session.usage().catch(() => undefined)
296 const now = await $.clock.now()
297 const decision = decide({
298 ...base,
299 ...(result ? { result } : {}),
300 ...(usage?.context?.tokens ? { contextTokens: usage.context.tokens } : {}),
301 now,
302 cacheTtlMs: ttlMs,
303 })
304
305 const kindOfFailure = decision.failure?.kind
306 if (kindOfFailure === "not-configured")
307 warnOnce($, "not-configured", `Clef router: set your Cloudflare account ID and API token with /plugin configure ${PLUGIN}, or set backend to local. Using the ${config.fallbackLevel} profile meanwhile.`)
308 else if (kindOfFailure === "auth")
309 warnOnce($, "auth", `Clef router: Cloudflare rejected the API token (${decision.failure?.message}). Falling back until it is fixed.`)
310 else if (kindOfFailure === "quota")
311 warnOnce($, `quota-${new Date(now).toISOString().slice(0, 10)}`, "Clef router: Workers AI's free daily allocation is used up; falling back until 00:00 UTC.")
312 else if (config.backend === "local" && (kindOfFailure === "network" || kindOfFailure === "timeout" || kindOfFailure === "bad-request"))
313 warnOnce($, "local-unreachable", `Clef router: no local Clef at ${config.localEndpoint} (${decision.failure?.message}). Start llama-server or Ollama; using the ${config.fallbackLevel} profile meanwhile.`)
314
315 if (decision.final) state.last = decision.final
316 const run: Run = { decision, prompt: text, passthrough: !decision.final, failedRewrite: false, steps: 0 }
317 const hash = await promptHash(text).catch(() => undefined)
318 if (hash) run.hash = hash
319 runs.set(turnId, run)
320 while (runs.size > RUN_LIMIT) runs.delete(runs.keys().next().value!)
321 state.history.push({ decision, prompt: text.slice(0, 200) })
322 while (state.history.length > HISTORY_LIMIT) state.history.shift()
323 announce($, decision)
324 await save($)
325}
326
327/** Watches Claude Code's own model and effort at a main-loop step, before any rewrite. */
328function noticeStep($: EngineInterface, run: Run, model: string, effort: string | number | undefined, index: number): void {
329 if (index === 0) {
330 run.engineModel = model
331 const tracked = trackNativeEffort({ baseline: state.effortBaseline, native: state.nativeEffort }, effort)
332 state.effortBaseline = tracked.baseline
333 if (tracked.native === undefined) delete state.nativeEffort
334 else state.nativeEffort = tracked.native
335 if (!config.pauseOnNativeChange) return
336 if (tracked.change === "set" && state.mode === "auto")
337 $.ui.toast(`Clef: using your effort ${String(tracked.native)}; Clef still picks the model. /clef auto hands effort back.`, { timeoutMs: 8000 })
338 // The person's /effort beats Clef's, not an explicit +target or pin.
339 const d = run.decision
340 if (state.nativeEffort !== undefined && d.final && (d.source === "clef" || d.source === "continuation" || d.source === "fallback")) {
341 const wanted = typeof state.nativeEffort === "string" && isEffort(state.nativeEffort) ? state.nativeEffort : undefined
342 const effortNow = wanted ? clampEffort(d.final.model, wanted) : undefined
343 if (effortNow !== d.final.effort) {
344 const final = { ...d.final }
345 if (effortNow) final.effort = effortNow
346 else delete final.effort
347 run.decision = {
348 ...d,
349 final,
350 adjustments: [...d.adjustments, { rule: "pinned-effort", from: d.final.effort ?? "default", to: effortNow ?? "default", reason: "your /effort" }],
351 }
352 state.last = final
353 announce($, run.decision)
354 }
355 }
356 } else if (run.engineModel && !sameModel(model, run.engineModel)) {
357 // Claude Code moved the turn to a fallback model (an error, or a safety
358 // classifier). That is never overridden.
359 run.passthrough = true
360 state.engineFallback = model
361 }
362}
363
364/** Records what a step's response says: failures of a routed model, usage, cache. */
365async function afterStep(
366 $: EngineInterface,
367 run: Run,
368 sent: { model: string; effort?: unknown },
369 rewritten: boolean,
370 result: { stopReason: string | null; usage: { model: string; input_tokens: number; output_tokens: number; cache_read_input_tokens: number; cache_creation_input_tokens: number } | null },
371 aborted: boolean,
372): Promise<void> {
373 const now = await $.clock.now()
374 if (rewritten) {
375 if (result.stopReason === null && !aborted) {
376 // The routed request got no response: stop routing this turn, and stop
377 // using the model after repeated failures.
378 run.failedRewrite = true
379 const n = (state.failures[sent.model] ?? 0) + 1
380 state.failures[sent.model] = n
381 if (n >= FAILURES_BEFORE_UNAVAILABLE && !state.unavailable.includes(sent.model)) {
382 state.unavailable.push(sent.model)
383 $.ui.toast(`Clef router: ${sent.model} failed ${n} times; not routing to it again this session.`, { timeoutMs: 8000 })
384 }
385 } else if (result.stopReason !== null) {
386 state.failures[sent.model] = 0
387 }
388 }
389 const u = result.usage
390 if (u) {
391 const a = run.answered ?? { model: u.model, inputTokens: 0, outputTokens: 0, cacheReadTokens: 0, cacheWriteTokens: 0 }
392 run.answered = {
393 model: u.model,
394 inputTokens: a.inputTokens + u.input_tokens,
395 outputTokens: a.outputTokens + u.output_tokens,
396 cacheReadTokens: a.cacheReadTokens + u.cache_read_input_tokens,
397 cacheWriteTokens: a.cacheWriteTokens + u.cache_creation_input_tokens,
398 }
399 const cache: CacheState = {
400 model: sent.model,
401 at: now,
402 promptTokens: u.input_tokens + u.cache_read_input_tokens + u.cache_creation_input_tokens,
403 }
404 if (typeof sent.effort === "string" && isEffort(sent.effort)) cache.effort = sent.effort
405 state.cache = cache
406 }
407 await save($)
408}
409
410async function completeTurn($: EngineInterface, run: Run, durationMs: number, reason: string): Promise<void> {
411 const ts = new Date(await $.clock.now()).toISOString()
412 const record = turnRecord({
413 decision: run.decision,
414 session: await $.session.id(),
415 ts,
416 promptText: run.prompt,
417 ...(run.hash ? { hash: run.hash } : {}),
418 logPrompts: config.logPrompts,
419 ...(run.answered ? { answered: run.answered } : {}),
420 steps: run.steps,
421 durationMs,
422 endReason: reason,
423 })
424 await appendLog($, JSON.stringify(record), ts)
425 const row = state.history.find((h) => h.decision.turnId === run.decision.turnId)
426 if (row) {
427 row.decision = run.decision
428 if (run.answered) row.answeredModel = run.answered.model
429 await save($)
430 }
431}
432
433async function registerCommand($: EngineInterface): Promise<void> {
434 try {
435 await $.command.register({
436 name: "clef",
437 description: "Clef router: status, history, stats, test, pin, auto, off, on, feedback",
438 argumentHint: "[status|history|stats|profiles|test|pin|auto|off|on|feedback|help]",
439 immediate: true,
440 })
441 } catch {
442 // Without the command the router still routes.
443 }
444 if (!config.enabled) showStatus($, undefined)
445 else if (!configured()) showStatus($, "Clef: not configured")
446 else {
447 // Always replace what an earlier load showed (a stale "not configured"
448 // survives a reload otherwise): the last route if there is one.
449 const last = state.history.at(-1)?.decision
450 showStatus($, (last && statusLine(last)) ?? "Clef: awaiting prompt")
451 }
452}
453
454async function onClear($: EngineInterface): Promise<void> {
455 // /clear starts a new conversation: nothing is cached and nothing continues.
456 delete state.last
457 delete state.cache
458 delete state.pendingOverride
459 state.history = []
460 runs.clear()
461 logFile = undefined
462 if (config.enabled && configured()) showStatus($, "Clef: awaiting prompt")
463 await save($)
464}
465
466async function setPendingOverride($: EngineInterface, override: Target | "off"): Promise<void> {
467 await load($)
468 state.pendingOverride = override
469 await save($)
470}
471
472async function statusText($: EngineInterface, now: number): Promise<string> {
473 const guard = await readGuard($, now)
474 const blocked = blockedReason(guard, { now, model: config.decisionModel, dailyNeuronBudget: neuronBudget() })
475 const last = state.history.at(-1)?.decision
476 return statusReport({
477 config,
478 configProblems: problems,
479 modelEnv,
480 mode: state.mode,
481 ...(state.pin ? { pin: state.pin } : {}),
482 ...(last ? { last } : {}),
483 guard: {
484 calls: guard.calls,
485 inputTokens: guard.inputTokens,
486 neurons: estimatedNeurons(guard, config.decisionModel),
487 ...(blocked ? { blocked: blocked.message } : {}),
488 },
489 ...(state.cache
490 ? {
491 cache: {
492 model: state.cache.model,
493 promptTokens: state.cache.promptTokens,
494 ageSeconds: Math.round((now - state.cache.at) / 1000),
495 ttlSeconds: Math.round(ttlMs / 1000),
496 },
497 }
498 : {}),
499 unavailable: state.unavailable,
500 logDir,
501 ...(advancedPath ? { advancedPath } : {}),
502 })
503}
504
505async function statsText($: EngineInterface, now: number, days: number): Promise<string> {
506 if (!logDir) return "No log directory."
507 const since = new Date(now - (days - 1) * 86_400_000).toISOString().slice(0, 10)
508 const entries = await $.fs.list(logDir).catch(() => [])
509 const files = entries.filter((f) => /^routing-\d{4}-\d{2}-\d{2}-/.test(f.name) && f.name.slice(8, 18) >= since)
510 const records = []
511 for (const f of files) {
512 const text = await $.fs.read(`${logDir}/${f.name}`).catch(() => "")
513 if (typeof text === "string") records.push(...parseLines(text))
514 }
515 return statsReport(aggregate(records), days, files.length)
516}
517
518async function testText($: EngineInterface, now: number, prompt: string): Promise<string> {
519 const result = await askClef($, prompt)
520 const d = decide({
521 turnId: "test",
522 kind: turnKind(prompt),
523 config,
524 modelEnv,
525 session: { ...state, mode: "auto" },
526 result,
527 now,
528 cacheTtlMs: ttlMs,
529 })
530 return [`Clef on: ${prompt.slice(0, 80)}`, ...explain(d), "", "(Not sent to Claude. Counts toward today's Clef usage.)"].join("\n")
531}
532
533async function feedbackText($: EngineInterface, now: number, cmd: Extract<ClefCommand, { kind: "feedback" }>): Promise<string> {
534 const last = state.history.at(-1)
535 const record: FeedbackRecord = {
536 v: 1,
537 type: "feedback",
538 ts: new Date(now).toISOString(),
539 session: await $.session.id(),
540 verdict: cmd.verdict,
541 ...(last ? { turn: last.decision.turnId } : {}),
542 ...(cmd.note ? { note: cmd.note } : {}),
543 }
544 await appendLog($, JSON.stringify(record), record.ts)
545 const route = last?.decision.final ? describe(last.decision.final) : "the last turn"
546 const verdict = cmd.verdict === "ok" ? "about right" : cmd.verdict === "under" ? "not capable enough" : "more than needed"
547 return `Noted: ${route} was ${verdict}.`
548}
549
550async function runCommand($: EngineInterface, args: string): Promise<string> {
551 await load($)
552 const cmd = parseCommand(args)
553 const now = await $.clock.now()
554 switch (cmd.kind) {
555 case "help":
556 return HELP
557 case "error":
558 return cmd.message
559 case "status":
560 return statusText($, now)
561 case "history":
562 return historyReport(state.history)
563 case "profiles":
564 return profilesReport(config, modelEnv, state.unavailable)
565 case "stats":
566 return statsText($, now, cmd.days)
567 case "test":
568 return testText($, now, cmd.prompt)
569 case "feedback":
570 return feedbackText($, now, cmd)
571 case "auto":
572 state.mode = "auto"
573 delete state.pin
574 delete state.nativeEffort
575 delete state.effortBaseline
576 await save($)
577 showStatus($, config.enabled ? `Clef auto · ${deciderName()}` : undefined)
578 return config.enabled ? "Routing is automatic again." : "Routing is disabled in the plugin config (enabled = false)."
579 case "on":
580 state.mode = "auto"
581 await save($)
582 showStatus($, `Clef on · ${deciderName()}`)
583 return state.pin ? `Routing on, still pinned to ${targetText(state.pin)} (/clef auto to unpin).` : "Routing on."
584 case "off":
585 state.mode = "off"
586 await save($)
587 showStatus($, "Clef off")
588 return "Routing off for this session: Claude Code's own model and effort apply. /clef on resumes."
589 case "pin":
590 state.pin = cmd.target
591 state.mode = "auto"
592 await save($)
593 showStatus($, `Pinned → ${targetText(cmd.target)}`)
594 return `Pinned to ${targetText(cmd.target)} for this session. /clef auto unpins.`
595 }
596}
597
598/** The request a step is sent with: the turn's route, effort fitted to the model. */
599function routed<E extends { model: string; effort?: unknown }>(e: E, route: Route): E {
600 const request = { ...e, model: route.model } as E & { effort?: unknown }
601 if (effortsFor(route.model) === null) delete request.effort
602 else if (route.effort) request.effort = route.effort
603 else if (typeof e.effort === "string" && isEffort(e.effort)) {
604 const clamped = clampEffort(route.model, e.effort)
605 if (clamped) request.effort = clamped
606 }
607 return request
608}
609
610export const register: Register = (on, pluginOptions) => {
611 options = pluginOptions
612 state = freshState()
613 hint = undefined
614 apiBase = undefined
615 loaded = false
616 logFile = undefined
617 logLines = []
618 runs.clear()
619
620 on("ui.render", { component: "PromptHint" }, async ($, e, next) => {
621 if (hint === undefined) return next(e)
622 const tail = e.props.tail ? `${e.props.tail} · ${hint}` : ` ${hint}`
623 return next({ ...e, props: { ...e.props, tail } })
624 })
625
626 on("session.start", async ($, e, next) => {
627 await load($)
628 await registerCommand($)
629 return next(e)
630 })
631
632 on("session.end", async ($, e, next) => {
633 if (e.reason === "clear") await onClear($)
634 return next(e)
635 })
636
637 on("prompt.submit", async ($, e, next) => {
638 // Only a prompt that starts a turn of its own; one typed into a running
639 // turn joins that turn, which keeps its route.
640 if (e.turnId !== undefined) return next(e)
641 const parsed = parsePrefix(e.text)
642 if (parsed.override === undefined) return next(e)
643 await setPendingOverride($, parsed.override)
644 return next({ ...e, text: parsed.text })
645 })
646
647 on("turn.start", async ($, e, next) => {
648 try {
649 await load($)
650 await routeTurn($, e.turnId, e.text)
651 } catch {
652 // Unrouted: Claude Code's own model and effort apply.
653 }
654 return next(e)
655 })
656
657 on("turn.step", async function* ($, e, next) {
658 // Subagents run on the model their definition or Claude Code gives them.
659 const run = e.agentId === undefined ? runs.get(e.turnId) : undefined
660 if (!run) return yield* next(e)
661 run.steps++
662 try {
663 noticeStep($, run, e.model, e.effort, e.index)
664 } catch {
665 run.passthrough = true
666 }
667 const route = run.decision.final
668 const rewrite = route !== undefined && !run.passthrough && !run.failedRewrite
669 const request = rewrite ? routed(e, route) : e
670 const result = yield* next(request)
671 try {
672 await afterStep($, run, request, rewrite, result, next.signal?.aborted === true)
673 } catch {
674 // Bookkeeping only.
675 }
676 return result
677 })
678
679 on("turn.complete", async ($, e, next) => {
680 const result = await next(e)
681 const run = e.agentId === undefined ? runs.get(e.turnId) : undefined
682 if (!run) return result
683 try {
684 await completeTurn($, run, e.durationMs, e.reason)
685 } catch {
686 // Logging only.
687 }
688 if (config.announce === "answer" || config.announce === "both") {
689 const line = answerLine(run.decision, run.decision.recommendation?.latencyMs)
690 if (line) return { ...result, text: line }
691 }
692 return result
693 })
694
695 on("command.run", { command: "clef" }, async ($, e) => ({ text: await runCommand($, e.args) }))
696}
697hooks/lib/clef.ts 267 lines1// Cloudflare Clef on Workers AI: the request, the response, and every way the
2// exchange can fail, turned into a normalised Recommendation or a classified
3// ProviderFailure. Nothing here throws.
4//
5// API (developers.cloudflare.com/workers-ai/models/clef-flash, Oct 2026):
6// POST https://api.cloudflare.com/client/v4/accounts/{account}/ai/run/@cf/cloudflare/{model}
7// Authorization: Bearer {token}
8// { "model": "clef-flash", "state": ..., "questions": { id: {type, instructions, criteria} } }
9// → { "result": { "model", "answers": { id: answer }, "usage": { input_tokens, output_tokens } },
10// "success": true, "errors": [], "messages": [] }
11
12import { buildQuestions, Q_DIFFICULTY, Q_FOLLOW_UP, type QuestionStyle, type Rubric } from "./rubric.ts"
13import { redact } from "./redact.ts"
14import { LEVELS, type DecisionProvider, type Level, type ProviderFailure, type ProviderResult, type Recommendation } from "./types.ts"
15
16export const CLEF_MODELS = ["clef-flash", "clef"] as const
17export type ClefModel = (typeof CLEF_MODELS)[number]
18
19export const DEFAULT_API_BASE = "https://api.cloudflare.com/client/v4"
20
21/** The slice of `$.http.fetch` (or the global fetch, in scripts) this needs. */
22export type HttpLike = (
23 url: string,
24 init: { method: string; headers: Record<string, string>; body: string },
25) => Promise<{ status: number; ok: boolean; text: string }>
26
27export type ClefOptions = {
28 accountId: string | undefined
29 apiToken: string | undefined
30 model: ClefModel
31 rubric: Rubric
32 style?: QuestionStyle
33 timeoutMs: number
34 maxPromptChars: number
35 apiBase?: string
36 fetch: HttpLike
37 /** Resolves after `ms` (rejects if `signal` aborts); the timeout races the request against it. */
38 sleep: (ms: number, signal: AbortSignal) => Promise<void>
39 now: () => number | Promise<number>
40}
41
42export function endpoint(accountId: string, model: ClefModel, apiBase = DEFAULT_API_BASE): string {
43 return `${apiBase.replace(/\/$/, "")}/accounts/${encodeURIComponent(accountId)}/ai/run/@cf/cloudflare/${model}`
44}
45
46/**
47 * Keeps the head and the tail of a long prompt. The request's intent is
48 * usually stated at one end; the middle of a long paste rarely changes how
49 * hard the task is, and every character sent is billed and leaves the machine.
50 */
51export function truncatePrompt(text: string, maxChars: number): string {
52 if (text.length <= maxChars) return text
53 const head = Math.floor(maxChars * 0.75)
54 const tail = maxChars - head
55 const omitted = text.length - head - tail
56 return `${text.slice(0, head)}\n[... ${omitted} characters omitted ...]\n${text.slice(text.length - tail)}`
57}
58
59export function buildRequestBody(
60 prompt: string,
61 opts: { model: string; rubric: Rubric; style?: QuestionStyle; maxPromptChars: number },
62): string {
63 return JSON.stringify({
64 model: opts.model,
65 state: truncatePrompt(prompt, opts.maxPromptChars),
66 questions: buildQuestions(opts.rubric, opts.style ?? "score"),
67 })
68}
69
70type Envelope = {
71 success?: unknown
72 result?: unknown
73 errors?: unknown
74}
75
76function errorsOf(envelope: Envelope): { code?: number; message: string }[] {
77 if (!Array.isArray(envelope.errors)) return []
78 return envelope.errors
79 .filter((e): e is Record<string, unknown> => typeof e === "object" && e !== null)
80 .map((e) => ({
81 code: typeof e.code === "number" ? e.code : undefined,
82 message: typeof e.message === "string" ? e.message : "",
83 }))
84}
85
86/** Maps an HTTP status and Cloudflare error envelope to a failure kind. */
87export function classifyHttpFailure(
88 status: number,
89 bodyText: string,
90 latencyMs: number,
91 secrets: readonly (string | undefined)[] = [],
92): ProviderFailure {
93 let envelope: Envelope = {}
94 try {
95 envelope = JSON.parse(bodyText) as Envelope
96 } catch {
97 // Not JSON (a proxy's HTML page, say); the status alone decides.
98 }
99 const errors = errorsOf(envelope)
100 const codes = errors.map((e) => e.code)
101 const text = errors.map((e) => (e.code === undefined ? e.message : `${e.code}: ${e.message}`)).join("; ")
102 const message = redact(text || `HTTP ${status}`, secrets).slice(0, 200)
103 const quotaText = /daily free allocation|neurons/i.test(text)
104
105 if (codes.includes(3036) || (status === 429 && quotaText)) return { kind: "quota", message, status, latencyMs }
106 if (status === 429) return { kind: "rate-limited", message, status, latencyMs }
107 if (status === 401 || status === 403 || codes.includes(10000)) return { kind: "auth", message, status, latencyMs }
108 if (status === 408 || codes.includes(3007)) return { kind: "timeout", message, status, latencyMs }
109 if (status >= 500) return { kind: "server", message, status, latencyMs }
110 if (status >= 400) return { kind: "bad-request", message, status, latencyMs }
111 return { kind: "malformed", message, status, latencyMs }
112}
113
114function num(value: unknown): number | undefined {
115 return typeof value === "number" && Number.isFinite(value) ? value : undefined
116}
117
118/**
119 * Reads a difficulty answer's per-level probabilities. A score answer keys
120 * them by level index ("0".."4", per the schema), a choice answer by option
121 * id (our level names). A 1-based index set is accepted too, defensively.
122 */
123export function levelProbabilities(answer: Record<string, unknown>, rubric: Rubric): Record<Level, number> | undefined {
124 const raw = answer.probabilities
125 if (typeof raw !== "object" || raw === null) return undefined
126 const entries = Object.entries(raw as Record<string, unknown>)
127 const out = Object.fromEntries(LEVELS.map((l) => [l, 0])) as Record<Level, number>
128 const numericKeys = entries.every(([k]) => /^\d+$/.test(k))
129 const base = numericKeys ? Math.min(...entries.map(([k]) => Number(k))) : 0
130 let matched = 0
131 for (const [key, value] of entries) {
132 const p = num(value)
133 if (p === undefined || p < 0 || p > 1.0001) return undefined
134 let index: number
135 if (numericKeys) index = Number(key) - (base === 1 && entries.length === LEVELS.length ? 1 : 0)
136 else if ((LEVELS as readonly string[]).includes(key)) index = LEVELS.indexOf(key as Level)
137 else index = rubric.levels.indexOf(key)
138 const level = LEVELS[index]
139 if (level === undefined) return undefined
140 out[level] += p
141 matched++
142 }
143 if (matched === 0) return undefined
144 const total = LEVELS.reduce((s, l) => s + out[l], 0)
145 if (total < 0.98 || total > 1.02) return undefined
146 return out
147}
148
149/** The most probable level; a tie goes to the more capable one. */
150export function topLevel(probabilities: Record<Level, number>): Level {
151 let best: Level = LEVELS[0]
152 for (const level of LEVELS) if (probabilities[level] >= probabilities[best]) best = level
153 return best
154}
155
156/**
157 * Turns an unwrapped SystemOne result (`{ model, answers, usage }`, as
158 * llama-server and Ollama return it and as Workers AI wraps it under
159 * `result`) into a Recommendation, or says why it cannot.
160 */
161export function parseResult(
162 result: unknown,
163 opts: { provider: string; rubric: Rubric; latencyMs: number },
164): ProviderResult {
165 const fail = (message: string): ProviderResult => ({
166 ok: false,
167 failure: { kind: "malformed", message, latencyMs: opts.latencyMs },
168 })
169 if (typeof result !== "object" || result === null) return fail("response has no result")
170 const r = result as Record<string, unknown>
171 const answers = r.answers as Record<string, unknown> | undefined
172 if (typeof answers !== "object" || answers === null) return fail("result has no answers")
173 const difficulty = answers[Q_DIFFICULTY] as Record<string, unknown> | undefined
174 if (typeof difficulty !== "object" || difficulty === null) return fail(`no answer for "${Q_DIFFICULTY}"`)
175
176 const probabilities = levelProbabilities(difficulty, opts.rubric)
177 if (probabilities === undefined) return fail("difficulty answer has no usable probabilities")
178 const level = topLevel(probabilities)
179 const confidence = num(difficulty.confidence)
180 const score = num(difficulty.score)
181
182 const followUp = answers[Q_FOLLOW_UP] as Record<string, unknown> | undefined
183 const contextDependent = followUp && typeof followUp === "object" ? num(followUp.noul) : undefined
184
185 const usage = r.usage as Record<string, unknown> | undefined
186 const recommendation: Recommendation = {
187 provider: opts.provider,
188 level,
189 confidence: probabilities[level],
190 probabilities,
191 latencyMs: opts.latencyMs,
192 }
193 if (confidence !== undefined && confidence >= 0 && confidence <= 1) recommendation.providerConfidence = confidence
194 if (score !== undefined) recommendation.score = score
195 if (contextDependent !== undefined && contextDependent >= 0 && contextDependent <= 1) recommendation.contextDependent = contextDependent
196 const inputTokens = usage ? num(usage.input_tokens) : undefined
197 if (inputTokens !== undefined) recommendation.inputTokens = inputTokens
198 return { ok: true, recommendation }
199}
200
201/**
202 * Turns a Workers AI success envelope into a Recommendation, or says why it
203 * cannot. Exported for tests and the calibration script.
204 */
205export function parseResponse(
206 bodyText: string,
207 opts: { provider: string; rubric: Rubric; latencyMs: number },
208): ProviderResult {
209 let envelope: Envelope
210 try {
211 envelope = JSON.parse(bodyText) as Envelope
212 } catch {
213 return { ok: false, failure: { kind: "malformed", message: "response is not JSON", latencyMs: opts.latencyMs } }
214 }
215 if (envelope.success === false) return { ok: false, failure: classifyHttpFailure(200, bodyText, opts.latencyMs) }
216 return parseResult(envelope.result, opts)
217}
218
219const TIMED_OUT: unique symbol = Symbol("timeout")
220
221/** The Clef provider. One HTTP request per call, no retries: a retry would
222 * only add latency in the interactive path, and the fallback is cheap. */
223export function clefProvider(opts: ClefOptions): DecisionProvider {
224 const secrets = [opts.apiToken, opts.accountId]
225 return {
226 name: opts.model,
227 async decide(prompt: string): Promise<ProviderResult> {
228 if (!opts.accountId || !opts.apiToken) {
229 return {
230 ok: false,
231 failure: { kind: "not-configured", message: "Cloudflare account ID or API token not set", latencyMs: 0 },
232 }
233 }
234 const started = await opts.now()
235 const elapsed = async () => Math.round((await opts.now()) - started)
236 let response: { status: number; ok: boolean; text: string } | typeof TIMED_OUT
237 const timer = new AbortController()
238 try {
239 const request = opts.fetch(endpoint(opts.accountId, opts.model, opts.apiBase), {
240 method: "POST",
241 headers: { Authorization: `Bearer ${opts.apiToken}`, "Content-Type": "application/json" },
242 body: buildRequestBody(prompt, opts),
243 })
244 // $.http.fetch takes no abort signal: on a timeout the request is left
245 // to finish on its own, and its answer is ignored.
246 request.catch(() => {})
247 const deadline = opts.sleep(opts.timeoutMs, timer.signal).then(
248 (): typeof TIMED_OUT => TIMED_OUT,
249 (): typeof TIMED_OUT => TIMED_OUT,
250 )
251 response = await Promise.race([request, deadline])
252 } catch (error) {
253 const message = redact(error instanceof Error ? error.message : String(error), secrets).slice(0, 200)
254 return { ok: false, failure: { kind: "network", message, latencyMs: await elapsed() } }
255 } finally {
256 timer.abort()
257 }
258 const latencyMs = await elapsed()
259 if (response === TIMED_OUT) {
260 return { ok: false, failure: { kind: "timeout", message: `no answer within ${opts.timeoutMs} ms`, latencyMs } }
261 }
262 if (!response.ok) return { ok: false, failure: classifyHttpFailure(response.status, response.text, latencyMs, secrets) }
263 return parseResponse(response.text, { provider: opts.model, rubric: opts.rubric, latencyMs })
264 },
265 }
266}
267hooks/lib/config.ts 237 lines1// Reads the plugin's `userConfig` values (set with /plugin configure or
2// /config) into a validated Config. Bad values fall back to defaults and are
3// reported, never thrown: a typo in one field must not stop Claude Code.
4
5import { isEffort } from "./models.ts"
6import { CLEF_MODELS, type ClefModel } from "./clef.ts"
7import { DEFAULT_LOCAL_ENDPOINT, DEFAULT_LOCAL_MODEL } from "./local.ts"
8import { EFFORTS, LEVELS, type Effort, type Level, type ProfileSpec } from "./types.ts"
9
10export const LOW_CONFIDENCE_POLICIES = ["upper-of-top-two", "bump", "hold", "fallback", "obey"] as const
11export type LowConfidencePolicy = (typeof LOW_CONFIDENCE_POLICIES)[number]
12
13export const ANNOUNCE_MODES = ["status", "answer", "both", "off"] as const
14export type AnnounceMode = (typeof ANNOUNCE_MODES)[number]
15
16export const BACKENDS = ["cloudflare", "local"] as const
17export type Backend = (typeof BACKENDS)[number]
18
19export type Config = {
20 enabled: boolean
21 accountId?: string
22 apiToken?: string
23 decisionModel: ClefModel
24 /** Where decisions come from: Workers AI, or a SystemOne server on this machine. */
25 backend: Backend
26 localEndpoint: string
27 localModel: string
28 localTimeoutMs: number
29 timeoutMs: number
30 profiles: Record<Level, ProfileSpec>
31 confidenceThreshold: number
32 lowConfidencePolicy: LowConfidencePolicy
33 fallbackLevel: Level
34 /** Honour /model (pause) and /effort (effort only) changes made mid-session. */
35 pauseOnNativeChange: boolean
36 /** P(follow-up) at or above which a prompt never routes below the last route. */
37 followUpThreshold: number
38 /** Hold a warm model on a downgrade when the context is at least this big; 0 = never hold. */
39 cacheHoldMinTokens: number
40 /** Prompt-cache TTL in minutes; 0 = work it out from the environment. */
41 cacheTtlMinutes: number
42 maxEffort?: Effort
43 dailyNeuronBudget: number
44 maxPromptChars: number
45 announce: AnnounceMode
46 logEnabled: boolean
47 logPrompts: boolean
48 logDir?: string
49 rubricFile?: string
50}
51
52export const DEFAULT_PROFILES: Record<Level, string> = {
53 trivial: "haiku",
54 simple: "sonnet:low",
55 standard: "sonnet:medium",
56 hard: "opus:high",
57 deep: "opus:xhigh",
58}
59
60export const DEFAULTS = {
61 decisionModel: "clef-flash" as ClefModel,
62 backend: "cloudflare" as Backend,
63 localEndpoint: DEFAULT_LOCAL_ENDPOINT,
64 localModel: DEFAULT_LOCAL_MODEL,
65 localTimeoutMs: 6000,
66 timeoutMs: 1500,
67 confidenceThreshold: 0.55,
68 lowConfidencePolicy: "upper-of-top-two" as LowConfidencePolicy,
69 fallbackLevel: "standard" as Level,
70 followUpThreshold: 0.6,
71 cacheHoldMinTokens: 40_000,
72 cacheTtlMinutes: 0,
73 dailyNeuronBudget: 9_000,
74 maxPromptChars: 6_000,
75 announce: "status" as AnnounceMode,
76}
77
78/** "opus:high" → { model: "opus", effort: "high" }; "haiku" → { model: "haiku" }. */
79export function parseProfile(level: Level, text: string): ProfileSpec | string {
80 const trimmed = text.trim()
81 if (trimmed === "") return `profile_${level} is empty`
82 const split = splitTarget(trimmed)
83 if (split.model === "") return `profile_${level} "${text}" names no model`
84 if (split.badEffort) return `profile_${level} "${text}": effort must be one of ${EFFORTS.join(", ")}`
85 return split.effort ? { level, model: split.model, effort: split.effort } : { level, model: split.model }
86}
87
88/**
89 * Splits "model:effort". A suffix that is not an effort name stays part of
90 * the model, since provider IDs carry colons ("...-v1:0" on Bedrock); a
91 * suffix that looks like a word but is no effort is reported.
92 */
93export function splitTarget(text: string): { model: string; effort?: Effort; badEffort?: boolean } {
94 const colon = text.lastIndexOf(":")
95 if (colon === -1) return { model: text.trim() }
96 const model = text.slice(0, colon).trim()
97 const suffix = text.slice(colon + 1).trim().toLowerCase()
98 if (suffix === "" || suffix === "default") return { model }
99 if (isEffort(suffix)) return { model, effort: suffix }
100 if (/^\d+$/.test(suffix)) return { model: text.trim() }
101 return { model, badEffort: true }
102}
103
104type Options = Readonly<Record<string, unknown>>
105
106function str(options: Options, key: string): string | undefined {
107 const v = options[key]
108 return typeof v === "string" && v.trim() !== "" ? v.trim() : undefined
109}
110
111function numIn(options: Options, key: string, min: number, max: number, fallback: number, problems: string[]): number {
112 const v = options[key]
113 if (v === undefined || v === "") return fallback
114 const n = typeof v === "number" ? v : Number(v)
115 if (!Number.isFinite(n) || n < min || n > max) {
116 problems.push(`${key} must be a number from ${min} to ${max}; using ${fallback}`)
117 return fallback
118 }
119 return n
120}
121
122function oneOf<T extends string>(options: Options, key: string, allowed: readonly T[], fallback: T, problems: string[]): T {
123 const v = str(options, key)
124 if (v === undefined) return fallback
125 if ((allowed as readonly string[]).includes(v)) return v as T
126 problems.push(`${key} must be one of ${allowed.join(", ")}; using ${fallback}`)
127 return fallback
128}
129
130function bool(options: Options, key: string, fallback: boolean): boolean {
131 const v = options[key]
132 if (typeof v === "boolean") return v
133 if (v === "true") return true
134 if (v === "false") return false
135 return fallback
136}
137
138function httpUrl(options: Options, key: string, fallback: string, problems: string[]): string {
139 const v = str(options, key)
140 if (v === undefined) return fallback
141 if (/^https?:\/\/\S+$/i.test(v)) return v
142 problems.push(`${key} must be an http:// or https:// URL; using ${fallback}`)
143 return fallback
144}
145
146/**
147 * Reads the optional advanced-settings file (`clef-model-router.json`): a
148 * JSON object with the same keys as the plugin options. Plugin options win
149 * over it; it wins over the defaults.
150 */
151export function parseAdvancedFile(text: string | undefined): { values: Options; problems: string[] } {
152 if (text === undefined) return { values: {}, problems: [] }
153 try {
154 const value = JSON.parse(text) as unknown
155 if (typeof value !== "object" || value === null || Array.isArray(value)) {
156 return { values: {}, problems: ["clef-model-router.json must hold a JSON object; ignoring it"] }
157 }
158 const values = { ...(value as Record<string, unknown>) }
159 // Credentials belong in the plugin's secure storage, not in a plain file.
160 const problems: string[] = []
161 if ("cloudflare_api_token" in values) {
162 delete values.cloudflare_api_token
163 problems.push("clef-model-router.json: cloudflare_api_token is ignored there; set it with /plugin configure")
164 }
165 return { values, problems }
166 } catch {
167 return { values: {}, problems: ["clef-model-router.json is not valid JSON; ignoring it"] }
168 }
169}
170
171/**
172 * Builds the Config from plugin options plus environment fallbacks for the
173 * credentials (CLOUDFLARE_ACCOUNT_ID / CLOUDFLARE_API_TOKEN), which the
174 * caller reads and passes in.
175 */
176export function parseConfig(
177 options: Options,
178 envFallback: { accountId?: string; apiToken?: string } = {},
179): { config: Config; problems: string[] } {
180 const problems: string[] = []
181 const profiles = {} as Record<Level, ProfileSpec>
182 // `profiles` is the five of them in one line, lowest first; a `profile_<level>`
183 // key (in the advanced file) names one.
184 const list = str(options, "profiles")?.split(",").map((s) => s.trim())
185 if (list && list.length !== LEVELS.length) {
186 problems.push(`profiles must list ${LEVELS.length} entries (${LEVELS.join(", ")}), got ${list.length}; using the defaults`)
187 }
188 for (const [i, level] of LEVELS.entries()) {
189 const fromList = list && list.length === LEVELS.length ? list[i] : undefined
190 const raw = str(options, `profile_${level}`) ?? fromList ?? DEFAULT_PROFILES[level]
191 const parsed = parseProfile(level, raw)
192 if (typeof parsed === "string") {
193 problems.push(`${parsed}; using "${DEFAULT_PROFILES[level]}"`)
194 profiles[level] = parseProfile(level, DEFAULT_PROFILES[level]) as ProfileSpec
195 } else profiles[level] = parsed
196 }
197 const maxEffortRaw = str(options, "max_effort")
198 let maxEffort: Effort | undefined
199 if (maxEffortRaw !== undefined && maxEffortRaw !== "none") {
200 if (isEffort(maxEffortRaw)) maxEffort = maxEffortRaw
201 else problems.push(`max_effort must be one of ${EFFORTS.join(", ")} or none; ignoring it`)
202 }
203
204 const config: Config = {
205 enabled: bool(options, "enabled", true),
206 decisionModel: oneOf(options, "decision_model", CLEF_MODELS, DEFAULTS.decisionModel, problems),
207 backend: oneOf(options, "backend", BACKENDS, DEFAULTS.backend, problems),
208 localEndpoint: httpUrl(options, "local_endpoint", DEFAULTS.localEndpoint, problems),
209 localModel: str(options, "local_model") ?? DEFAULTS.localModel,
210 localTimeoutMs: numIn(options, "local_timeout_ms", 100, 10_000, DEFAULTS.localTimeoutMs, problems),
211 timeoutMs: numIn(options, "timeout_ms", 100, 10_000, DEFAULTS.timeoutMs, problems),
212 profiles,
213 confidenceThreshold: numIn(options, "confidence_threshold", 0, 1, DEFAULTS.confidenceThreshold, problems),
214 lowConfidencePolicy: oneOf(options, "low_confidence_policy", LOW_CONFIDENCE_POLICIES, DEFAULTS.lowConfidencePolicy, problems),
215 fallbackLevel: oneOf(options, "fallback_profile", LEVELS, DEFAULTS.fallbackLevel, problems),
216 followUpThreshold: numIn(options, "follow_up_threshold", 0, 1, DEFAULTS.followUpThreshold, problems),
217 cacheHoldMinTokens: numIn(options, "cache_hold_min_tokens", 0, 10_000_000, DEFAULTS.cacheHoldMinTokens, problems),
218 cacheTtlMinutes: numIn(options, "cache_ttl_minutes", 0, 1440, DEFAULTS.cacheTtlMinutes, problems),
219 dailyNeuronBudget: numIn(options, "daily_neuron_budget", 0, 1_000_000_000, DEFAULTS.dailyNeuronBudget, problems),
220 maxPromptChars: numIn(options, "max_prompt_chars", 200, 200_000, DEFAULTS.maxPromptChars, problems),
221 announce: oneOf(options, "announce", ANNOUNCE_MODES, DEFAULTS.announce, problems),
222 pauseOnNativeChange: bool(options, "pause_on_native_change", true),
223 logEnabled: bool(options, "log_enabled", true),
224 logPrompts: bool(options, "log_prompts", false),
225 }
226 if (maxEffort) config.maxEffort = maxEffort
227 const accountId = str(options, "cloudflare_account_id") ?? envFallback.accountId
228 const apiToken = str(options, "cloudflare_api_token") ?? envFallback.apiToken
229 if (accountId) config.accountId = accountId
230 if (apiToken) config.apiToken = apiToken
231 const logDir = str(options, "log_dir")
232 if (logDir) config.logDir = logDir
233 const rubricFile = str(options, "rubric_file")
234 if (rubricFile) config.rubricFile = rubricFile
235 return { config, problems }
236}
237hooks/lib/local.ts 96 lines1// A Clef served on this machine (llama.cpp's llama-server, Ollama 0.35.1+,
2// or anything else that answers POST /v1/systemone): the same request body as
3// the Workers AI backend, no credentials, the raw SystemOne body back.
4// Nothing here throws.
5//
6// POST {endpoint}
7// { "model": "clef-flash", "state": ..., "questions": { id: {type, instructions, criteria} } }
8// → { "model", "answers": { id: answer }, "usage": { input_tokens, output_tokens } }
9
10import { buildRequestBody, parseResult, type HttpLike } from "./clef.ts"
11import { redact } from "./redact.ts"
12import type { QuestionStyle, Rubric } from "./rubric.ts"
13import type { DecisionProvider, ProviderFailure, ProviderResult } from "./types.ts"
14
15export const DEFAULT_LOCAL_ENDPOINT = "http://127.0.0.1:8080/v1/systemone"
16export const DEFAULT_LOCAL_MODEL = "clef-flash"
17
18export type LocalOptions = {
19 endpoint: string
20 /** The `model` field of the request: an alias llama-server was given, or an Ollama tag. */
21 model: string
22 rubric: Rubric
23 style?: QuestionStyle
24 timeoutMs: number
25 maxPromptChars: number
26 fetch: HttpLike
27 sleep: (ms: number, signal: AbortSignal) => Promise<void>
28 now: () => number | Promise<number>
29}
30
31/** Maps an HTTP status from a local server to a failure kind. */
32export function classifyLocalHttpFailure(status: number, bodyText: string, latencyMs: number): ProviderFailure {
33 const text = redact(bodyText.replace(/\s+/g, " ").trim()).slice(0, 160)
34 const message = text || `HTTP ${status}`
35 if (status === 404) {
36 return {
37 kind: "bad-request",
38 message: `${message} (no /v1/systemone here: needs llama.cpp b11379+ or Ollama 0.35.1+)`,
39 status,
40 latencyMs,
41 }
42 }
43 if (status === 408) return { kind: "timeout", message, status, latencyMs }
44 if (status === 429) return { kind: "rate-limited", message, status, latencyMs }
45 if (status >= 500) return { kind: "server", message, status, latencyMs }
46 if (status >= 400) return { kind: "bad-request", message, status, latencyMs }
47 return { kind: "malformed", message, status, latencyMs }
48}
49
50const TIMED_OUT: unique symbol = Symbol("timeout")
51
52export function localProvider(opts: LocalOptions): DecisionProvider {
53 const name = `local:${opts.model}`
54 return {
55 name,
56 async decide(prompt: string): Promise<ProviderResult> {
57 const started = await opts.now()
58 const elapsed = async () => Math.round((await opts.now()) - started)
59 let response: { status: number; ok: boolean; text: string } | typeof TIMED_OUT
60 const timer = new AbortController()
61 try {
62 const request = opts.fetch(opts.endpoint, {
63 method: "POST",
64 headers: { "Content-Type": "application/json" },
65 body: buildRequestBody(prompt, opts),
66 })
67 // $.http.fetch takes no abort signal: on a timeout the request is left
68 // to finish on its own, and its answer is ignored.
69 request.catch(() => {})
70 const deadline = opts.sleep(opts.timeoutMs, timer.signal).then(
71 (): typeof TIMED_OUT => TIMED_OUT,
72 (): typeof TIMED_OUT => TIMED_OUT,
73 )
74 response = await Promise.race([request, deadline])
75 } catch (error) {
76 const message = redact(error instanceof Error ? error.message : String(error)).slice(0, 200)
77 return { ok: false, failure: { kind: "network", message, latencyMs: await elapsed() } }
78 } finally {
79 timer.abort()
80 }
81 const latencyMs = await elapsed()
82 if (response === TIMED_OUT) {
83 return { ok: false, failure: { kind: "timeout", message: `no answer within ${opts.timeoutMs} ms`, latencyMs } }
84 }
85 if (!response.ok) return { ok: false, failure: classifyLocalHttpFailure(response.status, response.text, latencyMs) }
86 let body: unknown
87 try {
88 body = JSON.parse(response.text)
89 } catch {
90 return { ok: false, failure: { kind: "malformed", message: "response is not JSON", latencyMs } }
91 }
92 return parseResult(body, { provider: name, rubric: opts.rubric, latencyMs })
93 },
94 }
95}
96hooks/lib/env.ts 66 lines1// Turns the environment variables Claude Code itself honours into the facts
2// the policy needs. Pure: the hooks module reads the variables (by literal
3// name, as the engine requires) and passes them here.
4
5import type { ModelEnv } from "./models.ts"
6
7export type EnvValues = {
8 ANTHROPIC_DEFAULT_HAIKU_MODEL?: string
9 ANTHROPIC_DEFAULT_SONNET_MODEL?: string
10 ANTHROPIC_DEFAULT_OPUS_MODEL?: string
11 ANTHROPIC_DEFAULT_FABLE_MODEL?: string
12 CLAUDE_CODE_USE_BEDROCK?: string
13 CLAUDE_CODE_USE_VERTEX?: string
14 CLAUDE_CODE_USE_FOUNDRY?: string
15 CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS?: string
16 ANTHROPIC_API_KEY?: string
17 CLAUDE_CODE_PROMPT_CACHE_TTL?: string
18 FORCE_PROMPT_CACHING_5M?: string
19 ENABLE_PROMPT_CACHING_1H?: string
20}
21
22const truthy = (v: string | undefined) => v !== undefined && v !== "" && v !== "0" && v.toLowerCase() !== "false"
23
24export function modelEnvFrom(env: EnvValues): ModelEnv {
25 const defaults: ModelEnv["defaults"] = {}
26 if (env.ANTHROPIC_DEFAULT_HAIKU_MODEL) defaults.haiku = env.ANTHROPIC_DEFAULT_HAIKU_MODEL
27 if (env.ANTHROPIC_DEFAULT_SONNET_MODEL) defaults.sonnet = env.ANTHROPIC_DEFAULT_SONNET_MODEL
28 if (env.ANTHROPIC_DEFAULT_OPUS_MODEL) defaults.opus = env.ANTHROPIC_DEFAULT_OPUS_MODEL
29 if (env.ANTHROPIC_DEFAULT_FABLE_MODEL) defaults.fable = env.ANTHROPIC_DEFAULT_FABLE_MODEL
30 return {
31 defaults,
32 thirdParty: truthy(env.CLAUDE_CODE_USE_BEDROCK) || truthy(env.CLAUDE_CODE_USE_VERTEX) || truthy(env.CLAUDE_CODE_USE_FOUNDRY),
33 betasDisabled: truthy(env.CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS),
34 }
35}
36
37const FIVE_MIN = 5 * 60_000
38const ONE_HOUR = 60 * 60_000
39
40function ttlValue(v: string | undefined): number | undefined {
41 if (v === "5m") return FIVE_MIN
42 if (v === "1h") return ONE_HOUR
43 return undefined
44}
45
46/**
47 * The main conversation's prompt-cache TTL, resolved in the order Claude
48 * Code's prompt-caching docs give: FORCE_PROMPT_CACHING_5M, the TTL variable,
49 * the `promptCacheTtl` setting, ENABLE_PROMPT_CACHING_1H, then the default
50 * (one hour on a subscription, five minutes with an API key or a cloud
51 * provider). A subscription past its included usage drops to five minutes,
52 * which this cannot see; it then over-estimates warmth, which only makes the
53 * router hold a model it could have left.
54 */
55export function cacheTtlMs(configMinutes: number, env: EnvValues, settingsTtl?: unknown): number {
56 if (configMinutes > 0) return configMinutes * 60_000
57 if (truthy(env.FORCE_PROMPT_CACHING_5M)) return FIVE_MIN
58 const fromEnv = ttlValue(env.CLAUDE_CODE_PROMPT_CACHE_TTL)
59 if (fromEnv) return fromEnv
60 const fromSettings = typeof settingsTtl === "string" ? ttlValue(settingsTtl) : undefined
61 if (fromSettings) return fromSettings
62 if (truthy(env.ENABLE_PROMPT_CACHING_1H)) return ONE_HOUR
63 const viaKeyOrCloud = !!env.ANTHROPIC_API_KEY || modelEnvFrom(env).thirdParty
64 return viaKeyOrCloud ? FIVE_MIN : ONE_HOUR
65}
66hooks/lib/format.ts 215 lines1// Everything the router shows: the one-line status, and the text of each
2// /clef subcommand. Plain text, so it reads the same on every surface.
3
4import type { Config } from "./config.ts"
5import { displayName, resolveModel, type ModelEnv } from "./models.ts"
6import { describe, kTokens, pct, type RouterMode } from "./policy.ts"
7import { presence } from "./redact.ts"
8import type { Stats } from "./log.ts"
9import { LEVELS, type Decision, type Level, type Target } from "./types.ts"
10
11const RULE_LABEL: Record<string, string> = {
12 "low-confidence": "unsure",
13 "context-dependent": "follow-up",
14 unavailable: "unavailable",
15 "context-window": "window",
16 "cache-hold": "held for cache",
17 "effort-clamp": "effort clamped",
18 "effort-cap": "effort capped",
19 "pinned-effort": "pinned effort",
20}
21
22/** The one line under the prompt, e.g. "Clef → Sonnet · medium · 87%". */
23export function statusLine(d: Decision): string | undefined {
24 const route = d.final ? describe(d.final) : undefined
25 const notes = d.adjustments.map((a) => RULE_LABEL[a.rule] ?? a.rule)
26 const tail = notes.length > 0 ? ` (${[...new Set(notes)].join(", ")})` : ""
27 switch (d.source) {
28 case "clef": {
29 const conf = d.recommendation ? ` · ${pct(d.recommendation.confidence)}` : ""
30 return `Clef → ${route}${conf}${tail}`
31 }
32 case "fallback":
33 return `Clef ✕ ${d.failure?.kind ?? "no answer"} → ${route}${tail}`
34 case "continuation":
35 return route ? `Clef ↻ ${route}${tail}` : undefined
36 case "override":
37 return route ? `+ ${route}${tail}` : `+ ${d.note ?? "override"}`
38 case "pin":
39 return route ? `Pinned → ${route}${tail}` : `Pinned: ${d.note ?? ""}`
40 case "native":
41 return "Clef paused (you chose /model)"
42 case "disabled":
43 return d.note?.startsWith("+off") ? "Clef skipped this turn" : "Clef off"
44 }
45}
46
47/** A line under the answer, when `announce` asks for one. */
48export function answerLine(d: Decision, latencyMs?: number): string | undefined {
49 const line = statusLine(d)
50 if (!line) return undefined
51 return latencyMs !== undefined && d.source === "clef" ? `${line} · ${latencyMs} ms` : line
52}
53
54function bar(p: number, width = 20): string {
55 const filled = Math.round(p * width)
56 return "█".repeat(filled) + "·".repeat(width - filled)
57}
58
59export function distribution(probabilities: Record<Level, number>, mark?: Level, final?: Level): string[] {
60 return LEVELS.map((level) => {
61 const tags = [level === mark ? "← Clef" : "", level === final && final !== mark ? "← routed" : ""].filter(Boolean).join(" ")
62 return ` ${level.padEnd(9)} ${bar(probabilities[level])} ${pct(probabilities[level]).padStart(4)} ${tags}`.trimEnd()
63 })
64}
65
66export function explain(d: Decision): string[] {
67 const lines: string[] = []
68 lines.push(` route ${d.final ? `${describe(d.final)} (${d.final.model}${d.final.level ? `, profile ${d.final.level}` : ""})` : "untouched (Claude Code's own model and effort)"}`)
69 lines.push(` source ${d.source}${d.note ? ` — ${d.note}` : ""}`)
70 const rec = d.recommendation
71 if (rec) {
72 const extras = [
73 `${pct(rec.confidence)} on ${rec.level}`,
74 rec.providerConfidence !== undefined ? `clef confidence ${pct(rec.providerConfidence)}` : "",
75 rec.score !== undefined ? `score ${rec.score.toFixed(2)}/4` : "",
76 rec.contextDependent !== undefined ? `follow-up ${pct(rec.contextDependent)}` : "",
77 `${rec.latencyMs} ms`,
78 rec.inputTokens !== undefined ? `${rec.inputTokens} tokens` : "",
79 ].filter(Boolean)
80 lines.push(` clef ${rec.provider}: ${extras.join(" · ")}`)
81 lines.push(...distribution(rec.probabilities, rec.level, d.final?.level))
82 }
83 if (d.failure) lines.push(` failure ${d.failure.kind}: ${d.failure.message}${d.failure.latencyMs ? ` (${d.failure.latencyMs} ms)` : ""}`)
84 for (const a of d.adjustments) lines.push(` policy ${a.rule}: ${a.from} → ${a.to} — ${a.reason}`)
85 return lines
86}
87
88export type StatusArgs = {
89 config: Config
90 configProblems: readonly string[]
91 modelEnv: ModelEnv
92 mode: RouterMode
93 pin?: Target
94 last?: Decision
95 guard: { calls: number; inputTokens: number; neurons: number; blocked?: string }
96 cache?: { model: string; promptTokens: number; ageSeconds: number; ttlSeconds: number }
97 unavailable: readonly string[]
98 logDir: string | undefined
99 advancedPath?: string
100}
101
102export function targetText(t: Target): string {
103 return [t.level ?? t.model ?? "", t.effort ? `:${t.effort}` : ""].join("")
104}
105
106export function statusReport(a: StatusArgs): string {
107 const c = a.config
108 const lines = ["Clef router"]
109 const mode =
110 !c.enabled ? "disabled in plugin config" : a.mode === "auto" ? (a.pin ? `pinned to ${targetText(a.pin)} (/clef auto to unpin)` : "auto") : a.mode === "off" ? "off for this session (/clef on)" : "paused: you changed /model (/clef auto to resume)"
111 lines.push(` mode ${mode}`)
112 if (c.backend === "local") {
113 lines.push(` decider local · ${c.localModel} at ${c.localEndpoint} · timeout ${c.localTimeoutMs} ms`)
114 lines.push(` today ${a.guard.calls} Clef calls · ${kTokens(a.guard.inputTokens)} input tokens${a.guard.blocked ? ` · ${a.guard.blocked}` : ""}`)
115 } else {
116 lines.push(` decider ${c.decisionModel} · timeout ${c.timeoutMs} ms · account ${presence(c.accountId)} · token ${presence(c.apiToken)}`)
117 const budget = c.dailyNeuronBudget > 0 ? ` of ${c.dailyNeuronBudget} budget` : ""
118 lines.push(` today ${a.guard.calls} Clef calls · ${kTokens(a.guard.inputTokens)} input tokens · ~${Math.round(a.guard.neurons)} neurons${budget}${a.guard.blocked ? ` · ${a.guard.blocked}` : ""}`)
119 }
120 if (a.cache) {
121 const warm = a.cache.ageSeconds < a.cache.ttlSeconds
122 lines.push(` cache ${displayName(a.cache.model)} · ${kTokens(a.cache.promptTokens)} context · ${warm ? `warm (${a.cache.ageSeconds}s of ${a.cache.ttlSeconds}s)` : "cold"}`)
123 }
124 if (a.unavailable.length > 0) lines.push(` unusable ${a.unavailable.join(", ")}`)
125 for (const p of a.configProblems) lines.push(` config! ${p}`)
126 if (a.last) {
127 lines.push("", "Last turn")
128 lines.push(...explain(a.last))
129 } else {
130 lines.push("", "No turn routed yet this session.")
131 }
132 if (a.advancedPath) lines.push("", `Advanced settings: ${a.advancedPath} (optional)`)
133 lines.push(`Log: ${c.logEnabled ? (a.logDir ?? "(unavailable)") : "off"}${c.logEnabled && c.logPrompts ? " (with prompt text)" : ""}`)
134 lines.push("Commands: /clef history · stats · profiles · test <prompt> · pin <target> · auto · off · on · feedback under|ok|over")
135 return lines.join("\n")
136}
137
138export function profilesReport(config: Config, env: ModelEnv, unavailable: readonly string[]): string {
139 const lines = ["Profiles, lowest first. Change them with /config (Profiles) or /plugin configure."]
140 for (const level of LEVELS) {
141 const spec = config.profiles[level]
142 const id = resolveModel(spec.model, env)
143 const state = id === undefined ? "cannot resolve here (set a full model ID)" : unavailable.includes(id) ? `${id} (failed this session)` : id
144 lines.push(` ${level.padEnd(9)} ${`${spec.model}${spec.effort ? `:${spec.effort}` : ""}`.padEnd(16)} → ${state}`)
145 }
146 lines.push(` fallback ${config.fallbackLevel} · low confidence (< ${pct(config.confidenceThreshold)}): ${config.lowConfidencePolicy}`)
147 return lines.join("\n")
148}
149
150export type HistoryRow = { decision: Decision; prompt: string; answeredModel?: string }
151
152export function historyReport(rows: readonly HistoryRow[]): string {
153 if (rows.length === 0) return "No turns yet this session."
154 const lines = ["Turns this session, newest first"]
155 lines.push(" ms route conf source prompt")
156 for (const row of [...rows].reverse()) {
157 const d = row.decision
158 const ms = d.recommendation?.latencyMs ?? d.failure?.latencyMs
159 const conf = d.recommendation ? pct(d.recommendation.confidence) : "—"
160 const route = d.final ? describe(d.final) : "untouched"
161 const flag = d.adjustments.length > 0 ? "*" : " "
162 const prompt = row.prompt.replace(/\s+/g, " ").slice(0, 48)
163 lines.push(` ${String(ms ?? "—").padStart(5)} ${(route + flag).padEnd(22)} ${conf.padStart(4)} ${d.source.padEnd(12)} ${prompt}`)
164 for (const a of d.adjustments) lines.push(` ${a.rule}: ${a.from} → ${a.to}`)
165 if (d.failure) lines.push(` ${d.failure.kind}: ${d.failure.message}`)
166 if (row.answeredModel && d.final && !row.answeredModel.startsWith(d.final.model)) lines.push(` answered by ${row.answeredModel}`)
167 }
168 lines.push(" * policy changed Clef's recommendation")
169 return lines.join("\n")
170}
171
172function table(title: string, map: Record<string, number>, total: number): string[] {
173 const entries = Object.entries(map).sort((a, b) => b[1] - a[1])
174 if (entries.length === 0) return []
175 return [` ${title}`, ...entries.map(([k, v]) => ` ${k.padEnd(24)} ${String(v).padStart(5)} ${pct(total ? v / total : 0).padStart(4)}`)]
176}
177
178export function statsReport(s: Stats, days: number, files: number): string {
179 if (s.turns === 0 && Object.keys(s.feedback).length === 0) return `No routing log entries in the last ${days} day(s).`
180 const lines = [`Routing over the last ${days} day(s): ${s.turns} turns in ${files} log file(s)`]
181 lines.push(...table("by model", s.byModel, s.turns))
182 lines.push(...table("by effort", s.byEffort, s.turns))
183 lines.push(...table("by profile", s.byLevel, s.turns))
184 lines.push(...table("by source", s.bySource, s.turns))
185 lines.push(" clef")
186 lines.push(` calls ${s.clefCalls} · ${kTokens(s.clefInputTokens)} input tokens`)
187 if (s.latency) lines.push(` latency mean ${s.latency.mean} ms · p50 ${s.latency.p50} ms · p95 ${s.latency.p95} ms`)
188 if (s.meanConfidence !== undefined) lines.push(` mean probability of Clef's pick ${pct(s.meanConfidence)}`)
189 lines.push(` recommendation changed by policy: ${s.recommendationChanged} · cache holds: ${s.cacheHolds} · manual overrides: ${s.overrides}`)
190 const failures = Object.entries(s.failures)
191 if (failures.length > 0) lines.push(` fallbacks: ${failures.map(([k, v]) => `${k} ${v}`).join(", ")}`)
192 const fb = Object.entries(s.feedback)
193 if (fb.length > 0) lines.push(` your feedback: ${fb.map(([k, v]) => `${k} ${v}`).join(", ")}`)
194 return lines.join("\n")
195}
196
197export const HELP = [
198 "Clef router — picks the Claude model and effort for each turn.",
199 "",
200 " /clef status and the last decision, with Clef's probabilities",
201 " /clef history this session's turns",
202 " /clef stats [days] totals from the local log (default 7 days)",
203 " /clef profiles what each difficulty level runs on",
204 " /clef test <prompt> ask Clef about a prompt without sending it to Claude",
205 " /clef pin <target> use one target for the rest of the session",
206 " target: trivial|simple|standard|hard|deep, haiku|sonnet|opus|fable,",
207 " a model ID, with optional :effort; or :effort alone",
208 " /clef auto unpin, resume after /model, and hand effort back after /effort",
209 " /clef off | on stop or resume routing for this session",
210 " /clef feedback under|ok|over [note] rate the last route, for later analysis",
211 "",
212 "One turn only: start a prompt with +target, e.g. `+opus:max why does this deadlock?`,",
213 "or `+off ...` to leave that turn to Claude Code.",
214].join("\n")
215hooks/lib/guard.ts 131 lines1// Deterministic gates in front of the Clef call: a local daily budget that
2// keeps usage inside Workers AI's free allocation, a pause after the
3// allocation is reported exhausted, and a circuit breaker so a dead endpoint
4// costs one timeout, not one per prompt.
5//
6// The state is plain JSON kept in `$.store`, shared by every session on the
7// machine (best effort: two sessions writing at once may lose a count).
8
9import type { FailureKind, ProviderFailure } from "./types.ts"
10
11export type GuardState = {
12 /** UTC day (YYYY-MM-DD) the counters belong to. */
13 day: string
14 calls: number
15 inputTokens: number
16 /** Clef said the day's free allocation is used up; no calls until `day` changes. */
17 quotaExhausted?: boolean
18 consecutiveFailures: number
19 /** Epoch ms until which no call is made. */
20 pausedUntil?: number
21 pauseReason?: FailureKind
22}
23
24/**
25 * Neurons per million input tokens, derived from Cloudflare's published
26 * prices ($0.09/M for clef-flash, $0.24/M for clef) at $0.011 per 1,000
27 * neurons. Clef is not yet in the per-model neuron table; this is an
28 * estimate, and the Workers AI dashboard is the authority.
29 */
30export const NEURONS_PER_M_INPUT: Record<string, number> = {
31 "clef-flash": (0.09 / 0.011) * 1000,
32 clef: (0.24 / 0.011) * 1000,
33}
34
35export function utcDay(epochMs: number): string {
36 return new Date(epochMs).toISOString().slice(0, 10)
37}
38
39export function freshGuard(epochMs: number): GuardState {
40 return { day: utcDay(epochMs), calls: 0, inputTokens: 0, consecutiveFailures: 0 }
41}
42
43/** Rolls the counters over at 00:00 UTC, when Workers AI's allocation resets. */
44export function normaliseGuard(state: unknown, epochMs: number): GuardState {
45 const today = utcDay(epochMs)
46 if (typeof state !== "object" || state === null) return freshGuard(epochMs)
47 const s = state as Partial<GuardState>
48 if (s.day !== today) {
49 const next = freshGuard(epochMs)
50 if (typeof s.pausedUntil === "number" && s.pausedUntil > epochMs && s.pauseReason !== "quota") {
51 next.pausedUntil = s.pausedUntil
52 next.pauseReason = s.pauseReason
53 }
54 return next
55 }
56 return {
57 day: today,
58 calls: typeof s.calls === "number" ? s.calls : 0,
59 inputTokens: typeof s.inputTokens === "number" ? s.inputTokens : 0,
60 consecutiveFailures: typeof s.consecutiveFailures === "number" ? s.consecutiveFailures : 0,
61 ...(s.quotaExhausted ? { quotaExhausted: true } : {}),
62 ...(typeof s.pausedUntil === "number" ? { pausedUntil: s.pausedUntil } : {}),
63 ...(s.pauseReason ? { pauseReason: s.pauseReason } : {}),
64 }
65}
66
67export function estimatedNeurons(state: GuardState, model: string): number {
68 return (state.inputTokens / 1_000_000) * (NEURONS_PER_M_INPUT[model] ?? NEURONS_PER_M_INPUT.clef!)
69}
70
71/** Why no call should be made now, or undefined to go ahead. */
72export function blockedReason(
73 state: GuardState,
74 opts: { now: number; model: string; dailyNeuronBudget: number },
75): ProviderFailure | undefined {
76 if (state.quotaExhausted) {
77 return { kind: "quota", message: "Workers AI daily free allocation used up; resets 00:00 UTC", latencyMs: 0 }
78 }
79 if (opts.dailyNeuronBudget > 0 && estimatedNeurons(state, opts.model) >= opts.dailyNeuronBudget) {
80 return {
81 kind: "budget",
82 message: `local daily budget of ${opts.dailyNeuronBudget} neurons reached; resets 00:00 UTC`,
83 latencyMs: 0,
84 }
85 }
86 if (state.pausedUntil !== undefined && state.pausedUntil > opts.now) {
87 const seconds = Math.ceil((state.pausedUntil - opts.now) / 1000)
88 return {
89 kind: "circuit-open",
90 message: `paused ${seconds}s after ${state.pauseReason ?? "repeated failures"}`,
91 latencyMs: 0,
92 }
93 }
94 return undefined
95}
96
97/** How long to stop calling after a failure, in ms; 0 for none. */
98export const PAUSES = {
99 /** After this many failures in a row, stop calling for a while. */
100 breakerThreshold: 3,
101 breakerMs: 5 * 60_000,
102 rateLimitedMs: 60_000,
103 /** Auth and request errors need the user to fix configuration. */
104 configMs: 30 * 60_000,
105}
106
107export function recordSuccess(state: GuardState, inputTokens: number | undefined): GuardState {
108 const next: GuardState = {
109 ...state,
110 calls: state.calls + 1,
111 inputTokens: state.inputTokens + (inputTokens ?? 0),
112 consecutiveFailures: 0,
113 }
114 delete next.pausedUntil
115 delete next.pauseReason
116 return next
117}
118
119export function recordFailure(state: GuardState, failure: ProviderFailure, now: number): GuardState {
120 // Failures that never reached Cloudflare count toward nothing.
121 if (failure.kind === "not-configured" || failure.kind === "budget" || failure.kind === "circuit-open") return state
122 const next: GuardState = { ...state, calls: state.calls + 1, consecutiveFailures: state.consecutiveFailures + 1 }
123 if (failure.kind === "quota") return { ...next, quotaExhausted: true }
124 let pause = 0
125 if (failure.kind === "rate-limited") pause = PAUSES.rateLimitedMs
126 else if (failure.kind === "auth" || failure.kind === "bad-request") pause = PAUSES.configMs
127 else if (next.consecutiveFailures >= PAUSES.breakerThreshold) pause = PAUSES.breakerMs
128 if (pause > 0) return { ...next, pausedUntil: now + pause, pauseReason: failure.kind }
129 return next
130}
131hooks/lib/log.ts 229 lines1// The local routing log: one JSON object per line, one file per UTC day and
2// session, under the log directory. Nothing is sent anywhere. Prompt text is
3// left out unless `log_prompts` is on; a short SHA-256 prefix lets repeated
4// prompts be recognised without storing them.
5
6import type { Decision, Effort, Level, Source } from "./types.ts"
7
8export const LOG_VERSION = 1
9
10export type AnsweredUsage = {
11 model: string
12 inputTokens: number
13 outputTokens: number
14 cacheReadTokens: number
15 cacheWriteTokens: number
16}
17
18export type TurnRecord = {
19 v: number
20 type: "turn"
21 ts: string
22 session: string
23 turn: string
24 kind: Decision["kind"]
25 source: Source
26 promptHash?: string
27 promptChars: number
28 prompt?: string
29 provider?: string
30 recommendation?: {
31 level: Level
32 /** Probability of `level`: what the threshold compares. */
33 confidence: number
34 /** Clef's own `confidence` field (entropy-like; not a probability). */
35 clefConfidence?: number
36 probabilities: Record<Level, number>
37 score?: number
38 followUp?: number
39 }
40 latencyMs?: number
41 clefInputTokens?: number
42 proposed?: { level?: Level; model: string; effort?: Effort }
43 final?: { level?: Level; model: string; effort?: Effort }
44 adjustments: { rule: string; from: string; to: string; reason: string }[]
45 failure?: { kind: string; message: string; status?: number }
46 note?: string
47 /** What the API said answered, summed over the turn's main-loop requests. */
48 answered?: AnsweredUsage
49 steps?: number
50 durationMs?: number
51 endReason?: string
52}
53
54export type FeedbackRecord = {
55 v: number
56 type: "feedback"
57 ts: string
58 session: string
59 turn?: string
60 verdict: "under" | "ok" | "over"
61 note?: string
62}
63
64export type LogRecord = TurnRecord | FeedbackRecord
65
66export async function promptHash(text: string): Promise<string> {
67 const bytes = new TextEncoder().encode(text)
68 const digest = await crypto.subtle.digest("SHA-256", bytes)
69 return Array.from(new Uint8Array(digest).slice(0, 8), (b) => b.toString(16).padStart(2, "0")).join("")
70}
71
72export function turnRecord(args: {
73 decision: Decision
74 session: string
75 ts: string
76 promptText: string
77 hash?: string
78 logPrompts: boolean
79 answered?: AnsweredUsage
80 steps?: number
81 durationMs?: number
82 endReason?: string
83}): TurnRecord {
84 const { decision: d } = args
85 const record: TurnRecord = {
86 v: LOG_VERSION,
87 type: "turn",
88 ts: args.ts,
89 session: args.session,
90 turn: d.turnId,
91 kind: d.kind,
92 source: d.source,
93 promptChars: args.promptText.length,
94 adjustments: d.adjustments.map((a) => ({ ...a })),
95 }
96 if (args.hash) record.promptHash = args.hash
97 if (args.logPrompts) record.prompt = args.promptText
98 const rec = d.recommendation
99 if (rec) {
100 record.provider = rec.provider
101 record.recommendation = { level: rec.level, confidence: rec.confidence, probabilities: { ...rec.probabilities } }
102 if (rec.providerConfidence !== undefined) record.recommendation.clefConfidence = rec.providerConfidence
103 if (rec.score !== undefined) record.recommendation.score = rec.score
104 if (rec.contextDependent !== undefined) record.recommendation.followUp = rec.contextDependent
105 record.latencyMs = rec.latencyMs
106 if (rec.inputTokens !== undefined) record.clefInputTokens = rec.inputTokens
107 }
108 if (d.failure) {
109 record.failure = { kind: d.failure.kind, message: d.failure.message }
110 if (d.failure.status !== undefined) record.failure.status = d.failure.status
111 if (d.failure.latencyMs > 0) record.latencyMs = d.failure.latencyMs
112 }
113 if (d.proposed) record.proposed = { ...d.proposed }
114 if (d.final) record.final = { ...d.final }
115 if (d.note) record.note = d.note
116 if (args.answered) record.answered = { ...args.answered }
117 if (args.steps !== undefined) record.steps = args.steps
118 if (args.durationMs !== undefined) record.durationMs = args.durationMs
119 if (args.endReason) record.endReason = args.endReason
120 return record
121}
122
123export function logFileName(ts: string, session: string): string {
124 const safe = session.replace(/[^A-Za-z0-9_-]/g, "").slice(0, 12) || "session"
125 return `routing-${ts.slice(0, 10)}-${safe}.jsonl`
126}
127
128export function parseLines(text: string): LogRecord[] {
129 const out: LogRecord[] = []
130 for (const line of text.split("\n")) {
131 if (line.trim() === "") continue
132 try {
133 const value = JSON.parse(line) as LogRecord
134 if (value && typeof value === "object" && (value.type === "turn" || value.type === "feedback")) out.push(value)
135 } catch {
136 // A torn line from a crash mid-write; skip it.
137 }
138 }
139 return out
140}
141
142export type Stats = {
143 turns: number
144 bySource: Record<string, number>
145 byModel: Record<string, number>
146 byEffort: Record<string, number>
147 byLevel: Record<string, number>
148 clefCalls: number
149 latency: { mean: number; p50: number; p95: number } | undefined
150 meanConfidence: number | undefined
151 failures: Record<string, number>
152 overrides: number
153 cacheHolds: number
154 adjusted: number
155 /** Turns where Clef's raw level differs from the final route's level. */
156 recommendationChanged: number
157 feedback: Record<string, number>
158 clefInputTokens: number
159}
160
161function quantile(sorted: number[], q: number): number {
162 if (sorted.length === 0) return 0
163 const i = Math.min(sorted.length - 1, Math.max(0, Math.ceil(q * sorted.length) - 1))
164 return sorted[i]!
165}
166
167function bump(map: Record<string, number>, key: string) {
168 map[key] = (map[key] ?? 0) + 1
169}
170
171export function aggregate(records: readonly LogRecord[]): Stats {
172 const stats: Stats = {
173 turns: 0,
174 bySource: {},
175 byModel: {},
176 byEffort: {},
177 byLevel: {},
178 clefCalls: 0,
179 latency: undefined,
180 meanConfidence: undefined,
181 failures: {},
182 overrides: 0,
183 cacheHolds: 0,
184 adjusted: 0,
185 recommendationChanged: 0,
186 feedback: {},
187 clefInputTokens: 0,
188 }
189 const latencies: number[] = []
190 let confidenceSum = 0
191 let confidenceN = 0
192 for (const r of records) {
193 if (r.type === "feedback") {
194 bump(stats.feedback, r.verdict)
195 continue
196 }
197 stats.turns++
198 bump(stats.bySource, r.source)
199 const model = r.answered?.model ?? r.final?.model ?? "session default"
200 bump(stats.byModel, model.replace(/-\d{8}$/, ""))
201 bump(stats.byEffort, r.final ? (r.final.effort ?? "default") : "untouched")
202 if (r.final?.level) bump(stats.byLevel, r.final.level)
203 if (r.recommendation || r.failure) {
204 if (r.failure?.kind !== "not-configured" && r.failure?.kind !== "budget" && r.failure?.kind !== "circuit-open") stats.clefCalls++
205 }
206 if (r.recommendation) {
207 if (r.latencyMs !== undefined) latencies.push(r.latencyMs)
208 confidenceSum += r.recommendation.confidence
209 confidenceN++
210 if (r.final?.level !== r.recommendation.level || r.final?.model !== r.proposed?.model) stats.recommendationChanged++
211 }
212 if (r.failure) bump(stats.failures, r.failure.kind)
213 if (r.source === "override" || r.source === "pin") stats.overrides++
214 if (r.adjustments.some((a) => a.rule === "cache-hold")) stats.cacheHolds++
215 if (r.adjustments.length > 0) stats.adjusted++
216 stats.clefInputTokens += r.clefInputTokens ?? 0
217 }
218 if (latencies.length > 0) {
219 const sorted = [...latencies].sort((a, b) => a - b)
220 stats.latency = {
221 mean: Math.round(sorted.reduce((s, x) => s + x, 0) / sorted.length),
222 p50: quantile(sorted, 0.5),
223 p95: quantile(sorted, 0.95),
224 }
225 }
226 if (confidenceN > 0) stats.meanConfidence = confidenceSum / confidenceN
227 return stats
228}
229hooks/lib/models.ts 134 lines1// What this router knows about Claude models: how an alias resolves, which
2// effort levels a model takes, how big its window is, and how it ranks.
3//
4// Source: Claude Code model configuration docs (Claude Code 2.1.289,
5// October 2026). `turn.step` does not resolve aliases (a request for
6// "haiku" fails with unrecognized_model), so every route is resolved here to
7// a full ID before it is sent.
8
9import { EFFORTS, type Effort } from "./types.ts"
10
11export type Family = "haiku" | "sonnet" | "opus" | "fable"
12
13/** What each alias resolves to on the Anthropic API, as Claude Code does. */
14export const FIRST_PARTY_ALIASES: Record<Family, string> = {
15 haiku: "claude-haiku-4-5",
16 sonnet: "claude-sonnet-5-5",
17 opus: "claude-opus-5-5",
18 fable: "claude-fable-5-1",
19}
20
21/**
22 * The environment the resolution depends on: Claude Code's own
23 * ANTHROPIC_DEFAULT_*_MODEL pins, and whether a third-party provider is in use
24 * (where first-party IDs do not exist).
25 */
26export type ModelEnv = {
27 defaults: Partial<Record<Family, string>>
28 thirdParty: boolean
29 /** Prompt-cache beta features off (CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS). */
30 betasDisabled: boolean
31}
32
33export const FIRST_PARTY_ENV: ModelEnv = { defaults: {}, thirdParty: false, betasDisabled: false }
34
35export function isAlias(model: string): model is Family {
36 return model === "haiku" || model === "sonnet" || model === "opus" || model === "fable"
37}
38
39/** Resolves an alias or ID to a concrete model ID, or undefined if it cannot be. */
40export function resolveModel(model: string, env: ModelEnv): string | undefined {
41 const m = model.trim()
42 if (m === "") return undefined
43 if (!isAlias(m)) return m
44 const pinned = env.defaults[m]
45 if (pinned) return pinned
46 // On Bedrock, Vertex or Foundry, a first-party ID would fail.
47 return env.thirdParty ? undefined : FIRST_PARTY_ALIASES[m]
48}
49
50export function familyOf(modelId: string): Family | undefined {
51 const id = modelId.toLowerCase()
52 if (id.includes("haiku")) return "haiku"
53 if (id.includes("sonnet")) return "sonnet"
54 if (id.includes("fable")) return "fable"
55 if (id.includes("opus")) return "opus"
56 return undefined
57}
58
59const RANK: Record<Family, number> = { haiku: 0, sonnet: 1, opus: 2, fable: 3 }
60
61/** Capability/price rank for comparing two models; undefined if unknown. */
62export function rankOf(modelId: string): number | undefined {
63 const f = familyOf(modelId)
64 return f === undefined ? undefined : RANK[f]
65}
66
67/** Display name for the status line: "Sonnet", "Opus", or the raw ID. */
68export function displayName(modelId: string): string {
69 const f = familyOf(modelId)
70 return f === undefined ? modelId : f[0]!.toUpperCase() + f.slice(1)
71}
72
73/**
74 * The effort levels a model accepts; null when it takes no effort at all,
75 * undefined when the model is unknown (send what was asked; Claude Code
76 * clamps it).
77 */
78export function effortsFor(modelId: string): readonly Effort[] | null | undefined {
79 const id = modelId.toLowerCase()
80 if (id.includes("haiku")) return null
81 if (/(opus|sonnet)-4-6/.test(id)) return ["low", "medium", "high", "max"]
82 if (/fable|opus-5|sonnet-5|opus-4-[78]/.test(id)) return EFFORTS
83 return undefined
84}
85
86/**
87 * The highest supported level at or below the one asked, which is what
88 * Claude Code itself does; undefined when the model takes no effort.
89 */
90export function clampEffort(modelId: string, effort: Effort | undefined): Effort | undefined {
91 if (effort === undefined) return undefined
92 const supported = effortsFor(modelId)
93 if (supported === null) return undefined
94 if (supported === undefined) return effort
95 for (let i = EFFORTS.indexOf(effort); i >= 0; i--) {
96 const level = EFFORTS[i]!
97 if (supported.includes(level)) return level
98 }
99 return supported[0]
100}
101
102export function capEffort(effort: Effort | undefined, cap: Effort | undefined): Effort | undefined {
103 if (effort === undefined || cap === undefined) return effort
104 return EFFORTS.indexOf(effort) > EFFORTS.indexOf(cap) ? cap : effort
105}
106
107/** Context window in tokens where it is known to be smaller than 1M. */
108export function windowOf(modelId: string): number | undefined {
109 const id = modelId.toLowerCase()
110 if (id.includes("haiku")) return 200_000
111 return undefined
112}
113
114/**
115 * Whether changing effort between requests keeps the prompt cache. Per the
116 * Claude Code prompt-caching docs, it does on Opus 5.5, Sonnet 5.5 and
117 * Fable 5.1 with an API key or subscription, and not on Bedrock, Vertex, a
118 * gateway, or with experimental betas disabled. Elsewhere it is a full miss.
119 */
120export function effortChangeKeepsCache(modelId: string, env: ModelEnv): boolean {
121 if (env.thirdParty || env.betasDisabled) return false
122 return /claude-(opus|sonnet)-5-5|claude-fable-5-1/.test(modelId.toLowerCase())
123}
124
125/** Compares IDs ignoring a date suffix and a [1m] marker. */
126export function sameModel(a: string, b: string): boolean {
127 const norm = (id: string) => id.toLowerCase().replace(/\[1m\]$/, "").replace(/-\d{8}$/, "")
128 return norm(a) === norm(b)
129}
130
131export function isEffort(value: string): value is Effort {
132 return (EFFORTS as readonly string[]).includes(value)
133}
134hooks/lib/overrides.ts 183 lines1// Explicit user intent, parsed deterministically: the one-turn `+target`
2// prompt prefix, `/clef` command targets, and turns that only continue the
3// previous one. No model is involved in any of this.
4
5import { splitTarget } from "./config.ts"
6import { isAlias, isEffort } from "./models.ts"
7import { LEVELS, type Level, type Target, type TurnKind } from "./types.ts"
8
9/**
10 * Parses a target: a profile (`hard`), an alias or model ID with an optional
11 * effort (`opus`, `opus:max`, `claude-sonnet-5-5:low`), or an effort alone
12 * (`:high`). Undefined when the text is none of those.
13 */
14export function parseTarget(text: string): Target | undefined {
15 const t = text.trim()
16 if (t === "") return undefined
17 if (t.startsWith(":")) {
18 const effort = t.slice(1).toLowerCase()
19 return isEffort(effort) ? { effort } : undefined
20 }
21 const { model, effort, badEffort } = splitTarget(t)
22 if (badEffort || model === "") return undefined
23 const lower = model.toLowerCase()
24 const target: Target = {}
25 if ((LEVELS as readonly string[]).includes(lower)) target.level = lower as Level
26 else if (isAlias(lower)) target.model = lower
27 else if (/^claude-[a-z0-9.-]+(\[1m\])?$/i.test(model) || /anthropic\./i.test(model)) target.model = model
28 else return undefined
29 if (effort) target.effort = effort
30 return target
31}
32
33export type PrefixResult = { text: string; override?: Target | "off" }
34
35const PREFIX = /^\+(\S+)(?:\s+|$)/
36
37/**
38 * A prompt that starts with `+target ` routes that one turn to the target and
39 * reaches Claude without the prefix; `+off ` runs the turn as Claude Code
40 * would. Anything else (`+1`, `+x`) is left untouched.
41 */
42export function parsePrefix(text: string): PrefixResult {
43 const match = PREFIX.exec(text)
44 if (!match) return { text }
45 const token = match[1]!
46 const rest = text.slice(match[0].length)
47 if (token.toLowerCase() === "off" || token.toLowerCase() === "noroute") return { text: rest, override: "off" }
48 const target = parseTarget(token)
49 if (!target) return { text }
50 return { text: rest, override: target }
51}
52
53/** Words a go-ahead is made of ("yes, do it", "ok go ahead", "lgtm, ship it"). */
54const GO_AHEAD_WORDS = new Set([
55 "y", "ya", "yes", "yep", "yeah", "yup", "ok", "okay", "k", "sure", "alright", "fine", "cool", "great", "perfect",
56 "go", "ahead", "for", "it", "do", "that", "this", "continue", "proceed", "carry", "on", "keep", "going", "next",
57 "lgtm", "looks", "sounds", "good", "ship", "please", "approved", "approve", "confirm", "confirmed", "thanks",
58])
59/** A go-ahead says yes to something; "it", "on" or "good" alone do not. */
60const GO_AHEAD_ANCHORS = new Set([
61 "y", "ya", "yes", "yep", "yeah", "yup", "ok", "okay", "k", "sure", "alright", "go", "do", "continue", "proceed",
62 "carry", "keep", "next", "lgtm", "ship", "approved", "approve", "confirm", "confirmed", "sounds", "looks",
63])
64
65/** Classifies a turn's text before anything is asked of Clef. */
66export function turnKind(text: string): TurnKind {
67 const t = text.trim()
68 if (t === "") return "empty"
69 if (t.startsWith("<task-notification>")) return "notification"
70 if (t.length <= 40) {
71 const words = t.toLowerCase().replace(/[.,!;:'"]+/g, " ").split(/\s+/).filter(Boolean)
72 if (words.length > 0 && words.length <= 6 && words.every((w) => GO_AHEAD_WORDS.has(w)) && words.some((w) => GO_AHEAD_ANCHORS.has(w)))
73 return "go-ahead"
74 }
75 return "prompt"
76}
77
78export type ClefCommand =
79 | { kind: "status" }
80 | { kind: "history" }
81 | { kind: "stats"; days: number }
82 | { kind: "profiles" }
83 | { kind: "test"; prompt: string }
84 | { kind: "auto" }
85 | { kind: "on" }
86 | { kind: "off" }
87 | { kind: "pin"; target: Target }
88 | { kind: "feedback"; verdict: "under" | "ok" | "over"; note?: string }
89 | { kind: "help" }
90 | { kind: "error"; message: string }
91
92const FEEDBACK: Record<string, "under" | "ok" | "over"> = {
93 under: "under",
94 underpowered: "under",
95 weak: "under",
96 "too-weak": "under",
97 ok: "ok",
98 right: "ok",
99 good: "ok",
100 over: "over",
101 overpowered: "over",
102 strong: "over",
103 "too-strong": "over",
104}
105
106export function parseCommand(args: string): ClefCommand {
107 const trimmed = args.trim()
108 const space = trimmed.search(/\s/)
109 const head = (space === -1 ? trimmed : trimmed.slice(0, space)).toLowerCase()
110 const rest = space === -1 ? "" : trimmed.slice(space + 1).trim()
111 switch (head) {
112 case "":
113 case "status":
114 return { kind: "status" }
115 case "history":
116 case "log":
117 return { kind: "history" }
118 case "stats": {
119 const days = rest === "" ? 7 : Number(rest)
120 return Number.isInteger(days) && days > 0 && days <= 366
121 ? { kind: "stats", days }
122 : { kind: "error", message: "usage: /clef stats [days]" }
123 }
124 case "profiles":
125 return { kind: "profiles" }
126 case "test":
127 return rest === "" ? { kind: "error", message: "usage: /clef test <prompt>" } : { kind: "test", prompt: rest }
128 case "auto":
129 case "unpin":
130 return { kind: "auto" }
131 case "on":
132 return { kind: "on" }
133 case "off":
134 return { kind: "off" }
135 case "pin": {
136 const target = parseTarget(rest)
137 return target
138 ? { kind: "pin", target }
139 : { kind: "error", message: `usage: /clef pin <${LEVELS.join("|")}|haiku|sonnet|opus|fable|model-id>[:effort] or /clef pin :<effort>` }
140 }
141 case "feedback": {
142 const space2 = rest.search(/\s/)
143 const word = (space2 === -1 ? rest : rest.slice(0, space2)).toLowerCase()
144 const verdict = FEEDBACK[word]
145 if (!verdict) return { kind: "error", message: "usage: /clef feedback under|ok|over [note]" }
146 const note = space2 === -1 ? "" : rest.slice(space2 + 1).trim()
147 return note ? { kind: "feedback", verdict, note } : { kind: "feedback", verdict }
148 }
149 case "help":
150 return { kind: "help" }
151 default: {
152 // `/clef opus:high` as shorthand for `/clef pin opus:high`.
153 const target = parseTarget(trimmed)
154 return target ? { kind: "pin", target } : { kind: "error", message: `unknown subcommand "${head}"; try /clef help` }
155 }
156 }
157}
158
159export type NativeEffort = {
160 /** The effort Claude Code sent when the router started watching. */
161 baseline?: string | number
162 /** An effort the person set with /effort since, in force until /clef auto. */
163 native?: string | number
164}
165
166/**
167 * Tracks the effort Claude Code itself would send, seen at each turn's first
168 * request. A change from the baseline is the person's /effort (or a skill's
169 * `effort` for one turn): it is honoured, and dropped again when the engine's
170 * effort returns to the baseline, so a one-turn skill does not stick.
171 */
172export function trackNativeEffort(prev: NativeEffort, engine: string | number | undefined): NativeEffort & { change?: "set" | "cleared" } {
173 const keep = (): NativeEffort => ({
174 ...(prev.baseline !== undefined ? { baseline: prev.baseline } : {}),
175 ...(prev.native !== undefined ? { native: prev.native } : {}),
176 })
177 if (engine === undefined) return keep()
178 if (prev.baseline === undefined) return { baseline: engine }
179 if (engine === prev.native) return keep()
180 if (engine === prev.baseline) return prev.native === undefined ? keep() : { baseline: prev.baseline, change: "cleared" }
181 return { baseline: prev.baseline, native: engine, change: "set" }
182}
183hooks/lib/policy.ts 388 lines1// The routing policy: a pure function from what is known at the start of a
2// turn to the route its requests will use, with every change it made on the
3// way recorded. Clef makes the semantic judgment (how hard is this?); this
4// file makes the deterministic ones (what is allowed, what is safe, what is
5// worth the cache), in this order:
6//
7// 1. Routing off (config, +off; /clef off and a mid-session /model change, unless +model/+profile)
8// 2. An explicit choice: a +target prefix this turn, then a /clef pin
9// 3. A continuation (go-ahead, task notification, empty) reuses the last route
10// 4. Clef's recommendation, or the fallback when Clef did not answer
11// 5. Low confidence → the configured confidence policy
12// 6. A follow-up never routes below the route it follows
13// 7. A profile whose model is unavailable → the nearest available one, upward first
14// 8. A context too big for the model's window → the nearest profile that fits
15// 9. A downgrade that would throw away a large warm prompt cache → hold the model
16// 10. Effort is capped and clamped to what the model takes
17
18import type { Config } from "./config.ts"
19import {
20 capEffort,
21 clampEffort,
22 displayName,
23 effortChangeKeepsCache,
24 rankOf,
25 resolveModel,
26 sameModel,
27 windowOf,
28 type ModelEnv,
29} from "./models.ts"
30import {
31 EFFORTS,
32 LEVELS,
33 type Adjustment,
34 type Decision,
35 type Effort,
36 type Level,
37 type ProviderResult,
38 type Route,
39 type Source,
40 type Target,
41 type TurnKind,
42} from "./types.ts"
43
44export type RouterMode = "auto" | "off" | "paused-native"
45
46/** The last request the main loop sent, which is what the prompt cache holds. */
47export type CacheState = {
48 model: string
49 effort?: Effort
50 /** Epoch ms the response finished. */
51 at: number
52 /** Input + cache read + cache write tokens of that request. */
53 promptTokens: number
54}
55
56export type PolicySession = {
57 mode: RouterMode
58 pin?: Target
59 last?: Route
60 cache?: CacheState
61 /** Models that failed when routed to this session. */
62 unavailable: readonly string[]
63}
64
65export type PolicyInput = {
66 turnId: string
67 kind: TurnKind
68 config: Config
69 modelEnv: ModelEnv
70 session: PolicySession
71 /** A one-turn +target prefix; "off" leaves this turn alone. */
72 override?: Target | "off"
73 /** Clef's answer; absent when Clef was not asked. */
74 result?: ProviderResult
75 /** Tokens in the conversation now, when known. */
76 contextTokens?: number
77 now: number
78 cacheTtlMs: number
79}
80
81/** Headroom kept under a model's window for the reply and tool results. */
82export const WINDOW_HEADROOM = 20_000
83
84const idx = (level: Level) => LEVELS.indexOf(level)
85const atIdx = (i: number): Level => LEVELS[Math.max(0, Math.min(LEVELS.length - 1, i))]!
86const higher = (a: Level, b: Level): Level => (idx(a) >= idx(b) ? a : b)
87
88export function describe(route: Route | undefined): string {
89 if (!route) return "session default"
90 return route.effort ? `${displayName(route.model)} · ${route.effort}` : displayName(route.model)
91}
92
93/**
94 * Whether the policy should ask Clef at all for this turn. Explicit choices,
95 * continuations and a disabled router make no network call.
96 */
97export function needsClef(input: Omit<PolicyInput, "result" | "now" | "cacheTtlMs" | "contextTokens">): boolean {
98 const { config, session, override, kind } = input
99 if (!config.enabled || override === "off") return false
100 if (override && (override.level || override.model)) return false
101 if (session.mode !== "auto") return false
102 if (session.pin && (session.pin.level || session.pin.model)) return false
103 if (kind !== "prompt" && session.last) return false
104 if ((kind === "notification" || kind === "empty") && !session.last) return false
105 return true
106}
107
108function isUnavailable(model: string, session: PolicySession): boolean {
109 return session.unavailable.some((u) => sameModel(u, model))
110}
111
112/** The route for a profile, or the nearest usable profile (upward first). */
113function routeForLevel(
114 level: Level,
115 input: PolicyInput,
116 adjustments: Adjustment[],
117): Route | undefined {
118 const order = [idx(level), ...LEVELS.map((_, i) => i).filter((i) => i > idx(level)), ...LEVELS.map((_, i) => i).filter((i) => i < idx(level)).reverse()]
119 for (const i of order) {
120 const lv = atIdx(i)
121 const spec = input.config.profiles[lv]
122 const model = resolveModel(spec.model, input.modelEnv)
123 if (model === undefined || isUnavailable(model, input.session)) continue
124 const route: Route = { level: lv, model }
125 if (spec.effort) route.effort = spec.effort
126 if (lv !== level) {
127 adjustments.push({
128 rule: "unavailable",
129 from: level,
130 to: lv,
131 reason: `profile ${level} (${input.config.profiles[level].model}) has no usable model here`,
132 })
133 }
134 return route
135 }
136 return undefined
137}
138
139function routeForTarget(target: Target, input: PolicyInput, adjustments: Adjustment[]): Route | undefined {
140 if (target.level) {
141 const route = routeForLevel(target.level, input, adjustments)
142 if (route && target.effort) route.effort = target.effort
143 return route
144 }
145 if (target.model) {
146 const model = resolveModel(target.model, input.modelEnv)
147 if (model === undefined) return undefined
148 return target.effort ? { model, effort: target.effort } : { model }
149 }
150 return undefined
151}
152
153/** The more capable of the two most probable levels. */
154function upperOfTopTwo(probabilities: Record<Level, number>, top: Level): Level {
155 let second: Level | undefined
156 for (const level of LEVELS) {
157 if (level === top) continue
158 if (second === undefined || probabilities[level] >= probabilities[second]) second = level
159 }
160 return second === undefined ? top : higher(top, second)
161}
162
163export function decide(input: PolicyInput): Decision {
164 const { config, session, turnId, kind } = input
165 const adjustments: Adjustment[] = []
166 const base = { turnId, kind, adjustments }
167 const disabled = (source: Source, note: string): Decision => ({ ...base, source, note })
168
169 // 1. Off. A one-turn +model or +profile still applies while the session is
170 // off or paused: it is the most explicit request there is.
171 if (!config.enabled) return disabled("disabled", "routing disabled in plugin config")
172 if (input.override === "off") return disabled("disabled", "+off: this turn runs as Claude Code would")
173 const explicitTurn = input.override !== undefined && (input.override.level !== undefined || input.override.model !== undefined)
174 if (session.mode === "off" && !explicitTurn) return disabled("disabled", "routing off for this session (/clef on)")
175 if (session.mode === "paused-native" && !explicitTurn)
176 return disabled("native", "paused: you changed /model (/clef auto to resume)")
177
178 let source: Source
179 let route: Route | undefined
180 let proposed: Route | undefined
181 let note: string | undefined
182 const decision: Decision = { ...base, source: "clef" }
183
184 // 2. Explicit choice.
185 const explicit = input.override ?? (session.mode === "auto" ? session.pin : undefined)
186 const explicitSource: Source = input.override ? "override" : "pin"
187 if (explicit && (explicit.level || explicit.model)) {
188 source = explicitSource
189 route = routeForTarget(explicit, input, adjustments)
190 if (!route) return disabled(explicitSource, `cannot resolve ${explicit.model ?? explicit.level} here`)
191 proposed = { ...route }
192 } else if (kind !== "prompt" && session.last) {
193 // 3. Continuation.
194 source = "continuation"
195 route = { ...session.last }
196 proposed = { ...route }
197 note = kind === "go-ahead" ? "go-ahead continues the last route" : `${kind} continues the last route`
198 } else if ((kind === "notification" || kind === "empty") && !session.last) {
199 return disabled("continuation", "nothing to continue yet")
200 } else {
201 // 4. Clef, or the fallback.
202 const result = input.result
203 let level: Level
204 if (result?.ok) {
205 source = "clef"
206 const rec = result.recommendation
207 decision.recommendation = rec
208 level = rec.level
209 proposed = routeForLevel(level, input, [])
210
211 // 5. Low confidence.
212 if (rec.confidence < config.confidenceThreshold) {
213 let adjusted: Level = level
214 switch (config.lowConfidencePolicy) {
215 case "upper-of-top-two":
216 adjusted = upperOfTopTwo(rec.probabilities, level)
217 break
218 case "bump":
219 adjusted = atIdx(idx(level) + 1)
220 break
221 case "hold":
222 adjusted = session.last?.level ?? config.fallbackLevel
223 break
224 case "fallback":
225 adjusted = config.fallbackLevel
226 break
227 case "obey":
228 break
229 }
230 if (adjusted !== level) {
231 adjustments.push({
232 rule: "low-confidence",
233 from: level,
234 to: adjusted,
235 reason: `confidence ${pct(rec.confidence)} < ${pct(config.confidenceThreshold)} (${config.lowConfidencePolicy})`,
236 })
237 level = adjusted
238 }
239 }
240
241 // 6. Follow-up floor.
242 const lastLevel = session.last?.level
243 if (
244 rec.contextDependent !== undefined &&
245 rec.contextDependent >= config.followUpThreshold &&
246 lastLevel !== undefined &&
247 idx(lastLevel) > idx(level)
248 ) {
249 adjustments.push({
250 rule: "context-dependent",
251 from: level,
252 to: lastLevel,
253 reason: `follow-up (${pct(rec.contextDependent)}) to a ${lastLevel} turn`,
254 })
255 level = lastLevel
256 }
257 } else {
258 source = "fallback"
259 if (result && !result.ok) decision.failure = result.failure
260 const lastLevel = session.last?.level
261 level = lastLevel ? higher(lastLevel, config.fallbackLevel) : config.fallbackLevel
262 note = `${result && !result.ok ? result.failure.kind : "no answer"}: using ${level}`
263 }
264
265 // 7. Availability.
266 route = routeForLevel(level, input, adjustments)
267 if (!route) return { ...decision, ...disabled(source, "no configured profile resolves to a usable model here") }
268 if (source === "fallback") proposed = { ...route }
269 }
270
271 // 8. Context window.
272 const tokens = input.contextTokens ?? session.cache?.promptTokens
273 if (tokens !== undefined) {
274 const fits = (model: string) => {
275 const w = windowOf(model)
276 return w === undefined || tokens + WINDOW_HEADROOM <= w
277 }
278 if (!fits(route.model)) {
279 const start = route.level ? idx(route.level) : 0
280 let moved: Route | undefined
281 for (let i = start + 1; i < LEVELS.length && !moved; i++) {
282 const candidate = routeForLevel(atIdx(i), input, [])
283 if (candidate && fits(candidate.model)) moved = candidate
284 }
285 if (moved) {
286 adjustments.push({
287 rule: "context-window",
288 from: describe(route),
289 to: describe(moved),
290 reason: `${kTokens(tokens)} context does not fit ${displayName(route.model)}'s window`,
291 })
292 route = moved
293 }
294 }
295 }
296
297 // 9. Cache hold: only for routes the router chose itself.
298 const cache = session.cache
299 if ((source === "clef" || source === "continuation") && cache && config.cacheHoldMinTokens > 0) {
300 const warm = input.now - cache.at < input.cacheTtlMs
301 const big = cache.promptTokens >= config.cacheHoldMinTokens
302 const fromRank = rankOf(cache.model)
303 const toRank = rankOf(route.model)
304 if (warm && big && !isUnavailable(cache.model, session)) {
305 const ago = Math.round((input.now - cache.at) / 1000)
306 if (fromRank !== undefined && toRank !== undefined && toRank < fromRank) {
307 const keepsCache = effortChangeKeepsCache(cache.model, input.modelEnv)
308 const held: Route = { model: cache.model }
309 if (route.level) held.level = route.level
310 const effort = keepsCache ? (route.effort ?? cache.effort) : cache.effort
311 if (effort) held.effort = effort
312 adjustments.push({
313 rule: "cache-hold",
314 from: describe(route),
315 to: describe(held),
316 reason: `${kTokens(cache.promptTokens)} context is cached on ${displayName(cache.model)} (${ago}s ago); ${displayName(route.model)} would re-read it uncached`,
317 })
318 route = held
319 } else if (
320 sameModel(route.model, cache.model) &&
321 route.effort &&
322 cache.effort &&
323 EFFORTS.indexOf(route.effort) < EFFORTS.indexOf(cache.effort) &&
324 !effortChangeKeepsCache(cache.model, input.modelEnv)
325 ) {
326 const held: Route = { ...route, effort: cache.effort }
327 adjustments.push({
328 rule: "cache-hold",
329 from: describe(route),
330 to: describe(held),
331 reason: `changing effort on ${displayName(cache.model)} here invalidates the ${kTokens(cache.promptTokens)} cached context`,
332 })
333 route = held
334 }
335 }
336 }
337
338 // An effort-only pin applies to whatever model was chosen.
339 if (session.pin?.effort && !session.pin.level && !session.pin.model && !(input.override && input.override.effort)) {
340 if (route.effort !== session.pin.effort) {
341 adjustments.push({ rule: "pinned-effort", from: route.effort ?? "default", to: session.pin.effort, reason: "/clef pin" })
342 route = { ...route, effort: session.pin.effort }
343 }
344 }
345 if (input.override && !input.override.level && !input.override.model && input.override.effort) {
346 adjustments.push({ rule: "pinned-effort", from: route.effort ?? "default", to: input.override.effort, reason: "+ prefix" })
347 route = { ...route, effort: input.override.effort }
348 }
349
350 // 10. Effort cap and clamp.
351 if (route.effort) {
352 const capped = capEffort(route.effort, config.maxEffort)
353 if (capped !== route.effort) {
354 adjustments.push({ rule: "effort-cap", from: route.effort, to: capped ?? "none", reason: `max_effort is ${config.maxEffort}` })
355 }
356 const clamped = clampEffort(route.model, capped)
357 if (clamped !== capped) {
358 adjustments.push({
359 rule: "effort-clamp",
360 from: capped ?? "none",
361 to: clamped ?? "none",
362 reason: `${displayName(route.model)} ${clamped ? `tops out at ${clamped}` : "takes no effort setting"}`,
363 })
364 }
365 route = { ...route }
366 if (clamped) route.effort = clamped
367 else delete route.effort
368 }
369
370 const out: Decision = { ...decision, source, final: route }
371 if (proposed) out.proposed = proposed
372 if (note) out.note = note
373 return out
374}
375
376export function pct(p: number): string {
377 return `${Math.round(p * 100)}%`
378}
379
380export function kTokens(n: number): string {
381 return n >= 1000 ? `${Math.round(n / 1000)}k` : String(n)
382}
383
384/** Whether policy changed what Clef (or the override) proposed. */
385export function wasAdjusted(decision: Decision): boolean {
386 return decision.adjustments.length > 0
387}
388hooks/lib/rubric.ts 108 lines1// The questions Clef is asked. This is the calibration surface: change the
2// wording here (or point `rubric_file` at a JSON file with the same shape)
3// and nothing else needs to change. `npm run calibrate` shows the effect.
4//
5// Clef answers every question in one forward pass, so the second question
6// costs a few extra input tokens and no extra latency.
7
8import { LEVELS } from "./types.ts"
9
10export type Rubric = {
11 /** What Clef is asked to rate. The prompt itself is sent as the `state`. */
12 instructions: string
13 /** One description per level, lowest first; exactly five. */
14 levels: readonly string[]
15 /** The follow-up question; empty string to not ask it. */
16 followUp: string
17}
18
19export const DEFAULT_RUBRIC: Rubric = {
20 instructions:
21 "A developer sent this message to Claude Code, an AI coding agent working inside their software " +
22 "repository with tools to read, search, edit and run code. Rate how much model capability and " +
23 "reasoning effort the agent needs to do this well. Judge the work the message asks for, not the " +
24 "length of the message: a short request can be hard, and a long paste can still be a trivial task.",
25 levels: [
26 "Trivial: mechanical and obvious, no judgment. Fix a typo, rename one symbol, reformat, run a known " +
27 "command, answer a quick factual or yes/no question, find where something is defined.",
28 "Simple: a small, well-specified change or question in one place. A one-line fix with a clear cause, " +
29 "add a log line or one simple test, explain a short function, a small config edit.",
30 "Moderate: ordinary feature or bug work with a clear goal, a few files, following existing patterns. " +
31 "Add an endpoint or pagination, write tests for a module, fix a reproducible bug, a routine refactor.",
32 "Hard: tricky, ambiguous or multi-step work where a wrong answer is costly. Debug intermittent, " +
33 "concurrency or non-obvious failures, significant refactors, unfamiliar code, performance, " +
34 "security-sensitive changes, choosing between designs.",
35 "Very hard: open-ended investigation or design across a whole subsystem. Root-cause cascading or " +
36 "distributed failures, architecture or migration plans, weigh alternatives then implement and " +
37 "verify, long autonomous work.",
38 ],
39 followUp:
40 "Is this message a short follow-up whose actual task is defined by earlier conversation rather than " +
41 "by the message itself, such as 'yes do it', 'try again', 'go with option 2', 'that didn't work', " +
42 "or 'continue'?",
43}
44
45export type ClefQuestion =
46 | { type: "score"; instructions: string; criteria: string[] }
47 | { type: "choice"; instructions: string; criteria: Record<string, string> }
48 | { type: "noul"; instructions: string; criteria?: { true: string; false: string } }
49
50export type QuestionStyle = "score" | "choice"
51
52/** Question IDs, shared by the request builder and the response parser. */
53export const Q_DIFFICULTY = "difficulty"
54export const Q_FOLLOW_UP = "follow_up"
55
56/**
57 * Builds Clef's question map. `score` (the default) treats the levels as an
58 * ordered rubric, which is what they are; `choice` is kept so calibration can
59 * compare the two on the same corpus.
60 */
61export function buildQuestions(rubric: Rubric, style: QuestionStyle = "score"): Record<string, ClefQuestion> {
62 const questions: Record<string, ClefQuestion> = {}
63 questions[Q_DIFFICULTY] =
64 style === "score"
65 ? { type: "score", instructions: rubric.instructions, criteria: [...rubric.levels] }
66 : {
67 type: "choice",
68 instructions: rubric.instructions,
69 criteria: Object.fromEntries(LEVELS.map((level, i) => [level, rubric.levels[i]!])),
70 }
71 if (rubric.followUp.trim() !== "") {
72 questions[Q_FOLLOW_UP] = {
73 type: "noul",
74 instructions: rubric.followUp,
75 criteria: {
76 true: "The message only makes sense together with earlier conversation.",
77 false: "The message states a self-contained request.",
78 },
79 }
80 }
81 return questions
82}
83
84/** Validates a rubric loaded from a file; returns the problems found. */
85export function rubricProblems(value: unknown): string[] {
86 const problems: string[] = []
87 if (typeof value !== "object" || value === null) return ["rubric must be a JSON object"]
88 const r = value as Record<string, unknown>
89 if (typeof r.instructions !== "string" || r.instructions.trim() === "") problems.push("`instructions` must be a non-empty string")
90 if (!Array.isArray(r.levels) || r.levels.length !== LEVELS.length || !r.levels.every((l) => typeof l === "string" && l.trim() !== ""))
91 problems.push(`\`levels\` must be ${LEVELS.length} non-empty strings, lowest first`)
92 if (r.followUp !== undefined && typeof r.followUp !== "string") problems.push("`followUp` must be a string when present")
93 return problems
94}
95
96export function parseRubric(text: string): { rubric: Rubric } | { problems: string[] } {
97 let value: unknown
98 try {
99 value = JSON.parse(text)
100 } catch {
101 return { problems: ["rubric file is not valid JSON"] }
102 }
103 const problems = rubricProblems(value)
104 if (problems.length > 0) return { problems }
105 const r = value as { instructions: string; levels: string[]; followUp?: string }
106 return { rubric: { instructions: r.instructions, levels: r.levels, followUp: r.followUp ?? DEFAULT_RUBRIC.followUp } }
107}
108