Hands back the conversation itself in place of a compaction summary, less the tool calls and results a System One decision model judges no longer needed. The…

Compaction without a summary. A Claude Code mod for session.compact that hands the conversation back as it was, with only those tool calls and tool outputs removed that a System One decision model says the assistant can do without. What you and the assistant wrote stays word for word.
[!WARNING] The conversation text and the tool inputs are sent to the selected third-party API (TypeSafe, Cloudflare, OpenRouter, Codiv, Perplexity, decisions-api.dev, decisionapi.net or OpenAI) on every compaction. They are sent as they are: a secret you pasted into a prompt, or one that appears in a tool input such as a shell command, goes with them. Tool outputs are not sent: each is replaced by a note of its status and length.
With
providerDecisionon, a second provider can receive data in the same compaction. Every provider indecisionProviders(all eight by default) whose credentials resolve is a candidate to decide it, so a key you exported for another tool,OPENAI_API_KEYfor example, makes that vendor eligible to receive the conversation.decisionProvidersis the setting that narrows the set. The routing question, when there is a choice to make, goes to Cloudflare only whencloudflareis one of the providers considered (the configured provider, and those indecisionProviders) and its credentials are set; otherwise it goes to the configured provider. It carries a profile of the job that includes up to 2,000 characters of your prompts (the text typed after/compact, or your last three prompts).Do not enable this mod for a session whose content must not leave your machine.
session.compactturn.completeproviderDecisionFrom the marketplace, on Claude Code 2.1.275 or later:
claude plugin install decision-compaction --marketplace zchee/claude-code-mods
From a checkout, without installing:
claude --plugin-dir mods/decision-compaction
Then set a credential for the provider you want (see Providers), for example TYPESAFE_API_KEY in the environment, or the typesafeApiKey option in /config.
Options are read from settings under pluginConfigs and appear in /config. A mod loaded with --plugin-dir is keyed decision-compaction@inline (the bare decision-compaction is read too); an installed one is keyed by its plugin id, decision-compaction@<marketplace>.
session.compactOn session.compact (a /compact, the engine's own threshold, or this mod's trigger) the hook:
tool_use_id. The first message is never changed, and neither is any of the last preserveRecentMessages; a call that sits in one of those, or whose output does, is not judged. A tool_use_id that occurs on more than one call or output is not judged either, since a decision is applied by that id.keepThreshold. Either the call and its output both stay; or the call stays and the output is shortened to its opening truncateHeadChars characters, followed by a note saying how much was removed; or the call goes and its output goes with it. Output so short that shortening would save nothing is not touched.A message that had one of its blocks dropped or cut is rebuilt from its role, its text and its remaining tool blocks, and nothing else survives the rebuild. Two consequences: a result that is kept, but shares its message with a result that was dropped or cut, loses any image content it had; and an assistant message that made several calls, one of which was dropped, loses its thinking block.
It fails open. On any error (a missing key, a provider option that names no provider, an HTTP error, a request that is rejected, a compaction whose requests are not answered in time, a malformed answer, a conversation that cannot be fitted or would take too many requests) and when the decisions would remove less than minReductionRatio, it says why in one line and lets the built-in summary run. A precompute compaction is always left to the engine.
Every request that asks about a tool call goes to one provider: the configured one, or the one the provider decision picked. Nothing falls over to another provider. With providerDecision on there can be one more request before them, the routing question, and it may go to a different provider (see providerDecision).
turn.completeOn turn.complete, when the main conversation's context reaches compactAtPercent, the mod requests a compaction. What happens next depends on the usage that compaction leaves:
/compact of your own or a /clear).A turn you interrupted, and a subagent's turn, request nothing.
provider | Endpoint | Model by default | Credentials | Documented limits |
|---|---|---|---|---|
typesafe (default) | https://api.typesafe.ai/v1/systemone | jev-latest | TYPESAFE_API_KEY | 64,000 tokens a request; 32,000 for the state plus the longest question |
cloudflare | https://api.cloudflare.com/client/v4/accounts/<account id>/ai/run/@cf/cloudflare/<model> | clef (clef-flash is the other) | CLOUDFLARE_API_TOKEN and CLOUDFLARE_ACCOUNT_ID | 65,536 tokens; 64 questions a request |
openrouter | https://openrouter.ai/api/v1/systemone | ~typesafe/jev-latest | OPENROUTER_API_KEY | 32,000 tokens for the state plus the questions; 64 questions a request for any model other than Jev |
codiv | https://api.codiv.ai/v1/systemone | openjev-latest (OpenJev) | CODIV_API_KEY | 65,536 tokens for the state and the questions together (the vendor advises a state of about 60,000); no limit on questions |
perplexity | https://api.perplexity.ai/v1/decisions | pplx-decider-v1.1-27b | PERPLEXITY_API_KEY | under 262,144 tokens a request; 128 questions a request |
decisions-api-dev | https://decisions-api.dev/v1/systemone | jev-latest (Jev, through this gateway) | DECISIONS_API_KEY | 32 KiB a request body, counted in bytes, so text that is not ASCII uses more of it; 8 questions a request |
decisionapi-net | https://decisionapi.net/v1/systemone | jev-latest (Jev, through this gateway) | DECISIONAPI_API_KEY | 32 KiB a request body, counted in bytes; 8 questions a request |
openai | https://api.openai.com/v1/decisions | gpt-6-luna | OPENAI_API_KEY | not documented for the Decisions API; the mod holds a request to 32,000 tokens and 64 questions |
The provider option takes exactly these eight names. Any other value is not replaced by the default: nothing is sent anywhere, and every compaction uses the built-in summary with a line naming the value. Only an option left unset or empty means typesafe. For that reason /config shows the option as a text field rather than a list (a list would let the engine turn an unknown value into the default without saying so).
What differs between the providers beyond the table:
openrouter, decisions-api-dev and decisionapi-net are gateways to TypeSafe's Jev. codiv's OpenJev is a model of its own, not TypeSafe's Jev: at Codiv jev-latest is an alias of OpenJev.perplexity accepts only its own decider models. Any other name in model, jev-latest included, is refused with an HTTP 400 and the compaction uses the built-in summary.decisions-api-dev and decisionapi-net wrap the answer in a {code, message, data} envelope. A code other than 0, or an envelope without data.result, ends the compaction with the vendor's message quoted, as for any other failed request.decisions-api-dev accepts a question id only if it starts with a letter, uses letters, digits, _ and -, and is at most 64 characters. The mod's ids can hold . and run longer, so it names each question in a request by its place (q0, q1, …) and maps the answers back to the ids. An answer under a name the mod did not send is refused.openai takes a different request: the state goes as one JSON string in input and the questions as a list, each a predicate (yes/no) or a choice. The mod translates the request into that form and the answers back. OpenAI counts about 148 input tokens for each question (Codiv about 26 for the same one-line question), and the mod's size estimate adds that figure per question. If OpenAI refuses to answer a question (the line reads openai: refused to answer <question id>), or answers in a shape the mod does not read, the compaction ends with the built-in summary and the line names the question or the field.~typesafe/jev-latest is OpenRouter's alias for the newest Jev; typesafe/jev-1.13 pins a release. OpenRouter also takes TypeSafe's bare ids and maps them itself (jev-latest to ~typesafe/jev-latest, jev-1.13 to typesafe/jev-1.13). An id that already has an author prefix is used as written.
OpenRouter also serves Cloudflare's Clef (cloudflare/clef, cloudflare/clef-flash), and passes such a request on to Cloudflare, which refuses one with more than 64 questions. A model other than Jev on openrouter is therefore asked at most 64 questions (32 tool calls) a request, as on cloudflare, while the token limits stay OpenRouter's.
Each credential is read from its plugin option first, then from the environment, at the time of the compaction. The env block of the settings is read only when a credential a provider in play needs is still missing; if the settings cannot be read, that is logged and the compaction goes on with what it has. With providerDecision on and decisionProviders at its default of all eight providers, some key is usually unset, so the settings are read on most compactions.
Text that comes from a provider or from the host (an HTTP error body, a Cloudflare failure envelope, the reason the host gives for refusing a request or a read) is quoted in the toast and the log on one line of at most 160 characters. Before it is shown, every credential this mod resolved is replaced with [redacted], and so is the token after the word Bearer, whether it follows a space, a colon or an equals sign. A resolved value shorter than eight characters is not replaced where it stands on its own, since no provider issues a key that short and replacing it would garble ordinary words.
| Option | Default | Meaning |
|---|---|---|
provider | typesafe | Which provider decides; any other name sends nothing. |
model | provider default | Model name for the configured provider. |
providerDecision | false | Decide the provider per compaction; see providerDecision. |
decisionProviders | all eight providers | With providerDecision on, the providers that may be compared and so may receive the conversation: a list of provider names, or one text of names separated by commas. The configured provider is always considered. See providerDecision. |
typesafeApiKey | unset | TypeSafe API key (stored as a secret). |
openrouterApiKey | unset | OpenRouter API key (stored as a secret). |
cloudflareApiToken | unset | Cloudflare API token allowed to run Workers AI (stored as a secret). |
cloudflareAccountId | unset | Cloudflare account ID. |
codivApiKey | unset | Codiv API key (stored as a secret). |
perplexityApiKey | unset | Perplexity API key (stored as a secret). |
decisionsApiKey | unset | decisions-api.dev API key (stored as a secret). |
decisionapiApiKey | unset | decisionapi.net API key (stored as a secret). |
openaiApiKey | unset | OpenAI API key (stored as a secret). |
keepThreshold | 0.5 | Probability from which a call, or its full result, is kept. |
preserveRecentMessages | 6 | Newest messages that are never changed. |
compactAtPercent | 60 | Context usage at which a compaction is requested; 0 turns the trigger off. |
minReductionRatio | 0.25 | Smallest share of characters the decisions must remove. |
maxStateTokens | 25000 | Most estimated tokens the state may take, a whole number; see Limits. |
maxRequestTokens | 30000 | Most estimated tokens one request may take, state and questions together, a whole number; see Limits. |
truncateHeadChars | 300 | Characters kept of a shortened output; at 0 only the note remains. |
The bounds below are fixed; they are not options.
typesafe, 55,705 on cloudflare, 27,200 on openrouter and openai, 51,000 for the state plus the longest question (85% of the 60,000 the vendor advises) and 55,705 a request on codiv, 222,821 on perplexity, and 27,200 for the state plus the longest question and 54,400 a request on decisions-api-dev and decisionapi-net. Both serve TypeSafe's Jev 1.13 and state their own limit in bytes, so they are held to the token window TypeSafe documents for Jev, and the byte cap below is the bound that decides. openai documents no limit, so it is held to the tightest figures of the others, 32,000 tokens and 64 questions, before the margin. Where the margin leaves less than the default maxRequestTokens of 30,000 (openrouter and openai), the default is lowered there.decisions-api-dev and decisionapi-net also hold a request to the 85% share of their 32 KiB cap, 27,852 of 32,768 bytes, measured as UTF-8 of the encoded request body, since a body over the cap is refused whole. When one of them is the configured provider the state is fitted to that cap as well as to its tokens: the reductions above go on until the body, with the questions of one request set aside (4 tool calls, all that 8 questions hold), fits in 27,852 bytes. English prose takes more bytes for each estimated token than the token budget allows for, and text that is not ASCII more still, so on these two providers it is usually the bytes that decide how much of the conversation the model sees. Questions are counted at or above the bytes they are sent in (decisions-api-dev counts each under the longest name it may send, q7). A conversation still over the cap after every reduction is left to the built-in summary, and the line says how many bytes were left. With providerDecision on, the state is fitted for the configured provider, so one of these two offered as another route may have no room for a question within the cap; it is then ruled out, and the line says how many of the 27,852 bytes the request takes before any question.decisions-api-dev and decisionapi-net that limit is Jev's own 64,000 tokens, which a body within 32 KiB does not reach. On openai it is the mod's fallback of 32,000, not a window OpenAI states, so a reply counting that many is discarded even if the model read the state whole.providerDecision is one request more.providerDecision.At the default budgets the state holds about 25,000 estimated tokens after every reduction, and 16 requests ask about a bounded number of calls. Past either bound the result is always the built-in summary.
The table is a measurement, not a guarantee. It was taken with a fake provider on a synthetic session made only of Read calls with short file paths and 600-character outputs, between one opening and one closing prompt. A real session has longer inputs, more varied tools and more text, so its limits are lower, and differ from one session to the next; read the numbers as rough orders of size.
provider | Calls whose assistant message also carries a sentence of text | Call-only messages |
|---|---|---|
typesafe | roughly 400 (then the state no longer fits) | roughly 900 (then a 17th request) |
cloudflare | roughly 400 (then the state no longer fits) | 512 (32 calls a request) |
openrouter | roughly 290 (then a 17th request) | roughly 820 (then a 17th request) |
Longer tool inputs and longer tool names lower these numbers. Raising maxStateTokens and maxRequestTokens raises them, up to the margins above.
The table has no rows for codiv, perplexity, decisions-api-dev, decisionapi-net and openai; it was not measured on them. decisions-api-dev and decisionapi-net take 8 questions and 32 KiB a request, so of a long conversation the model sees only what is left once the state is shrunk to the cap. The cap counts bytes, so text that is not ASCII uses more of it. Eight questions are the questions of 4 tool calls, so a compaction asks about at most 64 calls over its 16 requests.
providerDecisionOff by default: the configured provider is always used, and no other provider is sent anything.
When on, the providers considered are the configured one and those in decisionProviders. Unset, that option is every provider, all eight. It takes a list of names, or one text of names separated by commas; a cleared field (an empty text) or an empty list names none, so only the configured provider is considered. A name that is no provider is dropped and named in one line of the transcript; the compaction goes on, because dropping a name can only narrow where data goes. A provider that is not considered is sent nothing, even when its credentials are set. A provider that is considered is a candidate as soon as its credentials resolve, so a key exported for another tool makes that vendor eligible. To keep the choice to the original three providers, set decisionProviders to typesafe,cloudflare,openrouter.
The mod first drops every route it can rule out in code: a provider whose credentials are unset, whose limits the fitted state exceeds, or that would need more than 16 requests. Cloudflare contributes two routes (clef and clef-flash). Among routes to the same model one is offered: the direct route when it is usable, otherwise the first gateway in the order openrouter, decisions-api-dev, decisionapi-net. Jev is reached directly through typesafe and through those three gateways; on OpenRouter an id that names Jev (jev, jev-…, alone or under typesafe/ or ~typesafe/) is Jev. Every other route to that model is left out, with the reason "the same model as" the one offered. Any other OpenRouter id, under typesafe/ included, is a model of its own. If two or more routes remain, one choice question is asked over a small profile of the job (message and call counts, state size, tool mix, error share, share of non-ASCII text, and the goal), not over the conversation. The goal is your own text: what you typed after /compact, cut to 2,000 characters, or your last three prompts, cut to 500 characters each. The question is asked on Cloudflare clef-flash only when cloudflare is one of the providers considered and its credentials are set, otherwise on the configured provider. A decisionProviders list wi
hooks/register.ts 295 lines1import type { EngineInterface, PluginOptions, Register } from 'claude-code'
2
3import { reductionOf } from './compact'
4import { configOf, eligibleOf } from './config'
5import {
6 briefOf,
7 credentialsFor,
8 credentialsOf,
9 cutOf,
10 environmentOf,
11 messageOf,
12 PROVIDERS,
13 redacted,
14 unresolvedOf,
15} from './providers'
16import type { Credentials, Environment, ProviderName } from './providers'
17import { decisionLinesOf, percentOf, summaryOf } from './report'
18import { run } from './run'
19import { armOf, isCompactionDue, settle } from './trigger'
20
21/**
22 * How long an outcome toast stays: long enough to read one full line.
23 */
24const TOAST_MS = 15_000
25
26/**
27 * The hook time that must remain after a retry's wait for the work that
28 * follows it: reading the answers and rebuilding the messages.
29 */
30const PAUSE_RESERVE_MS = 1500
31
32/**
33 * Says which names of the `decisionProviders` option are no provider. Each
34 * is quoted, and quoted only in part when long, as the `provider` option
35 * is: the field may hold anything that was pasted into it.
36 */
37function unknownLineOf(names: readonly string[]): string {
38 const quoted = names.map(name => cutOf(JSON.stringify(name)))
39
40 return (
41 `the decisionProviders option names ${quoted.join(', ')}, ` +
42 `${names.length === 1 ? 'which is' : 'which are'} none of ` +
43 `${PROVIDERS.join(', ')}; ignored`
44 )
45}
46
47/**
48 * Reads the credential variables from the process environment, and from the
49 * `env` block of the settings only while a credential one of the providers
50 * needs is still without a value: the settings are a second source, not a
51 * part of every compaction. Each name is spelled out because `$.env.get`
52 * takes a string literal only.
53 *
54 * Settings that cannot be read are no settings. The compaction goes on with
55 * what the options and the environment gave, and `problem` says what went
56 * wrong so that it can be logged.
57 */
58async function readEnvironment(
59 $: EngineInterface,
60 options: PluginOptions,
61 providers: readonly ProviderName[],
62): Promise<{ environment: Environment; problem?: string }> {
63 const read: Environment = {
64 TYPESAFE_API_KEY: await $.env.get('TYPESAFE_API_KEY'),
65 OPENROUTER_API_KEY: await $.env.get('OPENROUTER_API_KEY'),
66 CLOUDFLARE_API_TOKEN: await $.env.get('CLOUDFLARE_API_TOKEN'),
67 CLOUDFLARE_ACCOUNT_ID: await $.env.get('CLOUDFLARE_ACCOUNT_ID'),
68 CODIV_API_KEY: await $.env.get('CODIV_API_KEY'),
69 PERPLEXITY_API_KEY: await $.env.get('PERPLEXITY_API_KEY'),
70 DECISIONS_API_KEY: await $.env.get('DECISIONS_API_KEY'),
71 DECISIONAPI_API_KEY: await $.env.get('DECISIONAPI_API_KEY'),
72 OPENAI_API_KEY: await $.env.get('OPENAI_API_KEY'),
73 }
74
75 if (unresolvedOf(providers, options, read).length === 0) {
76 return { environment: environmentOf(read, undefined) }
77 }
78
79 try {
80 const settings = await $.settings.read()
81
82 return { environment: environmentOf(read, settings.env) }
83 } catch (error) {
84 return {
85 environment: environmentOf(read, undefined),
86 problem: messageOf(error),
87 }
88 }
89}
90
91/**
92 * Writes one line to the transcript. Every line this module shows goes
93 * through here or through `tell`, and so through `redacted`: whatever a
94 * line was built from, no credential leaves in it.
95 */
96function log($: EngineInterface, text: string, credentials: Credentials): void {
97 $.ui.log(redacted(text, credentials))
98}
99
100/**
101 * Says one line both where it stays (the transcript) and where it is seen
102 * at once (a toast), redacted like every other line.
103 */
104function tell(
105 $: EngineInterface,
106 text: string,
107 credentials: Credentials,
108): void {
109 const shown = redacted(text, credentials)
110
111 $.ui.log(shown)
112 $.ui.toast(shown, { timeoutMs: TOAST_MS })
113}
114
115/**
116 * Registers the two hooks.
117 *
118 * `session.compact` hands back the conversation itself in place of a
119 * summary, less the tool calls and results a decision model judged no
120 * longer needed. It
121 * fails open: on any error, and when too little would be removed, it says
122 * why in one line and hands the compaction on to the built-in summary.
123 *
124 * `turn.complete` requests a compaction once the main conversation's
125 * context reaches `compactAtPercent`, one request at a time, and not again
126 * while usage stays near what a compaction that could not get under the
127 * threshold left.
128 *
129 * @param on the engine's registrar
130 * @param options the plugin's options
131 */
132export const register: Register = (on, options) => {
133 const config = configOf(options)
134 const providers: readonly ProviderName[] = eligibleOf(config)
135
136 on('session.compact', async ($, e, next) => {
137 // A precompute installs nothing and its result is used only if the
138 // conversation is still the same when a real compaction comes. Deciding
139 // then would send the conversation to a third party, and pay for it,
140 // for a result that may be thrown away; the engine's own precompute is
141 // left to stand as the fallback this hook fails open to.
142 if (e.trigger === 'precompute') {
143 return next(e)
144 }
145
146 // Held outside the `try` so that the line written when it fails is
147 // redacted with whatever had been resolved by then.
148 let credentials: Credentials = {}
149 let result
150
151 try {
152 // Options that name no provider the person chose are acted on in no
153 // way: not even a credential is read for them.
154 if (config.refusal !== undefined) {
155 throw new Error(config.refusal)
156 }
157
158 // Credentials are read here, for this compaction, so a key set or
159 // rotated after the module loaded is the one that is used. What the
160 // host says when it refuses a read is its own text, so it is quoted
161 // short like a provider's.
162 const { environment, problem } = await readEnvironment(
163 $,
164 options,
165 providers,
166 ).catch(error => {
167 throw new Error(
168 `the environment could not be read: ${briefOf(messageOf(error), {})}`,
169 )
170 })
171
172 credentials = credentialsOf(options, environment)
173
174 if (config.unknownProviders !== undefined) {
175 log($, unknownLineOf(config.unknownProviders), credentials)
176 }
177
178 if (problem !== undefined) {
179 log(
180 $,
181 `the settings could not be read (${briefOf(problem, credentials)}); ` +
182 'credentials are taken from the options and the environment alone',
183 credentials,
184 )
185 }
186
187 result = await run({
188 messages: e.messages,
189 instructions: e.instructions,
190 config,
191 // Only the providers the person allows may be sent anything, so
192 // only their credentials go on. Every line shown is still redacted
193 // with all that was resolved: a key of another provider is no less
194 // a secret.
195 credentials: credentialsFor(providers, credentials),
196 ports: {
197 fetch: (url, init) => $.http.fetch(url, init),
198 pause: async (ms, signal) => {
199 if (next.budget.remainingMs < ms + PAUSE_RESERVE_MS) {
200 return false
201 }
202
203 await $.clock.sleep(ms, {
204 signal: AbortSignal.any([next.signal, signal]),
205 })
206
207 return true
208 },
209 after: (ms, fn) => $.clock.after(ms, fn),
210 signal: next.signal,
211 },
212 })
213 } catch (error) {
214 tell($, `built-in summary used instead: ${messageOf(error)}`, credentials)
215
216 return next(e)
217 }
218
219 if (result.routing !== undefined) {
220 log($, `provider decision: ${result.routing}`, credentials)
221 }
222
223 for (const line of decisionLinesOf(result.outcome.decisions)) {
224 log($, line, credentials)
225 }
226
227 const reduction = reductionOf(result.outcome)
228
229 if (reduction <= 0 || reduction < config.minReductionRatio) {
230 tell(
231 $,
232 'built-in summary used instead: the reduction is below the ' +
233 `${percentOf(config.minReductionRatio)} minimum (${summaryOf(result)})`,
234 credentials,
235 )
236
237 return next(e)
238 }
239
240 tell(
241 $,
242 `compacted without a summary: ${result.outcome.messages.length} of ` +
243 `${e.messages.length} messages remain (${summaryOf(result)})`,
244 credentials,
245 )
246
247 return { messages: result.outcome.messages }
248 })
249
250 const arm = armOf()
251 let isCompacting = false
252
253 on('turn.complete', async ($, e, next) => {
254 // A turn the person interrupted is not a moment to compact: they stopped
255 // the session to say something, and a compaction would hold it paused.
256 const isDue =
257 config.compactAtPercent > 0 &&
258 e.agentId === undefined &&
259 !e.isAborted &&
260 !isCompacting
261
262 if (isDue) {
263 isCompacting = true
264
265 try {
266 const { context } = await $.session.usage()
267
268 if (isCompactionDue(arm, context.percent, config.compactAtPercent)) {
269 await $.session.compact()
270
271 settle(
272 arm,
273 (await $.session.usage()).context.percent,
274 config.compactAtPercent,
275 )
276 }
277 } catch (error) {
278 // No credential is resolved on this path; the reason is the host's
279 // own text, so it is still quoted short and held to the rule that
280 // nothing after the word `Bearer` is shown.
281 log(
282 $,
283 `compaction at ${config.compactAtPercent}% of the context was ` +
284 `not run: ${briefOf(messageOf(error), {})}`,
285 {},
286 )
287 } finally {
288 isCompacting = false
289 }
290 }
291
292 return next(e)
293 })
294}
295hooks/compact.ts 845 lines1import type {
2 SessionMessage,
3 ToolResultSummary,
4 ToolUseSummary,
5} from 'claude-code'
6
7import type { Limits } from './providers'
8import { estimatedTokensOf, headOf } from './state'
9import type { Call, Fitted, Stage } from './state'
10import { bytesOf, noulOf } from './systemone'
11import type { Questions, Reply, Wire } from './systemone'
12
13/**
14 * Sends one batch of questions about a state and resolves with the reply.
15 */
16export type Ask = (state: unknown, questions: Questions) => Promise<Reply>
17
18/**
19 * What one request may hold, in estimated tokens: the person's budgets,
20 * held under what the provider documents, and the state's share of the
21 * request held to what a batch of questions leaves.
22 */
23export type Budget = {
24 stateTokens: number
25 requestTokens: number
26 questions?: number
27 /**
28 * For a provider that caps the request body in bytes: the bytes a body
29 * may take, already held to `LIMIT_SHARE` of the cap, and the bytes it
30 * takes besides its questions.
31 */
32 bytes?: { limit: number; state: number }
33}
34
35/**
36 * The two probabilities answered for one call.
37 */
38export type Scores = {
39 keepCall: number
40 keepResult: number
41}
42
43/**
44 * What happens to a call: `pinned` and `keep` leave it whole, `truncate`
45 * keeps the call and the head of its result, `drop` removes both.
46 */
47export type Action = 'pinned' | 'keep' | 'truncate' | 'drop'
48
49export type Decision = Scores & {
50 call: Call
51 action: Action
52}
53
54/**
55 * What the batches reported between them.
56 */
57type Reported = {
58 /**
59 * The largest `input_tokens` any batch reported. Every batch sends the
60 * same state, so the largest is the one to hold against the estimate.
61 */
62 inputTokens?: number
63 /**
64 * The costs the batches reported, summed.
65 */
66 cost?: number
67}
68
69/**
70 * A finished compaction: the conversation as it should read afterwards and
71 * the figures the report is written from.
72 */
73export type Outcome = {
74 messages: SessionMessage[]
75 decisions: Decision[]
76 charsBefore: number
77 charsAfter: number
78 stateTokens: number
79 /**
80 * The reduction the state needed; absent from an outcome that asked a
81 * provider nothing.
82 */
83 stage?: Stage
84 requests: number
85 reported: Reported
86 /**
87 * How many results were in fact cut to their head. It can be fewer than
88 * the `truncate` decisions: a result too short to gain from a cut is left
89 * whole, and its call then counts with the ones kept.
90 */
91 shortened: number
92}
93
94/**
95 * What the questions about a set of calls need of a request, whichever
96 * provider it goes to.
97 */
98export type Demand = {
99 /**
100 * The estimated size of the longest single question: what a limit stated
101 * for "the state plus the longest question" has to leave free.
102 */
103 longestQuestion: number
104 /**
105 * How many calls a request must at least be able to ask about.
106 */
107 batchCalls: number
108 /**
109 * The estimated size of the questions of that many calls, taking the
110 * calls whose questions are longest.
111 */
112 batchTokens: number
113 /**
114 * The UTF-8 bytes the questions of one request may take: those of the
115 * calls whose questions take the most bytes, `batchCalls` of them or as
116 * many as the provider's question cap lets one request hold, whichever
117 * is fewer. What a cap on the body in bytes has to leave free.
118 */
119 batchBytes: number
120}
121
122/**
123 * The questions asked about each call.
124 */
125const QUESTIONS_PER_CALL = 2
126
127/**
128 * What a request costs around its state and questions: the `model` field
129 * and the three key names, which are about the same in every wire format.
130 */
131const ENVELOPE_TOKENS = 32
132
133/**
134 * How many calls every request must have room to ask about, beside the
135 * state. Each request carries the whole state, so a state allowed to fill
136 * the request leaves room for one call's questions at a time, and the state
137 * is then paid for once per call instead of once per batch.
138 */
139const MIN_CALLS_PER_REQUEST = 16
140
141/**
142 * The most requests one compaction sends. The state rides in every one of
143 * them, so this bounds what a compaction can cost at this many times the
144 * state; a conversation that would need more is left to the built-in
145 * summary.
146 */
147const MAX_REQUESTS = 16
148
149/**
150 * How many requests are in flight at once. Waiting on a request is not
151 * charged to the hook, so there is no need to send every batch at the same
152 * moment, which is what a provider's rate limit is most likely to refuse.
153 */
154const CONCURRENT_REQUESTS = 3
155
156/**
157 * The share of a provider's documented limits a budget may reach. The token
158 * estimate is made without a tokenizer and its weights were fitted to one
159 * model, so a budget right at a limit could be over it by the provider's
160 * own count, and a provider may cut what is over without saying so.
161 */
162export const LIMIT_SHARE = 0.85
163
164/**
165 * A result is left whole unless cutting it saves more than the note that
166 * replaces its tail costs.
167 */
168const NOTE_ROOM = 120
169
170/**
171 * The scores a pinned call is reported with: it was not asked about, and it
172 * stays whole.
173 */
174const UNASKED: Scores = { keepCall: 1, keepResult: 1 }
175
176/**
177 * The id of the question whether a call should stay.
178 */
179function callQuestionId(call: Call): string {
180 return `call_${call.id}`
181}
182
183/**
184 * The id of the question whether a call's whole output should stay.
185 */
186function resultQuestionId(call: Call): string {
187 return `result_${call.id}`
188}
189
190/**
191 * What is asked about a single call, as two questions. They are separate
192 * because the answers differ in practice: that a file was read can matter
193 * long after its contents stopped mattering.
194 *
195 * The rubric they are read under is in the state's `context`, sent once.
196 *
197 * @returns its two `noul` questions, by id
198 */
199export function questionsOf(call: Call): Questions {
200 const named = `tool call ${call.id} (${call.tool})`
201 const outcome = call.isError ? 'an error, ' : ''
202
203 return {
204 [callQuestionId(call)]: {
205 type: 'noul',
206 instructions:
207 `Does the assistant still need to know that it made ${named}, and ` +
208 'with which input, to carry on with the goal?',
209 },
210 [resultQuestionId(call)]: {
211 type: 'noul',
212 instructions:
213 `Does the complete output of ${named}, ${outcome}` +
214 `${call.resultChars} characters long, have to stay word for word ` +
215 'because the assistant still relies on its content and running the ' +
216 'tool again would not bring it back?',
217 },
218 }
219}
220
221/**
222 * The estimated size of each of a call's questions and of the two together,
223 * as they ride in a request body written in the route's wire format, with
224 * what the provider counts for each question beyond its text; and the UTF-8
225 * bytes of the two, each with the comma after it.
226 */
227function questionTokensOf(
228 call: Call,
229 wire: Wire,
230): { both: number; longest: number; bytes: number } {
231 let both = 0
232 let longest = 0
233 let bytes = 0
234
235 for (const [id, question] of Object.entries(questionsOf(call))) {
236 const json = wire.questionJsonOf(id, question)
237 const tokens = estimatedTokensOf(json) + 1 + wire.tokensPerQuestion
238
239 both += tokens
240 longest = Math.max(longest, tokens)
241 bytes += bytesOf(json) + 1
242 }
243
244 return { both, longest, bytes }
245}
246
247/**
248 * The sum of the `count` largest of some sizes.
249 */
250function largestOf(sizes: readonly number[], count: number): number {
251 return [...sizes]
252 .sort((a, b) => b - a)
253 .slice(0, count)
254 .reduce((sum, size) => sum + size, 0)
255}
256
257/**
258 * Measures what the questions about the calls need of a request: the
259 * longest single question, and the questions of `MIN_CALLS_PER_REQUEST`
260 * calls, or of all of them when there are fewer. The calls with the longest
261 * questions are the ones counted, so the room set aside holds whichever
262 * calls end up in a batch together.
263 *
264 * @param calls the calls that will be asked about
265 * @param wire the wire format the questions are written in
266 * @param maxQuestions the provider's cap on the questions of one request,
267 * when it has one; it bounds only the bytes set aside
268 * @returns the demand, all zero when there are no calls
269 */
270export function demandOf(
271 calls: readonly Call[],
272 wire: Wire,
273 maxQuestions?: number,
274): Demand {
275 const sizes = calls.map(call => questionTokensOf(call, wire))
276 const batchCalls = Math.min(MIN_CALLS_PER_REQUEST, sizes.length)
277 const byteCalls =
278 maxQuestions === undefined
279 ? batchCalls
280 : Math.min(batchCalls, Math.floor(maxQuestions / QUESTIONS_PER_CALL))
281
282 return {
283 longestQuestion: sizes.reduce(
284 (longest, size) => Math.max(longest, size.longest),
285 0,
286 ),
287 batchCalls,
288 batchTokens: largestOf(
289 sizes.map(size => size.both),
290 batchCalls,
291 ),
292 batchBytes: largestOf(
293 sizes.map(size => size.bytes),
294 byteCalls,
295 ),
296 }
297}
298
299/**
300 * Holds the person's budgets against a provider's limits and against each
301 * other.
302 *
303 * A budget above `LIMIT_SHARE` of a documented limit is lowered to it rather
304 * than trusted. The state is further held to what the request leaves once
305 * the questions of a batch are set aside: a state that filled the request
306 * would be sent again for every single call.
307 *
308 * @param want the configured budgets
309 * @param limits the provider's documented limits
310 * @param demand what the questions need of a request
311 * @returns the budget one request is held to; throws when the request has
312 * no room for a state beside one batch of questions
313 */
314export function budgetOf(
315 want: { maxStateTokens: number; maxRequestTokens: number },
316 limits: Limits,
317 demand: Demand,
318): Budget {
319 const requestTokens = Math.min(
320 want.maxRequestTokens,
321 Math.floor(limits.maxRequestTokens * LIMIT_SHARE),
322 )
323 const beside = requestTokens - ENVELOPE_TOKENS - demand.batchTokens
324
325 if (beside < 1) {
326 throw new Error(
327 `a request of ${requestTokens} tokens has no room for a state beside ` +
328 `the questions of ${demand.batchCalls} ` +
329 `${demand.batchCalls === 1 ? 'call' : 'calls'}, which take about ` +
330 `${demand.batchTokens}`,
331 )
332 }
333
334 const budget: Budget = {
335 stateTokens: Math.min(
336 want.maxStateTokens,
337 Math.floor(limits.maxStateTokens * LIMIT_SHARE) - demand.longestQuestion,
338 beside,
339 ),
340 requestTokens,
341 }
342
343 if (limits.maxQuestions !== undefined) {
344 budget.questions = limits.maxQuestions
345 }
346
347 return budget
348}
349
350/**
351 * Splits the calls into batches, each of which fits one request beside the
352 * state under both the token budget and the provider's question count.
353 *
354 * The state is never split: every batch is asked against all of it, so the
355 * only thing that varies between requests is which calls they ask about.
356 *
357 * @param calls the calls to ask about, in order
358 * @param stateTokens the estimated size of the fitted state
359 * @param budget what one request may hold
360 * @param wire the wire format the questions are written in
361 * @returns the batches, in order; throws when the state leaves room for not
362 * even one call's questions, and when there would be more batches than
363 * `MAX_REQUESTS`
364 */
365export function batchesOf(
366 calls: readonly Call[],
367 stateTokens: number,
368 budget: Budget,
369 wire: Wire,
370): Call[][] {
371 const room = budget.requestTokens - stateTokens - ENVELOPE_TOKENS
372 const byteRoom =
373 budget.bytes === undefined
374 ? Infinity
375 : budget.bytes.limit - budget.bytes.state
376 const mostCalls =
377 budget.questions === undefined
378 ? Infinity
379 : Math.floor(budget.questions / QUESTIONS_PER_CALL)
380
381 if (mostCalls < 1) {
382 throw new Error(
383 `a request may hold ${budget.questions} questions, fewer than the ` +
384 `${QUESTIONS_PER_CALL} one call needs`,
385 )
386 }
387
388 const batches: Call[][] = []
389 let batch: Call[] = []
390 let used = 0
391 let usedBytes = 0
392
393 for (const call of calls) {
394 const { both, bytes } = questionTokensOf(call, wire)
395
396 if (both > room) {
397 throw new Error(
398 `no question fits beside the state: the state takes about ` +
399 `${stateTokens} of the ${budget.requestTokens} tokens a request ` +
400 'may hold',
401 )
402 }
403
404 if (bytes > byteRoom) {
405 throw new Error(
406 `no question fits beside the state: the request takes ` +
407 `${budget.bytes?.state} of the ${budget.bytes?.limit} bytes it ` +
408 'may hold before any question',
409 )
410 }
411
412 if (
413 batch.length >= mostCalls ||
414 used + both > room ||
415 usedBytes + bytes > byteRoom
416 ) {
417 batches.push(batch)
418 batch = []
419 used = 0
420 usedBytes = 0
421 }
422
423 batch.push(call)
424 used += both
425 usedBytes += bytes
426 }
427
428 if (batch.length > 0) {
429 batches.push(batch)
430 }
431
432 if (batches.length > MAX_REQUESTS) {
433 throw new Error(
434 `asking about ${calls.length} tool calls would take ` +
435 `${batches.length} requests, and one compaction sends at most ` +
436 `${MAX_REQUESTS}`,
437 )
438 }
439
440 return batches
441}
442
443/**
444 * Turns the two probabilities into what happens to the call. The result
445 * question is the stronger claim, so it is read first: a result worth
446 * keeping keeps its call with it.
447 *
448 * @param threshold the probability from which an item is kept
449 */
450export function decide(
451 call: Call,
452 scores: Scores,
453 threshold: number,
454): Decision {
455 if (call.isPinned) {
456 return { call, ...scores, action: 'pinned' }
457 }
458
459 if (scores.keepResult >= threshold) {
460 return { call, ...scores, action: 'keep' }
461 }
462
463 if (scores.keepCall >= threshold) {
464 return { call, ...scores, action: 'truncate' }
465 }
466
467 return { call, ...scores, action: 'drop' }
468}
469
470/**
471 * Says whether cutting a text of this length to its head removes more than
472 * the note that then stands for the rest adds.
473 */
474function isWorthCutting(chars: number, headChars: number): boolean {
475 return chars > headChars + NOTE_ROOM
476}
477
478/**
479 * A result text cut to its head, with one line saying what was done and by
480 * what, so the assistant reading it later knows the rest existed and how to
481 * get it back. A text too short for the cut to save anything is left as is.
482 */
483function cut(text: string, headChars: number, isError: boolean): string {
484 if (!isWorthCutting(text.length, headChars)) {
485 return text
486 }
487
488 const head = headOf(text, headChars)
489 const kind = isError ? 'error output' : 'output'
490
491 return (
492 `${head}${head === '' ? '' : '\n'}[decision-compaction removed ` +
493 `${text.length - head.length} more characters of this tool ${kind}; ` +
494 'run the tool again if they are needed]'
495 )
496}
497
498/**
499 * A result with its text cut, or the same object when nothing was cut. The
500 * stored record is left off a cut result: it holds the full output.
501 */
502function cutResult(
503 result: ToolResultSummary,
504 headChars: number,
505): ToolResultSummary {
506 const { tool_use_id, isError } = result
507 const text = cut(result.text, headChars, isError)
508
509 if (text === result.text) {
510 return result
511 }
512
513 return { tool_use_id, text, isError }
514}
515
516/**
517 * A tool use whose mirrored outcome is cut like its result, or the same
518 * object when nothing was cut. Only used inside a message that is rebuilt
519 * anyway, so the copy and the result it mirrors agree.
520 */
521function cutUse(use: ToolUseSummary, headChars: number): ToolUseSummary {
522 if (use.text === undefined) {
523 return use
524 }
525
526 const text = cut(use.text, headChars, use.isError === true)
527
528 if (text === use.text) {
529 return use
530 }
531
532 const { result: _record, ...kept } = use
533
534 return { ...kept, text }
535}
536
537/**
538 * Carries the decisions out on the conversation.
539 *
540 * A message nothing touched is returned as the very object it came in as, so
541 * it still carries the engine's `handle` and stands whole. A changed message
542 * is a new object with no `handle`, which the engine builds from its role,
543 * text and tool blocks; because that loses whatever else the original held,
544 * a message is rebuilt only when one of its own blocks changed. A dropped
545 * call takes its result with it, so no result is left without its call, and
546 * a message left with no text and no block is removed.
547 *
548 * @param messages the conversation
549 * @param decisions what was decided for each paired call
550 * @param headChars how much of a truncated result stays
551 * @returns the conversation afterwards, oldest first
552 */
553export function rebuild(
554 messages: readonly SessionMessage[],
555 decisions: readonly Decision[],
556 headChars: number,
557): SessionMessage[] {
558 const fate = new Map<string, 'truncate' | 'drop'>()
559
560 for (const { call, action } of decisions) {
561 if (action === 'truncate' || action === 'drop') {
562 fate.set(call.toolUseId, action)
563 }
564 }
565
566 const rebuilt: SessionMessage[] = []
567
568 for (const message of messages) {
569 const given = message.toolResults ?? []
570 const uses = message.toolUses.filter(
571 use => fate.get(use.tool_use_id) !== 'drop',
572 )
573 const results = given
574 .filter(result => fate.get(result.tool_use_id) !== 'drop')
575 .map(result =>
576 fate.get(result.tool_use_id) === 'truncate'
577 ? cutResult(result, headChars)
578 : result,
579 )
580 const isUntouched =
581 uses.length === message.toolUses.length &&
582 results.length === given.length &&
583 results.every((result, index) => result === given[index])
584
585 if (isUntouched) {
586 rebuilt.push(message)
587 continue
588 }
589
590 const isEmptied =
591 uses.length + results.length === 0 && message.text.trim() === ''
592
593 if (!isEmptied) {
594 rebuilt.push({
595 role: message.role,
596 text: message.text,
597 toolUses: uses.map(use =>
598 fate.get(use.tool_use_id) === 'truncate'
599 ? cutUse(use, headChars)
600 : use,
601 ),
602 ...(results.length > 0 ? { toolResults: results } : {}),
603 })
604 }
605 }
606
607 return rebuilt
608}
609
610/**
611 * The size of a tool input as the model reads it. An input that cannot be
612 * serialised has no size to compare, and counts as nothing.
613 */
614function inputCharsOf(use: ToolUseSummary): number {
615 try {
616 return (JSON.stringify(use.input) ?? '').length
617 } catch {
618 return 0
619 }
620}
621
622/**
623 * The characters a message holds for the model: its text, its tool inputs
624 * and its tool results. The outcome mirrored on a tool use is not counted,
625 * since the result block already is.
626 */
627function charsOf(message: SessionMessage): number {
628 const inputs = message.toolUses.reduce(
629 (sum, use) => sum + inputCharsOf(use),
630 0,
631 )
632 const results = (message.toolResults ?? []).reduce(
633 (sum, result) => sum + result.text.length,
634 0,
635 )
636
637 return message.text.length + inputs + results
638}
639
640function totalChars(messages: readonly SessionMessage[]): number {
641 return messages.reduce((chars, message) => chars + charsOf(message), 0)
642}
643
644/**
645 * The share of the conversation's characters a compaction removed.
646 *
647 * @returns a ratio from 0 to 1; 0 for an empty conversation
648 */
649export function reductionOf(
650 outcome: Pick<Outcome, 'charsBefore' | 'charsAfter'>,
651): number {
652 return outcome.charsBefore === 0
653 ? 0
654 : (outcome.charsBefore - outcome.charsAfter) / outcome.charsBefore
655}
656
657/**
658 * One batch as it came back: the reply and, for every call of the batch,
659 * the two probabilities read from it.
660 */
661type Judged = {
662 reply: Reply
663 scores: [Call, Scores][]
664}
665
666/**
667 * Puts one batch to the provider and reads both answers of every call in
668 * it. A reply that lacks one of them fails the batch: a call nobody answered
669 * for must not be decided by a default.
670 */
671async function judge(
672 batch: readonly Call[],
673 state: unknown,
674 ask: Ask,
675): Promise<Judged> {
676 const reply = await ask(
677 state,
678 Object.fromEntries(
679 batch.flatMap(call => Object.entries(questionsOf(call))),
680 ),
681 )
682
683 return {
684 reply,
685 scores: batch.map(call => [
686 call,
687 {
688 keepCall: noulOf(reply, callQuestionId(call)),
689 keepResult: noulOf(reply, resultQuestionId(call)),
690 },
691 ]),
692 }
693}
694
695/**
696 * The outcome when there is nothing to ask about: the conversation as it is.
697 *
698 * @param messages the conversation
699 * @param calls its paired calls, all of them pinned
700 * @returns an outcome with no request made
701 */
702export function untouched(
703 messages: readonly SessionMessage[],
704 calls: readonly Call[],
705): Outcome {
706 const chars = totalChars(messages)
707
708 return {
709 messages: [...messages],
710 decisions: calls.map(call => decide(call, UNASKED, 0)),
711 charsBefore: chars,
712 charsAfter: chars,
713 stateTokens: 0,
714 requests: 0,
715 reported: {},
716 shortened: 0,
717 }
718}
719
720/**
721 * Runs `work` over the items with at most `width` of them under way at a
722 * time, and resolves with the results in the items' order, whichever order
723 * they came back in. The first failure rejects the whole and starts nothing
724 * further; `onFailure` hears of it at once, so that the caller can stop what
725 * is still under way.
726 */
727async function inTurns<Item, Result>(
728 items: readonly Item[],
729 width: number,
730 work: (item: Item) => Promise<Result>,
731 onFailure: (reason: unknown) => void,
732): Promise<Result[]> {
733 const results: Result[] = []
734 let next = 0
735 let hasFailed = false
736
737 const worker = async () => {
738 while (!hasFailed && next < items.length) {
739 const index = next++
740
741 try {
742 results[index] = await work(items[index] as Item)
743 } catch (error) {
744 hasFailed = true
745 onFailure(error)
746
747 throw error
748 }
749 }
750 }
751
752 await Promise.all(
753 Array.from({ length: Math.min(width, items.length) }, worker),
754 )
755
756 return results
757}
758
759/**
760 * Asks about every candidate call and applies the answers to the
761 * conversation. The batches go out `CONCURRENT_REQUESTS` at a time, each
762 * with the whole state, and each call is decided by the answers of its own
763 * batch, in whatever order the batches come back. One batch failing fails
764 * the compaction, since half the answers decide nothing; no batch is sent
765 * after it, and `halt` is told at once so the others can be abandoned.
766 *
767 * @param messages the conversation
768 * @param calls its paired calls, pinned ones included
769 * @param fitted the state, already fitted to the budget, with its size as
770 * the route's wire format carries it
771 * @param ask how a batch is sent
772 * @param settings the budget, the route's wire format, the keep threshold,
773 * the truncation length, and `halt`, which hears the reason the moment a
774 * batch fails
775 */
776export async function compact(
777 messages: readonly SessionMessage[],
778 calls: readonly Call[],
779 fitted: Fitted,
780 ask: Ask,
781 settings: {
782 budget: Budget
783 wire: Wire
784 keepThreshold: number
785 truncateHeadChars: number
786 halt?: (reason: Error) => void
787 },
788): Promise<Outcome> {
789 const batches = batchesOf(
790 calls.filter(call => !call.isPinned),
791 fitted.tokens,
792 settings.budget,
793 settings.wire,
794 )
795 const answered = await inTurns(
796 batches,
797 CONCURRENT_REQUESTS,
798 batch => judge(batch, fitted.state, ask),
799 reason =>
800 settings.halt?.(
801 reason instanceof Error ? reason : new Error(String(reason)),
802 ),
803 )
804 const scores = new Map<Call, Scores>()
805 const reported: Reported = {}
806
807 for (const { scores: judged, reply } of answered) {
808 for (const [call, value] of judged) {
809 scores.set(call, value)
810 }
811
812 if (reply.usage.input_tokens !== undefined) {
813 reported.inputTokens = Math.max(
814 reported.inputTokens ?? 0,
815 reply.usage.input_tokens,
816 )
817 }
818
819 if (reply.usage.cost !== undefined) {
820 reported.cost = (reported.cost ?? 0) + reply.usage.cost
821 }
822 }
823
824 const decisions = calls.map(call =>
825 decide(call, scores.get(call) ?? UNASKED, settings.keepThreshold),
826 )
827 const kept = rebuild(messages, decisions, settings.truncateHeadChars)
828
829 return {
830 messages: kept,
831 decisions,
832 charsBefore: totalChars(messages),
833 charsAfter: totalChars(kept),
834 stateTokens: fitted.tokens,
835 stage: fitted.stage,
836 requests: batches.length,
837 reported,
838 shortened: decisions.filter(
839 ({ call, action }) =>
840 action === 'truncate' &&
841 isWorthCutting(call.resultChars, settings.truncateHeadChars),
842 ).length,
843 }
844}
845hooks/config.ts 212 lines1import type { PluginOptions } from 'claude-code'
2
3import { cutOf, providerOf, PROVIDERS } from './providers'
4import type { ProviderName } from './providers'
5
6/**
7 * The plugin's options as the hooks use them: every value present, typed
8 * and inside the range it is meaningful in.
9 */
10export type Config = {
11 provider: ProviderName
12 /**
13 * Why no compaction may be decided with these options, when there is such
14 * a reason. It is set for an option that says where the conversation is
15 * sent and cannot be read: a destination is never guessed.
16 */
17 refusal?: string
18 /**
19 * The model name for the configured provider; absent for its default.
20 */
21 model?: string
22 providerDecision: boolean
23 /**
24 * The providers the provider decision may pick besides the configured
25 * one, in `PROVIDERS` order.
26 */
27 decisionProviders: readonly ProviderName[]
28 /**
29 * The names in `decisionProviders` that are no provider, as given; absent
30 * when there are none. They are dropped, not refused: leaving a name out
31 * can only narrow where the conversation is sent.
32 */
33 unknownProviders?: string[]
34 keepThreshold: number
35 preserveRecentMessages: number
36 compactAtPercent: number
37 minReductionRatio: number
38 maxStateTokens: number
39 maxRequestTokens: number
40 truncateHeadChars: number
41}
42
43/**
44 * What an option reads as when it is unset or unusable; the manifest's
45 * `userConfig` states the same values to the person.
46 */
47const DEFAULTS = {
48 provider: 'typesafe',
49 providerDecision: false,
50 decisionProviders: PROVIDERS,
51 keepThreshold: 0.5,
52 preserveRecentMessages: 6,
53 compactAtPercent: 60,
54 minReductionRatio: 0.25,
55 maxStateTokens: 25_000,
56 maxRequestTokens: 30_000,
57 truncateHeadChars: 300,
58} as const satisfies Config
59
60/**
61 * A numeric option held to a range. A value that is not a finite number
62 * reads as the default; one outside the range is moved to its nearest end.
63 */
64function within(
65 value: unknown,
66 fallback: number,
67 low: number,
68 high: number,
69): number {
70 const number =
71 typeof value === 'number' && Number.isFinite(value) ? value : fallback
72
73 return Math.min(high, Math.max(low, number))
74}
75
76/**
77 * Reads the `decisionProviders` option: a list, or the names in one text
78 * separated by commas. Unset, it is every provider. The host fills in the
79 * manifest's default for a field that was never set, so an empty text is a
80 * field the person cleared, and like an empty list it names no provider:
81 * only the configured one is then eligible. Names are matched as `provider`
82 * is, exactly; a value of any other type names no provider.
83 */
84function decisionProvidersOf(value: unknown): {
85 named: readonly ProviderName[]
86 unknown: string[]
87} {
88 if (value === undefined || value === null) {
89 return { named: DEFAULTS.decisionProviders, unknown: [] }
90 }
91
92 const given = Array.isArray(value)
93 ? value
94 : typeof value === 'string'
95 ? value.split(',')
96 : [value]
97 const names = given
98 .map(name =>
99 typeof name === 'string' ? name.trim() : String(JSON.stringify(name)),
100 )
101 .filter(name => name !== '')
102
103 return {
104 named: PROVIDERS.filter(provider => names.includes(provider)),
105 unknown: [...new Set(names.filter(name => providerOf(name) === undefined))],
106 }
107}
108
109/**
110 * The providers a compaction may send anything to: the configured one, and
111 * with the provider decision on also those `decisionProviders` names, in
112 * `PROVIDERS` order. Everything that reads a credential or offers a route
113 * goes by this one list.
114 *
115 * @param config the configuration
116 * @returns the providers
117 */
118export function eligibleOf(config: Config): ProviderName[] {
119 return PROVIDERS.filter(
120 provider =>
121 provider === config.provider ||
122 (config.providerDecision && config.decisionProviders.includes(provider)),
123 )
124}
125
126/**
127 * Says whether an option was left unset: absent, or text with nothing in
128 * it, which is what a cleared field holds.
129 */
130function isUnset(value: unknown): boolean {
131 return (
132 value === undefined ||
133 value === null ||
134 (typeof value === 'string' && value.trim() === '')
135 )
136}
137
138/**
139 * Reads the plugin's options. Nothing here throws, because this runs while
140 * the module loads and a throw would take both hooks with it: an option the
141 * person mistyped reads as its default.
142 *
143 * The one exception is `provider`, which decides who is sent the
144 * conversation. Only an unset `provider` reads as the default. A value that
145 * names none of the providers is not replaced by one the person did not
146 * choose: it is reported in `refusal`, and every compaction is then left to
147 * the built-in summary.
148 *
149 * @param options what `register` received
150 * @returns the configuration
151 */
152export function configOf(options: PluginOptions): Config {
153 const named = providerOf(options.provider)
154 const deciding = decisionProvidersOf(options.decisionProviders)
155 const config: Config = {
156 provider: named ?? DEFAULTS.provider,
157 providerDecision: options.providerDecision === true,
158 decisionProviders: deciding.named,
159 keepThreshold: within(options.keepThreshold, DEFAULTS.keepThreshold, 0, 1),
160 preserveRecentMessages: Math.floor(
161 within(
162 options.preserveRecentMessages,
163 DEFAULTS.preserveRecentMessages,
164 0,
165 Infinity,
166 ),
167 ),
168 compactAtPercent: within(
169 options.compactAtPercent,
170 DEFAULTS.compactAtPercent,
171 0,
172 100,
173 ),
174 minReductionRatio: within(
175 options.minReductionRatio,
176 DEFAULTS.minReductionRatio,
177 0,
178 1,
179 ),
180 maxStateTokens: Math.floor(
181 within(options.maxStateTokens, DEFAULTS.maxStateTokens, 1, Infinity),
182 ),
183 maxRequestTokens: Math.floor(
184 within(options.maxRequestTokens, DEFAULTS.maxRequestTokens, 1, Infinity),
185 ),
186 truncateHeadChars: Math.floor(
187 within(
188 options.truncateHeadChars,
189 DEFAULTS.truncateHeadChars,
190 0,
191 Infinity,
192 ),
193 ),
194 }
195
196 if (named === undefined && !isUnset(options.provider)) {
197 config.refusal =
198 `the provider option is ${cutOf(JSON.stringify(options.provider))}, ` +
199 `which is none of ${PROVIDERS.join(', ')}`
200 }
201
202 if (deciding.unknown.length > 0) {
203 config.unknownProviders = deciding.unknown
204 }
205
206 if (typeof options.model === 'string' && options.model.trim() !== '') {
207 config.model = options.model.trim()
208 }
209
210 return config
211}
212hooks/providers.ts 1127 lines1import type { HttpResponse, PluginOptions } from 'claude-code'
2
3import { decodeOpenAI, OPENAI } from './openai'
4import { headOf } from './state'
5import { BARE, bodyOf, bytesOf, idsOf, isRecord, replyOf } from './systemone'
6import type { Question, Questions, Reply, Wire } from './systemone'
7
8/**
9 * The providers that serve the System One protocol, in the order the
10 * manifest lists them. Of two gateways to the same model the one listed
11 * first is the one the provider decision offers.
12 */
13export const PROVIDERS = [
14 'typesafe',
15 'cloudflare',
16 'openrouter',
17 'codiv',
18 'perplexity',
19 'decisions-api-dev',
20 'decisionapi-net',
21 'openai',
22] as const
23
24export type ProviderName = (typeof PROVIDERS)[number]
25
26/**
27 * One way to reach one model: the provider that is called and the model name
28 * as that provider's request body spells it.
29 */
30export type Route = {
31 provider: ProviderName
32 model: string
33}
34
35/**
36 * What a provider's documentation says a request may hold, in tokens as the
37 * provider counts them.
38 *
39 * `maxStateTokens` bounds the state together with the longest single
40 * question, which is how TypeSafe states its limit; `maxQuestions` is absent
41 * where no cap is documented. `maxRequestBytes` is a cap on the UTF-8 bytes
42 * of the whole request body, for a provider that states its limit in bytes.
43 */
44export type Limits = {
45 maxStateTokens: number
46 maxRequestTokens: number
47 maxQuestions?: number
48 maxRequestBytes?: number
49}
50
51/**
52 * The secrets and the account id the providers need, each present only when
53 * one was configured.
54 */
55export type Credentials = {
56 typesafeApiKey?: string
57 openrouterApiKey?: string
58 cloudflareApiToken?: string
59 cloudflareAccountId?: string
60 codivApiKey?: string
61 perplexityApiKey?: string
62 decisionsApiKey?: string
63 decisionapiApiKey?: string
64 openaiApiKey?: string
65}
66
67/**
68 * The environment variable each credential falls back to when its plugin
69 * option is unset.
70 */
71const ENV_NAMES = {
72 typesafeApiKey: 'TYPESAFE_API_KEY',
73 openrouterApiKey: 'OPENROUTER_API_KEY',
74 cloudflareApiToken: 'CLOUDFLARE_API_TOKEN',
75 cloudflareAccountId: 'CLOUDFLARE_ACCOUNT_ID',
76 codivApiKey: 'CODIV_API_KEY',
77 perplexityApiKey: 'PERPLEXITY_API_KEY',
78 decisionsApiKey: 'DECISIONS_API_KEY',
79 decisionapiApiKey: 'DECISIONAPI_API_KEY',
80 openaiApiKey: 'OPENAI_API_KEY',
81} as const satisfies Record<keyof Credentials, string>
82
83type EnvName = (typeof ENV_NAMES)[keyof Credentials]
84
85/**
86 * The values of the credential variables as one source holds them.
87 */
88export type Environment = Partial<Record<EnvName, string | undefined>>
89
90/**
91 * One request as `$.http.fetch` takes it.
92 */
93type Exchange = {
94 url: string
95 init: { method: 'POST'; headers: Record<string, string>; body: string }
96}
97
98/**
99 * Everything that differs between providers. The call path and the answer
100 * validation are shared, so a provider is this record and nothing else.
101 */
102type Descriptor = {
103 /**
104 * The provider's name as a sentence about it spells it.
105 */
106 title: string
107 defaultModel: string
108 /**
109 * The models the provider is offered with in the routing decision.
110 */
111 models: readonly string[]
112 /**
113 * The model family a model of this provider is, which is what two routes
114 * are compared by: the same Jev is served directly and through gateways.
115 */
116 family: (model: string) => string
117 /**
118 * True for a provider that forwards to a model another provider hosts.
119 */
120 isGateway: boolean
121 limits: Limits
122 needs: readonly (keyof Credentials)[]
123 bodyModelOf: (model: string) => string
124 urlOf: (model: string, credentials: Credentials) => string
125 tokenOf: (credentials: Credentials) => string | undefined
126 wire: Wire
127 /**
128 * Takes the response out of whatever the provider wraps it in and into
129 * the System One shape the readers of a reply take, given the questions
130 * the request asked.
131 */
132 unwrap: (
133 payload: unknown,
134 credentials: Credentials,
135 asked: Questions,
136 ) => unknown
137}
138
139const CLOUDFLARE_MODELS = ['clef', 'clef-flash'] as const
140
141/**
142 * What a request to a provider that documents no limit is held to: the
143 * tightest window and question cap among the documented ones (OpenRouter's
144 * 32,000 tokens for Jev, Clef's 64 questions).
145 */
146const UNDOCUMENTED: Limits = {
147 maxStateTokens: 32_000,
148 maxRequestTokens: 32_000,
149 maxQuestions: 64,
150}
151
152/**
153 * How many questions one request to either Decisions reseller may hold.
154 */
155const RESELLER_QUESTIONS = 8
156
157/**
158 * The two Decisions resellers state their cap in bytes: the text and the
159 * questions of a request within 32 KiB, and at most eight questions. Both
160 * serve TypeSafe's Jev 1.13, so the token figures are the window TypeSafe
161 * documents for it. They do not bind: 32 KiB of English is far fewer
162 * tokens, and the state is fitted to the byte cap as well. They stand
163 * against what the provider counts, so a reply that counts the whole window
164 * is still read as a state that was cut short.
165 */
166const RESELLER: Limits = {
167 maxStateTokens: 32_000,
168 maxRequestTokens: 64_000,
169 maxQuestions: RESELLER_QUESTIONS,
170 maxRequestBytes: 32_768,
171}
172
173/**
174 * The question ids decisions-api.dev accepts: a letter first, then letters,
175 * digits, `_` or `-`, 64 characters at most. The mod's ids may hold `.` and
176 * run to 100 characters, so the request names each question by its place
177 * instead.
178 */
179const ALIAS_OF = (index: number) => `q${index}`
180
181/**
182 * The longest alias a request can carry: the one of the last place the
183 * question cap allows. A question is measured under it, whatever its own
184 * id, since the alias is what is sent.
185 */
186const LONGEST_ALIAS = ALIAS_OF(RESELLER_QUESTIONS - 1)
187
188/**
189 * The System One body with every question id replaced by its place in the
190 * request, `q0`, `q1`, …. The ids are checked against the mod's own rule
191 * first, as for every provider; option names inside a `choice` are left as
192 * they are, since only question ids are restricted.
193 */
194const ALIASED: Wire = {
195 encode: (model, state, questions) =>
196 bodyOf(
197 model,
198 state,
199 Object.fromEntries(
200 idsOf(questions).map((id, index) => [
201 ALIAS_OF(index),
202 questions[id] as Question,
203 ]),
204 ),
205 ),
206 questionJsonOf: (_id, question) =>
207 BARE.questionJsonOf(LONGEST_ALIAS, question),
208 stateJsonOf: BARE.stateJsonOf,
209 tokensPerQuestion: 0,
210}
211
212/**
213 * A Jev id as TypeSafe names it: `jev`, `jev-latest`, `jev-1.13`.
214 */
215const JEV = /^jev(?:-|$)/
216
217/**
218 * The ids OpenRouter routes to Jev: a Jev id (`jev`, `jev-latest`,
219 * `jev-1.13`) either bare, which OpenRouter maps onto TypeSafe's namespace
220 * before routing, or under that namespace (`typesafe/`) or its alias
221 * (`~typesafe/`). The namespace alone does not make a model Jev: any other
222 * model TypeSafe publishes there is a model of its own.
223 */
224const OPENROUTER_JEV = /^(?:~?typesafe\/)?jev(?:-|$)/
225
226/**
227 * The Workers AI catalogue prefix; a Cloudflare model is `clef` in the body
228 * and this prefix plus `clef` in the URL.
229 */
230const CLOUDFLARE_PREFIX = '@cf/cloudflare/'
231
232/**
233 * How much is quoted of a text this mod did not write (an error body, a
234 * failure envelope, the host's reason for refusing a call): enough to read
235 * the reason, short enough for one toast line.
236 */
237export const REASON_CHARS = 160
238
239/**
240 * The shortest value treated as a secret when a text is redacted. A shorter
241 * one is not a key any of the providers issues, and replacing it wherever
242 * its few characters occur would garble the ordinary words around it.
243 */
244const MIN_SECRET_CHARS = 8
245
246/**
247 * What stands where a credential was. It holds no part of what it replaces,
248 * not even the word that announced it.
249 */
250const REDACTED = '[redacted]'
251
252/**
253 * An `Authorization` value as a response may echo it: the scheme and the
254 * token after it, whoever's token that is. The two are separated by white
255 * space in a header, and by a colon or an equals sign where the header was
256 * written out as a field (`Bearer: ...`, `bearer=...`).
257 */
258const BEARER = /bearer(?:\s*[:=]\s*|\s+)[^\s"'<>]+/gi
259
260function textOf(value: unknown): string | undefined {
261 return typeof value === 'string' && value.trim() !== ''
262 ? value.trim()
263 : undefined
264}
265
266/**
267 * Accepts a Cloudflare model by its short name or its catalogue id, and
268 * answers the short name. The body's `model` must be `clef` or `clef-flash`,
269 * so any other name is refused here rather than by a 4xx after the state was
270 * sent.
271 */
272function cloudflareModelOf(model: string): string {
273 const named = model.trim()
274 const short = named.startsWith(CLOUDFLARE_PREFIX)
275 ? named.slice(CLOUDFLARE_PREFIX.length)
276 : named
277
278 if (!CLOUDFLARE_MODELS.some(known => known === short)) {
279 throw new Error(
280 `cloudflare serves ${CLOUDFLARE_MODELS.join(' and ')}, not ` +
281 JSON.stringify(named),
282 )
283 }
284
285 return short
286}
287
288/**
289 * The first message of a Workers AI envelope's `errors` list.
290 */
291function envelopeErrorOf(payload: Record<string, unknown>): string | undefined {
292 const first = Array.isArray(payload.errors) ? payload.errors[0] : undefined
293
294 return isRecord(first) ? textOf(first.message) : undefined
295}
296
297/**
298 * The error for a response envelope that reports failure, quoting the
299 * reason it gives.
300 */
301function envelopeFailure(
302 reason: string | undefined,
303 credentials: Credentials,
304): Error {
305 return new Error(
306 'the response envelope reports failure: ' +
307 briefOf(reason ?? '', credentials),
308 )
309}
310
311/**
312 * Takes a Workers AI response out of its `{ result, success, errors }`
313 * envelope. A body that is already a bare reply is passed through, because
314 * the envelope is documented for the REST API in general and not for Clef.
315 * The reason a failed envelope gives is quoted like any other text of a
316 * provider's: redacted and held to one short line.
317 */
318function unwrapCloudflare(payload: unknown, credentials: Credentials): unknown {
319 if (!isRecord(payload)) {
320 return payload
321 }
322
323 if (payload.success === false) {
324 throw envelopeFailure(envelopeErrorOf(payload), credentials)
325 }
326
327 return isRecord(payload.result) ? payload.result : payload
328}
329
330/**
331 * Takes a response out of the `{ code, message, data: { result } }` envelope
332 * the two Decisions resellers answer with. A `code` other than 0, or no
333 * `result`, is a failure, and the envelope's `message` is quoted as its
334 * reason like any other text of a provider's: redacted and held to one
335 * short line.
336 */
337function unwrapDataEnvelope(
338 payload: unknown,
339 credentials: Credentials,
340): Record<string, unknown> {
341 const data = isRecord(payload) ? payload.data : undefined
342 const result = isRecord(data) ? data.result : undefined
343
344 if (!isRecord(payload) || payload.code !== 0 || !isRecord(result)) {
345 throw envelopeFailure(
346 isRecord(payload) ? textOf(payload.message) : undefined,
347 credentials,
348 )
349 }
350
351 return result
352}
353
354/**
355 * Reads a decisions-api.dev response: out of its envelope, and with every
356 * answer back under the id of the question asked in that place. An answer
357 * under a name that was not sent is refused: it answers nothing that was
358 * asked. Both maps have no prototype, so an id or a name such as
359 * `__proto__` is an ordinary key.
360 */
361function unwrapAliased(
362 payload: unknown,
363 credentials: Credentials,
364 asked: Questions,
365): unknown {
366 const result = unwrapDataEnvelope(payload, credentials)
367
368 if (!isRecord(result.answers)) {
369 return result
370 }
371
372 const idOf: Record<string, string> = Object.create(null)
373 const answers: Record<string, unknown> = Object.create(null)
374
375 Object.keys(asked).forEach((id, index) => {
376 idOf[ALIAS_OF(index)] = id
377 })
378
379 for (const [alias, answer] of Object.entries(result.answers)) {
380 const id = Object.hasOwn(idOf, alias) ? idOf[alias] : undefined
381
382 if (id === undefined) {
383 throw new Error(
384 `answered ${briefOf(alias, credentials)}, which was not asked`,
385 )
386 }
387
388 answers[id] = answer
389 }
390
391 return { ...result, answers }
392}
393
394/**
395 * A model name as configured, without the white space around it: how most
396 * providers' bodies spell it.
397 */
398const TRIMMED = (model: string) => model.trim()
399
400/**
401 * A response that already is a bare System One reply.
402 */
403const AS_SENT = (payload: unknown) => payload
404
405/**
406 * The family of a model at a reseller of TypeSafe's Jev: Jev for a Jev id,
407 * the model itself otherwise.
408 */
409const JEV_FAMILY = (model: string) => (JEV.test(model) ? 'jev' : model)
410
411const TABLE: Record<ProviderName, Descriptor> = {
412 typesafe: {
413 title: 'TypeSafe',
414 defaultModel: 'jev-latest',
415 models: ['jev-latest'],
416 family: () => 'jev',
417 isGateway: false,
418 limits: { maxStateTokens: 32_000, maxRequestTokens: 64_000 },
419 needs: ['typesafeApiKey'],
420 bodyModelOf: TRIMMED,
421 urlOf: () => 'https://api.typesafe.ai/v1/systemone',
422 tokenOf: credentials => credentials.typesafeApiKey,
423 wire: BARE,
424 unwrap: AS_SENT,
425 },
426 cloudflare: {
427 title: 'Cloudflare',
428 defaultModel: 'clef',
429 models: CLOUDFLARE_MODELS,
430 family: model => model,
431 isGateway: false,
432 limits: {
433 maxStateTokens: 65_536,
434 maxRequestTokens: 65_536,
435 maxQuestions: 64,
436 },
437 needs: ['cloudflareApiToken', 'cloudflareAccountId'],
438 bodyModelOf: cloudflareModelOf,
439 urlOf: (model, credentials) =>
440 'https://api.cloudflare.com/client/v4/accounts/' +
441 `${encodeURIComponent(credentials.cloudflareAccountId ?? '')}/ai/run/` +
442 `${CLOUDFLARE_PREFIX}${model}`,
443 tokenOf: credentials => credentials.cloudflareApiToken,
444 wire: BARE,
445 unwrap: unwrapCloudflare,
446 },
447 // OpenRouter documents Jev's window as 32,000 tokens for the state and the
448 // questions together, which is tighter than TypeSafe's own 64,000 a request.
449 // Its alias for the newest Jev is `~typesafe/jev-latest`. An id that has an
450 // author prefix is sent on as written, so the same name without the `~` is
451 // an id its documentation does not list.
452 openrouter: {
453 title: 'OpenRouter',
454 defaultModel: '~typesafe/jev-latest',
455 models: ['~typesafe/jev-latest'],
456 family: model => (OPENROUTER_JEV.test(model) ? 'jev' : model),
457 isGateway: true,
458 limits: { maxStateTokens: 32_000, maxRequestTokens: 32_000 },
459 needs: ['openrouterApiKey'],
460 bodyModelOf: TRIMMED,
461 urlOf: () => 'https://openrouter.ai/api/v1/systemone',
462 tokenOf: credentials => credentials.openrouterApiKey,
463 wire: BARE,
464 unwrap: AS_SENT,
465 },
466 // Codiv serves its own open Jev, which answers as `openjev-0.1` and is a
467 // model of its own, not TypeSafe's Jev; at Codiv `jev-latest` is an alias
468 // of it. Its window is 65,536 tokens for the state and the questions
469 // together, and it advises about 60,000 for the state. It documents no
470 // fixed question cap and answered 129 questions in one request.
471 codiv: {
472 title: 'Codiv',
473 defaultModel: 'openjev-latest',
474 models: ['openjev-latest'],
475 family: model => (/^(?:open)?jev(?:-|$)/.test(model) ? 'openjev' : model),
476 isGateway: false,
477 limits: { maxStateTokens: 60_000, maxRequestTokens: 65_536 },
478 needs: ['codivApiKey'],
479 bodyModelOf: TRIMMED,
480 urlOf: () => 'https://api.codiv.ai/v1/systemone',
481 tokenOf: credentials => credentials.codivApiKey,
482 wire: BARE,
483 unwrap: AS_SENT,
484 },
485 // Perplexity refuses any model but its own deciders (HTTP 400, "Invalid
486 // model"), Jev included. A request must stay under 262,144 input tokens,
487 // the state and every question counted, and holds 1 to 128 questions.
488 perplexity: {
489 title: 'Perplexity',
490 defaultModel: 'pplx-decider-v1.1-27b',
491 models: ['pplx-decider-v1.1-27b'],
492 family: model =>
493 /^pplx-decider(?:-|$)/.test(model) ? 'pplx-decider' : model,
494 isGateway: false,
495 limits: {
496 maxStateTokens: 262_143,
497 maxRequestTokens: 262_143,
498 maxQuestions: 128,
499 },
500 needs: ['perplexityApiKey'],
501 bodyModelOf: TRIMMED,
502 urlOf: () => 'https://api.perplexity.ai/v1/decisions',
503 tokenOf: credentials => credentials.perplexityApiKey,
504 wire: BARE,
505 unwrap: AS_SENT,
506 },
507 // The two Decisions resellers forward to TypeSafe's Jev and wrap the
508 // reply in an envelope of their own. decisions-api.dev alone restricts
509 // question ids, so its requests name questions by their place.
510 'decisions-api-dev': {
511 title: 'decisions-api.dev',
512 defaultModel: 'jev-latest',
513 models: ['jev-latest'],
514 family: JEV_FAMILY,
515 isGateway: true,
516 limits: RESELLER,
517 needs: ['decisionsApiKey'],
518 bodyModelOf: TRIMMED,
519 urlOf: () => 'https://decisions-api.dev/v1/systemone',
520 tokenOf: credentials => credentials.decisionsApiKey,
521 wire: ALIASED,
522 unwrap: unwrapAliased,
523 },
524 // decisionapi.net states its 32 KiB body limit only where it describes
525 // image input; it is the platform's one stated body limit, so text is
526 // held to it as well.
527 'decisionapi-net': {
528 title: 'decisionapi.net',
529 defaultModel: 'jev-latest',
530 models: ['jev-latest'],
531 family: JEV_FAMILY,
532 isGateway: true,
533 limits: RESELLER,
534 needs: ['decisionapiApiKey'],
535 bodyModelOf: TRIMMED,
536 urlOf: () => 'https://decisionapi.net/v1/systemone',
537 tokenOf: credentials => credentials.decisionapiApiKey,
538 wire: BARE,
539 unwrap: unwrapDataEnvelope,
540 },
541 // OpenAI takes the same questions in a body of its own shape and answers
542 // in a list. Its Decisions API documents no limits, so it is held to the
543 // fallback. Those figures are this mod's, not a window OpenAI states: a
544 // reply counting all of them is taken as a cut state although the model
545 // may have read it whole, which errs toward the built-in summary.
546 openai: {
547 title: 'OpenAI',
548 defaultModel: 'gpt-6-luna',
549 models: ['gpt-6-luna'],
550 family: model => model,
551 isGateway: false,
552 limits: UNDOCUMENTED,
553 needs: ['openaiApiKey'],
554 bodyModelOf: TRIMMED,
555 urlOf: () => 'https://api.openai.com/v1/decisions',
556 tokenOf: credentials => credentials.openaiApiKey,
557 wire: OPENAI,
558 unwrap: (payload, _credentials, asked) => decodeOpenAI(payload, asked),
559 },
560}
561
562/**
563 * Narrows a configured value to a provider name.
564 *
565 * @param value the value of the `provider` option
566 * @returns the provider, or undefined when the value names none
567 */
568export function providerOf(value: unknown): ProviderName | undefined {
569 return PROVIDERS.find(name => name === value)
570}
571
572/**
573 * The limits a request over a route must keep to. OpenRouter is a gateway:
574 * a request it forwards is also held to what the model's own host accepts.
575 * Its Jev window is documented, and no question cap with it. Any other model
576 * it serves is held to Cloudflare's 64 questions as well, because Clef, the
577 * one other System One model, is served there and refuses a request with
578 * more (HTTP 422, "Dictionary should have at most 64 items").
579 *
580 * @param route the provider and the model
581 * @returns its limits
582 */
583export function limitsOf(route: Route): Limits {
584 const limits = TABLE[route.provider].limits
585
586 if (route.provider === 'openrouter' && !OPENROUTER_JEV.test(route.model)) {
587 return {
588 ...limits,
589 maxQuestions: TABLE.cloudflare.limits.maxQuestions,
590 }
591 }
592
593 return limits
594}
595
596/**
597 * The models a provider is offered with in the routing decision.
598 *
599 * @param provider the provider
600 * @returns the model names as its body spells them
601 */
602export function modelsOf(provider: ProviderName): readonly string[] {
603 return TABLE[provider].models
604}
605
606/**
607 * The model family a route reaches: a gateway's route to Jev is the same
608 * Jev that TypeSafe serves directly.
609 *
610 * @param route the route
611 * @returns the family, the model name itself for a model of its own
612 */
613export function familyOf(route: Route): string {
614 return TABLE[route.provider].family(route.model)
615}
616
617/**
618 * Names a gateway as a sentence about it spells the name.
619 *
620 * @param provider the provider
621 * @returns the name, or undefined for a provider that hosts its models
622 */
623export function gatewayOf(provider: ProviderName): string | undefined {
624 const described = TABLE[provider]
625
626 return described.isGateway ? described.title : undefined
627}
628
629/**
630 * The UTF-8 bytes a request over a route takes besides its questions: the
631 * body as encoded around one question, less that question as the route's
632 * wire format measures it. The model name and the state are counted
633 * exactly as they are sent. What is left holds no bracket of the question
634 * map or list; each question is measured with brackets of its own and a
635 * comma after it, so a body counted this way is over by a few bytes a
636 * question, never under.
637 *
638 * @param route the route
639 * @param state what the questions are asked about
640 * @returns the byte count
641 */
642export function bytesBesideOf(route: Route, state: unknown): number {
643 const described = TABLE[route.provider]
644 const question: Question = { type: 'noul', instructions: '' }
645 const body = described.wire.encode(
646 described.bodyModelOf(route.model),
647 state,
648 { q: question },
649 )
650
651 return bytesOf(body) - bytesOf(described.wire.questionJsonOf('q', question))
652}
653
654/**
655 * How a request over a route is written, which is also how its size is
656 * estimated.
657 *
658 * @param route the route
659 * @returns the provider's wire format
660 */
661export function wireOf(route: Route): Wire {
662 return TABLE[route.provider].wire
663}
664
665/**
666 * Builds the route for a provider: the configured model when one is set, the
667 * provider's default otherwise, spelled as the provider's body expects it.
668 *
669 * @param provider the provider
670 * @param model the configured model name, when one is set
671 * @returns the route; throws when the provider does not serve the model
672 */
673export function routeOf(provider: ProviderName, model?: string): Route {
674 const described = TABLE[provider]
675
676 return {
677 provider,
678 model: described.bodyModelOf(textOf(model) ?? described.defaultModel),
679 }
680}
681
682/**
683 * Names the environment variables of the credentials a provider needs and
684 * does not have. The names are safe to show; the values never are.
685 *
686 * @param provider the provider
687 * @param credentials what was resolved
688 * @returns the variable names, empty when the provider can be called
689 */
690export function missingOf(
691 provider: ProviderName,
692 credentials: Credentials,
693): EnvName[] {
694 return TABLE[provider].needs
695 .filter(need => credentials[need] === undefined)
696 .map(need => ENV_NAMES[need])
697}
698
699/**
700 * Throws when a provider lacks a credential it needs, naming the variable
701 * to set. Checked before any work is done for a request that could not be
702 * sent.
703 *
704 * @param provider the provider
705 * @param credentials what was resolved
706 */
707export function assertConfigured(
708 provider: ProviderName,
709 credentials: Credentials,
710): void {
711 const missing = missingOf(provider, credentials)
712
713 if (missing.length > 0) {
714 throw new Error(
715 `${provider} is not configured: ${missing.join(' and ')} is unset`,
716 )
717 }
718}
719
720/**
721 * Names the credential variables that still have no value after the plugin
722 * options and one source of the environment were read, among those the given
723 * providers need. An empty answer means a further source would add nothing,
724 * so it need not be read.
725 *
726 * @param providers the providers a compaction may call
727 * @param options the plugin's options
728 * @param read the variables as that source holds them
729 * @returns the variable names, each once
730 */
731export function unresolvedOf(
732 providers: readonly ProviderName[],
733 options: PluginOptions,
734 read: Environment,
735): EnvName[] {
736 const credentials = credentialsOf(options, read)
737
738 return [
739 ...new Set(providers.flatMap(provider => missingOf(provider, credentials))),
740 ]
741}
742
743/**
744 * Keeps only the credentials the given providers need. A credential of a
745 * provider that may not be called is dropped here, so no later step can
746 * send anything with it.
747 *
748 * @param providers the providers a compaction may call
749 * @param credentials what was resolved
750 * @returns the credentials of those providers
751 */
752export function credentialsFor(
753 providers: readonly ProviderName[],
754 credentials: Credentials,
755): Credentials {
756 const needed = new Set(providers.flatMap(provider => TABLE[provider].needs))
757 const kept: Credentials = {}
758
759 for (const key of needed) {
760 const value = credentials[key]
761
762 if (value !== undefined) {
763 kept[key] = value
764 }
765 }
766
767 return kept
768}
769
770/**
771 * Lays one source of the credential variables under another: a value the
772 * first source holds wins, and the `env` block of the settings fills the
773 * rest. The block is read as untyped data, since settings are a plain object.
774 *
775 * @param read the variables as the process environment holds them
776 * @param settingsEnv the `env` key of the merged settings
777 * @returns the variables with every gap the settings can fill filled
778 */
779export function environmentOf(
780 read: Environment,
781 settingsEnv: unknown,
782): Environment {
783 const fallback = isRecord(settingsEnv) ? settingsEnv : {}
784 const merged: Environment = {}
785
786 for (const name of Object.values(ENV_NAMES)) {
787 merged[name] = textOf(read[name]) ?? textOf(fallback[name])
788 }
789
790 return merged
791}
792
793/**
794 * Resolves each credential: the plugin option when it is set, the
795 * environment otherwise.
796 *
797 * @param options the plugin's options
798 * @param environment the credential variables, already merged over settings
799 * @returns the credentials that have a value
800 */
801export function credentialsOf(
802 options: PluginOptions,
803 environment: Environment,
804): Credentials {
805 const credentials: Credentials = {}
806
807 for (const [key, name] of Object.entries(ENV_NAMES)) {
808 const value = textOf(options[key]) ?? textOf(environment[name])
809
810 if (value !== undefined) {
811 credentials[key as keyof Credentials] = value
812 }
813 }
814
815 return credentials
816}
817
818/**
819 * Takes every credential out of a text that is about to be shown.
820 *
821 * A provider's error body is quoted to the person, and a body may repeat
822 * what the request carried: the key itself, or the whole `Authorization`
823 * header. Each secret that was resolved is replaced wherever it stands, the
824 * longest first so that one secret containing another leaves no remainder,
825 * and then anything that follows the word `Bearer`, which covers a token
826 * this mod never held. A resolved value under `MIN_SECRET_CHARS` is not
827 * replaced by value; the `Bearer` rule still covers it where it follows
828 * that word. The account id is not a secret and is left alone.
829 *
830 * @param text the text
831 * @param credentials the resolved credentials
832 * @returns the text with `REDACTED` where a credential stood
833 */
834export function redacted(text: string, credentials: Credentials): string {
835 const secrets = PROVIDERS.map(provider =>
836 TABLE[provider].tokenOf(credentials),
837 )
838 .filter(
839 (secret): secret is string =>
840 secret !== undefined && secret.length >= MIN_SECRET_CHARS,
841 )
842 .sort((a, b) => b.length - a.length)
843 let shown = text
844
845 for (const secret of secrets) {
846 shown = shown.split(secret).join(REDACTED)
847 }
848
849 return shown.replace(BEARER, REDACTED)
850}
851
852/**
853 * The C0 and C1 control characters: a terminal acts on them rather than
854 * showing them, so a quoted text could move the cursor or recolour what
855 * follows it.
856 */
857const CONTROL = /[\u0000-\u001f\u007f-\u009f]/g
858
859/**
860 * A text this mod did not write, made fit to quote in a message: redacted,
861 * on one line, with no control character, and no longer than
862 * `REASON_CHARS`.
863 *
864 * The text is redacted before it is shortened: a cut that fell inside a
865 * credential would leave a part of it that no longer matches the whole. The
866 * cut never falls between the two halves of a surrogate pair.
867 *
868 * @param text what a provider or the host said
869 * @param credentials the resolved credentials
870 * @returns the line to quote; `no reason given` for a text that says nothing
871 */
872export function briefOf(text: string, credentials: Credentials): string {
873 const line = redacted(text, credentials)
874 .replace(CONTROL, ' ')
875 .replace(/\s+/g, ' ')
876 .trim()
877
878 return line === '' ? 'no reason given' : cutOf(line)
879}
880
881/**
882 * A text held to `REASON_CHARS`: cut there, never between the two halves of
883 * a surrogate pair, with `…` where it was cut.
884 *
885 * @param text the text
886 * @returns the text, or its first `REASON_CHARS` characters and `…`
887 */
888export function cutOf(text: string): string {
889 return text.length > REASON_CHARS ? `${headOf(text, REASON_CHARS)}…` : text
890}
891
892/**
893 * The message of a thrown value, which need not be an `Error`.
894 *
895 * @param error what was thrown
896 * @returns its message, or the value as text
897 */
898export function messageOf(error: unknown): string {
899 return error instanceof Error ? error.message : String(error)
900}
901
902/**
903 * Builds the request for one batch of questions. The credential goes into
904 * the `authorization` header and nowhere else, so nothing built from the
905 * returned URL or body can leak it.
906 *
907 * @param route which provider and model
908 * @param credentials the resolved credentials
909 * @param state what the questions are asked about
910 * @param questions the questions, by id
911 * @returns the URL and the init for `$.http.fetch`; throws when a credential
912 * the provider needs is missing, naming its environment variable
913 */
914export function exchangeOf(
915 route: Route,
916 credentials: Credentials,
917 state: unknown,
918 questions: Questions,
919): Exchange {
920 assertConfigured(route.provider, credentials)
921
922 const described = TABLE[route.provider]
923 const token = described.tokenOf(credentials) ?? ''
924 const model = described.bodyModelOf(route.model)
925
926 return {
927 url: described.urlOf(model, credentials),
928 init: {
929 method: 'POST',
930 headers: {
931 authorization: `Bearer ${token}`,
932 'content-type': 'application/json',
933 },
934 body: described.wire.encode(model, state, questions),
935 },
936 }
937}
938
939/**
940 * Says whether a status is worth another attempt: rate limiting, an
941 * overloaded provider and any server-side failure. A 4xx other than 429 is
942 * the request's own fault and would fail the same way again.
943 *
944 * @param status the HTTP status
945 * @returns true when a retry may succeed
946 */
947export function isRetryable(status: number): boolean {
948 return status === 429 || status >= 500
949}
950
951/**
952 * Finds the sentence an error body gives as its reason, across the shapes
953 * the providers use: `error.message`, `errors[0].message`, a bare
954 * `error`, `message` or `detail` string.
955 */
956function reasonIn(payload: unknown): string | undefined {
957 if (!isRecord(payload)) {
958 return undefined
959 }
960
961 const nested = isRecord(payload.error)
962 ? textOf(payload.error.message)
963 : undefined
964
965 return (
966 nested ??
967 envelopeErrorOf(payload) ??
968 textOf(payload.error) ??
969 textOf(payload.message) ??
970 textOf(payload.detail)
971 )
972}
973
974/**
975 * The field errors of a validation failure, as `field: message` sentences:
976 * the `error.details.fieldErrors` map Workers AI answers a body it refuses
977 * with, which names what was wrong where the reason only says that it was.
978 */
979function fieldErrorsIn(payload: unknown): string | undefined {
980 const error = isRecord(payload) ? payload.error : undefined
981 const details = isRecord(error) ? error.details : undefined
982 const fields = isRecord(details) ? details.fieldErrors : undefined
983
984 if (!isRecord(fields)) {
985 return undefined
986 }
987
988 const sentences = Object.entries(fields).flatMap(([field, messages]) =>
989 Array.isArray(messages)
990 ? messages.filter(m => typeof m === 'string').map(m => `${field}: ${m}`)
991 : [],
992 )
993
994 return sentences.length > 0 ? sentences.join('; ') : undefined
995}
996
997/**
998 * How deep a reason is followed into the bodies quoted inside it.
999 */
1000const REASON_DEPTH = 4
1001
1002/**
1003 * A reason with the bodies it quotes unwrapped. A gateway passes on the
1004 * failure of the host it forwarded to as text: OpenRouter's reason is
1005 * `HTTP 422: ` and the host's body, whose own reason is `AiError: AiError: `
1006 * and another body. Only the innermost sentence says what was refused, and
1007 * it would not fit a toast line behind the wrappers, so every `HTTP nnn:`
1008 * and `AiError:` prefix is dropped and a quoted body is read for its own
1009 * reason and field errors.
1010 */
1011function innermostReason(reason: string, depth: number): string {
1012 const bare = reason.replace(/^(?:\s*(?:HTTP \d{3}|AiError):)+\s*/i, '')
1013 const start = bare.indexOf('{')
1014 const end = bare.lastIndexOf('}')
1015
1016 if (depth >= REASON_DEPTH || start !== 0 || end < start) {
1017 return bare
1018 }
1019
1020 let quoted: unknown
1021
1022 try {
1023 quoted = JSON.parse(bare.slice(start, end + 1))
1024 } catch {
1025 return bare
1026 }
1027
1028 return statedIn(quoted, depth + 1) ?? bare
1029}
1030
1031/**
1032 * What an error body states: its reason with the bodies quoted in it
1033 * unwrapped, and its field errors after it, or either alone.
1034 */
1035function statedIn(payload: unknown, depth: number): string | undefined {
1036 const reason = reasonIn(payload)
1037 const fields = fieldErrorsIn(payload)
1038
1039 return reason === undefined
1040 ? fields
1041 : fields === undefined
1042 ? innermostReason(reason, depth)
1043 : `${innermostReason(reason, depth)}: ${fields}`
1044}
1045
1046/**
1047 * The reason a failed response gives, as one short line. The response body
1048 * is the only source: the sentence it names as its reason when it is JSON
1049 * of a known shape, the body itself otherwise.
1050 */
1051function reasonOf(text: string, credentials: Credentials): string {
1052 let payload: unknown
1053
1054 try {
1055 payload = JSON.parse(text)
1056 } catch {
1057 payload = undefined
1058 }
1059
1060 return briefOf(statedIn(payload, 0) ?? text, credentials)
1061}
1062
1063/**
1064 * Turns a provider's HTTP response into a reply, or throws saying why not: a
1065 * failed status with the provider's own reason, a body that is not JSON, an
1066 * envelope reporting failure, a body without answers, or a count
1067 * of input tokens that shows the request was not read whole.
1068 *
1069 * Every message that quotes the response is redacted here, where it is
1070 * built, so no caller can show one that is not.
1071 *
1072 * @param route the route the request went over
1073 * @param response what `$.http.fetch` resolved with
1074 * @param credentials the resolved credentials, to keep out of the messages
1075 * @param asked the questions the request asked, for a provider whose answers
1076 * name them otherwise
1077 * @returns the reply
1078 */
1079export function replyFrom(
1080 route: Route,
1081 response: HttpResponse,
1082 credentials: Credentials,
1083 asked: Questions,
1084): Reply {
1085 if (!response.ok) {
1086 throw new Error(
1087 `${route.provider} answered HTTP ${response.status}: ` +
1088 reasonOf(response.text, credentials),
1089 )
1090 }
1091
1092 let payload: unknown
1093
1094 try {
1095 payload = JSON.parse(response.text)
1096 } catch {
1097 throw new Error(`${route.provider} answered a body that is not JSON`)
1098 }
1099
1100 let reply: Reply
1101
1102 try {
1103 reply = replyOf(TABLE[route.provider].unwrap(payload, credentials, asked))
1104 } catch (error) {
1105 throw new Error(
1106 `${route.provider}: ${briefOf(messageOf(error), credentials)}`,
1107 )
1108 }
1109
1110 // The size of a request is only estimated here, and a provider may cut a
1111 // state that is too long without saying so. A count that is at the limit
1112 // is what a request cut to the limit reports, and answers about a state
1113 // the model saw part of are no ground for removing anything.
1114 const counted = reply.usage.input_tokens
1115 const { maxRequestTokens } = TABLE[route.provider].limits
1116
1117 if (counted !== undefined && counted >= maxRequestTokens) {
1118 throw new Error(
1119 `${route.provider} counted ${counted} input tokens, all that a ` +
1120 `request of its may hold (${maxRequestTokens}): the state was ` +
1121 'probably cut short, so the answers decide nothing',
1122 )
1123 }
1124
1125 return reply
1126}
1127hooks/report.ts 151 lines1import { reductionOf } from './compact'
2import type { Action, Decision } from './compact'
3import { labelOf } from './route'
4import type { Result } from './run'
5
6/**
7 * The longest line `$.ui.log` is given; a longer decisions list is split.
8 */
9const LOG_LINE_CHARS = 4096
10
11/**
12 * What every line of the per-call verdicts starts with.
13 */
14const VERDICTS = 'per-call verdicts'
15
16/**
17 * Room kept in every line for the lead, the longest of which is
18 * `per-call verdicts, part 12 of 34: `.
19 */
20const LEAD_ROOM = 40
21
22/**
23 * What stands between two entries on a line.
24 */
25const BETWEEN = '; '
26
27/**
28 * A ratio as a whole percentage.
29 *
30 * @param ratio the ratio, 0 to 1
31 * @returns for example `42%`
32 */
33export function percentOf(ratio: number): string {
34 return `${Math.round(ratio * 100)}%`
35}
36
37function countOf(decisions: readonly Decision[], action: Action): number {
38 return decisions.filter(decision => decision.action === action).length
39}
40
41/**
42 * One line saying what a compaction did and what it cost to decide.
43 *
44 * It names the provider and model that answered, and puts the state's
45 * estimated size beside the input tokens the provider itself counted. The
46 * two disagreeing shows how far the estimate drifts; the provider's count
47 * falling well below the estimate shows a state the provider cut short.
48 *
49 * The calls are counted by what happened to them, not by what was decided:
50 * a call whose result was to be cut, but was too short to gain from it,
51 * stands whole and is counted so.
52 *
53 * @param result the compaction
54 * @returns the line, without a lead
55 */
56export function summaryOf(result: Result): string {
57 const { outcome, route } = result
58 const uncut = countOf(outcome.decisions, 'truncate') - outcome.shortened
59 const counts = (
60 [
61 ['left whole', countOf(outcome.decisions, 'keep') + uncut],
62 ['with the result cut short', outcome.shortened],
63 ['removed', countOf(outcome.decisions, 'drop')],
64 ['not judged', countOf(outcome.decisions, 'pinned')],
65 ] as const
66 )
67 .filter(([, count]) => count > 0)
68 .map(([said, count]) => `${count} ${said}`)
69 const parts = [
70 `${percentOf(reductionOf(outcome))} smaller`,
71 counts.length > 0
72 ? `tool calls: ${counts.join(', ')}`
73 : 'no answered tool call',
74 ]
75
76 if (outcome.requests > 0) {
77 const counted =
78 outcome.reported.inputTokens === undefined
79 ? ''
80 : `, ${outcome.reported.inputTokens} input tokens counted by the provider`
81 const cost =
82 outcome.reported.cost === undefined
83 ? ''
84 : `, cost $${outcome.reported.cost.toFixed(6)}`
85
86 parts.push(
87 `${labelOf(route)} in ${outcome.requests} ` +
88 `${outcome.requests === 1 ? 'request' : 'requests'}`,
89 `state about ${outcome.stateTokens} tokens estimated ` +
90 `(${outcome.stage ?? 'whole'})${counted}${cost}`,
91 )
92 }
93
94 return parts.join('; ')
95}
96
97/**
98 * The decisions as log lines: one entry per call that was asked about, with
99 * both probabilities, split so that no line exceeds what one log line holds.
100 *
101 * @param decisions every decision of the compaction
102 * @param maxChars the longest a line may be
103 * @returns the lines, in call order
104 */
105export function decisionLinesOf(
106 decisions: readonly Decision[],
107 maxChars: number = LOG_LINE_CHARS,
108): string[] {
109 const entries = decisions
110 .filter(decision => decision.action !== 'pinned')
111 .map(
112 ({ call, action, keepCall, keepResult }) =>
113 `${call.id} ${call.tool} -> ${action} ` +
114 `(call ${keepCall.toFixed(2)}, result ${keepResult.toFixed(2)})`,
115 )
116
117 // Each line is a group of entries, filled by length: an entry joins the
118 // newest group while the group, with it and the separator before it,
119 // stays within the room, and opens a group of its own otherwise. An entry
120 // longer than the room therefore still gets a line.
121 const room = Math.max(1, maxChars - LEAD_ROOM)
122 const groups: string[][] = []
123 let used = 0
124
125 for (const entry of entries) {
126 const newest = groups.at(-1)
127 const added = BETWEEN.length + entry.length
128
129 if (newest === undefined || used + added > room) {
130 groups.push([entry])
131 used = entry.length
132 } else {
133 newest.push(entry)
134 used += added
135 }
136 }
137
138 if (groups.length === 0) {
139 return [`${VERDICTS}: no tool call was asked about`]
140 }
141
142 return groups.map((group, index) => {
143 const lead =
144 groups.length === 1
145 ? VERDICTS
146 : `${VERDICTS}, part ${index + 1} of ${groups.length}`
147
148 return `${lead}: ${group.join(BETWEEN)}`
149 })
150}
151hooks/run.ts 214 lines1import type { SessionMessage } from 'claude-code'
2
3import { askerOf, attemptOf } from './ask'
4import type { Ports } from './ask'
5import {
6 batchesOf,
7 budgetOf,
8 compact,
9 demandOf,
10 LIMIT_SHARE,
11 untouched,
12} from './compact'
13import type { Budget, Outcome } from './compact'
14import { eligibleOf } from './config'
15import type { Config } from './config'
16import {
17 assertConfigured,
18 bytesBesideOf,
19 limitsOf,
20 messageOf,
21 routeOf,
22 wireOf,
23} from './providers'
24import type { Credentials, Route } from './providers'
25import { chooseRoute, profileOf } from './route'
26import { estimatedTokensOf, goalOf, pairCalls, stateWithin } from './state'
27import type { Fitted } from './state'
28
29/**
30 * One compaction to carry out.
31 */
32export type Job = {
33 messages: readonly SessionMessage[]
34 /**
35 * The instructions given with the compaction, when there were any.
36 */
37 instructions?: string
38 config: Config
39 credentials: Credentials
40 ports: Ports
41}
42
43/**
44 * A compaction carried out: the outcome, the route it went over, and how
45 * the route was picked when the decision was on.
46 */
47export type Result = {
48 outcome: Outcome
49 route: Route
50 routing?: string
51}
52
53/**
54 * Carries out one compaction from the transcript to the rebuilt messages.
55 *
56 * The state is fitted once, against the configured provider's limits and
57 * wire format, a cap on the body in bytes included, and every request that
58 * asks about a call then goes over one route: the configured one, or the
59 * one the provider decision picked. A route whose wire format carries the
60 * state differently is held to the state's size in that format. A failure
61 * anywhere throws; the caller falls back to the built-in summary.
62 *
63 * Everything sent to a provider, the routing question included, is sent
64 * inside one attempt: one deadline, armed here once the state is fitted and
65 * released when the outcome is known. Past it, on the engine's signal, and
66 * at the first request about tool calls that fails for good, whatever is
67 * still out is abandoned, nothing further is sent or waited for, and this
68 * throws the one reason. The routing question failing is not such an end:
69 * the configured route decides instead. Once the outcome is known, whatever
70 * is still out starts nothing more.
71 *
72 * Options that carry a refusal are not acted on at all: they name no
73 * provider the person chose, so nothing may be sent anywhere.
74 *
75 * @param job the transcript, the configuration, the credentials, the ports
76 * @returns the result
77 */
78export async function run(job: Job): Promise<Result> {
79 const { messages, config, credentials, ports } = job
80
81 if (config.refusal !== undefined) {
82 throw new Error(config.refusal)
83 }
84
85 const configured = routeOf(config.provider, config.model)
86
87 if (!config.providerDecision) {
88 assertConfigured(configured.provider, credentials)
89 }
90
91 const calls = pairCalls(messages, config.preserveRecentMessages)
92 const unpinned = calls.filter(call => !call.isPinned)
93
94 if (unpinned.length === 0) {
95 return { outcome: untouched(messages, calls), route: configured }
96 }
97
98 const demandFor = (route: Route) =>
99 demandOf(unpinned, wireOf(route), limitsOf(route).maxQuestions)
100 const budgetFor = (route: Route) =>
101 budgetOf(config, limitsOf(route), demandFor(route))
102 const goal = goalOf(messages, job.instructions)
103 const { maxRequestBytes } = limitsOf(configured)
104 const fitted = stateWithin(messages, calls, {
105 maxStateTokens: budgetFor(configured).stateTokens,
106 preserveRecentMessages: config.preserveRecentMessages,
107 goal,
108 wire: wireOf(configured),
109 // Held to the same share of a byte cap as the request is below, with
110 // the bytes of one request's questions left free.
111 ...(maxRequestBytes === undefined
112 ? {}
113 : {
114 bytes: {
115 limit:
116 Math.floor(maxRequestBytes * LIMIT_SHARE) -
117 demandFor(configured).batchBytes,
118 besideOf: (state: unknown) => bytesBesideOf(configured, state),
119 },
120 }),
121 })
122 const stateJson = JSON.stringify(fitted.state)
123 const fittedFor = (route: Route): Fitted =>
124 wireOf(route) === wireOf(configured)
125 ? fitted
126 : {
127 ...fitted,
128 tokens: estimatedTokensOf(wireOf(route).stateJsonOf(stateJson)),
129 }
130 // A provider that caps the body in bytes is held to the bytes the state
131 // takes in its request as written, counted exactly, and to the same share
132 // of its cap as of its token limits: what a question will take is still
133 // an estimate, and a body over the cap is refused whole.
134 const budgetOn = (route: Route): Budget => {
135 const budget = budgetFor(route)
136 const { maxRequestBytes } = limitsOf(route)
137
138 return maxRequestBytes === undefined
139 ? budget
140 : {
141 ...budget,
142 bytes: {
143 limit: Math.floor(maxRequestBytes * LIMIT_SHARE),
144 state: bytesBesideOf(route, fitted.state),
145 },
146 }
147 }
148 const attempt = attemptOf(ports)
149
150 try {
151 const result: Pick<Result, 'route' | 'routing'> = { route: configured }
152
153 if (config.providerDecision) {
154 const routed = await chooseRoute(
155 configured,
156 eligibleOf(config),
157 credentials,
158 route => {
159 try {
160 const budget = budgetOn(route)
161 const { tokens } = fittedFor(route)
162
163 if (tokens > budget.stateTokens) {
164 return (
165 `the state is about ${tokens} tokens and it takes ` +
166 `${budget.stateTokens}`
167 )
168 }
169
170 batchesOf(unpinned, tokens, budget, wireOf(route))
171 } catch (error) {
172 return messageOf(error)
173 }
174
175 return undefined
176 },
177 profileOf(messages, unpinned, fitted.tokens, goal),
178 // The routing question is one the compaction can do without: when
179 // it fails, the configured route decides.
180 route => askerOf(route, credentials, attempt, false),
181 )
182 const failure = attempt.failure()
183
184 // A routing question cut off by the deadline or the engine's signal
185 // is not a reason to go on with the configured route: the attempt is
186 // over, and what ended it is the reason to give.
187 if (failure !== undefined) {
188 throw failure
189 }
190
191 result.route = routed.route
192 result.routing = routed.why
193 }
194
195 const outcome = await compact(
196 messages,
197 calls,
198 fittedFor(result.route),
199 askerOf(result.route, credentials, attempt, true),
200 {
201 budget: budgetOn(result.route),
202 wire: wireOf(result.route),
203 keepThreshold: config.keepThreshold,
204 truncateHeadChars: config.truncateHeadChars,
205 halt: attempt.fail,
206 },
207 )
208
209 return { ...result, outcome }
210 } finally {
211 attempt.close()
212 }
213}
214hooks/trigger.ts 108 lines1/**
2 * What the `turn.complete` trigger remembers between turns.
3 */
4type Arm = {
5 /**
6 * The usage the last requested compaction left, while that is still at or
7 * above the threshold: the level the next request has to rise from.
8 */
9 floor?: number
10 /**
11 * True from a requested compaction until usage is next known. Usage is
12 * absent right after a compaction (it comes with the next response), so
13 * the floor is then taken from the first reading that follows.
14 */
15 isFloorPending: boolean
16}
17
18/**
19 * How far usage has to rise above the floor before another compaction is
20 * requested, in points of the context window.
21 */
22const REARM_POINTS = 10
23
24/**
25 * A trigger that has requested nothing yet.
26 *
27 * @returns the initial state
28 */
29export function armOf(): Arm {
30 return { isFloorPending: false }
31}
32
33/**
34 * Says whether a compaction should be requested at this reading of usage,
35 * and updates what is remembered.
36 *
37 * A reading at or above the threshold asks for one, unless an earlier
38 * request left usage near where it is now: a compaction that could not get
39 * usage under the threshold would otherwise be requested again at the end
40 * of every turn. The next request then waits until usage has risen
41 * `REARM_POINTS` above that floor. A reading under the threshold forgets
42 * the floor. When the answer is yes the floor is set to the reading at
43 * once, so a request that then fails is not repeated on the next turn.
44 *
45 * @param arm what is remembered, changed in place
46 * @param percent the context usage in percent, absent when it is not known
47 * @param threshold the `compactAtPercent` option
48 * @returns true when a compaction should be requested now
49 */
50export function isCompactionDue(
51 arm: Arm,
52 percent: number | undefined,
53 threshold: number,
54): boolean {
55 if (percent === undefined) {
56 return false
57 }
58
59 if (percent < threshold) {
60 delete arm.floor
61 arm.isFloorPending = false
62
63 return false
64 }
65
66 if (arm.isFloorPending) {
67 arm.floor = percent
68 arm.isFloorPending = false
69 }
70
71 if (arm.floor !== undefined && percent < arm.floor + REARM_POINTS) {
72 return false
73 }
74
75 arm.floor = percent
76
77 return true
78}
79
80/**
81 * Records what a requested compaction left behind.
82 *
83 * A reading at or above the threshold becomes the floor: the compaction
84 * could not get usage under it, so the next request has to wait for a rise.
85 * A reading under the threshold leaves no floor, because the compaction
86 * worked and the next crossing of the threshold is a new one; a floor kept
87 * there would hold the next request back until usage was `REARM_POINTS`
88 * above a level that was never a problem, which a high threshold can put
89 * past the top of the window. Usage that is not known yet marks the floor as
90 * pending.
91 *
92 * @param arm what is remembered, changed in place
93 * @param percent the context usage read after the compaction
94 * @param threshold the `compactAtPercent` option
95 */
96export function settle(
97 arm: Arm,
98 percent: number | undefined,
99 threshold: number,
100): void {
101 delete arm.floor
102 arm.isFloorPending = percent === undefined
103
104 if (percent !== undefined && percent >= threshold) {
105 arm.floor = percent
106 }
107}
108hooks/state.ts 779 lines1import type { SessionMessage, ToolUseSummary } from 'claude-code'
2
3import { bytesOf } from './systemone'
4import type { Wire } from './systemone'
5
6/**
7 * One tool call with the result that answers it, found by `tool_use_id`.
8 */
9export type Call = {
10 /**
11 * A name of a few characters (`t1`, `t2`, ...) that stands for the call in
12 * the state and in the question ids, where the engine's own `tool_use_id`
13 * would be paid for, character by character, in every request.
14 */
15 id: string
16 toolUseId: string
17 tool: string
18 input: Record<string, unknown>
19 /**
20 * Where in the conversation the assistant made the call.
21 */
22 callIndex: number
23 /**
24 * Where in the conversation the answer to the call arrived.
25 */
26 resultIndex: number
27 resultChars: number
28 isError: boolean
29 /**
30 * True when the call is shown to the model but never asked about, so it
31 * stays exactly as it is. That holds when either of its two messages is
32 * one that is never changed, and when its `tool_use_id` is not unique in
33 * the conversation: a decision is applied by that id, so one taken for
34 * such a call would also reach the other blocks that carry it.
35 */
36 isPinned: boolean
37}
38
39/**
40 * A call as the state shows it: its input, and a note in place of its output.
41 */
42export type CallNote = {
43 id: string
44 tool: string
45 input: string
46 result: string
47}
48
49/**
50 * One message of the conversation as the state shows it. `calls` holds
51 * notes, or one line per call once the state had to shrink that far.
52 */
53type Entry = {
54 at: number
55 role: SessionMessage['role']
56 text: string
57 calls?: CallNote[] | string[]
58}
59
60/**
61 * What every request sends as `state`.
62 */
63export type State = {
64 context: string
65 goal: string
66 conversation: Entry[]
67}
68
69/**
70 * How far the state had to be reduced before it fitted, for the report.
71 */
72export type Stage =
73 | 'whole'
74 | 'inputs shortened'
75 | 'long texts cut in the middle'
76 | 'old texts replaced by their length'
77 | 'old calls on one line'
78 | 'old call-free messages omitted'
79 | 'rows of old calls joined'
80
81/**
82 * A state that fits its budget.
83 */
84export type Fitted = {
85 state: State
86 /**
87 * The estimated size of the serialised state, as the wire format it was
88 * fitted for carries it.
89 */
90 tokens: number
91 stage: Stage
92}
93
94/**
95 * What the fitting needs to know.
96 */
97type FitNeed = {
98 maxStateTokens: number
99 preserveRecentMessages: number
100 goal: string
101 /**
102 * The wire format the state is sent in. It is measured as that format
103 * carries it: a format that escapes the state inside a string makes it
104 * larger than its own JSON.
105 */
106 wire: Wire
107 /**
108 * For a route that caps the request body in bytes: the UTF-8 bytes the
109 * body may take besides the questions it must leave room for, and the
110 * bytes it takes around a state, questions left out. The state is then
111 * held to both budgets; absent, it is fitted in tokens alone.
112 */
113 bytes?: { limit: number; besideOf: (state: State) => number }
114}
115
116/**
117 * The frame every question is read in. It carries the rubric once, so the
118 * per-call questions can stay short: they are repeated for every call.
119 */
120const CONTEXT =
121 'This is the transcript of a coding assistant session whose context ' +
122 'window is nearly full. To free space, old tool calls and their outputs ' +
123 'are being removed; nothing is summarised or reworded. `conversation` ' +
124 'lists the messages in order, oldest first. A tool call appears under ' +
125 '`calls` with its `id` and its input. Its output is not shown: `result` ' +
126 'only says whether it succeeded and how long it was. Long texts may be ' +
127 'shortened and old messages may be missing. `goal` says what the ' +
128 'assistant is working on. Each question is about one tool call, named by ' +
129 'its `id`. What is removed is gone for good, but the assistant can run ' +
130 'any tool again and read any file again.'
131
132/**
133 * The successive caps on a call's serialised input, in characters.
134 */
135const INPUT_CAPS = [1000, 240, 64] as const
136
137/**
138 * What stays of a long text when it is abridged: its start and its end.
139 */
140const KEPT_START_CHARS = 400
141const KEPT_END_CHARS = 160
142
143/**
144 * A text is abridged only past this length, so the note that replaces its
145 * middle is always shorter than what it replaces.
146 */
147const LONG_TEXT = KEPT_START_CHARS + KEPT_END_CHARS + 48
148
149/**
150 * The longest the note standing for a collapsed text gets; a text no longer
151 * than this is left alone, since the note would not be shorter.
152 */
153const COLLAPSED_NOTE_CHARS = 48
154
155const GOAL_CHARS = 2000
156const PROMPT_CHARS = 500
157const GOAL_PROMPTS = 3
158
159/**
160 * Letters one token is assumed to cover inside a word.
161 */
162const LETTERS_PER_TOKEN = 6
163
164function isLetter(code: number): boolean {
165 return (code >= 65 && code <= 90) || (code >= 97 && code <= 122)
166}
167
168function isDigit(code: number): boolean {
169 return code >= 48 && code <= 57
170}
171
172function isHighSurrogate(code: number): boolean {
173 return code >= 0xd800 && code <= 0xdbff
174}
175
176/**
177 * Estimates how many tokens a provider will count for a text, with no
178 * tokenizer: a hooks module cannot load one.
179 *
180 * Each character is weighed by its class:
181 *
182 * - ASCII letters: 1 token for each started group of `LETTERS_PER_TOKEN`
183 * in an unbroken run of them;
184 * - ASCII digits: 0.5 each;
185 * - other visible ASCII characters: 0.9 each;
186 * - characters beyond ASCII: 1 each, which is about what CJK text comes to;
187 * - white space: nothing.
188 *
189 * The weights follow the design this mod is modelled on, which reports them
190 * landing a little above Jev's own count; they are an estimate, so the
191 * budgets they are compared with must leave headroom.
192 */
193export function estimatedTokensOf(text: string): number {
194 // Counted in tenths of a token so the sum stays in whole numbers: adding
195 // 0.9 repeatedly drifts, and the drift would move the rounded result.
196 let tenths = 0
197 let letters = 0
198
199 for (let index = 0; index < text.length; index++) {
200 const code = text.charCodeAt(index)
201
202 if (isLetter(code)) {
203 letters++
204 continue
205 }
206
207 tenths += Math.ceil(letters / LETTERS_PER_TOKEN) * 10
208 letters = 0
209
210 if (isDigit(code)) {
211 tenths += 5
212 } else if (code > 127) {
213 tenths += 10
214 } else if (code > 32) {
215 tenths += 9
216 }
217 }
218
219 tenths += Math.ceil(letters / LETTERS_PER_TOKEN) * 10
220
221 return Math.ceil(tenths / 10)
222}
223
224/**
225 * The first `count` characters of a text, never ending between the two
226 * halves of a surrogate pair: a lone half does not survive JSON transport.
227 *
228 * @param count how many characters at most
229 * @returns the start of the text
230 */
231export function headOf(text: string, count: number): string {
232 const end = Math.max(0, Math.min(count, text.length))
233
234 return end > 0 &&
235 end < text.length &&
236 isHighSurrogate(text.charCodeAt(end - 1))
237 ? text.slice(0, end - 1)
238 : text.slice(0, end)
239}
240
241/**
242 * The last `count` characters of a text, never starting between the two
243 * halves of a surrogate pair.
244 */
245function tailOf(text: string, count: number): string {
246 const start = Math.max(0, text.length - count)
247
248 return start > 0 && isHighSurrogate(text.charCodeAt(start - 1))
249 ? text.slice(start + 1)
250 : text.slice(start)
251}
252
253/**
254 * A text cut to a length, with a mark where it was cut.
255 */
256function clipped(text: string, limit: number): string {
257 if (text.length > limit) {
258 return `${headOf(text, limit - 1)}…`
259 }
260
261 return text
262}
263
264/**
265 * A long text with its middle replaced by a note of how much is missing.
266 */
267function abridged(text: string): string {
268 const head = headOf(text, KEPT_START_CHARS)
269 const tail = tailOf(text, KEPT_END_CHARS)
270 const missing = text.length - head.length - tail.length
271
272 return `${head}\n[${missing} characters left out]\n${tail}`
273}
274
275/**
276 * Says whether a message is one that is never changed: the first, which
277 * usually states the task, and the newest `recent`, which the assistant is
278 * still working from.
279 *
280 * @param index the message's index
281 * @param total how many messages the conversation has
282 * @param recent how many of the newest are kept as they are
283 * @returns true when the message is pinned
284 */
285function isPinnedIndex(
286 index: number,
287 total: number,
288 recent: number,
289): boolean {
290 // 1 for the newest message, 2 for the one before it, and so on.
291 const fromNewest = total - index
292
293 return index === 0 || fromNewest <= recent
294}
295
296/**
297 * One tool use and where it stands: the message that holds it, and its
298 * place among all the tool uses of the conversation.
299 */
300type Site = {
301 use: ToolUseSummary
302 at: number
303 order: number
304}
305
306/**
307 * Lists the calls that have been answered, in the order they were made,
308 * each joined to its answer through the `tool_use_id` the two share. A call
309 * still waiting for its answer is not listed: removing it would leave the
310 * answer, when it arrives, with nothing to belong to.
311 *
312 * An id is expected on one tool use and on one result. Where it stands on
313 * more than one of either, every call that carries it is marked as not to
314 * be judged: what is decided for a call is applied by its id, so it would
315 * reach each block with that id, one in a pinned message included. Such a
316 * call is joined to the first result with its id.
317 *
318 * @param messages the conversation
319 * @param recent how many of the newest messages are pinned
320 * @returns the paired calls, the ones not to be judged included
321 */
322export function pairCalls(
323 messages: readonly SessionMessage[],
324 recent: number,
325): Call[] {
326 const sites = new Map<string, Site[]>()
327 let order = 0
328
329 for (const [at, message] of messages.entries()) {
330 for (const use of message.toolUses) {
331 const site: Site = { use, at, order: order++ }
332 const sharing = sites.get(use.tool_use_id)
333
334 if (sharing === undefined) {
335 sites.set(use.tool_use_id, [site])
336 } else {
337 sharing.push(site)
338 }
339 }
340 }
341
342 const answered = new Set<string>()
343 const repeated = new Set<string>()
344 const pairs: {
345 site: Site
346 resultIndex: number
347 resultChars: number
348 isError: boolean
349 }[] = []
350
351 for (const [resultIndex, message] of messages.entries()) {
352 for (const { tool_use_id: id, text, isError } of message.toolResults ??
353 []) {
354 if (answered.has(id)) {
355 repeated.add(id)
356 continue
357 }
358
359 answered.add(id)
360
361 for (const site of sites.get(id) ?? []) {
362 pairs.push({ site, resultIndex, resultChars: text.length, isError })
363 }
364 }
365 }
366
367 const isKept = (at: number) => isPinnedIndex(at, messages.length, recent)
368 const isShared = (id: string) =>
369 repeated.has(id) || (sites.get(id)?.length ?? 0) > 1
370
371 return pairs
372 .sort((a, b) => a.site.order - b.site.order)
373 .map(({ site, resultIndex, resultChars, isError }, index) => ({
374 id: `t${index + 1}`,
375 toolUseId: site.use.tool_use_id,
376 tool: site.use.tool,
377 input: site.use.input,
378 callIndex: site.at,
379 resultIndex,
380 resultChars,
381 isError,
382 isPinned:
383 isShared(site.use.tool_use_id) ||
384 isKept(site.at) ||
385 isKept(resultIndex),
386 }))
387}
388
389/**
390 * Says whether a message is something the person typed: a user message
391 * with text that brings no tool result.
392 */
393function isPrompt(message: SessionMessage): boolean {
394 return (
395 message.role === 'user' &&
396 (message.toolResults?.length ?? 0) === 0 &&
397 message.text.trim() !== ''
398 )
399}
400
401/**
402 * Says what the assistant is working on: the instructions given with the
403 * compaction when there are any (the text typed after `/compact`, or a
404 * plugin's), since they were written for exactly this; otherwise the last
405 * prompts the person typed.
406 *
407 * @param messages the conversation
408 * @param instructions the compaction's instructions, when given
409 * @returns the goal text, possibly empty
410 */
411export function goalOf(
412 messages: readonly SessionMessage[],
413 instructions?: string,
414): string {
415 const stated = instructions?.trim() ?? ''
416
417 if (stated !== '') {
418 return clipped(stated, GOAL_CHARS)
419 }
420
421 // Read from the newest message back, and stop at the first few prompts:
422 // the conversation before them has no part in the goal.
423 const prompts: string[] = []
424
425 for (
426 let at = messages.length - 1;
427 at >= 0 && prompts.length < GOAL_PROMPTS;
428 at--
429 ) {
430 const message = messages[at]
431
432 if (message !== undefined && isPrompt(message)) {
433 prompts.unshift(clipped(message.text.trim(), PROMPT_CHARS))
434 }
435 }
436
437 return prompts.join('\n')
438}
439
440function statusOf(call: Call): string {
441 return call.isError ? 'error' : 'ok'
442}
443
444/**
445 * A call's input as JSON text. An input that cannot be serialised (a cycle)
446 * is named rather than allowed to fail the whole compaction.
447 */
448function inputJsonOf(call: Call): string {
449 try {
450 return JSON.stringify(call.input) ?? '{}'
451 } catch {
452 return '[input that cannot be shown]'
453 }
454}
455
456/**
457 * One argument of a call as `name=value` on a single line: a string as it
458 * is, any other value as its JSON, every stretch of white space as one
459 * space.
460 */
461function argumentOf(name: string, value: unknown): string {
462 const shown =
463 typeof value === 'string' ? value : String(JSON.stringify(value))
464
465 return `${name}=${shown.replace(/\s+/g, ' ')}`
466}
467
468/**
469 * A call on one line, for when a structured note costs too much. Its parts,
470 * a space apart: the id, the tool with its arguments in brackets (held to
471 * the tightest input cap), the outcome, and the length of the result.
472 */
473function lineOf(call: Call): string {
474 const said = Object.keys(call.input)
475 .map(name => argumentOf(name, call.input[name]))
476 .join(' ')
477
478 return [
479 call.id,
480 `${call.tool}(${clipped(said, INPUT_CAPS[2])})`,
481 statusOf(call),
482 `${call.resultChars}ch`,
483 ].join(' ')
484}
485
486/**
487 * One entry of the working list: the entry, what it costs, and whether it
488 * may be reduced.
489 */
490type Slot = {
491 entry: Entry
492 tokens: number
493 /**
494 * What the entry takes in UTF-8 bytes, comma included; 0 when the state
495 * has no byte budget.
496 */
497 bytes: number
498 isOld: boolean
499 isOut: boolean
500}
501
502/**
503 * One reduction: the stage it reports, the slots it visits in order, which
504 * of them it applies to, and what it does to one.
505 */
506type Pass = {
507 stage: Stage
508 over: readonly Slot[]
509 applies: (slot: Slot) => boolean
510 reduce: (slot: Slot) => void
511}
512
513/**
514 * Where several entries in a row hold nothing but one-line calls, gives the
515 * lines of all of them to the first and drops the rest. The keys of an
516 * entry are then paid for once for the whole row, and every line still
517 * starts with the id its questions name. `reweigh` measures a slot again
518 * once its entry has grown.
519 */
520function folded(
521 slots: readonly Slot[],
522 reweigh: (slot: Slot) => void,
523): Slot[] {
524 const isFoldable = (slot: Slot) =>
525 slot.isOld &&
526 slot.entry.text === '' &&
527 typeof slot.entry.calls?.[0] === 'string'
528 const kept: Slot[] = []
529 const grown = new Set<Slot>()
530
531 for (const slot of slots) {
532 const last = kept.at(-1)
533
534 if (
535 last !== undefined &&
536 isFoldable(last) &&
537 isFoldable(slot) &&
538 last.entry.role === slot.entry.role
539 ) {
540 // The first merge of a run gives the entry a list of its own; every
541 // later one appends to it. Copying the list or weighing the entry per
542 // merge would cost the square of the run's length, and a session of
543 // nothing but tool calls is one long run.
544 if (!grown.has(last)) {
545 last.entry.calls = [...(last.entry.calls as string[])]
546 grown.add(last)
547 }
548
549 ;(last.entry.calls as string[]).push(...(slot.entry.calls as string[]))
550 continue
551 }
552
553 kept.push(slot)
554 }
555
556 for (const slot of grown) {
557 reweigh(slot)
558 }
559
560 return kept
561}
562
563/**
564 * Describes the conversation to the decision model within `maxStateTokens`.
565 *
566 * The model judges better the more of the session it can see, so the state
567 * starts as everything and gives up detail only while its estimate is over
568 * the budget, cheapest loss first, and stops the moment it is under. Each
569 * rung below is tried only when the one above was not enough:
570 *
571 * 1. a cap on the JSON of each call's input, tightened twice;
572 * 2. the middle of every long text, in old messages before pinned ones;
573 * 3. the text of an old message, of which only its length is then stated;
574 * 4. the structure of an old call, written as a single line instead;
575 * 5. old messages in which no call was made;
576 * 6. the separate entries of neighbouring old messages that hold only calls.
577 *
578 * The estimate is checked here, whichever provider is asked, because a
579 * provider may cut a state that is too long without saying so. Where the
580 * route caps the body in bytes, the state is held to that budget too, on
581 * the same rungs: text with many letters to a token takes more bytes than
582 * the token budget assumes. A conversation still over either budget on the
583 * last rung is refused by throwing.
584 *
585 * @param messages the conversation
586 * @param calls its paired calls, as `pairCalls` answered them
587 * @param need the budget, the pinning, the goal and the wire format
588 * @returns the fitted state, its estimated size and the stage reached
589 */
590export function stateWithin(
591 messages: readonly SessionMessage[],
592 calls: readonly Call[],
593 need: FitNeed,
594): Fitted {
595 const frame = (conversation: Entry[]): State => ({
596 context: CONTEXT,
597 goal: need.goal,
598 conversation,
599 })
600 const tokensOf = (state: State) =>
601 estimatedTokensOf(need.wire.stateJsonOf(JSON.stringify(state)))
602 const frameTokens = tokensOf(frame([]))
603 const { bytes } = need
604 const frameBytes = bytes === undefined ? 0 : bytes.besideOf(frame([]))
605 // What an entry costs inside the list: its own JSON as the wire format
606 // carries it and the comma after it, in tokens and in UTF-8 bytes. The
607 // last entry has no comma, so a sum of these is over by one byte at most.
608 const weigh = (entry: Entry): Pick<Slot, 'tokens' | 'bytes'> => {
609 const json = need.wire.stateJsonOf(JSON.stringify(entry))
610
611 return {
612 tokens: estimatedTokensOf(json) + 1,
613 bytes: bytes === undefined ? 0 : bytesOf(json) + 1,
614 }
615 }
616 const inputs = new Map(calls.map(call => [call, inputJsonOf(call)]))
617 const made = new Map<number, Call[]>()
618
619 for (const call of calls) {
620 made.set(call.callIndex, [...(made.get(call.callIndex) ?? []), call])
621 }
622
623 let slots: Slot[] = []
624 let total = 0
625 let totalBytes = 0
626
627 const tally = () => {
628 total = frameTokens
629 totalBytes = frameBytes
630
631 for (const slot of slots) {
632 total += slot.tokens
633 totalBytes += slot.bytes
634 }
635 }
636
637 const lay = (inputCap: number) => {
638 slots = []
639
640 messages.forEach((message, at) => {
641 const own = made.get(at) ?? []
642
643 if (message.text.trim() === '' && own.length === 0) {
644 return
645 }
646
647 const entry: Entry = {
648 at,
649 role: message.role,
650 text: message.text.trim() === '' ? '' : message.text,
651 }
652
653 if (own.length > 0) {
654 entry.calls = own.map(call => ({
655 id: call.id,
656 tool: call.tool,
657 input: clipped(inputs.get(call) ?? '', inputCap),
658 result: `${statusOf(call)}, ${call.resultChars} characters, not shown`,
659 }))
660 }
661
662 slots.push({
663 entry,
664 ...weigh(entry),
665 isOld: !isPinnedIndex(at, messages.length, need.preserveRecentMessages),
666 isOut: false,
667 })
668 })
669
670 tally()
671 }
672
673 const fits = () =>
674 total <= need.maxStateTokens &&
675 (bytes === undefined || totalBytes <= bytes.limit)
676
677 const fitted = (stage: Stage): Fitted => {
678 const state = frame(
679 slots.filter(slot => !slot.isOut).map(slot => slot.entry),
680 )
681
682 return { state, tokens: tokensOf(state), stage }
683 }
684
685 for (const cap of INPUT_CAPS) {
686 lay(cap)
687
688 if (fits()) {
689 return fitted(cap === INPUT_CAPS[0] ? 'whole' : 'inputs shortened')
690 }
691 }
692
693 const old = slots.filter(slot => slot.isOld)
694 const oldFirst = [...old, ...slots.filter(slot => !slot.isOld)]
695
696 const passes: readonly Pass[] = [
697 {
698 stage: 'long texts cut in the middle',
699 over: oldFirst,
700 applies: slot => slot.entry.text.length > LONG_TEXT,
701 reduce: slot => {
702 slot.entry.text = abridged(slot.entry.text)
703 },
704 },
705 {
706 stage: 'old texts replaced by their length',
707 over: old,
708 applies: slot => slot.entry.text.length > COLLAPSED_NOTE_CHARS,
709 reduce: slot => {
710 const length = messages[slot.entry.at]?.text.length ?? 0
711
712 slot.entry.text = `[${length} characters of text left out]`
713 },
714 },
715 {
716 stage: 'old calls on one line',
717 over: old,
718 applies: slot => slot.entry.calls !== undefined,
719 reduce: slot => {
720 slot.entry.calls = (made.get(slot.entry.at) ?? []).map(lineOf)
721 },
722 },
723 {
724 stage: 'old call-free messages omitted',
725 over: old,
726 applies: slot => slot.entry.calls === undefined,
727 reduce: slot => {
728 slot.isOut = true
729 },
730 },
731 ]
732
733 for (const pass of passes) {
734 for (const slot of pass.over) {
735 if (!pass.applies(slot)) {
736 continue
737 }
738
739 pass.reduce(slot)
740
741 const now = slot.isOut ? { tokens: 0, bytes: 0 } : weigh(slot.entry)
742
743 total += now.tokens - slot.tokens
744 totalBytes += now.bytes - slot.bytes
745 slot.tokens = now.tokens
746 slot.bytes = now.bytes
747
748 if (fits()) {
749 return fitted(pass.stage)
750 }
751 }
752 }
753
754 slots = folded(
755 slots.filter(slot => !slot.isOut),
756 slot => {
757 Object.assign(slot, weigh(slot.entry))
758 },
759 )
760 tally()
761
762 if (fits()) {
763 return fitted('rows of old calls joined')
764 }
765
766 if (total <= need.maxStateTokens && bytes !== undefined) {
767 throw new Error(
768 'the conversation does not fit the request body: about ' +
769 `${totalBytes} bytes are left after every reduction and ` +
770 `${bytes.limit} are allowed`,
771 )
772 }
773
774 throw new Error(
775 `the conversation does not fit the state budget: about ${total} tokens ` +
776 `are left after every reduction and ${need.maxStateTokens} are allowed`,
777 )
778}
779hooks/systemone.ts 244 lines1/**
2 * A yes/no question; the answer is the probability of yes.
3 */
4type NoulQuestion = {
5 type: 'noul'
6 instructions: string
7 criteria?: { true?: string; false?: string }
8}
9
10/**
11 * A pick of one option; `criteria` maps each option to what it stands for.
12 */
13export type ChoiceQuestion = {
14 type: 'choice'
15 instructions: string
16 criteria: Record<string, string | null>
17}
18
19export type Question = NoulQuestion | ChoiceQuestion
20
21/**
22 * The questions of one request, by the id their answers come back under.
23 */
24export type Questions = Record<string, Question>
25
26/**
27 * What a provider reported a request cost. Every field is optional because
28 * the eight providers do not report the same fields.
29 */
30type Usage = {
31 input_tokens?: number
32 cost?: number
33}
34
35/**
36 * One response, unwrapped from whatever envelope its provider put around it.
37 */
38export type Reply = {
39 model?: string
40 answers: Record<string, unknown>
41 usage: Usage
42}
43
44/**
45 * The picked option of a `choice` answer and how sure the model was.
46 */
47type Picked = {
48 choice: string
49 confidence?: number
50}
51
52/**
53 * The rule every question id must pass before any provider-specific alias
54 * is applied: letters, digits, `_`, `.` and `-`, 100 characters at most. A
55 * provider with a stricter rule is sent aliases in place of the ids.
56 */
57const QUESTION_ID = /^[A-Za-z0-9_.-]{1,100}$/
58
59/**
60 * Whether a decoded JSON value is an object, as opposed to an array, `null`
61 * or a scalar.
62 */
63export function isRecord(value: unknown): value is Record<string, unknown> {
64 return typeof value === 'object' && value !== null && !Array.isArray(value)
65}
66
67function numberOf(value: unknown): number | undefined {
68 return typeof value === 'number' && Number.isFinite(value) ? value : undefined
69}
70
71/**
72 * How one provider's request body is written, and how big what it carries
73 * comes out once written: the token estimate is taken of the text that is
74 * sent, not of the objects it was written from.
75 */
76export type Wire = {
77 /**
78 * Serialises one request.
79 */
80 encode: (model: string, state: unknown, questions: Questions) => string
81 /**
82 * The JSON one question takes in the body, its id included.
83 */
84 questionJsonOf: (id: string, question: Question) => string
85 /**
86 * The JSON of the state, or of a part of it, as it stands in the body:
87 * as it is, or escaped inside a string.
88 */
89 stateJsonOf: (json: string) => string
90 /**
91 * The tokens a provider counts for each question beyond its own text.
92 */
93 tokensPerQuestion: number
94}
95
96const UTF8 = new TextEncoder()
97
98/**
99 * The size of a text once sent, in UTF-8 bytes: what a cap stated in bytes
100 * is held against. A character beyond ASCII takes two to four of them.
101 */
102export function bytesOf(text: string): number {
103 return UTF8.encode(text).length
104}
105
106/**
107 * Checks the ids of one request. An id a provider would reject is refused
108 * before anything is sent: a 422 on one id would otherwise cost the whole
109 * batch it rides in.
110 *
111 * @param questions the questions, by id
112 * @returns the ids; throws when there is none or one is not a valid id
113 */
114export function idsOf(questions: Questions): string[] {
115 const ids = Object.keys(questions)
116
117 if (ids.length === 0) {
118 throw new Error('a System One request needs at least one question')
119 }
120
121 for (const id of ids) {
122 if (!QUESTION_ID.test(id)) {
123 throw new Error(`question id ${JSON.stringify(id)} is not a valid id`)
124 }
125 }
126
127 return ids
128}
129
130/**
131 * Serialises one request in the System One form: the three fields every
132 * provider that speaks the protocol as published takes.
133 *
134 * @param model the model name as the provider's body spells it
135 * @param state what every question is asked about
136 * @param questions the questions, by id
137 * @returns the JSON text of the request body
138 */
139export function bodyOf(
140 model: string,
141 state: unknown,
142 questions: Questions,
143): string {
144 idsOf(questions)
145
146 return JSON.stringify({ model, state, questions })
147}
148
149/**
150 * The System One body: the state as it is and the questions as a map by id.
151 */
152export const BARE: Wire = {
153 encode: bodyOf,
154 questionJsonOf: (id, question) => JSON.stringify({ [id]: question }),
155 stateJsonOf: json => json,
156 tokensPerQuestion: 0,
157}
158
159/**
160 * Reads a decoded response body as a reply. Only the `answers` map is
161 * required; each answer is checked when it is read, by the reader that knows
162 * which type the question had.
163 *
164 * @param payload the decoded body, already out of any provider envelope
165 */
166export function replyOf(payload: unknown): Reply {
167 if (!isRecord(payload) || !isRecord(payload.answers)) {
168 throw new Error('the response holds no answers object')
169 }
170
171 const reported = isRecord(payload.usage) ? payload.usage : {}
172 const usage: Usage = {}
173
174 for (const field of ['input_tokens', 'cost'] as const) {
175 const value = numberOf(reported[field])
176
177 if (value !== undefined) {
178 usage[field] = value
179 }
180 }
181
182 const reply: Reply = { answers: payload.answers, usage }
183
184 if (typeof payload.model === 'string') {
185 reply.model = payload.model
186 }
187
188 return reply
189}
190
191/**
192 * Reads the answer to a `noul` question: a probability, so a finite number
193 * from 0 to 1. Anything else throws, because a decision taken on a missing or
194 * out-of-range value would delete or keep a tool result for no reason.
195 *
196 * @param reply the reply the answer is in
197 * @param id the question's id
198 * @returns the probability of yes
199 */
200export function noulOf(reply: Reply, id: string): number {
201 const answer = reply.answers[id]
202 const value = isRecord(answer) ? numberOf(answer.noul) : undefined
203
204 if (value === undefined || value < 0 || value > 1) {
205 throw new Error(`no probability from 0 to 1 was answered for ${id}`)
206 }
207
208 return value
209}
210
211/**
212 * Reads the answer to a `choice` question. The pick must be one of the
213 * options that were offered: a name outside them cannot be acted on.
214 *
215 * @param reply the reply the answer is in
216 * @param id the question's id
217 * @param options the option names the question offered
218 * @returns the picked option and the confidence, when one was reported
219 */
220export function choiceOf(
221 reply: Reply,
222 id: string,
223 options: readonly string[],
224): Picked {
225 const answer = reply.answers[id]
226
227 if (
228 !isRecord(answer) ||
229 typeof answer.choice !== 'string' ||
230 !options.includes(answer.choice)
231 ) {
232 throw new Error(`no offered option was answered for ${id}`)
233 }
234
235 const picked: Picked = { choice: answer.choice }
236 const confidence = numberOf(answer.confidence)
237
238 if (confidence !== undefined) {
239 picked.confidence = confidence
240 }
241
242 return picked
243}
244hooks/openai.ts 201 lines1import { idsOf, isRecord } from './systemone'
2import type { Question, Questions, Wire } from './systemone'
3
4/**
5 * One question as OpenAI's Decisions API takes it: a list entry that names
6 * itself, where System One keys a map by the id.
7 */
8type OpenAIQuestion =
9 | { name: string; type: 'predicate'; instructions: string }
10 | {
11 name: string
12 type: 'choice'
13 instructions: string
14 choices: { value: string; description?: string }[]
15 }
16
17/**
18 * The longest JSON of a value OpenAI sent that an error quotes.
19 */
20const QUOTED_CHARS = 64
21
22/**
23 * A value OpenAI sent, made fit to quote in an error: its JSON, which
24 * escapes every control character, when that is short, and only its length
25 * otherwise. It is never cut: the message is redacted where it is shown,
26 * and a credential cut short would no longer be found there.
27 */
28function quoted(value: unknown): string {
29 const json = JSON.stringify(value)
30
31 if (json === undefined) {
32 return 'missing'
33 }
34
35 return json.length <= QUOTED_CHARS
36 ? json
37 : `a value of ${json.length} characters`
38}
39
40/**
41 * A text made to end a sentence, so that two of them read as two sentences
42 * once appended to the instructions.
43 */
44function sentenceOf(text: string): string {
45 const trimmed = text.trim()
46
47 return /[.!?]$/.test(trimmed) ? trimmed : `${trimmed}.`
48}
49
50/**
51 * A question in OpenAI's form. A `predicate` has no field for what yes and
52 * no stand for, so a `noul` question's criteria are appended to its
53 * instructions, one sentence each; a `choice` option's description is
54 * omitted where System One gives `null`.
55 */
56function questionOf(name: string, question: Question): OpenAIQuestion {
57 if (question.type === 'choice') {
58 return {
59 name,
60 type: 'choice',
61 instructions: question.instructions,
62 choices: Object.entries(question.criteria).map(([value, description]) =>
63 description === null ? { value } : { value, description },
64 ),
65 }
66 }
67
68 const said = [question.instructions]
69 const criteria = question.criteria ?? {}
70
71 if (criteria.true !== undefined) {
72 said.push(`True means: ${sentenceOf(criteria.true)}`)
73 }
74
75 if (criteria.false !== undefined) {
76 said.push(`False means: ${sentenceOf(criteria.false)}`)
77 }
78
79 return { name, type: 'predicate', instructions: said.join(' ') }
80}
81
82/**
83 * Serialises one request for OpenAI: the state as a JSON string in `input`,
84 * which takes text and no object, and the questions as a list.
85 *
86 * @param model the model name
87 * @param state what every question is asked about
88 * @param questions the questions, by id
89 * @returns the JSON text of the request body
90 */
91function encodeOpenAI(
92 model: string,
93 state: unknown,
94 questions: Questions,
95): string {
96 return JSON.stringify({
97 model,
98 input: JSON.stringify(state),
99 questions: idsOf(questions).map(id =>
100 questionOf(id, questions[id] as Question),
101 ),
102 })
103}
104
105/**
106 * Reads OpenAI's list of answers as the System One map by id, which is what
107 * the readers of a reply take: a `predicate`'s probability as `noul`, a
108 * `choice`'s pick and confidence as they are. Only the fields read here are
109 * relied on, so a field OpenAI adds changes nothing; one it renames throws,
110 * naming the field. A refusal to answer throws, naming the question: every
111 * question was sent with a name, so every answer has one.
112 *
113 * An answer is taken only under the name of a question that was asked, and
114 * only once: any other name answers nothing the request asked, and a second
115 * answer to the same question would leave which one counts to chance. The
116 * map has no prototype, so a name such as `__proto__` is an ordinary key.
117 *
118 * @param payload the decoded response body
119 * @param asked the questions the request asked, by id
120 * @returns the body in the System One shape
121 */
122export function decodeOpenAI(payload: unknown, asked: Questions): unknown {
123 if (!isRecord(payload) || !Array.isArray(payload.answers)) {
124 throw new Error('the response holds no answers list')
125 }
126
127 const answers: Record<string, unknown> = Object.create(null)
128
129 payload.answers.forEach((answer: unknown, index) => {
130 const at = `answers[${index}]`
131
132 if (!isRecord(answer)) {
133 throw new Error(`${at} is not an object`)
134 }
135
136 if (typeof answer.name !== 'string' || answer.name === '') {
137 throw new Error(`${at}.name is missing`)
138 }
139
140 if (!Object.hasOwn(asked, answer.name)) {
141 throw new Error(
142 `${at}.name is ${quoted(answer.name)}, which was not asked`,
143 )
144 }
145
146 if (Object.hasOwn(answers, answer.name)) {
147 throw new Error(`${at} answers ${answer.name} a second time`)
148 }
149
150 if (answer.type === 'refusal') {
151 throw new Error(`refused to answer ${answer.name}`)
152 }
153
154 if (answer.type === 'predicate') {
155 answers[answer.name] = { noul: answer.probability }
156 } else if (answer.type === 'choice') {
157 // An option's value may be a boolean in OpenAI's schema. Read as text
158 // it can only match an option that was offered as that same text;
159 // any other pick is refused by the reader of the answer.
160 answers[answer.name] = {
161 choice:
162 typeof answer.choice === 'boolean'
163 ? String(answer.choice)
164 : answer.choice,
165 confidence: answer.confidence,
166 }
167 } else {
168 throw new Error(
169 `${at}.type is ${quoted(answer.type)}, neither predicate nor choice`,
170 )
171 }
172 })
173
174 const reported = isRecord(payload.usage) ? payload.usage : {}
175 const decoded: Record<string, unknown> = {
176 answers,
177 usage: { input_tokens: reported.input_tokens },
178 }
179
180 if (typeof payload.model === 'string') {
181 decoded.model = payload.model
182 }
183
184 return decoded
185}
186
187/**
188 * The body OpenAI's Decisions API takes. The state rides as a string, so its
189 * JSON is escaped once more and is measured in that form.
190 *
191 * OpenAI counts far more input tokens per question than its text: 129
192 * one-line predicates were counted as 19,139 input tokens, about 148 each,
193 * where Codiv counted about 26 each for the same questions.
194 */
195export const OPENAI: Wire = {
196 encode: encodeOpenAI,
197 questionJsonOf: (id, question) => JSON.stringify(questionOf(id, question)),
198 stateJsonOf: json => JSON.stringify(json).slice(1, -1),
199 tokensPerQuestion: 148,
200}
201hooks/route.ts 428 lines1import type { SessionMessage } from 'claude-code'
2
3import type { Ask } from './compact'
4import {
5 briefOf,
6 familyOf,
7 gatewayOf,
8 messageOf,
9 missingOf,
10 modelsOf,
11} from './providers'
12import type { Credentials, ProviderName, Route } from './providers'
13import type { Call } from './state'
14import { choiceOf } from './systemone'
15import type { ChoiceQuestion } from './systemone'
16
17/**
18 * A route the decision may pick: the name it is offered under and what is
19 * documented about the model behind it.
20 */
21type Candidate = {
22 key: string
23 route: Route
24 about: string
25}
26
27/**
28 * The job in a few numbers: what the routing question is asked about. It
29 * stands in for the full state, which would cost a second full-size request
30 * just to choose where to send the first.
31 */
32export type Profile = {
33 context: string
34 goal: string
35 messages: number
36 candidate_calls: number
37 state_tokens: number
38 tools: Record<string, number>
39 error_share: number
40 non_ascii_share: number
41}
42
43/**
44 * Says why the fitted state and its questions do not fit a route's limits,
45 * or nothing when they do.
46 */
47type SizeCheck = (route: Route) => string | undefined
48
49/**
50 * The route to use and one line saying how it was arrived at.
51 */
52type Routed = {
53 route: Route
54 why: string
55}
56
57/**
58 * The id of the one routing question.
59 */
60const ROUTE_QUESTION = 'route'
61
62/**
63 * How many tools the profile names; the rest are counted together.
64 */
65const TOOLS_SHOWN = 8
66
67/**
68 * What the providers' own documentation says each model is for. The
69 * decision is only as good as these lines, so each states published facts
70 * and nothing inferred.
71 */
72const ABOUT: Record<string, string> = {
73 clef:
74 'Clef, a 27B model. Cloudflare recommends it for highest-precision ' +
75 'decisions. It is slower and costs more per token than Clef-flash.',
76 'clef-flash':
77 'Clef-flash, a 9B model. Cloudflare recommends it for latency-critical ' +
78 'decisions on a hot path. It is the fastest of the models offered here.',
79 jev:
80 "Jev, TypeSafe's flagship System One model. It reads text only, and " +
81 'English is its primary training language: TypeSafe reports lower ' +
82 'accuracy on other languages, CJK scripts included.',
83 openjev:
84 'OpenJev, DiffusionGemma 26B-A4B made into a System One model; it ' +
85 'reads a distribution for every answer in one denoising step.',
86 'pplx-decider':
87 "Perplexity's 27B decision model. It reads text and images and " +
88 'returns typed answers with probabilities.',
89 'gpt-6-luna':
90 'GPT-6 Luna, which OpenAI describes as its most efficient model for ' +
91 'focused, high-volume tasks, served through its Decisions API.',
92}
93
94const PROFILE_CONTEXT =
95 'A coding assistant session is about to be compacted: a decision model ' +
96 'will be asked, for each old tool call, whether the call and its output ' +
97 'still have to stay in the transcript. This object describes that job: ' +
98 'how many messages and candidate tool calls there are, the estimated ' +
99 'size in tokens of the state every request will carry, which tools were ' +
100 'called how often, the share of calls that ended in an error, and the ' +
101 'share of the conversation text that is not ASCII. `goal` is what the ' +
102 'assistant is working on.'
103
104/**
105 * A route as a line of text shows it.
106 *
107 * @param route the route
108 * @returns `provider/model`
109 */
110export function labelOf(route: Route): string {
111 return `${route.provider}/${route.model}`
112}
113
114function keyOf(route: Route): string {
115 return `${route.provider}.${route.model.replace(/[^A-Za-z0-9_.-]/g, '-')}`
116}
117
118function aboutOf(route: Route): string {
119 const known = ABOUT[familyOf(route)]
120 const gateway = gatewayOf(route.provider)
121 const via = gateway === undefined ? '' : ` Reached through ${gateway}.`
122
123 return known === undefined
124 ? `The model ${route.model} at ${route.provider}; nothing is documented ` +
125 'here about what it is best at.'
126 : `${known}${via}`
127}
128
129/**
130 * Says why a route cannot take the job, or nothing when it can. Both checks
131 * are ones code can make with certainty, which is why they are made here
132 * and not put to a model.
133 */
134function refusalOf(
135 route: Route,
136 credentials: Credentials,
137 sizeRefusalOf: SizeCheck,
138): string | undefined {
139 const missing = missingOf(route.provider, credentials)
140
141 return missing.length > 0
142 ? `${missing.join(' and ')} unset`
143 : sizeRefusalOf(route)
144}
145
146/**
147 * Lists the routes that can take the job.
148 *
149 * Every eligible provider is considered with each model it offers; the
150 * configured provider with the configured model in place of its default. A
151 * route is left out when a credential it needs is unset or the fitted state
152 * exceeds its documented limit. A model family is offered once: over the
153 * route that reaches it directly whenever there is one, else over the
154 * first gateway to it in `PROVIDERS` order, and every other route to it is
155 * left out of the choice.
156 *
157 * @param configured the route the options name
158 * @param eligible the providers that may be sent anything, in `PROVIDERS`
159 * order
160 * @param credentials the resolved credentials
161 * @param sizeRefusalOf the check of the fitted state against a route's limits
162 * @returns the candidates, for every route not offered the reason, and
163 * whether the configured provider is among the routes that can take the job
164 */
165export function candidatesOf(
166 configured: Route,
167 eligible: readonly ProviderName[],
168 credentials: Credentials,
169 sizeRefusalOf: SizeCheck,
170): { candidates: Candidate[]; refused: string[]; isConfiguredUsable: boolean } {
171 const usable: Route[] = []
172 const refused: string[] = []
173
174 for (const provider of eligible) {
175 const offered = modelsOf(provider)
176 const models =
177 provider === configured.provider && !offered.includes(configured.model)
178 ? [configured.model]
179 : offered
180
181 for (const model of models) {
182 const route = { provider, model }
183 const refusal = refusalOf(route, credentials, sizeRefusalOf)
184
185 if (refusal === undefined) {
186 usable.push(route)
187 } else {
188 refused.push(`${labelOf(route)}: ${refusal}`)
189 }
190 }
191 }
192
193 const offered = new Map<string, Route>()
194
195 for (const route of [
196 ...usable.filter(route => gatewayOf(route.provider) === undefined),
197 ...usable.filter(route => gatewayOf(route.provider) !== undefined),
198 ]) {
199 if (!offered.has(familyOf(route))) {
200 offered.set(familyOf(route), route)
201 }
202 }
203
204 const candidates: Candidate[] = []
205
206 for (const route of usable) {
207 const same = offered.get(familyOf(route))
208
209 if (same !== undefined && same !== route) {
210 refused.push(`${labelOf(route)}: the same model as ${labelOf(same)}`)
211 continue
212 }
213
214 candidates.push({ key: keyOf(route), route, about: aboutOf(route) })
215 }
216
217 return {
218 candidates,
219 refused,
220 isConfiguredUsable: usable.some(
221 route => route.provider === configured.provider,
222 ),
223 }
224}
225
226function shareOf(part: number, whole: number): number {
227 return whole === 0 ? 0 : Math.round((part / whole) * 100) / 100
228}
229
230/**
231 * Describes the job for the routing question.
232 *
233 * @param messages the conversation
234 * @param calls the calls that will be asked about
235 * @param stateTokens the estimated size of the fitted state
236 * @param goal what the assistant is working on
237 * @returns the profile
238 */
239export function profileOf(
240 messages: readonly SessionMessage[],
241 calls: readonly Call[],
242 stateTokens: number,
243 goal: string,
244): Profile {
245 const counts = new Map<string, number>()
246
247 for (const call of calls) {
248 counts.set(call.tool, (counts.get(call.tool) ?? 0) + 1)
249 }
250
251 const ranked = [...counts].sort(([, a], [, b]) => b - a)
252 const tools = Object.fromEntries(ranked.slice(0, TOOLS_SHOWN))
253 const rest = ranked
254 .slice(TOOLS_SHOWN)
255 .reduce((sum, [, count]) => sum + count, 0)
256
257 if (rest > 0) {
258 tools.other = rest
259 }
260
261 let chars = 0
262 let nonAscii = 0
263
264 for (const { text } of messages) {
265 chars += text.length
266
267 for (let index = 0; index < text.length; index++) {
268 if (text.charCodeAt(index) > 127) {
269 nonAscii++
270 }
271 }
272 }
273
274 return {
275 context: PROFILE_CONTEXT,
276 goal,
277 messages: messages.length,
278 candidate_calls: calls.length,
279 state_tokens: stateTokens,
280 tools,
281 error_share: shareOf(
282 calls.filter(call => call.isError).length,
283 calls.length,
284 ),
285 non_ascii_share: shareOf(nonAscii, chars),
286 }
287}
288
289/**
290 * The one question the decision asks: which of the candidates should judge
291 * this job, each described by what is documented about it.
292 *
293 * @param candidates the routes on offer, two or more
294 * @returns the `choice` question
295 */
296function routeQuestionOf(
297 candidates: readonly Candidate[],
298): ChoiceQuestion {
299 return {
300 type: 'choice',
301 instructions:
302 'Which decision model should judge this compaction job? A wrong ' +
303 '"remove" loses output the assistant may still need, a wrong "keep" ' +
304 'only wastes context, and the answers are waited for while the ' +
305 'session is paused. Pick the model whose documented strengths fit ' +
306 'the job the state describes.',
307 criteria: Object.fromEntries(
308 candidates.map(candidate => [candidate.key, candidate.about]),
309 ),
310 }
311}
312
313/**
314 * Decides which route a compaction goes over.
315 *
316 * First by what code can check: `candidatesOf` leaves out every route that
317 * cannot take the job. Only when two or more remain is a model asked, with
318 * one `choice` question over the job's profile, on the fastest route there
319 * is: Cloudflare's Clef-flash when Cloudflare is eligible and configured,
320 * the configured route otherwise. The question carries a profile with the
321 * person's prompts in it, so it goes to no provider outside the eligible
322 * ones. With one candidate or a failed routing request the
323 * configured route is used, so the decision can never make a compaction
324 * fail that would have worked without it; the exception is a configured
325 * route that cannot take the job at all, where the one that can is used.
326 * The one candidate may be the direct route to the model a configured
327 * gateway also reaches: there is then one model and nothing to choose, and
328 * the configured gateway stays in use.
329 *
330 * With no candidate at all there is nothing to send the job to, the
331 * configured route included, so this throws, and the reason lists why each
332 * route was ruled out: the person has to see which credential or limit to
333 * change.
334 *
335 * @param configured the route the options name
336 * @param eligible the providers that may be sent anything, in `PROVIDERS`
337 * order
338 * @param credentials the resolved credentials
339 * @param sizeRefusalOf the check of the fitted state against a route's limits
340 * @param profile the job's profile
341 * @param askOn how a question is sent over a given route
342 * @returns the route and why it was picked; throws when no route can take
343 * the job
344 */
345export async function chooseRoute(
346 configured: Route,
347 eligible: readonly ProviderName[],
348 credentials: Credentials,
349 sizeRefusalOf: SizeCheck,
350 profile: Profile,
351 askOn: (route: Route) => Ask,
352): Promise<Routed> {
353 const { candidates, refused, isConfiguredUsable } = candidatesOf(
354 configured,
355 eligible,
356 credentials,
357 sizeRefusalOf,
358 )
359 const [first, second] = candidates
360 const left = refused.length === 0 ? '' : ` (${refused.join('; ')})`
361
362 if (first === undefined) {
363 throw new Error(`no route can take this job: ${refused.join('; ')}`)
364 }
365
366 if (second === undefined) {
367 if (!isConfiguredUsable) {
368 return {
369 route: first.route,
370 why: `only ${labelOf(first.route)} can take this job${left}`,
371 }
372 }
373
374 const found =
375 first.route.provider === configured.provider
376 ? 'one route is available'
377 : 'one model is available'
378
379 return {
380 route: configured,
381 why: `${found}${left}; using the configured ${labelOf(configured)}`,
382 }
383 }
384
385 const asked: Route =
386 eligible.includes('cloudflare') &&
387 missingOf('cloudflare', credentials).length === 0
388 ? { provider: 'cloudflare', model: 'clef-flash' }
389 : configured
390
391 try {
392 const reply = await askOn(asked)(profile, {
393 [ROUTE_QUESTION]: routeQuestionOf(candidates),
394 })
395 const picked = choiceOf(
396 reply,
397 ROUTE_QUESTION,
398 candidates.map(candidate => candidate.key),
399 )
400 const chosen = candidates.find(candidate => candidate.key === picked.choice)
401
402 if (chosen === undefined) {
403 throw new Error('the picked option is not a candidate')
404 }
405
406 const sure =
407 picked.confidence === undefined
408 ? ''
409 : `, confidence ${picked.confidence.toFixed(2)}`
410
411 return {
412 route: chosen.route,
413 why: `${labelOf(asked)} picked ${labelOf(chosen.route)}${sure}`,
414 }
415 } catch (error) {
416 // The reason may quote a provider's text, and it reaches the report
417 // line, so it is quoted like any other: redacted and on one short line.
418 const reason = briefOf(messageOf(error), credentials)
419
420 return {
421 route: configured,
422 why:
423 `the routing request failed (${reason}); using the configured ` +
424 labelOf(configured),
425 }
426 }
427}
428hooks/ask.ts 280 lines1import type { HttpInit, HttpResponse } from 'claude-code'
2
3import type { Ask } from './compact'
4import {
5 briefOf,
6 exchangeOf,
7 isRetryable,
8 messageOf,
9 replyFrom,
10} from './providers'
11import type { Credentials, Route } from './providers'
12
13/**
14 * What a request needs from the engine, handed in as functions because the
15 * engine interface itself cannot leave the hooks module's own file.
16 */
17export type Ports = {
18 /**
19 * `$.http.fetch`.
20 */
21 fetch: (url: string, init: HttpInit) => Promise<HttpResponse>
22 /**
23 * Waits before a retry. Resolves false, without waiting, when the hook has
24 * no time left to spend on it. `signal` aborts when the attempt the wait
25 * belongs to is over, and the wait must end with it then: a wait that
26 * outlived its attempt would hold the hook for nothing.
27 */
28 pause: (ms: number, signal: AbortSignal) => Promise<boolean>
29 /**
30 * `$.clock.after`: runs `fn` once after `ms` milliseconds unless the
31 * answer's `cancel` is called first.
32 */
33 after: (ms: number, fn: () => void) => { cancel: () => void }
34 /**
35 * `next.signal`: aborts when the compaction this hook serves is given up.
36 */
37 signal: AbortSignal
38}
39
40/**
41 * How long to wait before the second and the third attempt of one request.
42 * The hook's own code and its waits are charged to its ten seconds; the time
43 * a request is out is not. The two waits of one request stay under two
44 * seconds together, but that bound is per request, not per compaction: with
45 * `MAX_REQUESTS` batches going out `CONCURRENT_REQUESTS` at a time, as many
46 * rounds of batches as that makes (six, at sixteen and three) can each wait
47 * in turn. What bounds the waits of a compaction together is the `pause`
48 * port, which declines a wait when the hook has too little of its time left.
49 */
50const BACKOFF_MS = [400, 1200] as const
51
52/**
53 * How long one compaction may wait on its providers, from the moment its
54 * first request can go out: the routing question, every batch, every retry
55 * and every wait between retries together. A request has no timeout of its
56 * own, and the time it takes is not charged to the hook, so without a bound
57 * a provider that accepts a request and never answers would hold the
58 * compaction open for good, and the built-in summary would never get its
59 * turn. The bound is the longest the person waits on the providers, with the
60 * session paused, before that summary starts.
61 */
62const COMPACTION_DEADLINE_MS = 30_000
63
64/**
65 * The requests and waits of one compaction, under the one bound they share.
66 *
67 * The compaction is over, as far as its providers go, at the first of: its
68 * deadline passing, the engine giving the compaction up, `fail` being
69 * called, or `close` being called. From then on `fetch` and `pause` start
70 * nothing and reject with the reason. When it is over for any reason but
71 * `close`, whatever they had under way rejects with it at once; however it
72 * ends, a wait under way is aborted. A request that is out cannot be
73 * withdrawn, so its late answer is ignored.
74 */
75type Attempt = {
76 /**
77 * `Ports.fetch`, started only while the compaction is not over. A request
78 * the host refuses by throwing, instead of by rejecting, rejects all the
79 * same.
80 */
81 fetch: Ports['fetch']
82 /**
83 * `Ports.pause`, started only while the compaction is not over, with a
84 * signal that aborts when it is.
85 */
86 pause: (ms: number) => Promise<boolean>
87 /**
88 * Why the compaction is over, or undefined while it is not.
89 */
90 failure: () => Error | undefined
91 /**
92 * Ends the compaction with a reason. Only the first reason counts.
93 */
94 fail: (reason: Error) => void
95 /**
96 * Ends the compaction, when nothing ended it before, and releases the
97 * deadline's timer and the listener on the engine's signal. Called once
98 * the compaction has its outcome, whichever that is: a request still out
99 * then has nobody to answer, and must not retry or wait.
100 */
101 close: () => void
102}
103
104/**
105 * Puts one compaction's requests under one deadline, armed here and nowhere
106 * else, and under the engine's signal.
107 *
108 * @param ports the engine calls
109 * @returns the attempt; the caller closes it when the compaction is decided
110 */
111export function attemptOf(ports: Ports): Attempt {
112 let failure: Error | undefined
113 let disarm = () => {}
114 let reject: (reason: Error) => void = () => {}
115 // Aborted when the attempt is over, however it ends: the waits listen to
116 // it, so that none of them outlives the attempt.
117 const ended = new AbortController()
118
119 const cut = new Promise<never>((_resolve, rejectCut) => {
120 reject = rejectCut
121 })
122
123 // An end that comes while nothing is under way has nobody to reach, and
124 // must not surface as a rejection nobody handled.
125 cut.catch(() => {})
126
127 const fail = (reason: Error) => {
128 if (failure !== undefined) {
129 return
130 }
131
132 failure = reason
133 disarm()
134 reject(reason)
135 ended.abort(reason)
136 }
137
138 const interrupted = () =>
139 new Error(
140 'the compaction was interrupted before its requests were answered',
141 )
142
143 if (ports.signal.aborted) {
144 fail(interrupted())
145 } else {
146 const onAbort = () => fail(interrupted())
147
148 // Neither is left set up without the other: the listener goes on first
149 // and comes off again when the timer cannot be armed, so an attempt that
150 // could not be made leaves nothing on the engine's signal or its clock.
151 ports.signal.addEventListener('abort', onAbort, { once: true })
152
153 let timer: { cancel: () => void }
154
155 try {
156 timer = ports.after(COMPACTION_DEADLINE_MS, () =>
157 fail(
158 new Error(
159 'the requests of this compaction were not all answered within ' +
160 `${COMPACTION_DEADLINE_MS / 1000} seconds`,
161 ),
162 ),
163 )
164 } catch (error) {
165 ports.signal.removeEventListener('abort', onAbort)
166
167 throw error
168 }
169
170 disarm = () => {
171 disarm = () => {}
172 timer.cancel()
173 ports.signal.removeEventListener('abort', onAbort)
174 }
175 }
176
177 // The executor turns a `start` that throws into a rejection, so a call
178 // refused on the spot takes the same path as one refused later.
179 const within = <Value>(start: () => Promise<Value>): Promise<Value> =>
180 failure === undefined
181 ? Promise.race([new Promise<Value>(resolve => resolve(start())), cut])
182 : Promise.reject(failure)
183
184 return {
185 fetch: (url, init) => within(() => ports.fetch(url, init)),
186 pause: ms => within(() => ports.pause(ms, ended.signal)),
187 failure: () => failure,
188 fail,
189 // The shared cut is not rejected here: whatever is still out when the
190 // outcome is known is waited on by nobody and is left to settle on its
191 // own. Only what it would start from here on is refused.
192 close: () => {
193 failure ??= new Error('the compaction already has its outcome')
194 disarm()
195 ended.abort(failure)
196 },
197 }
198}
199
200/**
201 * Builds the function that sends one set of questions over a route.
202 *
203 * A response that may succeed on another attempt (rate limiting, an
204 * overloaded or failing server) is retried after a short wait, at most
205 * twice; any other failure is final at once. Only this one route is ever
206 * tried: there is no falling over to another provider.
207 *
208 * A failure of a decisive request ends the attempt on the spot, in the same
209 * step that found it, so that no other request of the compaction retries or
210 * waits after it. Once the attempt is over, every request fails with the
211 * attempt's reason, whatever its own would have been.
212 *
213 * What the host says when it refuses a call is quoted redacted and short:
214 * it may repeat the request's headers.
215 *
216 * @param route the provider and model
217 * @param credentials the resolved credentials
218 * @param attempt the compaction's requests and their shared bound
219 * @param isDecisive true for a request the compaction cannot do without (a
220 * batch of questions about calls), false for one it can (the routing
221 * question)
222 * @returns the function `compact` and the routing decision ask through
223 */
224export function askerOf(
225 route: Route,
226 credentials: Credentials,
227 attempt: Attempt,
228 isDecisive: boolean,
229): Ask {
230 const hosted = <Value>(what: string, call: Promise<Value>): Promise<Value> =>
231 call.catch(error => {
232 throw (
233 attempt.failure() ??
234 new Error(`${what} failed: ${briefOf(messageOf(error), credentials)}`)
235 )
236 })
237
238 return async (state, questions) => {
239 try {
240 const exchange = exchangeOf(route, credentials, state, questions)
241 const send = () =>
242 hosted(
243 `the request to ${route.provider}`,
244 attempt.fetch(exchange.url, exchange.init),
245 )
246
247 let response = await send()
248
249 for (const wait of BACKOFF_MS) {
250 if (response.ok || !isRetryable(response.status)) {
251 break
252 }
253
254 const waited = hosted(
255 `the wait before asking ${route.provider} again`,
256 attempt.pause(wait),
257 )
258
259 if (!(await waited)) {
260 break
261 }
262
263 response = await send()
264 }
265
266 return replyFrom(route, response, credentials, questions)
267 } catch (error) {
268 const reason =
269 attempt.failure() ??
270 (error instanceof Error ? error : new Error(String(error)))
271
272 if (isDecisive) {
273 attempt.fail(reason)
274 }
275
276 throw reason
277 }
278 }
279}
280