SLOPSHOPPER

decision-compaction

Hands back the conversation itself in place of a compaction summary, less the tool calls and results a System One decision model judges no longer needed. The…

newtoastnetworktimer
★ 1v0.1.0Apache-2.0updated 2026-10-09zchee/claude-code-mods/mods/decision-compaction
A shopper browsing a rack in a slop shop
README

decision-compaction

Compaction without a summary. A Claude Code mod for session.compact that hands the conversation back as it was, with only those tool calls and tool outputs removed that a System One decision model says the assistant can do without. What you and the assistant wrote stays word for word.

License: Apache-2.0 Plugin API: Claude Code 2.1.292 Providers: 8

[!WARNING] The conversation text and the tool inputs are sent to the selected third-party API (TypeSafe, Cloudflare, OpenRouter, Codiv, Perplexity, decisions-api.dev, decisionapi.net or OpenAI) on every compaction. They are sent as they are: a secret you pasted into a prompt, or one that appears in a tool input such as a shell command, goes with them. Tool outputs are not sent: each is replaced by a note of its status and length.

With providerDecision on, a second provider can receive data in the same compaction. Every provider in decisionProviders (all eight by default) whose credentials resolve is a candidate to decide it, so a key you exported for another tool, OPENAI_API_KEY for example, makes that vendor eligible to receive the conversation. decisionProviders is the setting that narrows the set. The routing question, when there is a choice to make, goes to Cloudflare only when cloudflare is one of the providers considered (the configured provider, and those in decisionProviders) and its credentials are set; otherwise it goes to the configured provider. It carries a profile of the job that includes up to 2,000 characters of your prompts (the text typed after /compact, or your last three prompts).

Do not enable this mod for a session whose content must not leave your machine.

Table of contents

Install

From the marketplace, on Claude Code 2.1.275 or later:

claude plugin install decision-compaction --marketplace zchee/claude-code-mods

From a checkout, without installing:

claude --plugin-dir mods/decision-compaction

Then set a credential for the provider you want (see Providers), for example TYPESAFE_API_KEY in the environment, or the typesafeApiKey option in /config.

Options are read from settings under pluginConfigs and appear in /config. A mod loaded with --plugin-dir is keyed decision-compaction@inline (the bare decision-compaction is read too); an installed one is keyed by its plugin id, decision-compaction@<marketplace>.

How it works

On session.compact

On session.compact (a /compact, the engine's own threshold, or this mod's trigger) the hook:

  1. Matches each tool call to its output by tool_use_id. The first message is never changed, and neither is any of the last preserveRecentMessages; a call that sits in one of those, or whose output does, is not judged. A tool_use_id that occurs on more than one call or output is not judged either, since a decision is applied by that id.
  2. Describes the whole conversation in one state and shrinks it only as far as its budget requires. The reductions are tried in a fixed order and the first that is enough ends the search: shorter tool inputs, the middle cut out of long texts (old messages before the protected ones), old texts replaced by their length, old calls written on one line, old messages without a call omitted, and neighbouring old messages that hold nothing but calls merged into one entry. For a provider that caps the request body in bytes, the same reductions go on until the state also fits that cap with room left for one request's questions. Sizes are estimates; no tokenizer is used.
  3. Puts two yes/no questions to the provider for every call left to judge: is the call itself still needed, and is its complete output still needed word for word. A request holds as many of these questions as the token budget, the provider's own limit on questions and, for a provider that caps the body in bytes, that cap allow, and each request carries the full state. See Limits for the bounds.
  4. Compares each answer with keepThreshold. Either the call and its output both stay; or the call stays and the output is shortened to its opening truncateHeadChars characters, followed by a note saying how much was removed; or the call goes and its output goes with it. Output so short that shortening would save nothing is not touched.
  5. Hands back the new list of messages. A message nothing touched goes back exactly as the engine gave it, and an output never outlives its call.

A message that had one of its blocks dropped or cut is rebuilt from its role, its text and its remaining tool blocks, and nothing else survives the rebuild. Two consequences: a result that is kept, but shares its message with a result that was dropped or cut, loses any image content it had; and an assistant message that made several calls, one of which was dropped, loses its thinking block.

It fails open. On any error (a missing key, a provider option that names no provider, an HTTP error, a request that is rejected, a compaction whose requests are not answered in time, a malformed answer, a conversation that cannot be fitted or would take too many requests) and when the decisions would remove less than minReductionRatio, it says why in one line and lets the built-in summary run. A precompute compaction is always left to the engine.

Every request that asks about a tool call goes to one provider: the configured one, or the one the provider decision picked. Nothing falls over to another provider. With providerDecision on there can be one more request before them, the routing question, and it may go to a different provider (see providerDecision).

On turn.complete

On turn.complete, when the main conversation's context reaches compactAtPercent, the mod requests a compaction. What happens next depends on the usage that compaction leaves:

  • Under the threshold: the next time usage reaches the threshold, another compaction is requested at once.
  • At or above the threshold: that level becomes a floor, and no further compaction is requested until usage has risen ten points above it. A floor above 90% cannot be risen from, so after a compaction that leaves usage there, no automatic compaction is requested again until usage has first dropped under the threshold (for example after a /compact of your own or a /clear).
  • Unknown: usage is not known right after a compaction (it arrives with the next response). The first later reading at or above the threshold then becomes the floor, and the next request waits for a rise of ten points above that reading.

A turn you interrupted, and a subagent's turn, request nothing.

Providers

providerEndpointModel by defaultCredentialsDocumented limits
typesafe (default)https://api.typesafe.ai/v1/systemonejev-latestTYPESAFE_API_KEY64,000 tokens a request; 32,000 for the state plus the longest question
cloudflarehttps://api.cloudflare.com/client/v4/accounts/<account id>/ai/run/@cf/cloudflare/<model>clef (clef-flash is the other)CLOUDFLARE_API_TOKEN and CLOUDFLARE_ACCOUNT_ID65,536 tokens; 64 questions a request
openrouterhttps://openrouter.ai/api/v1/systemone~typesafe/jev-latestOPENROUTER_API_KEY32,000 tokens for the state plus the questions; 64 questions a request for any model other than Jev
codivhttps://api.codiv.ai/v1/systemoneopenjev-latest (OpenJev)CODIV_API_KEY65,536 tokens for the state and the questions together (the vendor advises a state of about 60,000); no limit on questions
perplexityhttps://api.perplexity.ai/v1/decisionspplx-decider-v1.1-27bPERPLEXITY_API_KEYunder 262,144 tokens a request; 128 questions a request
decisions-api-devhttps://decisions-api.dev/v1/systemonejev-latest (Jev, through this gateway)DECISIONS_API_KEY32 KiB a request body, counted in bytes, so text that is not ASCII uses more of it; 8 questions a request
decisionapi-nethttps://decisionapi.net/v1/systemonejev-latest (Jev, through this gateway)DECISIONAPI_API_KEY32 KiB a request body, counted in bytes; 8 questions a request
openaihttps://api.openai.com/v1/decisionsgpt-6-lunaOPENAI_API_KEYnot documented for the Decisions API; the mod holds a request to 32,000 tokens and 64 questions

The provider option takes exactly these eight names. Any other value is not replaced by the default: nothing is sent anywhere, and every compaction uses the built-in summary with a line naming the value. Only an option left unset or empty means typesafe. For that reason /config shows the option as a text field rather than a list (a list would let the engine turn an unknown value into the default without saying so).

What differs between the providers beyond the table:

  • openrouter, decisions-api-dev and decisionapi-net are gateways to TypeSafe's Jev. codiv's OpenJev is a model of its own, not TypeSafe's Jev: at Codiv jev-latest is an alias of OpenJev.
  • perplexity accepts only its own decider models. Any other name in model, jev-latest included, is refused with an HTTP 400 and the compaction uses the built-in summary.
  • decisions-api-dev and decisionapi-net wrap the answer in a {code, message, data} envelope. A code other than 0, or an envelope without data.result, ends the compaction with the vendor's message quoted, as for any other failed request.
  • decisions-api-dev accepts a question id only if it starts with a letter, uses letters, digits, _ and -, and is at most 64 characters. The mod's ids can hold . and run longer, so it names each question in a request by its place (q0, q1, …) and maps the answers back to the ids. An answer under a name the mod did not send is refused.
  • openai takes a different request: the state goes as one JSON string in input and the questions as a list, each a predicate (yes/no) or a choice. The mod translates the request into that form and the answers back. OpenAI counts about 148 input tokens for each question (Codiv about 26 for the same one-line question), and the mod's size estimate adds that figure per question. If OpenAI refuses to answer a question (the line reads openai: refused to answer <question id>), or answers in a shape the mod does not read, the compaction ends with the built-in summary and the line names the question or the field.

~typesafe/jev-latest is OpenRouter's alias for the newest Jev; typesafe/jev-1.13 pins a release. OpenRouter also takes TypeSafe's bare ids and maps them itself (jev-latest to ~typesafe/jev-latest, jev-1.13 to typesafe/jev-1.13). An id that already has an author prefix is used as written.

OpenRouter also serves Cloudflare's Clef (cloudflare/clef, cloudflare/clef-flash), and passes such a request on to Cloudflare, which refuses one with more than 64 questions. A model other than Jev on openrouter is therefore asked at most 64 questions (32 tool calls) a request, as on cloudflare, while the token limits stay OpenRouter's.

Credentials

Each credential is read from its plugin option first, then from the environment, at the time of the compaction. The env block of the settings is read only when a credential a provider in play needs is still missing; if the settings cannot be read, that is logged and the compaction goes on with what it has. With providerDecision on and decisionProviders at its default of all eight providers, some key is usually unset, so the settings are read on most compactions.

Redaction

Text that comes from a provider or from the host (an HTTP error body, a Cloudflare failure envelope, the reason the host gives for refusing a request or a read) is quoted in the toast and the log on one line of at most 160 characters. Before it is shown, every credential this mod resolved is replaced with [redacted], and so is the token after the word Bearer, whether it follows a space, a colon or an equals sign. A resolved value shorter than eight characters is not replaced where it stands on its own, since no provider issues a key that short and replacing it would garble ordinary words.

Options

OptionDefaultMeaning
providertypesafeWhich provider decides; any other name sends nothing.
modelprovider defaultModel name for the configured provider.
providerDecisionfalseDecide the provider per compaction; see providerDecision.
decisionProvidersall eight providersWith providerDecision on, the providers that may be compared and so may receive the conversation: a list of provider names, or one text of names separated by commas. The configured provider is always considered. See providerDecision.
typesafeApiKeyunsetTypeSafe API key (stored as a secret).
openrouterApiKeyunsetOpenRouter API key (stored as a secret).
cloudflareApiTokenunsetCloudflare API token allowed to run Workers AI (stored as a secret).
cloudflareAccountIdunsetCloudflare account ID.
codivApiKeyunsetCodiv API key (stored as a secret).
perplexityApiKeyunsetPerplexity API key (stored as a secret).
decisionsApiKeyunsetdecisions-api.dev API key (stored as a secret).
decisionapiApiKeyunsetdecisionapi.net API key (stored as a secret).
openaiApiKeyunsetOpenAI API key (stored as a secret).
keepThreshold0.5Probability from which a call, or its full result, is kept.
preserveRecentMessages6Newest messages that are never changed.
compactAtPercent60Context usage at which a compaction is requested; 0 turns the trigger off.
minReductionRatio0.25Smallest share of characters the decisions must remove.
maxStateTokens25000Most estimated tokens the state may take, a whole number; see Limits.
maxRequestTokens30000Most estimated tokens one request may take, state and questions together, a whole number; see Limits.
truncateHeadChars300Characters kept of a shortened output; at 0 only the note remains.

Limits

The bounds below are fixed; they are not options.

  • Margin under the documented limits. The token estimate is made without a tokenizer and its weights were fitted to Jev, and Cloudflare cuts a state that is too long without saying so. A budget is therefore held to 85% of the provider's documented limit: 27,200 tokens for the state plus the longest question and 54,400 a request on typesafe, 55,705 on cloudflare, 27,200 on openrouter and openai, 51,000 for the state plus the longest question (85% of the 60,000 the vendor advises) and 55,705 a request on codiv, 222,821 on perplexity, and 27,200 for the state plus the longest question and 54,400 a request on decisions-api-dev and decisionapi-net. Both serve TypeSafe's Jev 1.13 and state their own limit in bytes, so they are held to the token window TypeSafe documents for Jev, and the byte cap below is the bound that decides. openai documents no limit, so it is held to the tightest figures of the others, 32,000 tokens and 64 questions, before the margin. Where the margin leaves less than the default maxRequestTokens of 30,000 (openrouter and openai), the default is lowered there.
  • A byte cap. decisions-api-dev and decisionapi-net also hold a request to the 85% share of their 32 KiB cap, 27,852 of 32,768 bytes, measured as UTF-8 of the encoded request body, since a body over the cap is refused whole. When one of them is the configured provider the state is fitted to that cap as well as to its tokens: the reductions above go on until the body, with the questions of one request set aside (4 tool calls, all that 8 questions hold), fits in 27,852 bytes. English prose takes more bytes for each estimated token than the token budget allows for, and text that is not ASCII more still, so on these two providers it is usually the bytes that decide how much of the conversation the model sees. Questions are counted at or above the bytes they are sent in (decisions-api-dev counts each under the longest name it may send, q7). A conversation still over the cap after every reduction is left to the built-in summary, and the line says how many bytes were left. With providerDecision on, the state is fitted for the configured provider, so one of these two offered as another route may have no room for a question within the cap; it is then ruled out, and the line says how many of the 27,852 bytes the request takes before any question.
  • A reply that shows a cut. If a provider reports as many input tokens as the request limit the mod holds for it (before the margin), the state was probably cut short. The answers are discarded and the built-in summary runs. On decisions-api-dev and decisionapi-net that limit is Jev's own 64,000 tokens, which a body within 32 KiB does not reach. On openai it is the mod's fallback of 32,000, not a window OpenAI states, so a reply counting that many is discarded even if the model read the state whole.
  • Room for a batch. The state may take at most the request budget less the questions of 16 tool calls (of all of them, when there are fewer). Every request carries the whole state, so without this a state that filled the request would be sent once per call.
  • At most 16 requests ask about tool calls in one compaction, three at a time. A conversation that would need a 17th is left to the built-in summary, and the line says how many it would have taken. The routing question of providerDecision is one request more.
  • Retries. A 429 or a 5xx is retried after 0.4 s and again after 1.2 s; nothing else is. A retry is skipped when the hook has too little of its ten seconds left to wait and still finish (the wait plus 1.5 s); the last 429 or 5xx then decides, as a failure.
  • One failure ends the compaction. Once a request about tool calls has failed for good (an answer that is not retried, or the last retry), no request is sent and no retry is waited for after it, by any batch; the requests still out are abandoned. A failed routing question is the one exception: the configured provider then decides, as described under providerDecision.
  • 30 seconds for the whole compaction. One deadline covers everything a compaction sends: the routing question, every batch, every retry and the waits before retries. It starts when the first request can go out. When it passes, every request still out is abandoned, nothing further is sent, and the built-in summary runs; that is the longest the session stays paused waiting on the providers. A request cannot be withdrawn, so a late answer is ignored.

How large a conversation can be decided

At the default budgets the state holds about 25,000 estimated tokens after every reduction, and 16 requests ask about a bounded number of calls. Past either bound the result is always the built-in summary.

The table is a measurement, not a guarantee. It was taken with a fake provider on a synthetic session made only of Read calls with short file paths and 600-character outputs, between one opening and one closing prompt. A real session has longer inputs, more varied tools and more text, so its limits are lower, and differ from one session to the next; read the numbers as rough orders of size.

providerCalls whose assistant message also carries a sentence of textCall-only messages
typesaferoughly 400 (then the state no longer fits)roughly 900 (then a 17th request)
cloudflareroughly 400 (then the state no longer fits)512 (32 calls a request)
openrouterroughly 290 (then a 17th request)roughly 820 (then a 17th request)

Longer tool inputs and longer tool names lower these numbers. Raising maxStateTokens and maxRequestTokens raises them, up to the margins above.

The table has no rows for codiv, perplexity, decisions-api-dev, decisionapi-net and openai; it was not measured on them. decisions-api-dev and decisionapi-net take 8 questions and 32 KiB a request, so of a long conversation the model sees only what is left once the state is shrunk to the cap. The cap counts bytes, so text that is not ASCII uses more of it. Eight questions are the questions of 4 tool calls, so a compaction asks about at most 64 calls over its 16 requests.

providerDecision

Off by default: the configured provider is always used, and no other provider is sent anything.

When on, the providers considered are the configured one and those in decisionProviders. Unset, that option is every provider, all eight. It takes a list of names, or one text of names separated by commas; a cleared field (an empty text) or an empty list names none, so only the configured provider is considered. A name that is no provider is dropped and named in one line of the transcript; the compaction goes on, because dropping a name can only narrow where data goes. A provider that is not considered is sent nothing, even when its credentials are set. A provider that is considered is a candidate as soon as its credentials resolve, so a key exported for another tool makes that vendor eligible. To keep the choice to the original three providers, set decisionProviders to typesafe,cloudflare,openrouter.

The mod first drops every route it can rule out in code: a provider whose credentials are unset, whose limits the fitted state exceeds, or that would need more than 16 requests. Cloudflare contributes two routes (clef and clef-flash). Among routes to the same model one is offered: the direct route when it is usable, otherwise the first gateway in the order openrouter, decisions-api-dev, decisionapi-net. Jev is reached directly through typesafe and through those three gateways; on OpenRouter an id that names Jev (jev, jev-…, alone or under typesafe/ or ~typesafe/) is Jev. Every other route to that model is left out, with the reason "the same model as" the one offered. Any other OpenRouter id, under typesafe/ included, is a model of its own. If two or more routes remain, one choice question is asked over a small profile of the job (message and call counts, state size, tool mix, error share, share of non-ASCII text, and the goal), not over the conversation. The goal is your own text: what you typed after /compact, cut to 2,000 characters, or your last three prompts, cut to 500 characters each. The question is asked on Cloudflare clef-flash only when cloudflare is one of the providers considered and its credentials are set, otherwise on the configured provider. A decisionProviders list wi

Source 12 files
hooks/register.ts 295 lines
1import type { EngineInterface, PluginOptions, Register } from 'claude-code'
2
3import { reductionOf } from './compact'
4import { configOf, eligibleOf } from './config'
5import {
6  briefOf,
7  credentialsFor,
8  credentialsOf,
9  cutOf,
10  environmentOf,
11  messageOf,
12  PROVIDERS,
13  redacted,
14  unresolvedOf,
15} from './providers'
16import type { Credentials, Environment, ProviderName } from './providers'
17import { decisionLinesOf, percentOf, summaryOf } from './report'
18import { run } from './run'
19import { armOf, isCompactionDue, settle } from './trigger'
20
21/**
22 * How long an outcome toast stays: long enough to read one full line.
23 */
24const TOAST_MS = 15_000
25
26/**
27 * The hook time that must remain after a retry's wait for the work that
28 * follows it: reading the answers and rebuilding the messages.
29 */
30const PAUSE_RESERVE_MS = 1500
31
32/**
33 * Says which names of the `decisionProviders` option are no provider. Each
34 * is quoted, and quoted only in part when long, as the `provider` option
35 * is: the field may hold anything that was pasted into it.
36 */
37function unknownLineOf(names: readonly string[]): string {
38  const quoted = names.map(name => cutOf(JSON.stringify(name)))
39
40  return (
41    `the decisionProviders option names ${quoted.join(', ')}, ` +
42    `${names.length === 1 ? 'which is' : 'which are'} none of ` +
43    `${PROVIDERS.join(', ')}; ignored`
44  )
45}
46
47/**
48 * Reads the credential variables from the process environment, and from the
49 * `env` block of the settings only while a credential one of the providers
50 * needs is still without a value: the settings are a second source, not a
51 * part of every compaction. Each name is spelled out because `$.env.get`
52 * takes a string literal only.
53 *
54 * Settings that cannot be read are no settings. The compaction goes on with
55 * what the options and the environment gave, and `problem` says what went
56 * wrong so that it can be logged.
57 */
58async function readEnvironment(
59  $: EngineInterface,
60  options: PluginOptions,
61  providers: readonly ProviderName[],
62): Promise<{ environment: Environment; problem?: string }> {
63  const read: Environment = {
64    TYPESAFE_API_KEY: await $.env.get('TYPESAFE_API_KEY'),
65    OPENROUTER_API_KEY: await $.env.get('OPENROUTER_API_KEY'),
66    CLOUDFLARE_API_TOKEN: await $.env.get('CLOUDFLARE_API_TOKEN'),
67    CLOUDFLARE_ACCOUNT_ID: await $.env.get('CLOUDFLARE_ACCOUNT_ID'),
68    CODIV_API_KEY: await $.env.get('CODIV_API_KEY'),
69    PERPLEXITY_API_KEY: await $.env.get('PERPLEXITY_API_KEY'),
70    DECISIONS_API_KEY: await $.env.get('DECISIONS_API_KEY'),
71    DECISIONAPI_API_KEY: await $.env.get('DECISIONAPI_API_KEY'),
72    OPENAI_API_KEY: await $.env.get('OPENAI_API_KEY'),
73  }
74
75  if (unresolvedOf(providers, options, read).length === 0) {
76    return { environment: environmentOf(read, undefined) }
77  }
78
79  try {
80    const settings = await $.settings.read()
81
82    return { environment: environmentOf(read, settings.env) }
83  } catch (error) {
84    return {
85      environment: environmentOf(read, undefined),
86      problem: messageOf(error),
87    }
88  }
89}
90
91/**
92 * Writes one line to the transcript. Every line this module shows goes
93 * through here or through `tell`, and so through `redacted`: whatever a
94 * line was built from, no credential leaves in it.
95 */
96function log($: EngineInterface, text: string, credentials: Credentials): void {
97  $.ui.log(redacted(text, credentials))
98}
99
100/**
101 * Says one line both where it stays (the transcript) and where it is seen
102 * at once (a toast), redacted like every other line.
103 */
104function tell(
105  $: EngineInterface,
106  text: string,
107  credentials: Credentials,
108): void {
109  const shown = redacted(text, credentials)
110
111  $.ui.log(shown)
112  $.ui.toast(shown, { timeoutMs: TOAST_MS })
113}
114
115/**
116 * Registers the two hooks.
117 *
118 * `session.compact` hands back the conversation itself in place of a
119 * summary, less the tool calls and results a decision model judged no
120 * longer needed. It
121 * fails open: on any error, and when too little would be removed, it says
122 * why in one line and hands the compaction on to the built-in summary.
123 *
124 * `turn.complete` requests a compaction once the main conversation's
125 * context reaches `compactAtPercent`, one request at a time, and not again
126 * while usage stays near what a compaction that could not get under the
127 * threshold left.
128 *
129 * @param on the engine's registrar
130 * @param options the plugin's options
131 */
132export const register: Register = (on, options) => {
133  const config = configOf(options)
134  const providers: readonly ProviderName[] = eligibleOf(config)
135
136  on('session.compact', async ($, e, next) => {
137    // A precompute installs nothing and its result is used only if the
138    // conversation is still the same when a real compaction comes. Deciding
139    // then would send the conversation to a third party, and pay for it,
140    // for a result that may be thrown away; the engine's own precompute is
141    // left to stand as the fallback this hook fails open to.
142    if (e.trigger === 'precompute') {
143      return next(e)
144    }
145
146    // Held outside the `try` so that the line written when it fails is
147    // redacted with whatever had been resolved by then.
148    let credentials: Credentials = {}
149    let result
150
151    try {
152      // Options that name no provider the person chose are acted on in no
153      // way: not even a credential is read for them.
154      if (config.refusal !== undefined) {
155        throw new Error(config.refusal)
156      }
157
158      // Credentials are read here, for this compaction, so a key set or
159      // rotated after the module loaded is the one that is used. What the
160      // host says when it refuses a read is its own text, so it is quoted
161      // short like a provider's.
162      const { environment, problem } = await readEnvironment(
163        $,
164        options,
165        providers,
166      ).catch(error => {
167        throw new Error(
168          `the environment could not be read: ${briefOf(messageOf(error), {})}`,
169        )
170      })
171
172      credentials = credentialsOf(options, environment)
173
174      if (config.unknownProviders !== undefined) {
175        log($, unknownLineOf(config.unknownProviders), credentials)
176      }
177
178      if (problem !== undefined) {
179        log(
180          $,
181          `the settings could not be read (${briefOf(problem, credentials)}); ` +
182            'credentials are taken from the options and the environment alone',
183          credentials,
184        )
185      }
186
187      result = await run({
188        messages: e.messages,
189        instructions: e.instructions,
190        config,
191        // Only the providers the person allows may be sent anything, so
192        // only their credentials go on. Every line shown is still redacted
193        // with all that was resolved: a key of another provider is no less
194        // a secret.
195        credentials: credentialsFor(providers, credentials),
196        ports: {
197          fetch: (url, init) => $.http.fetch(url, init),
198          pause: async (ms, signal) => {
199            if (next.budget.remainingMs < ms + PAUSE_RESERVE_MS) {
200              return false
201            }
202
203            await $.clock.sleep(ms, {
204              signal: AbortSignal.any([next.signal, signal]),
205            })
206
207            return true
208          },
209          after: (ms, fn) => $.clock.after(ms, fn),
210          signal: next.signal,
211        },
212      })
213    } catch (error) {
214      tell($, `built-in summary used instead: ${messageOf(error)}`, credentials)
215
216      return next(e)
217    }
218
219    if (result.routing !== undefined) {
220      log($, `provider decision: ${result.routing}`, credentials)
221    }
222
223    for (const line of decisionLinesOf(result.outcome.decisions)) {
224      log($, line, credentials)
225    }
226
227    const reduction = reductionOf(result.outcome)
228
229    if (reduction <= 0 || reduction < config.minReductionRatio) {
230      tell(
231        $,
232        'built-in summary used instead: the reduction is below the ' +
233          `${percentOf(config.minReductionRatio)} minimum (${summaryOf(result)})`,
234        credentials,
235      )
236
237      return next(e)
238    }
239
240    tell(
241      $,
242      `compacted without a summary: ${result.outcome.messages.length} of ` +
243        `${e.messages.length} messages remain (${summaryOf(result)})`,
244      credentials,
245    )
246
247    return { messages: result.outcome.messages }
248  })
249
250  const arm = armOf()
251  let isCompacting = false
252
253  on('turn.complete', async ($, e, next) => {
254    // A turn the person interrupted is not a moment to compact: they stopped
255    // the session to say something, and a compaction would hold it paused.
256    const isDue =
257      config.compactAtPercent > 0 &&
258      e.agentId === undefined &&
259      !e.isAborted &&
260      !isCompacting
261
262    if (isDue) {
263      isCompacting = true
264
265      try {
266        const { context } = await $.session.usage()
267
268        if (isCompactionDue(arm, context.percent, config.compactAtPercent)) {
269          await $.session.compact()
270
271          settle(
272            arm,
273            (await $.session.usage()).context.percent,
274            config.compactAtPercent,
275          )
276        }
277      } catch (error) {
278        // No credential is resolved on this path; the reason is the host's
279        // own text, so it is still quoted short and held to the rule that
280        // nothing after the word `Bearer` is shown.
281        log(
282          $,
283          `compaction at ${config.compactAtPercent}% of the context was ` +
284            `not run: ${briefOf(messageOf(error), {})}`,
285          {},
286        )
287      } finally {
288        isCompacting = false
289      }
290    }
291
292    return next(e)
293  })
294}
295
hooks/compact.ts 845 lines
1import type {
2  SessionMessage,
3  ToolResultSummary,
4  ToolUseSummary,
5} from 'claude-code'
6
7import type { Limits } from './providers'
8import { estimatedTokensOf, headOf } from './state'
9import type { Call, Fitted, Stage } from './state'
10import { bytesOf, noulOf } from './systemone'
11import type { Questions, Reply, Wire } from './systemone'
12
13/**
14 * Sends one batch of questions about a state and resolves with the reply.
15 */
16export type Ask = (state: unknown, questions: Questions) => Promise<Reply>
17
18/**
19 * What one request may hold, in estimated tokens: the person's budgets,
20 * held under what the provider documents, and the state's share of the
21 * request held to what a batch of questions leaves.
22 */
23export type Budget = {
24  stateTokens: number
25  requestTokens: number
26  questions?: number
27  /**
28   * For a provider that caps the request body in bytes: the bytes a body
29   * may take, already held to `LIMIT_SHARE` of the cap, and the bytes it
30   * takes besides its questions.
31   */
32  bytes?: { limit: number; state: number }
33}
34
35/**
36 * The two probabilities answered for one call.
37 */
38export type Scores = {
39  keepCall: number
40  keepResult: number
41}
42
43/**
44 * What happens to a call: `pinned` and `keep` leave it whole, `truncate`
45 * keeps the call and the head of its result, `drop` removes both.
46 */
47export type Action = 'pinned' | 'keep' | 'truncate' | 'drop'
48
49export type Decision = Scores & {
50  call: Call
51  action: Action
52}
53
54/**
55 * What the batches reported between them.
56 */
57type Reported = {
58  /**
59   * The largest `input_tokens` any batch reported. Every batch sends the
60   * same state, so the largest is the one to hold against the estimate.
61   */
62  inputTokens?: number
63  /**
64   * The costs the batches reported, summed.
65   */
66  cost?: number
67}
68
69/**
70 * A finished compaction: the conversation as it should read afterwards and
71 * the figures the report is written from.
72 */
73export type Outcome = {
74  messages: SessionMessage[]
75  decisions: Decision[]
76  charsBefore: number
77  charsAfter: number
78  stateTokens: number
79  /**
80   * The reduction the state needed; absent from an outcome that asked a
81   * provider nothing.
82   */
83  stage?: Stage
84  requests: number
85  reported: Reported
86  /**
87   * How many results were in fact cut to their head. It can be fewer than
88   * the `truncate` decisions: a result too short to gain from a cut is left
89   * whole, and its call then counts with the ones kept.
90   */
91  shortened: number
92}
93
94/**
95 * What the questions about a set of calls need of a request, whichever
96 * provider it goes to.
97 */
98export type Demand = {
99  /**
100   * The estimated size of the longest single question: what a limit stated
101   * for "the state plus the longest question" has to leave free.
102   */
103  longestQuestion: number
104  /**
105   * How many calls a request must at least be able to ask about.
106   */
107  batchCalls: number
108  /**
109   * The estimated size of the questions of that many calls, taking the
110   * calls whose questions are longest.
111   */
112  batchTokens: number
113  /**
114   * The UTF-8 bytes the questions of one request may take: those of the
115   * calls whose questions take the most bytes, `batchCalls` of them or as
116   * many as the provider's question cap lets one request hold, whichever
117   * is fewer. What a cap on the body in bytes has to leave free.
118   */
119  batchBytes: number
120}
121
122/**
123 * The questions asked about each call.
124 */
125const QUESTIONS_PER_CALL = 2
126
127/**
128 * What a request costs around its state and questions: the `model` field
129 * and the three key names, which are about the same in every wire format.
130 */
131const ENVELOPE_TOKENS = 32
132
133/**
134 * How many calls every request must have room to ask about, beside the
135 * state. Each request carries the whole state, so a state allowed to fill
136 * the request leaves room for one call's questions at a time, and the state
137 * is then paid for once per call instead of once per batch.
138 */
139const MIN_CALLS_PER_REQUEST = 16
140
141/**
142 * The most requests one compaction sends. The state rides in every one of
143 * them, so this bounds what a compaction can cost at this many times the
144 * state; a conversation that would need more is left to the built-in
145 * summary.
146 */
147const MAX_REQUESTS = 16
148
149/**
150 * How many requests are in flight at once. Waiting on a request is not
151 * charged to the hook, so there is no need to send every batch at the same
152 * moment, which is what a provider's rate limit is most likely to refuse.
153 */
154const CONCURRENT_REQUESTS = 3
155
156/**
157 * The share of a provider's documented limits a budget may reach. The token
158 * estimate is made without a tokenizer and its weights were fitted to one
159 * model, so a budget right at a limit could be over it by the provider's
160 * own count, and a provider may cut what is over without saying so.
161 */
162export const LIMIT_SHARE = 0.85
163
164/**
165 * A result is left whole unless cutting it saves more than the note that
166 * replaces its tail costs.
167 */
168const NOTE_ROOM = 120
169
170/**
171 * The scores a pinned call is reported with: it was not asked about, and it
172 * stays whole.
173 */
174const UNASKED: Scores = { keepCall: 1, keepResult: 1 }
175
176/**
177 * The id of the question whether a call should stay.
178 */
179function callQuestionId(call: Call): string {
180  return `call_${call.id}`
181}
182
183/**
184 * The id of the question whether a call's whole output should stay.
185 */
186function resultQuestionId(call: Call): string {
187  return `result_${call.id}`
188}
189
190/**
191 * What is asked about a single call, as two questions. They are separate
192 * because the answers differ in practice: that a file was read can matter
193 * long after its contents stopped mattering.
194 *
195 * The rubric they are read under is in the state's `context`, sent once.
196 *
197 * @returns its two `noul` questions, by id
198 */
199export function questionsOf(call: Call): Questions {
200  const named = `tool call ${call.id} (${call.tool})`
201  const outcome = call.isError ? 'an error, ' : ''
202
203  return {
204    [callQuestionId(call)]: {
205      type: 'noul',
206      instructions:
207        `Does the assistant still need to know that it made ${named}, and ` +
208        'with which input, to carry on with the goal?',
209    },
210    [resultQuestionId(call)]: {
211      type: 'noul',
212      instructions:
213        `Does the complete output of ${named}, ${outcome}` +
214        `${call.resultChars} characters long, have to stay word for word ` +
215        'because the assistant still relies on its content and running the ' +
216        'tool again would not bring it back?',
217    },
218  }
219}
220
221/**
222 * The estimated size of each of a call's questions and of the two together,
223 * as they ride in a request body written in the route's wire format, with
224 * what the provider counts for each question beyond its text; and the UTF-8
225 * bytes of the two, each with the comma after it.
226 */
227function questionTokensOf(
228  call: Call,
229  wire: Wire,
230): { both: number; longest: number; bytes: number } {
231  let both = 0
232  let longest = 0
233  let bytes = 0
234
235  for (const [id, question] of Object.entries(questionsOf(call))) {
236    const json = wire.questionJsonOf(id, question)
237    const tokens = estimatedTokensOf(json) + 1 + wire.tokensPerQuestion
238
239    both += tokens
240    longest = Math.max(longest, tokens)
241    bytes += bytesOf(json) + 1
242  }
243
244  return { both, longest, bytes }
245}
246
247/**
248 * The sum of the `count` largest of some sizes.
249 */
250function largestOf(sizes: readonly number[], count: number): number {
251  return [...sizes]
252    .sort((a, b) => b - a)
253    .slice(0, count)
254    .reduce((sum, size) => sum + size, 0)
255}
256
257/**
258 * Measures what the questions about the calls need of a request: the
259 * longest single question, and the questions of `MIN_CALLS_PER_REQUEST`
260 * calls, or of all of them when there are fewer. The calls with the longest
261 * questions are the ones counted, so the room set aside holds whichever
262 * calls end up in a batch together.
263 *
264 * @param calls the calls that will be asked about
265 * @param wire the wire format the questions are written in
266 * @param maxQuestions the provider's cap on the questions of one request,
267 * when it has one; it bounds only the bytes set aside
268 * @returns the demand, all zero when there are no calls
269 */
270export function demandOf(
271  calls: readonly Call[],
272  wire: Wire,
273  maxQuestions?: number,
274): Demand {
275  const sizes = calls.map(call => questionTokensOf(call, wire))
276  const batchCalls = Math.min(MIN_CALLS_PER_REQUEST, sizes.length)
277  const byteCalls =
278    maxQuestions === undefined
279      ? batchCalls
280      : Math.min(batchCalls, Math.floor(maxQuestions / QUESTIONS_PER_CALL))
281
282  return {
283    longestQuestion: sizes.reduce(
284      (longest, size) => Math.max(longest, size.longest),
285      0,
286    ),
287    batchCalls,
288    batchTokens: largestOf(
289      sizes.map(size => size.both),
290      batchCalls,
291    ),
292    batchBytes: largestOf(
293      sizes.map(size => size.bytes),
294      byteCalls,
295    ),
296  }
297}
298
299/**
300 * Holds the person's budgets against a provider's limits and against each
301 * other.
302 *
303 * A budget above `LIMIT_SHARE` of a documented limit is lowered to it rather
304 * than trusted. The state is further held to what the request leaves once
305 * the questions of a batch are set aside: a state that filled the request
306 * would be sent again for every single call.
307 *
308 * @param want the configured budgets
309 * @param limits the provider's documented limits
310 * @param demand what the questions need of a request
311 * @returns the budget one request is held to; throws when the request has
312 * no room for a state beside one batch of questions
313 */
314export function budgetOf(
315  want: { maxStateTokens: number; maxRequestTokens: number },
316  limits: Limits,
317  demand: Demand,
318): Budget {
319  const requestTokens = Math.min(
320    want.maxRequestTokens,
321    Math.floor(limits.maxRequestTokens * LIMIT_SHARE),
322  )
323  const beside = requestTokens - ENVELOPE_TOKENS - demand.batchTokens
324
325  if (beside < 1) {
326    throw new Error(
327      `a request of ${requestTokens} tokens has no room for a state beside ` +
328        `the questions of ${demand.batchCalls} ` +
329        `${demand.batchCalls === 1 ? 'call' : 'calls'}, which take about ` +
330        `${demand.batchTokens}`,
331    )
332  }
333
334  const budget: Budget = {
335    stateTokens: Math.min(
336      want.maxStateTokens,
337      Math.floor(limits.maxStateTokens * LIMIT_SHARE) - demand.longestQuestion,
338      beside,
339    ),
340    requestTokens,
341  }
342
343  if (limits.maxQuestions !== undefined) {
344    budget.questions = limits.maxQuestions
345  }
346
347  return budget
348}
349
350/**
351 * Splits the calls into batches, each of which fits one request beside the
352 * state under both the token budget and the provider's question count.
353 *
354 * The state is never split: every batch is asked against all of it, so the
355 * only thing that varies between requests is which calls they ask about.
356 *
357 * @param calls the calls to ask about, in order
358 * @param stateTokens the estimated size of the fitted state
359 * @param budget what one request may hold
360 * @param wire the wire format the questions are written in
361 * @returns the batches, in order; throws when the state leaves room for not
362 * even one call's questions, and when there would be more batches than
363 * `MAX_REQUESTS`
364 */
365export function batchesOf(
366  calls: readonly Call[],
367  stateTokens: number,
368  budget: Budget,
369  wire: Wire,
370): Call[][] {
371  const room = budget.requestTokens - stateTokens - ENVELOPE_TOKENS
372  const byteRoom =
373    budget.bytes === undefined
374      ? Infinity
375      : budget.bytes.limit - budget.bytes.state
376  const mostCalls =
377    budget.questions === undefined
378      ? Infinity
379      : Math.floor(budget.questions / QUESTIONS_PER_CALL)
380
381  if (mostCalls < 1) {
382    throw new Error(
383      `a request may hold ${budget.questions} questions, fewer than the ` +
384        `${QUESTIONS_PER_CALL} one call needs`,
385    )
386  }
387
388  const batches: Call[][] = []
389  let batch: Call[] = []
390  let used = 0
391  let usedBytes = 0
392
393  for (const call of calls) {
394    const { both, bytes } = questionTokensOf(call, wire)
395
396    if (both > room) {
397      throw new Error(
398        `no question fits beside the state: the state takes about ` +
399          `${stateTokens} of the ${budget.requestTokens} tokens a request ` +
400          'may hold',
401      )
402    }
403
404    if (bytes > byteRoom) {
405      throw new Error(
406        `no question fits beside the state: the request takes ` +
407          `${budget.bytes?.state} of the ${budget.bytes?.limit} bytes it ` +
408          'may hold before any question',
409      )
410    }
411
412    if (
413      batch.length >= mostCalls ||
414      used + both > room ||
415      usedBytes + bytes > byteRoom
416    ) {
417      batches.push(batch)
418      batch = []
419      used = 0
420      usedBytes = 0
421    }
422
423    batch.push(call)
424    used += both
425    usedBytes += bytes
426  }
427
428  if (batch.length > 0) {
429    batches.push(batch)
430  }
431
432  if (batches.length > MAX_REQUESTS) {
433    throw new Error(
434      `asking about ${calls.length} tool calls would take ` +
435        `${batches.length} requests, and one compaction sends at most ` +
436        `${MAX_REQUESTS}`,
437    )
438  }
439
440  return batches
441}
442
443/**
444 * Turns the two probabilities into what happens to the call. The result
445 * question is the stronger claim, so it is read first: a result worth
446 * keeping keeps its call with it.
447 *
448 * @param threshold the probability from which an item is kept
449 */
450export function decide(
451  call: Call,
452  scores: Scores,
453  threshold: number,
454): Decision {
455  if (call.isPinned) {
456    return { call, ...scores, action: 'pinned' }
457  }
458
459  if (scores.keepResult >= threshold) {
460    return { call, ...scores, action: 'keep' }
461  }
462
463  if (scores.keepCall >= threshold) {
464    return { call, ...scores, action: 'truncate' }
465  }
466
467  return { call, ...scores, action: 'drop' }
468}
469
470/**
471 * Says whether cutting a text of this length to its head removes more than
472 * the note that then stands for the rest adds.
473 */
474function isWorthCutting(chars: number, headChars: number): boolean {
475  return chars > headChars + NOTE_ROOM
476}
477
478/**
479 * A result text cut to its head, with one line saying what was done and by
480 * what, so the assistant reading it later knows the rest existed and how to
481 * get it back. A text too short for the cut to save anything is left as is.
482 */
483function cut(text: string, headChars: number, isError: boolean): string {
484  if (!isWorthCutting(text.length, headChars)) {
485    return text
486  }
487
488  const head = headOf(text, headChars)
489  const kind = isError ? 'error output' : 'output'
490
491  return (
492    `${head}${head === '' ? '' : '\n'}[decision-compaction removed ` +
493    `${text.length - head.length} more characters of this tool ${kind}; ` +
494    'run the tool again if they are needed]'
495  )
496}
497
498/**
499 * A result with its text cut, or the same object when nothing was cut. The
500 * stored record is left off a cut result: it holds the full output.
501 */
502function cutResult(
503  result: ToolResultSummary,
504  headChars: number,
505): ToolResultSummary {
506  const { tool_use_id, isError } = result
507  const text = cut(result.text, headChars, isError)
508
509  if (text === result.text) {
510    return result
511  }
512
513  return { tool_use_id, text, isError }
514}
515
516/**
517 * A tool use whose mirrored outcome is cut like its result, or the same
518 * object when nothing was cut. Only used inside a message that is rebuilt
519 * anyway, so the copy and the result it mirrors agree.
520 */
521function cutUse(use: ToolUseSummary, headChars: number): ToolUseSummary {
522  if (use.text === undefined) {
523    return use
524  }
525
526  const text = cut(use.text, headChars, use.isError === true)
527
528  if (text === use.text) {
529    return use
530  }
531
532  const { result: _record, ...kept } = use
533
534  return { ...kept, text }
535}
536
537/**
538 * Carries the decisions out on the conversation.
539 *
540 * A message nothing touched is returned as the very object it came in as, so
541 * it still carries the engine's `handle` and stands whole. A changed message
542 * is a new object with no `handle`, which the engine builds from its role,
543 * text and tool blocks; because that loses whatever else the original held,
544 * a message is rebuilt only when one of its own blocks changed. A dropped
545 * call takes its result with it, so no result is left without its call, and
546 * a message left with no text and no block is removed.
547 *
548 * @param messages the conversation
549 * @param decisions what was decided for each paired call
550 * @param headChars how much of a truncated result stays
551 * @returns the conversation afterwards, oldest first
552 */
553export function rebuild(
554  messages: readonly SessionMessage[],
555  decisions: readonly Decision[],
556  headChars: number,
557): SessionMessage[] {
558  const fate = new Map<string, 'truncate' | 'drop'>()
559
560  for (const { call, action } of decisions) {
561    if (action === 'truncate' || action === 'drop') {
562      fate.set(call.toolUseId, action)
563    }
564  }
565
566  const rebuilt: SessionMessage[] = []
567
568  for (const message of messages) {
569    const given = message.toolResults ?? []
570    const uses = message.toolUses.filter(
571      use => fate.get(use.tool_use_id) !== 'drop',
572    )
573    const results = given
574      .filter(result => fate.get(result.tool_use_id) !== 'drop')
575      .map(result =>
576        fate.get(result.tool_use_id) === 'truncate'
577          ? cutResult(result, headChars)
578          : result,
579      )
580    const isUntouched =
581      uses.length === message.toolUses.length &&
582      results.length === given.length &&
583      results.every((result, index) => result === given[index])
584
585    if (isUntouched) {
586      rebuilt.push(message)
587      continue
588    }
589
590    const isEmptied =
591      uses.length + results.length === 0 && message.text.trim() === ''
592
593    if (!isEmptied) {
594      rebuilt.push({
595        role: message.role,
596        text: message.text,
597        toolUses: uses.map(use =>
598          fate.get(use.tool_use_id) === 'truncate'
599            ? cutUse(use, headChars)
600            : use,
601        ),
602        ...(results.length > 0 ? { toolResults: results } : {}),
603      })
604    }
605  }
606
607  return rebuilt
608}
609
610/**
611 * The size of a tool input as the model reads it. An input that cannot be
612 * serialised has no size to compare, and counts as nothing.
613 */
614function inputCharsOf(use: ToolUseSummary): number {
615  try {
616    return (JSON.stringify(use.input) ?? '').length
617  } catch {
618    return 0
619  }
620}
621
622/**
623 * The characters a message holds for the model: its text, its tool inputs
624 * and its tool results. The outcome mirrored on a tool use is not counted,
625 * since the result block already is.
626 */
627function charsOf(message: SessionMessage): number {
628  const inputs = message.toolUses.reduce(
629    (sum, use) => sum + inputCharsOf(use),
630    0,
631  )
632  const results = (message.toolResults ?? []).reduce(
633    (sum, result) => sum + result.text.length,
634    0,
635  )
636
637  return message.text.length + inputs + results
638}
639
640function totalChars(messages: readonly SessionMessage[]): number {
641  return messages.reduce((chars, message) => chars + charsOf(message), 0)
642}
643
644/**
645 * The share of the conversation's characters a compaction removed.
646 *
647 * @returns a ratio from 0 to 1; 0 for an empty conversation
648 */
649export function reductionOf(
650  outcome: Pick<Outcome, 'charsBefore' | 'charsAfter'>,
651): number {
652  return outcome.charsBefore === 0
653    ? 0
654    : (outcome.charsBefore - outcome.charsAfter) / outcome.charsBefore
655}
656
657/**
658 * One batch as it came back: the reply and, for every call of the batch,
659 * the two probabilities read from it.
660 */
661type Judged = {
662  reply: Reply
663  scores: [Call, Scores][]
664}
665
666/**
667 * Puts one batch to the provider and reads both answers of every call in
668 * it. A reply that lacks one of them fails the batch: a call nobody answered
669 * for must not be decided by a default.
670 */
671async function judge(
672  batch: readonly Call[],
673  state: unknown,
674  ask: Ask,
675): Promise<Judged> {
676  const reply = await ask(
677    state,
678    Object.fromEntries(
679      batch.flatMap(call => Object.entries(questionsOf(call))),
680    ),
681  )
682
683  return {
684    reply,
685    scores: batch.map(call => [
686      call,
687      {
688        keepCall: noulOf(reply, callQuestionId(call)),
689        keepResult: noulOf(reply, resultQuestionId(call)),
690      },
691    ]),
692  }
693}
694
695/**
696 * The outcome when there is nothing to ask about: the conversation as it is.
697 *
698 * @param messages the conversation
699 * @param calls its paired calls, all of them pinned
700 * @returns an outcome with no request made
701 */
702export function untouched(
703  messages: readonly SessionMessage[],
704  calls: readonly Call[],
705): Outcome {
706  const chars = totalChars(messages)
707
708  return {
709    messages: [...messages],
710    decisions: calls.map(call => decide(call, UNASKED, 0)),
711    charsBefore: chars,
712    charsAfter: chars,
713    stateTokens: 0,
714    requests: 0,
715    reported: {},
716    shortened: 0,
717  }
718}
719
720/**
721 * Runs `work` over the items with at most `width` of them under way at a
722 * time, and resolves with the results in the items' order, whichever order
723 * they came back in. The first failure rejects the whole and starts nothing
724 * further; `onFailure` hears of it at once, so that the caller can stop what
725 * is still under way.
726 */
727async function inTurns<Item, Result>(
728  items: readonly Item[],
729  width: number,
730  work: (item: Item) => Promise<Result>,
731  onFailure: (reason: unknown) => void,
732): Promise<Result[]> {
733  const results: Result[] = []
734  let next = 0
735  let hasFailed = false
736
737  const worker = async () => {
738    while (!hasFailed && next < items.length) {
739      const index = next++
740
741      try {
742        results[index] = await work(items[index] as Item)
743      } catch (error) {
744        hasFailed = true
745        onFailure(error)
746
747        throw error
748      }
749    }
750  }
751
752  await Promise.all(
753    Array.from({ length: Math.min(width, items.length) }, worker),
754  )
755
756  return results
757}
758
759/**
760 * Asks about every candidate call and applies the answers to the
761 * conversation. The batches go out `CONCURRENT_REQUESTS` at a time, each
762 * with the whole state, and each call is decided by the answers of its own
763 * batch, in whatever order the batches come back. One batch failing fails
764 * the compaction, since half the answers decide nothing; no batch is sent
765 * after it, and `halt` is told at once so the others can be abandoned.
766 *
767 * @param messages the conversation
768 * @param calls its paired calls, pinned ones included
769 * @param fitted the state, already fitted to the budget, with its size as
770 * the route's wire format carries it
771 * @param ask how a batch is sent
772 * @param settings the budget, the route's wire format, the keep threshold,
773 * the truncation length, and `halt`, which hears the reason the moment a
774 * batch fails
775 */
776export async function compact(
777  messages: readonly SessionMessage[],
778  calls: readonly Call[],
779  fitted: Fitted,
780  ask: Ask,
781  settings: {
782    budget: Budget
783    wire: Wire
784    keepThreshold: number
785    truncateHeadChars: number
786    halt?: (reason: Error) => void
787  },
788): Promise<Outcome> {
789  const batches = batchesOf(
790    calls.filter(call => !call.isPinned),
791    fitted.tokens,
792    settings.budget,
793    settings.wire,
794  )
795  const answered = await inTurns(
796    batches,
797    CONCURRENT_REQUESTS,
798    batch => judge(batch, fitted.state, ask),
799    reason =>
800      settings.halt?.(
801        reason instanceof Error ? reason : new Error(String(reason)),
802      ),
803  )
804  const scores = new Map<Call, Scores>()
805  const reported: Reported = {}
806
807  for (const { scores: judged, reply } of answered) {
808    for (const [call, value] of judged) {
809      scores.set(call, value)
810    }
811
812    if (reply.usage.input_tokens !== undefined) {
813      reported.inputTokens = Math.max(
814        reported.inputTokens ?? 0,
815        reply.usage.input_tokens,
816      )
817    }
818
819    if (reply.usage.cost !== undefined) {
820      reported.cost = (reported.cost ?? 0) + reply.usage.cost
821    }
822  }
823
824  const decisions = calls.map(call =>
825    decide(call, scores.get(call) ?? UNASKED, settings.keepThreshold),
826  )
827  const kept = rebuild(messages, decisions, settings.truncateHeadChars)
828
829  return {
830    messages: kept,
831    decisions,
832    charsBefore: totalChars(messages),
833    charsAfter: totalChars(kept),
834    stateTokens: fitted.tokens,
835    stage: fitted.stage,
836    requests: batches.length,
837    reported,
838    shortened: decisions.filter(
839      ({ call, action }) =>
840        action === 'truncate' &&
841        isWorthCutting(call.resultChars, settings.truncateHeadChars),
842    ).length,
843  }
844}
845
hooks/config.ts 212 lines
1import type { PluginOptions } from 'claude-code'
2
3import { cutOf, providerOf, PROVIDERS } from './providers'
4import type { ProviderName } from './providers'
5
6/**
7 * The plugin's options as the hooks use them: every value present, typed
8 * and inside the range it is meaningful in.
9 */
10export type Config = {
11  provider: ProviderName
12  /**
13   * Why no compaction may be decided with these options, when there is such
14   * a reason. It is set for an option that says where the conversation is
15   * sent and cannot be read: a destination is never guessed.
16   */
17  refusal?: string
18  /**
19   * The model name for the configured provider; absent for its default.
20   */
21  model?: string
22  providerDecision: boolean
23  /**
24   * The providers the provider decision may pick besides the configured
25   * one, in `PROVIDERS` order.
26   */
27  decisionProviders: readonly ProviderName[]
28  /**
29   * The names in `decisionProviders` that are no provider, as given; absent
30   * when there are none. They are dropped, not refused: leaving a name out
31   * can only narrow where the conversation is sent.
32   */
33  unknownProviders?: string[]
34  keepThreshold: number
35  preserveRecentMessages: number
36  compactAtPercent: number
37  minReductionRatio: number
38  maxStateTokens: number
39  maxRequestTokens: number
40  truncateHeadChars: number
41}
42
43/**
44 * What an option reads as when it is unset or unusable; the manifest's
45 * `userConfig` states the same values to the person.
46 */
47const DEFAULTS = {
48  provider: 'typesafe',
49  providerDecision: false,
50  decisionProviders: PROVIDERS,
51  keepThreshold: 0.5,
52  preserveRecentMessages: 6,
53  compactAtPercent: 60,
54  minReductionRatio: 0.25,
55  maxStateTokens: 25_000,
56  maxRequestTokens: 30_000,
57  truncateHeadChars: 300,
58} as const satisfies Config
59
60/**
61 * A numeric option held to a range. A value that is not a finite number
62 * reads as the default; one outside the range is moved to its nearest end.
63 */
64function within(
65  value: unknown,
66  fallback: number,
67  low: number,
68  high: number,
69): number {
70  const number =
71    typeof value === 'number' && Number.isFinite(value) ? value : fallback
72
73  return Math.min(high, Math.max(low, number))
74}
75
76/**
77 * Reads the `decisionProviders` option: a list, or the names in one text
78 * separated by commas. Unset, it is every provider. The host fills in the
79 * manifest's default for a field that was never set, so an empty text is a
80 * field the person cleared, and like an empty list it names no provider:
81 * only the configured one is then eligible. Names are matched as `provider`
82 * is, exactly; a value of any other type names no provider.
83 */
84function decisionProvidersOf(value: unknown): {
85  named: readonly ProviderName[]
86  unknown: string[]
87} {
88  if (value === undefined || value === null) {
89    return { named: DEFAULTS.decisionProviders, unknown: [] }
90  }
91
92  const given = Array.isArray(value)
93    ? value
94    : typeof value === 'string'
95      ? value.split(',')
96      : [value]
97  const names = given
98    .map(name =>
99      typeof name === 'string' ? name.trim() : String(JSON.stringify(name)),
100    )
101    .filter(name => name !== '')
102
103  return {
104    named: PROVIDERS.filter(provider => names.includes(provider)),
105    unknown: [...new Set(names.filter(name => providerOf(name) === undefined))],
106  }
107}
108
109/**
110 * The providers a compaction may send anything to: the configured one, and
111 * with the provider decision on also those `decisionProviders` names, in
112 * `PROVIDERS` order. Everything that reads a credential or offers a route
113 * goes by this one list.
114 *
115 * @param config the configuration
116 * @returns the providers
117 */
118export function eligibleOf(config: Config): ProviderName[] {
119  return PROVIDERS.filter(
120    provider =>
121      provider === config.provider ||
122      (config.providerDecision && config.decisionProviders.includes(provider)),
123  )
124}
125
126/**
127 * Says whether an option was left unset: absent, or text with nothing in
128 * it, which is what a cleared field holds.
129 */
130function isUnset(value: unknown): boolean {
131  return (
132    value === undefined ||
133    value === null ||
134    (typeof value === 'string' && value.trim() === '')
135  )
136}
137
138/**
139 * Reads the plugin's options. Nothing here throws, because this runs while
140 * the module loads and a throw would take both hooks with it: an option the
141 * person mistyped reads as its default.
142 *
143 * The one exception is `provider`, which decides who is sent the
144 * conversation. Only an unset `provider` reads as the default. A value that
145 * names none of the providers is not replaced by one the person did not
146 * choose: it is reported in `refusal`, and every compaction is then left to
147 * the built-in summary.
148 *
149 * @param options what `register` received
150 * @returns the configuration
151 */
152export function configOf(options: PluginOptions): Config {
153  const named = providerOf(options.provider)
154  const deciding = decisionProvidersOf(options.decisionProviders)
155  const config: Config = {
156    provider: named ?? DEFAULTS.provider,
157    providerDecision: options.providerDecision === true,
158    decisionProviders: deciding.named,
159    keepThreshold: within(options.keepThreshold, DEFAULTS.keepThreshold, 0, 1),
160    preserveRecentMessages: Math.floor(
161      within(
162        options.preserveRecentMessages,
163        DEFAULTS.preserveRecentMessages,
164        0,
165        Infinity,
166      ),
167    ),
168    compactAtPercent: within(
169      options.compactAtPercent,
170      DEFAULTS.compactAtPercent,
171      0,
172      100,
173    ),
174    minReductionRatio: within(
175      options.minReductionRatio,
176      DEFAULTS.minReductionRatio,
177      0,
178      1,
179    ),
180    maxStateTokens: Math.floor(
181      within(options.maxStateTokens, DEFAULTS.maxStateTokens, 1, Infinity),
182    ),
183    maxRequestTokens: Math.floor(
184      within(options.maxRequestTokens, DEFAULTS.maxRequestTokens, 1, Infinity),
185    ),
186    truncateHeadChars: Math.floor(
187      within(
188        options.truncateHeadChars,
189        DEFAULTS.truncateHeadChars,
190        0,
191        Infinity,
192      ),
193    ),
194  }
195
196  if (named === undefined && !isUnset(options.provider)) {
197    config.refusal =
198      `the provider option is ${cutOf(JSON.stringify(options.provider))}, ` +
199      `which is none of ${PROVIDERS.join(', ')}`
200  }
201
202  if (deciding.unknown.length > 0) {
203    config.unknownProviders = deciding.unknown
204  }
205
206  if (typeof options.model === 'string' && options.model.trim() !== '') {
207    config.model = options.model.trim()
208  }
209
210  return config
211}
212
hooks/providers.ts 1127 lines
1import type { HttpResponse, PluginOptions } from 'claude-code'
2
3import { decodeOpenAI, OPENAI } from './openai'
4import { headOf } from './state'
5import { BARE, bodyOf, bytesOf, idsOf, isRecord, replyOf } from './systemone'
6import type { Question, Questions, Reply, Wire } from './systemone'
7
8/**
9 * The providers that serve the System One protocol, in the order the
10 * manifest lists them. Of two gateways to the same model the one listed
11 * first is the one the provider decision offers.
12 */
13export const PROVIDERS = [
14  'typesafe',
15  'cloudflare',
16  'openrouter',
17  'codiv',
18  'perplexity',
19  'decisions-api-dev',
20  'decisionapi-net',
21  'openai',
22] as const
23
24export type ProviderName = (typeof PROVIDERS)[number]
25
26/**
27 * One way to reach one model: the provider that is called and the model name
28 * as that provider's request body spells it.
29 */
30export type Route = {
31  provider: ProviderName
32  model: string
33}
34
35/**
36 * What a provider's documentation says a request may hold, in tokens as the
37 * provider counts them.
38 *
39 * `maxStateTokens` bounds the state together with the longest single
40 * question, which is how TypeSafe states its limit; `maxQuestions` is absent
41 * where no cap is documented. `maxRequestBytes` is a cap on the UTF-8 bytes
42 * of the whole request body, for a provider that states its limit in bytes.
43 */
44export type Limits = {
45  maxStateTokens: number
46  maxRequestTokens: number
47  maxQuestions?: number
48  maxRequestBytes?: number
49}
50
51/**
52 * The secrets and the account id the providers need, each present only when
53 * one was configured.
54 */
55export type Credentials = {
56  typesafeApiKey?: string
57  openrouterApiKey?: string
58  cloudflareApiToken?: string
59  cloudflareAccountId?: string
60  codivApiKey?: string
61  perplexityApiKey?: string
62  decisionsApiKey?: string
63  decisionapiApiKey?: string
64  openaiApiKey?: string
65}
66
67/**
68 * The environment variable each credential falls back to when its plugin
69 * option is unset.
70 */
71const ENV_NAMES = {
72  typesafeApiKey: 'TYPESAFE_API_KEY',
73  openrouterApiKey: 'OPENROUTER_API_KEY',
74  cloudflareApiToken: 'CLOUDFLARE_API_TOKEN',
75  cloudflareAccountId: 'CLOUDFLARE_ACCOUNT_ID',
76  codivApiKey: 'CODIV_API_KEY',
77  perplexityApiKey: 'PERPLEXITY_API_KEY',
78  decisionsApiKey: 'DECISIONS_API_KEY',
79  decisionapiApiKey: 'DECISIONAPI_API_KEY',
80  openaiApiKey: 'OPENAI_API_KEY',
81} as const satisfies Record<keyof Credentials, string>
82
83type EnvName = (typeof ENV_NAMES)[keyof Credentials]
84
85/**
86 * The values of the credential variables as one source holds them.
87 */
88export type Environment = Partial<Record<EnvName, string | undefined>>
89
90/**
91 * One request as `$.http.fetch` takes it.
92 */
93type Exchange = {
94  url: string
95  init: { method: 'POST'; headers: Record<string, string>; body: string }
96}
97
98/**
99 * Everything that differs between providers. The call path and the answer
100 * validation are shared, so a provider is this record and nothing else.
101 */
102type Descriptor = {
103  /**
104   * The provider's name as a sentence about it spells it.
105   */
106  title: string
107  defaultModel: string
108  /**
109   * The models the provider is offered with in the routing decision.
110   */
111  models: readonly string[]
112  /**
113   * The model family a model of this provider is, which is what two routes
114   * are compared by: the same Jev is served directly and through gateways.
115   */
116  family: (model: string) => string
117  /**
118   * True for a provider that forwards to a model another provider hosts.
119   */
120  isGateway: boolean
121  limits: Limits
122  needs: readonly (keyof Credentials)[]
123  bodyModelOf: (model: string) => string
124  urlOf: (model: string, credentials: Credentials) => string
125  tokenOf: (credentials: Credentials) => string | undefined
126  wire: Wire
127  /**
128   * Takes the response out of whatever the provider wraps it in and into
129   * the System One shape the readers of a reply take, given the questions
130   * the request asked.
131   */
132  unwrap: (
133    payload: unknown,
134    credentials: Credentials,
135    asked: Questions,
136  ) => unknown
137}
138
139const CLOUDFLARE_MODELS = ['clef', 'clef-flash'] as const
140
141/**
142 * What a request to a provider that documents no limit is held to: the
143 * tightest window and question cap among the documented ones (OpenRouter's
144 * 32,000 tokens for Jev, Clef's 64 questions).
145 */
146const UNDOCUMENTED: Limits = {
147  maxStateTokens: 32_000,
148  maxRequestTokens: 32_000,
149  maxQuestions: 64,
150}
151
152/**
153 * How many questions one request to either Decisions reseller may hold.
154 */
155const RESELLER_QUESTIONS = 8
156
157/**
158 * The two Decisions resellers state their cap in bytes: the text and the
159 * questions of a request within 32 KiB, and at most eight questions. Both
160 * serve TypeSafe's Jev 1.13, so the token figures are the window TypeSafe
161 * documents for it. They do not bind: 32 KiB of English is far fewer
162 * tokens, and the state is fitted to the byte cap as well. They stand
163 * against what the provider counts, so a reply that counts the whole window
164 * is still read as a state that was cut short.
165 */
166const RESELLER: Limits = {
167  maxStateTokens: 32_000,
168  maxRequestTokens: 64_000,
169  maxQuestions: RESELLER_QUESTIONS,
170  maxRequestBytes: 32_768,
171}
172
173/**
174 * The question ids decisions-api.dev accepts: a letter first, then letters,
175 * digits, `_` or `-`, 64 characters at most. The mod's ids may hold `.` and
176 * run to 100 characters, so the request names each question by its place
177 * instead.
178 */
179const ALIAS_OF = (index: number) => `q${index}`
180
181/**
182 * The longest alias a request can carry: the one of the last place the
183 * question cap allows. A question is measured under it, whatever its own
184 * id, since the alias is what is sent.
185 */
186const LONGEST_ALIAS = ALIAS_OF(RESELLER_QUESTIONS - 1)
187
188/**
189 * The System One body with every question id replaced by its place in the
190 * request, `q0`, `q1`, …. The ids are checked against the mod's own rule
191 * first, as for every provider; option names inside a `choice` are left as
192 * they are, since only question ids are restricted.
193 */
194const ALIASED: Wire = {
195  encode: (model, state, questions) =>
196    bodyOf(
197      model,
198      state,
199      Object.fromEntries(
200        idsOf(questions).map((id, index) => [
201          ALIAS_OF(index),
202          questions[id] as Question,
203        ]),
204      ),
205    ),
206  questionJsonOf: (_id, question) =>
207    BARE.questionJsonOf(LONGEST_ALIAS, question),
208  stateJsonOf: BARE.stateJsonOf,
209  tokensPerQuestion: 0,
210}
211
212/**
213 * A Jev id as TypeSafe names it: `jev`, `jev-latest`, `jev-1.13`.
214 */
215const JEV = /^jev(?:-|$)/
216
217/**
218 * The ids OpenRouter routes to Jev: a Jev id (`jev`, `jev-latest`,
219 * `jev-1.13`) either bare, which OpenRouter maps onto TypeSafe's namespace
220 * before routing, or under that namespace (`typesafe/`) or its alias
221 * (`~typesafe/`). The namespace alone does not make a model Jev: any other
222 * model TypeSafe publishes there is a model of its own.
223 */
224const OPENROUTER_JEV = /^(?:~?typesafe\/)?jev(?:-|$)/
225
226/**
227 * The Workers AI catalogue prefix; a Cloudflare model is `clef` in the body
228 * and this prefix plus `clef` in the URL.
229 */
230const CLOUDFLARE_PREFIX = '@cf/cloudflare/'
231
232/**
233 * How much is quoted of a text this mod did not write (an error body, a
234 * failure envelope, the host's reason for refusing a call): enough to read
235 * the reason, short enough for one toast line.
236 */
237export const REASON_CHARS = 160
238
239/**
240 * The shortest value treated as a secret when a text is redacted. A shorter
241 * one is not a key any of the providers issues, and replacing it wherever
242 * its few characters occur would garble the ordinary words around it.
243 */
244const MIN_SECRET_CHARS = 8
245
246/**
247 * What stands where a credential was. It holds no part of what it replaces,
248 * not even the word that announced it.
249 */
250const REDACTED = '[redacted]'
251
252/**
253 * An `Authorization` value as a response may echo it: the scheme and the
254 * token after it, whoever's token that is. The two are separated by white
255 * space in a header, and by a colon or an equals sign where the header was
256 * written out as a field (`Bearer: ...`, `bearer=...`).
257 */
258const BEARER = /bearer(?:\s*[:=]\s*|\s+)[^\s"'<>]+/gi
259
260function textOf(value: unknown): string | undefined {
261  return typeof value === 'string' && value.trim() !== ''
262    ? value.trim()
263    : undefined
264}
265
266/**
267 * Accepts a Cloudflare model by its short name or its catalogue id, and
268 * answers the short name. The body's `model` must be `clef` or `clef-flash`,
269 * so any other name is refused here rather than by a 4xx after the state was
270 * sent.
271 */
272function cloudflareModelOf(model: string): string {
273  const named = model.trim()
274  const short = named.startsWith(CLOUDFLARE_PREFIX)
275    ? named.slice(CLOUDFLARE_PREFIX.length)
276    : named
277
278  if (!CLOUDFLARE_MODELS.some(known => known === short)) {
279    throw new Error(
280      `cloudflare serves ${CLOUDFLARE_MODELS.join(' and ')}, not ` +
281        JSON.stringify(named),
282    )
283  }
284
285  return short
286}
287
288/**
289 * The first message of a Workers AI envelope's `errors` list.
290 */
291function envelopeErrorOf(payload: Record<string, unknown>): string | undefined {
292  const first = Array.isArray(payload.errors) ? payload.errors[0] : undefined
293
294  return isRecord(first) ? textOf(first.message) : undefined
295}
296
297/**
298 * The error for a response envelope that reports failure, quoting the
299 * reason it gives.
300 */
301function envelopeFailure(
302  reason: string | undefined,
303  credentials: Credentials,
304): Error {
305  return new Error(
306    'the response envelope reports failure: ' +
307      briefOf(reason ?? '', credentials),
308  )
309}
310
311/**
312 * Takes a Workers AI response out of its `{ result, success, errors }`
313 * envelope. A body that is already a bare reply is passed through, because
314 * the envelope is documented for the REST API in general and not for Clef.
315 * The reason a failed envelope gives is quoted like any other text of a
316 * provider's: redacted and held to one short line.
317 */
318function unwrapCloudflare(payload: unknown, credentials: Credentials): unknown {
319  if (!isRecord(payload)) {
320    return payload
321  }
322
323  if (payload.success === false) {
324    throw envelopeFailure(envelopeErrorOf(payload), credentials)
325  }
326
327  return isRecord(payload.result) ? payload.result : payload
328}
329
330/**
331 * Takes a response out of the `{ code, message, data: { result } }` envelope
332 * the two Decisions resellers answer with. A `code` other than 0, or no
333 * `result`, is a failure, and the envelope's `message` is quoted as its
334 * reason like any other text of a provider's: redacted and held to one
335 * short line.
336 */
337function unwrapDataEnvelope(
338  payload: unknown,
339  credentials: Credentials,
340): Record<string, unknown> {
341  const data = isRecord(payload) ? payload.data : undefined
342  const result = isRecord(data) ? data.result : undefined
343
344  if (!isRecord(payload) || payload.code !== 0 || !isRecord(result)) {
345    throw envelopeFailure(
346      isRecord(payload) ? textOf(payload.message) : undefined,
347      credentials,
348    )
349  }
350
351  return result
352}
353
354/**
355 * Reads a decisions-api.dev response: out of its envelope, and with every
356 * answer back under the id of the question asked in that place. An answer
357 * under a name that was not sent is refused: it answers nothing that was
358 * asked. Both maps have no prototype, so an id or a name such as
359 * `__proto__` is an ordinary key.
360 */
361function unwrapAliased(
362  payload: unknown,
363  credentials: Credentials,
364  asked: Questions,
365): unknown {
366  const result = unwrapDataEnvelope(payload, credentials)
367
368  if (!isRecord(result.answers)) {
369    return result
370  }
371
372  const idOf: Record<string, string> = Object.create(null)
373  const answers: Record<string, unknown> = Object.create(null)
374
375  Object.keys(asked).forEach((id, index) => {
376    idOf[ALIAS_OF(index)] = id
377  })
378
379  for (const [alias, answer] of Object.entries(result.answers)) {
380    const id = Object.hasOwn(idOf, alias) ? idOf[alias] : undefined
381
382    if (id === undefined) {
383      throw new Error(
384        `answered ${briefOf(alias, credentials)}, which was not asked`,
385      )
386    }
387
388    answers[id] = answer
389  }
390
391  return { ...result, answers }
392}
393
394/**
395 * A model name as configured, without the white space around it: how most
396 * providers' bodies spell it.
397 */
398const TRIMMED = (model: string) => model.trim()
399
400/**
401 * A response that already is a bare System One reply.
402 */
403const AS_SENT = (payload: unknown) => payload
404
405/**
406 * The family of a model at a reseller of TypeSafe's Jev: Jev for a Jev id,
407 * the model itself otherwise.
408 */
409const JEV_FAMILY = (model: string) => (JEV.test(model) ? 'jev' : model)
410
411const TABLE: Record<ProviderName, Descriptor> = {
412  typesafe: {
413    title: 'TypeSafe',
414    defaultModel: 'jev-latest',
415    models: ['jev-latest'],
416    family: () => 'jev',
417    isGateway: false,
418    limits: { maxStateTokens: 32_000, maxRequestTokens: 64_000 },
419    needs: ['typesafeApiKey'],
420    bodyModelOf: TRIMMED,
421    urlOf: () => 'https://api.typesafe.ai/v1/systemone',
422    tokenOf: credentials => credentials.typesafeApiKey,
423    wire: BARE,
424    unwrap: AS_SENT,
425  },
426  cloudflare: {
427    title: 'Cloudflare',
428    defaultModel: 'clef',
429    models: CLOUDFLARE_MODELS,
430    family: model => model,
431    isGateway: false,
432    limits: {
433      maxStateTokens: 65_536,
434      maxRequestTokens: 65_536,
435      maxQuestions: 64,
436    },
437    needs: ['cloudflareApiToken', 'cloudflareAccountId'],
438    bodyModelOf: cloudflareModelOf,
439    urlOf: (model, credentials) =>
440      'https://api.cloudflare.com/client/v4/accounts/' +
441      `${encodeURIComponent(credentials.cloudflareAccountId ?? '')}/ai/run/` +
442      `${CLOUDFLARE_PREFIX}${model}`,
443    tokenOf: credentials => credentials.cloudflareApiToken,
444    wire: BARE,
445    unwrap: unwrapCloudflare,
446  },
447  // OpenRouter documents Jev's window as 32,000 tokens for the state and the
448  // questions together, which is tighter than TypeSafe's own 64,000 a request.
449  // Its alias for the newest Jev is `~typesafe/jev-latest`. An id that has an
450  // author prefix is sent on as written, so the same name without the `~` is
451  // an id its documentation does not list.
452  openrouter: {
453    title: 'OpenRouter',
454    defaultModel: '~typesafe/jev-latest',
455    models: ['~typesafe/jev-latest'],
456    family: model => (OPENROUTER_JEV.test(model) ? 'jev' : model),
457    isGateway: true,
458    limits: { maxStateTokens: 32_000, maxRequestTokens: 32_000 },
459    needs: ['openrouterApiKey'],
460    bodyModelOf: TRIMMED,
461    urlOf: () => 'https://openrouter.ai/api/v1/systemone',
462    tokenOf: credentials => credentials.openrouterApiKey,
463    wire: BARE,
464    unwrap: AS_SENT,
465  },
466  // Codiv serves its own open Jev, which answers as `openjev-0.1` and is a
467  // model of its own, not TypeSafe's Jev; at Codiv `jev-latest` is an alias
468  // of it. Its window is 65,536 tokens for the state and the questions
469  // together, and it advises about 60,000 for the state. It documents no
470  // fixed question cap and answered 129 questions in one request.
471  codiv: {
472    title: 'Codiv',
473    defaultModel: 'openjev-latest',
474    models: ['openjev-latest'],
475    family: model => (/^(?:open)?jev(?:-|$)/.test(model) ? 'openjev' : model),
476    isGateway: false,
477    limits: { maxStateTokens: 60_000, maxRequestTokens: 65_536 },
478    needs: ['codivApiKey'],
479    bodyModelOf: TRIMMED,
480    urlOf: () => 'https://api.codiv.ai/v1/systemone',
481    tokenOf: credentials => credentials.codivApiKey,
482    wire: BARE,
483    unwrap: AS_SENT,
484  },
485  // Perplexity refuses any model but its own deciders (HTTP 400, "Invalid
486  // model"), Jev included. A request must stay under 262,144 input tokens,
487  // the state and every question counted, and holds 1 to 128 questions.
488  perplexity: {
489    title: 'Perplexity',
490    defaultModel: 'pplx-decider-v1.1-27b',
491    models: ['pplx-decider-v1.1-27b'],
492    family: model =>
493      /^pplx-decider(?:-|$)/.test(model) ? 'pplx-decider' : model,
494    isGateway: false,
495    limits: {
496      maxStateTokens: 262_143,
497      maxRequestTokens: 262_143,
498      maxQuestions: 128,
499    },
500    needs: ['perplexityApiKey'],
501    bodyModelOf: TRIMMED,
502    urlOf: () => 'https://api.perplexity.ai/v1/decisions',
503    tokenOf: credentials => credentials.perplexityApiKey,
504    wire: BARE,
505    unwrap: AS_SENT,
506  },
507  // The two Decisions resellers forward to TypeSafe's Jev and wrap the
508  // reply in an envelope of their own. decisions-api.dev alone restricts
509  // question ids, so its requests name questions by their place.
510  'decisions-api-dev': {
511    title: 'decisions-api.dev',
512    defaultModel: 'jev-latest',
513    models: ['jev-latest'],
514    family: JEV_FAMILY,
515    isGateway: true,
516    limits: RESELLER,
517    needs: ['decisionsApiKey'],
518    bodyModelOf: TRIMMED,
519    urlOf: () => 'https://decisions-api.dev/v1/systemone',
520    tokenOf: credentials => credentials.decisionsApiKey,
521    wire: ALIASED,
522    unwrap: unwrapAliased,
523  },
524  // decisionapi.net states its 32 KiB body limit only where it describes
525  // image input; it is the platform's one stated body limit, so text is
526  // held to it as well.
527  'decisionapi-net': {
528    title: 'decisionapi.net',
529    defaultModel: 'jev-latest',
530    models: ['jev-latest'],
531    family: JEV_FAMILY,
532    isGateway: true,
533    limits: RESELLER,
534    needs: ['decisionapiApiKey'],
535    bodyModelOf: TRIMMED,
536    urlOf: () => 'https://decisionapi.net/v1/systemone',
537    tokenOf: credentials => credentials.decisionapiApiKey,
538    wire: BARE,
539    unwrap: unwrapDataEnvelope,
540  },
541  // OpenAI takes the same questions in a body of its own shape and answers
542  // in a list. Its Decisions API documents no limits, so it is held to the
543  // fallback. Those figures are this mod's, not a window OpenAI states: a
544  // reply counting all of them is taken as a cut state although the model
545  // may have read it whole, which errs toward the built-in summary.
546  openai: {
547    title: 'OpenAI',
548    defaultModel: 'gpt-6-luna',
549    models: ['gpt-6-luna'],
550    family: model => model,
551    isGateway: false,
552    limits: UNDOCUMENTED,
553    needs: ['openaiApiKey'],
554    bodyModelOf: TRIMMED,
555    urlOf: () => 'https://api.openai.com/v1/decisions',
556    tokenOf: credentials => credentials.openaiApiKey,
557    wire: OPENAI,
558    unwrap: (payload, _credentials, asked) => decodeOpenAI(payload, asked),
559  },
560}
561
562/**
563 * Narrows a configured value to a provider name.
564 *
565 * @param value the value of the `provider` option
566 * @returns the provider, or undefined when the value names none
567 */
568export function providerOf(value: unknown): ProviderName | undefined {
569  return PROVIDERS.find(name => name === value)
570}
571
572/**
573 * The limits a request over a route must keep to. OpenRouter is a gateway:
574 * a request it forwards is also held to what the model's own host accepts.
575 * Its Jev window is documented, and no question cap with it. Any other model
576 * it serves is held to Cloudflare's 64 questions as well, because Clef, the
577 * one other System One model, is served there and refuses a request with
578 * more (HTTP 422, "Dictionary should have at most 64 items").
579 *
580 * @param route the provider and the model
581 * @returns its limits
582 */
583export function limitsOf(route: Route): Limits {
584  const limits = TABLE[route.provider].limits
585
586  if (route.provider === 'openrouter' && !OPENROUTER_JEV.test(route.model)) {
587    return {
588      ...limits,
589      maxQuestions: TABLE.cloudflare.limits.maxQuestions,
590    }
591  }
592
593  return limits
594}
595
596/**
597 * The models a provider is offered with in the routing decision.
598 *
599 * @param provider the provider
600 * @returns the model names as its body spells them
601 */
602export function modelsOf(provider: ProviderName): readonly string[] {
603  return TABLE[provider].models
604}
605
606/**
607 * The model family a route reaches: a gateway's route to Jev is the same
608 * Jev that TypeSafe serves directly.
609 *
610 * @param route the route
611 * @returns the family, the model name itself for a model of its own
612 */
613export function familyOf(route: Route): string {
614  return TABLE[route.provider].family(route.model)
615}
616
617/**
618 * Names a gateway as a sentence about it spells the name.
619 *
620 * @param provider the provider
621 * @returns the name, or undefined for a provider that hosts its models
622 */
623export function gatewayOf(provider: ProviderName): string | undefined {
624  const described = TABLE[provider]
625
626  return described.isGateway ? described.title : undefined
627}
628
629/**
630 * The UTF-8 bytes a request over a route takes besides its questions: the
631 * body as encoded around one question, less that question as the route's
632 * wire format measures it. The model name and the state are counted
633 * exactly as they are sent. What is left holds no bracket of the question
634 * map or list; each question is measured with brackets of its own and a
635 * comma after it, so a body counted this way is over by a few bytes a
636 * question, never under.
637 *
638 * @param route the route
639 * @param state what the questions are asked about
640 * @returns the byte count
641 */
642export function bytesBesideOf(route: Route, state: unknown): number {
643  const described = TABLE[route.provider]
644  const question: Question = { type: 'noul', instructions: '' }
645  const body = described.wire.encode(
646    described.bodyModelOf(route.model),
647    state,
648    { q: question },
649  )
650
651  return bytesOf(body) - bytesOf(described.wire.questionJsonOf('q', question))
652}
653
654/**
655 * How a request over a route is written, which is also how its size is
656 * estimated.
657 *
658 * @param route the route
659 * @returns the provider's wire format
660 */
661export function wireOf(route: Route): Wire {
662  return TABLE[route.provider].wire
663}
664
665/**
666 * Builds the route for a provider: the configured model when one is set, the
667 * provider's default otherwise, spelled as the provider's body expects it.
668 *
669 * @param provider the provider
670 * @param model the configured model name, when one is set
671 * @returns the route; throws when the provider does not serve the model
672 */
673export function routeOf(provider: ProviderName, model?: string): Route {
674  const described = TABLE[provider]
675
676  return {
677    provider,
678    model: described.bodyModelOf(textOf(model) ?? described.defaultModel),
679  }
680}
681
682/**
683 * Names the environment variables of the credentials a provider needs and
684 * does not have. The names are safe to show; the values never are.
685 *
686 * @param provider the provider
687 * @param credentials what was resolved
688 * @returns the variable names, empty when the provider can be called
689 */
690export function missingOf(
691  provider: ProviderName,
692  credentials: Credentials,
693): EnvName[] {
694  return TABLE[provider].needs
695    .filter(need => credentials[need] === undefined)
696    .map(need => ENV_NAMES[need])
697}
698
699/**
700 * Throws when a provider lacks a credential it needs, naming the variable
701 * to set. Checked before any work is done for a request that could not be
702 * sent.
703 *
704 * @param provider the provider
705 * @param credentials what was resolved
706 */
707export function assertConfigured(
708  provider: ProviderName,
709  credentials: Credentials,
710): void {
711  const missing = missingOf(provider, credentials)
712
713  if (missing.length > 0) {
714    throw new Error(
715      `${provider} is not configured: ${missing.join(' and ')} is unset`,
716    )
717  }
718}
719
720/**
721 * Names the credential variables that still have no value after the plugin
722 * options and one source of the environment were read, among those the given
723 * providers need. An empty answer means a further source would add nothing,
724 * so it need not be read.
725 *
726 * @param providers the providers a compaction may call
727 * @param options the plugin's options
728 * @param read the variables as that source holds them
729 * @returns the variable names, each once
730 */
731export function unresolvedOf(
732  providers: readonly ProviderName[],
733  options: PluginOptions,
734  read: Environment,
735): EnvName[] {
736  const credentials = credentialsOf(options, read)
737
738  return [
739    ...new Set(providers.flatMap(provider => missingOf(provider, credentials))),
740  ]
741}
742
743/**
744 * Keeps only the credentials the given providers need. A credential of a
745 * provider that may not be called is dropped here, so no later step can
746 * send anything with it.
747 *
748 * @param providers the providers a compaction may call
749 * @param credentials what was resolved
750 * @returns the credentials of those providers
751 */
752export function credentialsFor(
753  providers: readonly ProviderName[],
754  credentials: Credentials,
755): Credentials {
756  const needed = new Set(providers.flatMap(provider => TABLE[provider].needs))
757  const kept: Credentials = {}
758
759  for (const key of needed) {
760    const value = credentials[key]
761
762    if (value !== undefined) {
763      kept[key] = value
764    }
765  }
766
767  return kept
768}
769
770/**
771 * Lays one source of the credential variables under another: a value the
772 * first source holds wins, and the `env` block of the settings fills the
773 * rest. The block is read as untyped data, since settings are a plain object.
774 *
775 * @param read the variables as the process environment holds them
776 * @param settingsEnv the `env` key of the merged settings
777 * @returns the variables with every gap the settings can fill filled
778 */
779export function environmentOf(
780  read: Environment,
781  settingsEnv: unknown,
782): Environment {
783  const fallback = isRecord(settingsEnv) ? settingsEnv : {}
784  const merged: Environment = {}
785
786  for (const name of Object.values(ENV_NAMES)) {
787    merged[name] = textOf(read[name]) ?? textOf(fallback[name])
788  }
789
790  return merged
791}
792
793/**
794 * Resolves each credential: the plugin option when it is set, the
795 * environment otherwise.
796 *
797 * @param options the plugin's options
798 * @param environment the credential variables, already merged over settings
799 * @returns the credentials that have a value
800 */
801export function credentialsOf(
802  options: PluginOptions,
803  environment: Environment,
804): Credentials {
805  const credentials: Credentials = {}
806
807  for (const [key, name] of Object.entries(ENV_NAMES)) {
808    const value = textOf(options[key]) ?? textOf(environment[name])
809
810    if (value !== undefined) {
811      credentials[key as keyof Credentials] = value
812    }
813  }
814
815  return credentials
816}
817
818/**
819 * Takes every credential out of a text that is about to be shown.
820 *
821 * A provider's error body is quoted to the person, and a body may repeat
822 * what the request carried: the key itself, or the whole `Authorization`
823 * header. Each secret that was resolved is replaced wherever it stands, the
824 * longest first so that one secret containing another leaves no remainder,
825 * and then anything that follows the word `Bearer`, which covers a token
826 * this mod never held. A resolved value under `MIN_SECRET_CHARS` is not
827 * replaced by value; the `Bearer` rule still covers it where it follows
828 * that word. The account id is not a secret and is left alone.
829 *
830 * @param text the text
831 * @param credentials the resolved credentials
832 * @returns the text with `REDACTED` where a credential stood
833 */
834export function redacted(text: string, credentials: Credentials): string {
835  const secrets = PROVIDERS.map(provider =>
836    TABLE[provider].tokenOf(credentials),
837  )
838    .filter(
839      (secret): secret is string =>
840        secret !== undefined && secret.length >= MIN_SECRET_CHARS,
841    )
842    .sort((a, b) => b.length - a.length)
843  let shown = text
844
845  for (const secret of secrets) {
846    shown = shown.split(secret).join(REDACTED)
847  }
848
849  return shown.replace(BEARER, REDACTED)
850}
851
852/**
853 * The C0 and C1 control characters: a terminal acts on them rather than
854 * showing them, so a quoted text could move the cursor or recolour what
855 * follows it.
856 */
857const CONTROL = /[\u0000-\u001f\u007f-\u009f]/g
858
859/**
860 * A text this mod did not write, made fit to quote in a message: redacted,
861 * on one line, with no control character, and no longer than
862 * `REASON_CHARS`.
863 *
864 * The text is redacted before it is shortened: a cut that fell inside a
865 * credential would leave a part of it that no longer matches the whole. The
866 * cut never falls between the two halves of a surrogate pair.
867 *
868 * @param text what a provider or the host said
869 * @param credentials the resolved credentials
870 * @returns the line to quote; `no reason given` for a text that says nothing
871 */
872export function briefOf(text: string, credentials: Credentials): string {
873  const line = redacted(text, credentials)
874    .replace(CONTROL, ' ')
875    .replace(/\s+/g, ' ')
876    .trim()
877
878  return line === '' ? 'no reason given' : cutOf(line)
879}
880
881/**
882 * A text held to `REASON_CHARS`: cut there, never between the two halves of
883 * a surrogate pair, with `…` where it was cut.
884 *
885 * @param text the text
886 * @returns the text, or its first `REASON_CHARS` characters and `…`
887 */
888export function cutOf(text: string): string {
889  return text.length > REASON_CHARS ? `${headOf(text, REASON_CHARS)}…` : text
890}
891
892/**
893 * The message of a thrown value, which need not be an `Error`.
894 *
895 * @param error what was thrown
896 * @returns its message, or the value as text
897 */
898export function messageOf(error: unknown): string {
899  return error instanceof Error ? error.message : String(error)
900}
901
902/**
903 * Builds the request for one batch of questions. The credential goes into
904 * the `authorization` header and nowhere else, so nothing built from the
905 * returned URL or body can leak it.
906 *
907 * @param route which provider and model
908 * @param credentials the resolved credentials
909 * @param state what the questions are asked about
910 * @param questions the questions, by id
911 * @returns the URL and the init for `$.http.fetch`; throws when a credential
912 * the provider needs is missing, naming its environment variable
913 */
914export function exchangeOf(
915  route: Route,
916  credentials: Credentials,
917  state: unknown,
918  questions: Questions,
919): Exchange {
920  assertConfigured(route.provider, credentials)
921
922  const described = TABLE[route.provider]
923  const token = described.tokenOf(credentials) ?? ''
924  const model = described.bodyModelOf(route.model)
925
926  return {
927    url: described.urlOf(model, credentials),
928    init: {
929      method: 'POST',
930      headers: {
931        authorization: `Bearer ${token}`,
932        'content-type': 'application/json',
933      },
934      body: described.wire.encode(model, state, questions),
935    },
936  }
937}
938
939/**
940 * Says whether a status is worth another attempt: rate limiting, an
941 * overloaded provider and any server-side failure. A 4xx other than 429 is
942 * the request's own fault and would fail the same way again.
943 *
944 * @param status the HTTP status
945 * @returns true when a retry may succeed
946 */
947export function isRetryable(status: number): boolean {
948  return status === 429 || status >= 500
949}
950
951/**
952 * Finds the sentence an error body gives as its reason, across the shapes
953 * the providers use: `error.message`, `errors[0].message`, a bare
954 * `error`, `message` or `detail` string.
955 */
956function reasonIn(payload: unknown): string | undefined {
957  if (!isRecord(payload)) {
958    return undefined
959  }
960
961  const nested = isRecord(payload.error)
962    ? textOf(payload.error.message)
963    : undefined
964
965  return (
966    nested ??
967    envelopeErrorOf(payload) ??
968    textOf(payload.error) ??
969    textOf(payload.message) ??
970    textOf(payload.detail)
971  )
972}
973
974/**
975 * The field errors of a validation failure, as `field: message` sentences:
976 * the `error.details.fieldErrors` map Workers AI answers a body it refuses
977 * with, which names what was wrong where the reason only says that it was.
978 */
979function fieldErrorsIn(payload: unknown): string | undefined {
980  const error = isRecord(payload) ? payload.error : undefined
981  const details = isRecord(error) ? error.details : undefined
982  const fields = isRecord(details) ? details.fieldErrors : undefined
983
984  if (!isRecord(fields)) {
985    return undefined
986  }
987
988  const sentences = Object.entries(fields).flatMap(([field, messages]) =>
989    Array.isArray(messages)
990      ? messages.filter(m => typeof m === 'string').map(m => `${field}: ${m}`)
991      : [],
992  )
993
994  return sentences.length > 0 ? sentences.join('; ') : undefined
995}
996
997/**
998 * How deep a reason is followed into the bodies quoted inside it.
999 */
1000const REASON_DEPTH = 4
1001
1002/**
1003 * A reason with the bodies it quotes unwrapped. A gateway passes on the
1004 * failure of the host it forwarded to as text: OpenRouter's reason is
1005 * `HTTP 422: ` and the host's body, whose own reason is `AiError: AiError: `
1006 * and another body. Only the innermost sentence says what was refused, and
1007 * it would not fit a toast line behind the wrappers, so every `HTTP nnn:`
1008 * and `AiError:` prefix is dropped and a quoted body is read for its own
1009 * reason and field errors.
1010 */
1011function innermostReason(reason: string, depth: number): string {
1012  const bare = reason.replace(/^(?:\s*(?:HTTP \d{3}|AiError):)+\s*/i, '')
1013  const start = bare.indexOf('{')
1014  const end = bare.lastIndexOf('}')
1015
1016  if (depth >= REASON_DEPTH || start !== 0 || end < start) {
1017    return bare
1018  }
1019
1020  let quoted: unknown
1021
1022  try {
1023    quoted = JSON.parse(bare.slice(start, end + 1))
1024  } catch {
1025    return bare
1026  }
1027
1028  return statedIn(quoted, depth + 1) ?? bare
1029}
1030
1031/**
1032 * What an error body states: its reason with the bodies quoted in it
1033 * unwrapped, and its field errors after it, or either alone.
1034 */
1035function statedIn(payload: unknown, depth: number): string | undefined {
1036  const reason = reasonIn(payload)
1037  const fields = fieldErrorsIn(payload)
1038
1039  return reason === undefined
1040    ? fields
1041    : fields === undefined
1042      ? innermostReason(reason, depth)
1043      : `${innermostReason(reason, depth)}: ${fields}`
1044}
1045
1046/**
1047 * The reason a failed response gives, as one short line. The response body
1048 * is the only source: the sentence it names as its reason when it is JSON
1049 * of a known shape, the body itself otherwise.
1050 */
1051function reasonOf(text: string, credentials: Credentials): string {
1052  let payload: unknown
1053
1054  try {
1055    payload = JSON.parse(text)
1056  } catch {
1057    payload = undefined
1058  }
1059
1060  return briefOf(statedIn(payload, 0) ?? text, credentials)
1061}
1062
1063/**
1064 * Turns a provider's HTTP response into a reply, or throws saying why not: a
1065 * failed status with the provider's own reason, a body that is not JSON, an
1066 * envelope reporting failure, a body without answers, or a count
1067 * of input tokens that shows the request was not read whole.
1068 *
1069 * Every message that quotes the response is redacted here, where it is
1070 * built, so no caller can show one that is not.
1071 *
1072 * @param route the route the request went over
1073 * @param response what `$.http.fetch` resolved with
1074 * @param credentials the resolved credentials, to keep out of the messages
1075 * @param asked the questions the request asked, for a provider whose answers
1076 * name them otherwise
1077 * @returns the reply
1078 */
1079export function replyFrom(
1080  route: Route,
1081  response: HttpResponse,
1082  credentials: Credentials,
1083  asked: Questions,
1084): Reply {
1085  if (!response.ok) {
1086    throw new Error(
1087      `${route.provider} answered HTTP ${response.status}: ` +
1088        reasonOf(response.text, credentials),
1089    )
1090  }
1091
1092  let payload: unknown
1093
1094  try {
1095    payload = JSON.parse(response.text)
1096  } catch {
1097    throw new Error(`${route.provider} answered a body that is not JSON`)
1098  }
1099
1100  let reply: Reply
1101
1102  try {
1103    reply = replyOf(TABLE[route.provider].unwrap(payload, credentials, asked))
1104  } catch (error) {
1105    throw new Error(
1106      `${route.provider}: ${briefOf(messageOf(error), credentials)}`,
1107    )
1108  }
1109
1110  // The size of a request is only estimated here, and a provider may cut a
1111  // state that is too long without saying so. A count that is at the limit
1112  // is what a request cut to the limit reports, and answers about a state
1113  // the model saw part of are no ground for removing anything.
1114  const counted = reply.usage.input_tokens
1115  const { maxRequestTokens } = TABLE[route.provider].limits
1116
1117  if (counted !== undefined && counted >= maxRequestTokens) {
1118    throw new Error(
1119      `${route.provider} counted ${counted} input tokens, all that a ` +
1120        `request of its may hold (${maxRequestTokens}): the state was ` +
1121        'probably cut short, so the answers decide nothing',
1122    )
1123  }
1124
1125  return reply
1126}
1127
hooks/report.ts 151 lines
1import { reductionOf } from './compact'
2import type { Action, Decision } from './compact'
3import { labelOf } from './route'
4import type { Result } from './run'
5
6/**
7 * The longest line `$.ui.log` is given; a longer decisions list is split.
8 */
9const LOG_LINE_CHARS = 4096
10
11/**
12 * What every line of the per-call verdicts starts with.
13 */
14const VERDICTS = 'per-call verdicts'
15
16/**
17 * Room kept in every line for the lead, the longest of which is
18 * `per-call verdicts, part 12 of 34: `.
19 */
20const LEAD_ROOM = 40
21
22/**
23 * What stands between two entries on a line.
24 */
25const BETWEEN = '; '
26
27/**
28 * A ratio as a whole percentage.
29 *
30 * @param ratio the ratio, 0 to 1
31 * @returns for example `42%`
32 */
33export function percentOf(ratio: number): string {
34  return `${Math.round(ratio * 100)}%`
35}
36
37function countOf(decisions: readonly Decision[], action: Action): number {
38  return decisions.filter(decision => decision.action === action).length
39}
40
41/**
42 * One line saying what a compaction did and what it cost to decide.
43 *
44 * It names the provider and model that answered, and puts the state's
45 * estimated size beside the input tokens the provider itself counted. The
46 * two disagreeing shows how far the estimate drifts; the provider's count
47 * falling well below the estimate shows a state the provider cut short.
48 *
49 * The calls are counted by what happened to them, not by what was decided:
50 * a call whose result was to be cut, but was too short to gain from it,
51 * stands whole and is counted so.
52 *
53 * @param result the compaction
54 * @returns the line, without a lead
55 */
56export function summaryOf(result: Result): string {
57  const { outcome, route } = result
58  const uncut = countOf(outcome.decisions, 'truncate') - outcome.shortened
59  const counts = (
60    [
61      ['left whole', countOf(outcome.decisions, 'keep') + uncut],
62      ['with the result cut short', outcome.shortened],
63      ['removed', countOf(outcome.decisions, 'drop')],
64      ['not judged', countOf(outcome.decisions, 'pinned')],
65    ] as const
66  )
67    .filter(([, count]) => count > 0)
68    .map(([said, count]) => `${count} ${said}`)
69  const parts = [
70    `${percentOf(reductionOf(outcome))} smaller`,
71    counts.length > 0
72      ? `tool calls: ${counts.join(', ')}`
73      : 'no answered tool call',
74  ]
75
76  if (outcome.requests > 0) {
77    const counted =
78      outcome.reported.inputTokens === undefined
79        ? ''
80        : `, ${outcome.reported.inputTokens} input tokens counted by the provider`
81    const cost =
82      outcome.reported.cost === undefined
83        ? ''
84        : `, cost $${outcome.reported.cost.toFixed(6)}`
85
86    parts.push(
87      `${labelOf(route)} in ${outcome.requests} ` +
88        `${outcome.requests === 1 ? 'request' : 'requests'}`,
89      `state about ${outcome.stateTokens} tokens estimated ` +
90        `(${outcome.stage ?? 'whole'})${counted}${cost}`,
91    )
92  }
93
94  return parts.join('; ')
95}
96
97/**
98 * The decisions as log lines: one entry per call that was asked about, with
99 * both probabilities, split so that no line exceeds what one log line holds.
100 *
101 * @param decisions every decision of the compaction
102 * @param maxChars the longest a line may be
103 * @returns the lines, in call order
104 */
105export function decisionLinesOf(
106  decisions: readonly Decision[],
107  maxChars: number = LOG_LINE_CHARS,
108): string[] {
109  const entries = decisions
110    .filter(decision => decision.action !== 'pinned')
111    .map(
112      ({ call, action, keepCall, keepResult }) =>
113        `${call.id} ${call.tool} -> ${action} ` +
114        `(call ${keepCall.toFixed(2)}, result ${keepResult.toFixed(2)})`,
115    )
116
117  // Each line is a group of entries, filled by length: an entry joins the
118  // newest group while the group, with it and the separator before it,
119  // stays within the room, and opens a group of its own otherwise. An entry
120  // longer than the room therefore still gets a line.
121  const room = Math.max(1, maxChars - LEAD_ROOM)
122  const groups: string[][] = []
123  let used = 0
124
125  for (const entry of entries) {
126    const newest = groups.at(-1)
127    const added = BETWEEN.length + entry.length
128
129    if (newest === undefined || used + added > room) {
130      groups.push([entry])
131      used = entry.length
132    } else {
133      newest.push(entry)
134      used += added
135    }
136  }
137
138  if (groups.length === 0) {
139    return [`${VERDICTS}: no tool call was asked about`]
140  }
141
142  return groups.map((group, index) => {
143    const lead =
144      groups.length === 1
145        ? VERDICTS
146        : `${VERDICTS}, part ${index + 1} of ${groups.length}`
147
148    return `${lead}: ${group.join(BETWEEN)}`
149  })
150}
151
hooks/run.ts 214 lines
1import type { SessionMessage } from 'claude-code'
2
3import { askerOf, attemptOf } from './ask'
4import type { Ports } from './ask'
5import {
6  batchesOf,
7  budgetOf,
8  compact,
9  demandOf,
10  LIMIT_SHARE,
11  untouched,
12} from './compact'
13import type { Budget, Outcome } from './compact'
14import { eligibleOf } from './config'
15import type { Config } from './config'
16import {
17  assertConfigured,
18  bytesBesideOf,
19  limitsOf,
20  messageOf,
21  routeOf,
22  wireOf,
23} from './providers'
24import type { Credentials, Route } from './providers'
25import { chooseRoute, profileOf } from './route'
26import { estimatedTokensOf, goalOf, pairCalls, stateWithin } from './state'
27import type { Fitted } from './state'
28
29/**
30 * One compaction to carry out.
31 */
32export type Job = {
33  messages: readonly SessionMessage[]
34  /**
35   * The instructions given with the compaction, when there were any.
36   */
37  instructions?: string
38  config: Config
39  credentials: Credentials
40  ports: Ports
41}
42
43/**
44 * A compaction carried out: the outcome, the route it went over, and how
45 * the route was picked when the decision was on.
46 */
47export type Result = {
48  outcome: Outcome
49  route: Route
50  routing?: string
51}
52
53/**
54 * Carries out one compaction from the transcript to the rebuilt messages.
55 *
56 * The state is fitted once, against the configured provider's limits and
57 * wire format, a cap on the body in bytes included, and every request that
58 * asks about a call then goes over one route: the configured one, or the
59 * one the provider decision picked. A route whose wire format carries the
60 * state differently is held to the state's size in that format. A failure
61 * anywhere throws; the caller falls back to the built-in summary.
62 *
63 * Everything sent to a provider, the routing question included, is sent
64 * inside one attempt: one deadline, armed here once the state is fitted and
65 * released when the outcome is known. Past it, on the engine's signal, and
66 * at the first request about tool calls that fails for good, whatever is
67 * still out is abandoned, nothing further is sent or waited for, and this
68 * throws the one reason. The routing question failing is not such an end:
69 * the configured route decides instead. Once the outcome is known, whatever
70 * is still out starts nothing more.
71 *
72 * Options that carry a refusal are not acted on at all: they name no
73 * provider the person chose, so nothing may be sent anywhere.
74 *
75 * @param job the transcript, the configuration, the credentials, the ports
76 * @returns the result
77 */
78export async function run(job: Job): Promise<Result> {
79  const { messages, config, credentials, ports } = job
80
81  if (config.refusal !== undefined) {
82    throw new Error(config.refusal)
83  }
84
85  const configured = routeOf(config.provider, config.model)
86
87  if (!config.providerDecision) {
88    assertConfigured(configured.provider, credentials)
89  }
90
91  const calls = pairCalls(messages, config.preserveRecentMessages)
92  const unpinned = calls.filter(call => !call.isPinned)
93
94  if (unpinned.length === 0) {
95    return { outcome: untouched(messages, calls), route: configured }
96  }
97
98  const demandFor = (route: Route) =>
99    demandOf(unpinned, wireOf(route), limitsOf(route).maxQuestions)
100  const budgetFor = (route: Route) =>
101    budgetOf(config, limitsOf(route), demandFor(route))
102  const goal = goalOf(messages, job.instructions)
103  const { maxRequestBytes } = limitsOf(configured)
104  const fitted = stateWithin(messages, calls, {
105    maxStateTokens: budgetFor(configured).stateTokens,
106    preserveRecentMessages: config.preserveRecentMessages,
107    goal,
108    wire: wireOf(configured),
109    // Held to the same share of a byte cap as the request is below, with
110    // the bytes of one request's questions left free.
111    ...(maxRequestBytes === undefined
112      ? {}
113      : {
114          bytes: {
115            limit:
116              Math.floor(maxRequestBytes * LIMIT_SHARE) -
117              demandFor(configured).batchBytes,
118            besideOf: (state: unknown) => bytesBesideOf(configured, state),
119          },
120        }),
121  })
122  const stateJson = JSON.stringify(fitted.state)
123  const fittedFor = (route: Route): Fitted =>
124    wireOf(route) === wireOf(configured)
125      ? fitted
126      : {
127          ...fitted,
128          tokens: estimatedTokensOf(wireOf(route).stateJsonOf(stateJson)),
129        }
130  // A provider that caps the body in bytes is held to the bytes the state
131  // takes in its request as written, counted exactly, and to the same share
132  // of its cap as of its token limits: what a question will take is still
133  // an estimate, and a body over the cap is refused whole.
134  const budgetOn = (route: Route): Budget => {
135    const budget = budgetFor(route)
136    const { maxRequestBytes } = limitsOf(route)
137
138    return maxRequestBytes === undefined
139      ? budget
140      : {
141          ...budget,
142          bytes: {
143            limit: Math.floor(maxRequestBytes * LIMIT_SHARE),
144            state: bytesBesideOf(route, fitted.state),
145          },
146        }
147  }
148  const attempt = attemptOf(ports)
149
150  try {
151    const result: Pick<Result, 'route' | 'routing'> = { route: configured }
152
153    if (config.providerDecision) {
154      const routed = await chooseRoute(
155        configured,
156        eligibleOf(config),
157        credentials,
158        route => {
159          try {
160            const budget = budgetOn(route)
161            const { tokens } = fittedFor(route)
162
163            if (tokens > budget.stateTokens) {
164              return (
165                `the state is about ${tokens} tokens and it takes ` +
166                `${budget.stateTokens}`
167              )
168            }
169
170            batchesOf(unpinned, tokens, budget, wireOf(route))
171          } catch (error) {
172            return messageOf(error)
173          }
174
175          return undefined
176        },
177        profileOf(messages, unpinned, fitted.tokens, goal),
178        // The routing question is one the compaction can do without: when
179        // it fails, the configured route decides.
180        route => askerOf(route, credentials, attempt, false),
181      )
182      const failure = attempt.failure()
183
184      // A routing question cut off by the deadline or the engine's signal
185      // is not a reason to go on with the configured route: the attempt is
186      // over, and what ended it is the reason to give.
187      if (failure !== undefined) {
188        throw failure
189      }
190
191      result.route = routed.route
192      result.routing = routed.why
193    }
194
195    const outcome = await compact(
196      messages,
197      calls,
198      fittedFor(result.route),
199      askerOf(result.route, credentials, attempt, true),
200      {
201        budget: budgetOn(result.route),
202        wire: wireOf(result.route),
203        keepThreshold: config.keepThreshold,
204        truncateHeadChars: config.truncateHeadChars,
205        halt: attempt.fail,
206      },
207    )
208
209    return { ...result, outcome }
210  } finally {
211    attempt.close()
212  }
213}
214
hooks/trigger.ts 108 lines
1/**
2 * What the `turn.complete` trigger remembers between turns.
3 */
4type Arm = {
5  /**
6   * The usage the last requested compaction left, while that is still at or
7   * above the threshold: the level the next request has to rise from.
8   */
9  floor?: number
10  /**
11   * True from a requested compaction until usage is next known. Usage is
12   * absent right after a compaction (it comes with the next response), so
13   * the floor is then taken from the first reading that follows.
14   */
15  isFloorPending: boolean
16}
17
18/**
19 * How far usage has to rise above the floor before another compaction is
20 * requested, in points of the context window.
21 */
22const REARM_POINTS = 10
23
24/**
25 * A trigger that has requested nothing yet.
26 *
27 * @returns the initial state
28 */
29export function armOf(): Arm {
30  return { isFloorPending: false }
31}
32
33/**
34 * Says whether a compaction should be requested at this reading of usage,
35 * and updates what is remembered.
36 *
37 * A reading at or above the threshold asks for one, unless an earlier
38 * request left usage near where it is now: a compaction that could not get
39 * usage under the threshold would otherwise be requested again at the end
40 * of every turn. The next request then waits until usage has risen
41 * `REARM_POINTS` above that floor. A reading under the threshold forgets
42 * the floor. When the answer is yes the floor is set to the reading at
43 * once, so a request that then fails is not repeated on the next turn.
44 *
45 * @param arm what is remembered, changed in place
46 * @param percent the context usage in percent, absent when it is not known
47 * @param threshold the `compactAtPercent` option
48 * @returns true when a compaction should be requested now
49 */
50export function isCompactionDue(
51  arm: Arm,
52  percent: number | undefined,
53  threshold: number,
54): boolean {
55  if (percent === undefined) {
56    return false
57  }
58
59  if (percent < threshold) {
60    delete arm.floor
61    arm.isFloorPending = false
62
63    return false
64  }
65
66  if (arm.isFloorPending) {
67    arm.floor = percent
68    arm.isFloorPending = false
69  }
70
71  if (arm.floor !== undefined && percent < arm.floor + REARM_POINTS) {
72    return false
73  }
74
75  arm.floor = percent
76
77  return true
78}
79
80/**
81 * Records what a requested compaction left behind.
82 *
83 * A reading at or above the threshold becomes the floor: the compaction
84 * could not get usage under it, so the next request has to wait for a rise.
85 * A reading under the threshold leaves no floor, because the compaction
86 * worked and the next crossing of the threshold is a new one; a floor kept
87 * there would hold the next request back until usage was `REARM_POINTS`
88 * above a level that was never a problem, which a high threshold can put
89 * past the top of the window. Usage that is not known yet marks the floor as
90 * pending.
91 *
92 * @param arm what is remembered, changed in place
93 * @param percent the context usage read after the compaction
94 * @param threshold the `compactAtPercent` option
95 */
96export function settle(
97  arm: Arm,
98  percent: number | undefined,
99  threshold: number,
100): void {
101  delete arm.floor
102  arm.isFloorPending = percent === undefined
103
104  if (percent !== undefined && percent >= threshold) {
105    arm.floor = percent
106  }
107}
108
hooks/state.ts 779 lines
1import type { SessionMessage, ToolUseSummary } from 'claude-code'
2
3import { bytesOf } from './systemone'
4import type { Wire } from './systemone'
5
6/**
7 * One tool call with the result that answers it, found by `tool_use_id`.
8 */
9export type Call = {
10  /**
11   * A name of a few characters (`t1`, `t2`, ...) that stands for the call in
12   * the state and in the question ids, where the engine's own `tool_use_id`
13   * would be paid for, character by character, in every request.
14   */
15  id: string
16  toolUseId: string
17  tool: string
18  input: Record<string, unknown>
19  /**
20   * Where in the conversation the assistant made the call.
21   */
22  callIndex: number
23  /**
24   * Where in the conversation the answer to the call arrived.
25   */
26  resultIndex: number
27  resultChars: number
28  isError: boolean
29  /**
30   * True when the call is shown to the model but never asked about, so it
31   * stays exactly as it is. That holds when either of its two messages is
32   * one that is never changed, and when its `tool_use_id` is not unique in
33   * the conversation: a decision is applied by that id, so one taken for
34   * such a call would also reach the other blocks that carry it.
35   */
36  isPinned: boolean
37}
38
39/**
40 * A call as the state shows it: its input, and a note in place of its output.
41 */
42export type CallNote = {
43  id: string
44  tool: string
45  input: string
46  result: string
47}
48
49/**
50 * One message of the conversation as the state shows it. `calls` holds
51 * notes, or one line per call once the state had to shrink that far.
52 */
53type Entry = {
54  at: number
55  role: SessionMessage['role']
56  text: string
57  calls?: CallNote[] | string[]
58}
59
60/**
61 * What every request sends as `state`.
62 */
63export type State = {
64  context: string
65  goal: string
66  conversation: Entry[]
67}
68
69/**
70 * How far the state had to be reduced before it fitted, for the report.
71 */
72export type Stage =
73  | 'whole'
74  | 'inputs shortened'
75  | 'long texts cut in the middle'
76  | 'old texts replaced by their length'
77  | 'old calls on one line'
78  | 'old call-free messages omitted'
79  | 'rows of old calls joined'
80
81/**
82 * A state that fits its budget.
83 */
84export type Fitted = {
85  state: State
86  /**
87   * The estimated size of the serialised state, as the wire format it was
88   * fitted for carries it.
89   */
90  tokens: number
91  stage: Stage
92}
93
94/**
95 * What the fitting needs to know.
96 */
97type FitNeed = {
98  maxStateTokens: number
99  preserveRecentMessages: number
100  goal: string
101  /**
102   * The wire format the state is sent in. It is measured as that format
103   * carries it: a format that escapes the state inside a string makes it
104   * larger than its own JSON.
105   */
106  wire: Wire
107  /**
108   * For a route that caps the request body in bytes: the UTF-8 bytes the
109   * body may take besides the questions it must leave room for, and the
110   * bytes it takes around a state, questions left out. The state is then
111   * held to both budgets; absent, it is fitted in tokens alone.
112   */
113  bytes?: { limit: number; besideOf: (state: State) => number }
114}
115
116/**
117 * The frame every question is read in. It carries the rubric once, so the
118 * per-call questions can stay short: they are repeated for every call.
119 */
120const CONTEXT =
121  'This is the transcript of a coding assistant session whose context ' +
122  'window is nearly full. To free space, old tool calls and their outputs ' +
123  'are being removed; nothing is summarised or reworded. `conversation` ' +
124  'lists the messages in order, oldest first. A tool call appears under ' +
125  '`calls` with its `id` and its input. Its output is not shown: `result` ' +
126  'only says whether it succeeded and how long it was. Long texts may be ' +
127  'shortened and old messages may be missing. `goal` says what the ' +
128  'assistant is working on. Each question is about one tool call, named by ' +
129  'its `id`. What is removed is gone for good, but the assistant can run ' +
130  'any tool again and read any file again.'
131
132/**
133 * The successive caps on a call's serialised input, in characters.
134 */
135const INPUT_CAPS = [1000, 240, 64] as const
136
137/**
138 * What stays of a long text when it is abridged: its start and its end.
139 */
140const KEPT_START_CHARS = 400
141const KEPT_END_CHARS = 160
142
143/**
144 * A text is abridged only past this length, so the note that replaces its
145 * middle is always shorter than what it replaces.
146 */
147const LONG_TEXT = KEPT_START_CHARS + KEPT_END_CHARS + 48
148
149/**
150 * The longest the note standing for a collapsed text gets; a text no longer
151 * than this is left alone, since the note would not be shorter.
152 */
153const COLLAPSED_NOTE_CHARS = 48
154
155const GOAL_CHARS = 2000
156const PROMPT_CHARS = 500
157const GOAL_PROMPTS = 3
158
159/**
160 * Letters one token is assumed to cover inside a word.
161 */
162const LETTERS_PER_TOKEN = 6
163
164function isLetter(code: number): boolean {
165  return (code >= 65 && code <= 90) || (code >= 97 && code <= 122)
166}
167
168function isDigit(code: number): boolean {
169  return code >= 48 && code <= 57
170}
171
172function isHighSurrogate(code: number): boolean {
173  return code >= 0xd800 && code <= 0xdbff
174}
175
176/**
177 * Estimates how many tokens a provider will count for a text, with no
178 * tokenizer: a hooks module cannot load one.
179 *
180 * Each character is weighed by its class:
181 *
182 * - ASCII letters: 1 token for each started group of `LETTERS_PER_TOKEN`
183 *   in an unbroken run of them;
184 * - ASCII digits: 0.5 each;
185 * - other visible ASCII characters: 0.9 each;
186 * - characters beyond ASCII: 1 each, which is about what CJK text comes to;
187 * - white space: nothing.
188 *
189 * The weights follow the design this mod is modelled on, which reports them
190 * landing a little above Jev's own count; they are an estimate, so the
191 * budgets they are compared with must leave headroom.
192 */
193export function estimatedTokensOf(text: string): number {
194  // Counted in tenths of a token so the sum stays in whole numbers: adding
195  // 0.9 repeatedly drifts, and the drift would move the rounded result.
196  let tenths = 0
197  let letters = 0
198
199  for (let index = 0; index < text.length; index++) {
200    const code = text.charCodeAt(index)
201
202    if (isLetter(code)) {
203      letters++
204      continue
205    }
206
207    tenths += Math.ceil(letters / LETTERS_PER_TOKEN) * 10
208    letters = 0
209
210    if (isDigit(code)) {
211      tenths += 5
212    } else if (code > 127) {
213      tenths += 10
214    } else if (code > 32) {
215      tenths += 9
216    }
217  }
218
219  tenths += Math.ceil(letters / LETTERS_PER_TOKEN) * 10
220
221  return Math.ceil(tenths / 10)
222}
223
224/**
225 * The first `count` characters of a text, never ending between the two
226 * halves of a surrogate pair: a lone half does not survive JSON transport.
227 *
228 * @param count how many characters at most
229 * @returns the start of the text
230 */
231export function headOf(text: string, count: number): string {
232  const end = Math.max(0, Math.min(count, text.length))
233
234  return end > 0 &&
235    end < text.length &&
236    isHighSurrogate(text.charCodeAt(end - 1))
237    ? text.slice(0, end - 1)
238    : text.slice(0, end)
239}
240
241/**
242 * The last `count` characters of a text, never starting between the two
243 * halves of a surrogate pair.
244 */
245function tailOf(text: string, count: number): string {
246  const start = Math.max(0, text.length - count)
247
248  return start > 0 && isHighSurrogate(text.charCodeAt(start - 1))
249    ? text.slice(start + 1)
250    : text.slice(start)
251}
252
253/**
254 * A text cut to a length, with a mark where it was cut.
255 */
256function clipped(text: string, limit: number): string {
257  if (text.length > limit) {
258    return `${headOf(text, limit - 1)}…`
259  }
260
261  return text
262}
263
264/**
265 * A long text with its middle replaced by a note of how much is missing.
266 */
267function abridged(text: string): string {
268  const head = headOf(text, KEPT_START_CHARS)
269  const tail = tailOf(text, KEPT_END_CHARS)
270  const missing = text.length - head.length - tail.length
271
272  return `${head}\n[${missing} characters left out]\n${tail}`
273}
274
275/**
276 * Says whether a message is one that is never changed: the first, which
277 * usually states the task, and the newest `recent`, which the assistant is
278 * still working from.
279 *
280 * @param index the message's index
281 * @param total how many messages the conversation has
282 * @param recent how many of the newest are kept as they are
283 * @returns true when the message is pinned
284 */
285function isPinnedIndex(
286  index: number,
287  total: number,
288  recent: number,
289): boolean {
290  // 1 for the newest message, 2 for the one before it, and so on.
291  const fromNewest = total - index
292
293  return index === 0 || fromNewest <= recent
294}
295
296/**
297 * One tool use and where it stands: the message that holds it, and its
298 * place among all the tool uses of the conversation.
299 */
300type Site = {
301  use: ToolUseSummary
302  at: number
303  order: number
304}
305
306/**
307 * Lists the calls that have been answered, in the order they were made,
308 * each joined to its answer through the `tool_use_id` the two share. A call
309 * still waiting for its answer is not listed: removing it would leave the
310 * answer, when it arrives, with nothing to belong to.
311 *
312 * An id is expected on one tool use and on one result. Where it stands on
313 * more than one of either, every call that carries it is marked as not to
314 * be judged: what is decided for a call is applied by its id, so it would
315 * reach each block with that id, one in a pinned message included. Such a
316 * call is joined to the first result with its id.
317 *
318 * @param messages the conversation
319 * @param recent how many of the newest messages are pinned
320 * @returns the paired calls, the ones not to be judged included
321 */
322export function pairCalls(
323  messages: readonly SessionMessage[],
324  recent: number,
325): Call[] {
326  const sites = new Map<string, Site[]>()
327  let order = 0
328
329  for (const [at, message] of messages.entries()) {
330    for (const use of message.toolUses) {
331      const site: Site = { use, at, order: order++ }
332      const sharing = sites.get(use.tool_use_id)
333
334      if (sharing === undefined) {
335        sites.set(use.tool_use_id, [site])
336      } else {
337        sharing.push(site)
338      }
339    }
340  }
341
342  const answered = new Set<string>()
343  const repeated = new Set<string>()
344  const pairs: {
345    site: Site
346    resultIndex: number
347    resultChars: number
348    isError: boolean
349  }[] = []
350
351  for (const [resultIndex, message] of messages.entries()) {
352    for (const { tool_use_id: id, text, isError } of message.toolResults ??
353      []) {
354      if (answered.has(id)) {
355        repeated.add(id)
356        continue
357      }
358
359      answered.add(id)
360
361      for (const site of sites.get(id) ?? []) {
362        pairs.push({ site, resultIndex, resultChars: text.length, isError })
363      }
364    }
365  }
366
367  const isKept = (at: number) => isPinnedIndex(at, messages.length, recent)
368  const isShared = (id: string) =>
369    repeated.has(id) || (sites.get(id)?.length ?? 0) > 1
370
371  return pairs
372    .sort((a, b) => a.site.order - b.site.order)
373    .map(({ site, resultIndex, resultChars, isError }, index) => ({
374      id: `t${index + 1}`,
375      toolUseId: site.use.tool_use_id,
376      tool: site.use.tool,
377      input: site.use.input,
378      callIndex: site.at,
379      resultIndex,
380      resultChars,
381      isError,
382      isPinned:
383        isShared(site.use.tool_use_id) ||
384        isKept(site.at) ||
385        isKept(resultIndex),
386    }))
387}
388
389/**
390 * Says whether a message is something the person typed: a user message
391 * with text that brings no tool result.
392 */
393function isPrompt(message: SessionMessage): boolean {
394  return (
395    message.role === 'user' &&
396    (message.toolResults?.length ?? 0) === 0 &&
397    message.text.trim() !== ''
398  )
399}
400
401/**
402 * Says what the assistant is working on: the instructions given with the
403 * compaction when there are any (the text typed after `/compact`, or a
404 * plugin's), since they were written for exactly this; otherwise the last
405 * prompts the person typed.
406 *
407 * @param messages the conversation
408 * @param instructions the compaction's instructions, when given
409 * @returns the goal text, possibly empty
410 */
411export function goalOf(
412  messages: readonly SessionMessage[],
413  instructions?: string,
414): string {
415  const stated = instructions?.trim() ?? ''
416
417  if (stated !== '') {
418    return clipped(stated, GOAL_CHARS)
419  }
420
421  // Read from the newest message back, and stop at the first few prompts:
422  // the conversation before them has no part in the goal.
423  const prompts: string[] = []
424
425  for (
426    let at = messages.length - 1;
427    at >= 0 && prompts.length < GOAL_PROMPTS;
428    at--
429  ) {
430    const message = messages[at]
431
432    if (message !== undefined && isPrompt(message)) {
433      prompts.unshift(clipped(message.text.trim(), PROMPT_CHARS))
434    }
435  }
436
437  return prompts.join('\n')
438}
439
440function statusOf(call: Call): string {
441  return call.isError ? 'error' : 'ok'
442}
443
444/**
445 * A call's input as JSON text. An input that cannot be serialised (a cycle)
446 * is named rather than allowed to fail the whole compaction.
447 */
448function inputJsonOf(call: Call): string {
449  try {
450    return JSON.stringify(call.input) ?? '{}'
451  } catch {
452    return '[input that cannot be shown]'
453  }
454}
455
456/**
457 * One argument of a call as `name=value` on a single line: a string as it
458 * is, any other value as its JSON, every stretch of white space as one
459 * space.
460 */
461function argumentOf(name: string, value: unknown): string {
462  const shown =
463    typeof value === 'string' ? value : String(JSON.stringify(value))
464
465  return `${name}=${shown.replace(/\s+/g, ' ')}`
466}
467
468/**
469 * A call on one line, for when a structured note costs too much. Its parts,
470 * a space apart: the id, the tool with its arguments in brackets (held to
471 * the tightest input cap), the outcome, and the length of the result.
472 */
473function lineOf(call: Call): string {
474  const said = Object.keys(call.input)
475    .map(name => argumentOf(name, call.input[name]))
476    .join(' ')
477
478  return [
479    call.id,
480    `${call.tool}(${clipped(said, INPUT_CAPS[2])})`,
481    statusOf(call),
482    `${call.resultChars}ch`,
483  ].join(' ')
484}
485
486/**
487 * One entry of the working list: the entry, what it costs, and whether it
488 * may be reduced.
489 */
490type Slot = {
491  entry: Entry
492  tokens: number
493  /**
494   * What the entry takes in UTF-8 bytes, comma included; 0 when the state
495   * has no byte budget.
496   */
497  bytes: number
498  isOld: boolean
499  isOut: boolean
500}
501
502/**
503 * One reduction: the stage it reports, the slots it visits in order, which
504 * of them it applies to, and what it does to one.
505 */
506type Pass = {
507  stage: Stage
508  over: readonly Slot[]
509  applies: (slot: Slot) => boolean
510  reduce: (slot: Slot) => void
511}
512
513/**
514 * Where several entries in a row hold nothing but one-line calls, gives the
515 * lines of all of them to the first and drops the rest. The keys of an
516 * entry are then paid for once for the whole row, and every line still
517 * starts with the id its questions name. `reweigh` measures a slot again
518 * once its entry has grown.
519 */
520function folded(
521  slots: readonly Slot[],
522  reweigh: (slot: Slot) => void,
523): Slot[] {
524  const isFoldable = (slot: Slot) =>
525    slot.isOld &&
526    slot.entry.text === '' &&
527    typeof slot.entry.calls?.[0] === 'string'
528  const kept: Slot[] = []
529  const grown = new Set<Slot>()
530
531  for (const slot of slots) {
532    const last = kept.at(-1)
533
534    if (
535      last !== undefined &&
536      isFoldable(last) &&
537      isFoldable(slot) &&
538      last.entry.role === slot.entry.role
539    ) {
540      // The first merge of a run gives the entry a list of its own; every
541      // later one appends to it. Copying the list or weighing the entry per
542      // merge would cost the square of the run's length, and a session of
543      // nothing but tool calls is one long run.
544      if (!grown.has(last)) {
545        last.entry.calls = [...(last.entry.calls as string[])]
546        grown.add(last)
547      }
548
549      ;(last.entry.calls as string[]).push(...(slot.entry.calls as string[]))
550      continue
551    }
552
553    kept.push(slot)
554  }
555
556  for (const slot of grown) {
557    reweigh(slot)
558  }
559
560  return kept
561}
562
563/**
564 * Describes the conversation to the decision model within `maxStateTokens`.
565 *
566 * The model judges better the more of the session it can see, so the state
567 * starts as everything and gives up detail only while its estimate is over
568 * the budget, cheapest loss first, and stops the moment it is under. Each
569 * rung below is tried only when the one above was not enough:
570 *
571 * 1. a cap on the JSON of each call's input, tightened twice;
572 * 2. the middle of every long text, in old messages before pinned ones;
573 * 3. the text of an old message, of which only its length is then stated;
574 * 4. the structure of an old call, written as a single line instead;
575 * 5. old messages in which no call was made;
576 * 6. the separate entries of neighbouring old messages that hold only calls.
577 *
578 * The estimate is checked here, whichever provider is asked, because a
579 * provider may cut a state that is too long without saying so. Where the
580 * route caps the body in bytes, the state is held to that budget too, on
581 * the same rungs: text with many letters to a token takes more bytes than
582 * the token budget assumes. A conversation still over either budget on the
583 * last rung is refused by throwing.
584 *
585 * @param messages the conversation
586 * @param calls its paired calls, as `pairCalls` answered them
587 * @param need the budget, the pinning, the goal and the wire format
588 * @returns the fitted state, its estimated size and the stage reached
589 */
590export function stateWithin(
591  messages: readonly SessionMessage[],
592  calls: readonly Call[],
593  need: FitNeed,
594): Fitted {
595  const frame = (conversation: Entry[]): State => ({
596    context: CONTEXT,
597    goal: need.goal,
598    conversation,
599  })
600  const tokensOf = (state: State) =>
601    estimatedTokensOf(need.wire.stateJsonOf(JSON.stringify(state)))
602  const frameTokens = tokensOf(frame([]))
603  const { bytes } = need
604  const frameBytes = bytes === undefined ? 0 : bytes.besideOf(frame([]))
605  // What an entry costs inside the list: its own JSON as the wire format
606  // carries it and the comma after it, in tokens and in UTF-8 bytes. The
607  // last entry has no comma, so a sum of these is over by one byte at most.
608  const weigh = (entry: Entry): Pick<Slot, 'tokens' | 'bytes'> => {
609    const json = need.wire.stateJsonOf(JSON.stringify(entry))
610
611    return {
612      tokens: estimatedTokensOf(json) + 1,
613      bytes: bytes === undefined ? 0 : bytesOf(json) + 1,
614    }
615  }
616  const inputs = new Map(calls.map(call => [call, inputJsonOf(call)]))
617  const made = new Map<number, Call[]>()
618
619  for (const call of calls) {
620    made.set(call.callIndex, [...(made.get(call.callIndex) ?? []), call])
621  }
622
623  let slots: Slot[] = []
624  let total = 0
625  let totalBytes = 0
626
627  const tally = () => {
628    total = frameTokens
629    totalBytes = frameBytes
630
631    for (const slot of slots) {
632      total += slot.tokens
633      totalBytes += slot.bytes
634    }
635  }
636
637  const lay = (inputCap: number) => {
638    slots = []
639
640    messages.forEach((message, at) => {
641      const own = made.get(at) ?? []
642
643      if (message.text.trim() === '' && own.length === 0) {
644        return
645      }
646
647      const entry: Entry = {
648        at,
649        role: message.role,
650        text: message.text.trim() === '' ? '' : message.text,
651      }
652
653      if (own.length > 0) {
654        entry.calls = own.map(call => ({
655          id: call.id,
656          tool: call.tool,
657          input: clipped(inputs.get(call) ?? '', inputCap),
658          result: `${statusOf(call)}, ${call.resultChars} characters, not shown`,
659        }))
660      }
661
662      slots.push({
663        entry,
664        ...weigh(entry),
665        isOld: !isPinnedIndex(at, messages.length, need.preserveRecentMessages),
666        isOut: false,
667      })
668    })
669
670    tally()
671  }
672
673  const fits = () =>
674    total <= need.maxStateTokens &&
675    (bytes === undefined || totalBytes <= bytes.limit)
676
677  const fitted = (stage: Stage): Fitted => {
678    const state = frame(
679      slots.filter(slot => !slot.isOut).map(slot => slot.entry),
680    )
681
682    return { state, tokens: tokensOf(state), stage }
683  }
684
685  for (const cap of INPUT_CAPS) {
686    lay(cap)
687
688    if (fits()) {
689      return fitted(cap === INPUT_CAPS[0] ? 'whole' : 'inputs shortened')
690    }
691  }
692
693  const old = slots.filter(slot => slot.isOld)
694  const oldFirst = [...old, ...slots.filter(slot => !slot.isOld)]
695
696  const passes: readonly Pass[] = [
697    {
698      stage: 'long texts cut in the middle',
699      over: oldFirst,
700      applies: slot => slot.entry.text.length > LONG_TEXT,
701      reduce: slot => {
702        slot.entry.text = abridged(slot.entry.text)
703      },
704    },
705    {
706      stage: 'old texts replaced by their length',
707      over: old,
708      applies: slot => slot.entry.text.length > COLLAPSED_NOTE_CHARS,
709      reduce: slot => {
710        const length = messages[slot.entry.at]?.text.length ?? 0
711
712        slot.entry.text = `[${length} characters of text left out]`
713      },
714    },
715    {
716      stage: 'old calls on one line',
717      over: old,
718      applies: slot => slot.entry.calls !== undefined,
719      reduce: slot => {
720        slot.entry.calls = (made.get(slot.entry.at) ?? []).map(lineOf)
721      },
722    },
723    {
724      stage: 'old call-free messages omitted',
725      over: old,
726      applies: slot => slot.entry.calls === undefined,
727      reduce: slot => {
728        slot.isOut = true
729      },
730    },
731  ]
732
733  for (const pass of passes) {
734    for (const slot of pass.over) {
735      if (!pass.applies(slot)) {
736        continue
737      }
738
739      pass.reduce(slot)
740
741      const now = slot.isOut ? { tokens: 0, bytes: 0 } : weigh(slot.entry)
742
743      total += now.tokens - slot.tokens
744      totalBytes += now.bytes - slot.bytes
745      slot.tokens = now.tokens
746      slot.bytes = now.bytes
747
748      if (fits()) {
749        return fitted(pass.stage)
750      }
751    }
752  }
753
754  slots = folded(
755    slots.filter(slot => !slot.isOut),
756    slot => {
757      Object.assign(slot, weigh(slot.entry))
758    },
759  )
760  tally()
761
762  if (fits()) {
763    return fitted('rows of old calls joined')
764  }
765
766  if (total <= need.maxStateTokens && bytes !== undefined) {
767    throw new Error(
768      'the conversation does not fit the request body: about ' +
769        `${totalBytes} bytes are left after every reduction and ` +
770        `${bytes.limit} are allowed`,
771    )
772  }
773
774  throw new Error(
775    `the conversation does not fit the state budget: about ${total} tokens ` +
776      `are left after every reduction and ${need.maxStateTokens} are allowed`,
777  )
778}
779
hooks/systemone.ts 244 lines
1/**
2 * A yes/no question; the answer is the probability of yes.
3 */
4type NoulQuestion = {
5  type: 'noul'
6  instructions: string
7  criteria?: { true?: string; false?: string }
8}
9
10/**
11 * A pick of one option; `criteria` maps each option to what it stands for.
12 */
13export type ChoiceQuestion = {
14  type: 'choice'
15  instructions: string
16  criteria: Record<string, string | null>
17}
18
19export type Question = NoulQuestion | ChoiceQuestion
20
21/**
22 * The questions of one request, by the id their answers come back under.
23 */
24export type Questions = Record<string, Question>
25
26/**
27 * What a provider reported a request cost. Every field is optional because
28 * the eight providers do not report the same fields.
29 */
30type Usage = {
31  input_tokens?: number
32  cost?: number
33}
34
35/**
36 * One response, unwrapped from whatever envelope its provider put around it.
37 */
38export type Reply = {
39  model?: string
40  answers: Record<string, unknown>
41  usage: Usage
42}
43
44/**
45 * The picked option of a `choice` answer and how sure the model was.
46 */
47type Picked = {
48  choice: string
49  confidence?: number
50}
51
52/**
53 * The rule every question id must pass before any provider-specific alias
54 * is applied: letters, digits, `_`, `.` and `-`, 100 characters at most. A
55 * provider with a stricter rule is sent aliases in place of the ids.
56 */
57const QUESTION_ID = /^[A-Za-z0-9_.-]{1,100}$/
58
59/**
60 * Whether a decoded JSON value is an object, as opposed to an array, `null`
61 * or a scalar.
62 */
63export function isRecord(value: unknown): value is Record<string, unknown> {
64  return typeof value === 'object' && value !== null && !Array.isArray(value)
65}
66
67function numberOf(value: unknown): number | undefined {
68  return typeof value === 'number' && Number.isFinite(value) ? value : undefined
69}
70
71/**
72 * How one provider's request body is written, and how big what it carries
73 * comes out once written: the token estimate is taken of the text that is
74 * sent, not of the objects it was written from.
75 */
76export type Wire = {
77  /**
78   * Serialises one request.
79   */
80  encode: (model: string, state: unknown, questions: Questions) => string
81  /**
82   * The JSON one question takes in the body, its id included.
83   */
84  questionJsonOf: (id: string, question: Question) => string
85  /**
86   * The JSON of the state, or of a part of it, as it stands in the body:
87   * as it is, or escaped inside a string.
88   */
89  stateJsonOf: (json: string) => string
90  /**
91   * The tokens a provider counts for each question beyond its own text.
92   */
93  tokensPerQuestion: number
94}
95
96const UTF8 = new TextEncoder()
97
98/**
99 * The size of a text once sent, in UTF-8 bytes: what a cap stated in bytes
100 * is held against. A character beyond ASCII takes two to four of them.
101 */
102export function bytesOf(text: string): number {
103  return UTF8.encode(text).length
104}
105
106/**
107 * Checks the ids of one request. An id a provider would reject is refused
108 * before anything is sent: a 422 on one id would otherwise cost the whole
109 * batch it rides in.
110 *
111 * @param questions the questions, by id
112 * @returns the ids; throws when there is none or one is not a valid id
113 */
114export function idsOf(questions: Questions): string[] {
115  const ids = Object.keys(questions)
116
117  if (ids.length === 0) {
118    throw new Error('a System One request needs at least one question')
119  }
120
121  for (const id of ids) {
122    if (!QUESTION_ID.test(id)) {
123      throw new Error(`question id ${JSON.stringify(id)} is not a valid id`)
124    }
125  }
126
127  return ids
128}
129
130/**
131 * Serialises one request in the System One form: the three fields every
132 * provider that speaks the protocol as published takes.
133 *
134 * @param model the model name as the provider's body spells it
135 * @param state what every question is asked about
136 * @param questions the questions, by id
137 * @returns the JSON text of the request body
138 */
139export function bodyOf(
140  model: string,
141  state: unknown,
142  questions: Questions,
143): string {
144  idsOf(questions)
145
146  return JSON.stringify({ model, state, questions })
147}
148
149/**
150 * The System One body: the state as it is and the questions as a map by id.
151 */
152export const BARE: Wire = {
153  encode: bodyOf,
154  questionJsonOf: (id, question) => JSON.stringify({ [id]: question }),
155  stateJsonOf: json => json,
156  tokensPerQuestion: 0,
157}
158
159/**
160 * Reads a decoded response body as a reply. Only the `answers` map is
161 * required; each answer is checked when it is read, by the reader that knows
162 * which type the question had.
163 *
164 * @param payload the decoded body, already out of any provider envelope
165 */
166export function replyOf(payload: unknown): Reply {
167  if (!isRecord(payload) || !isRecord(payload.answers)) {
168    throw new Error('the response holds no answers object')
169  }
170
171  const reported = isRecord(payload.usage) ? payload.usage : {}
172  const usage: Usage = {}
173
174  for (const field of ['input_tokens', 'cost'] as const) {
175    const value = numberOf(reported[field])
176
177    if (value !== undefined) {
178      usage[field] = value
179    }
180  }
181
182  const reply: Reply = { answers: payload.answers, usage }
183
184  if (typeof payload.model === 'string') {
185    reply.model = payload.model
186  }
187
188  return reply
189}
190
191/**
192 * Reads the answer to a `noul` question: a probability, so a finite number
193 * from 0 to 1. Anything else throws, because a decision taken on a missing or
194 * out-of-range value would delete or keep a tool result for no reason.
195 *
196 * @param reply the reply the answer is in
197 * @param id the question's id
198 * @returns the probability of yes
199 */
200export function noulOf(reply: Reply, id: string): number {
201  const answer = reply.answers[id]
202  const value = isRecord(answer) ? numberOf(answer.noul) : undefined
203
204  if (value === undefined || value < 0 || value > 1) {
205    throw new Error(`no probability from 0 to 1 was answered for ${id}`)
206  }
207
208  return value
209}
210
211/**
212 * Reads the answer to a `choice` question. The pick must be one of the
213 * options that were offered: a name outside them cannot be acted on.
214 *
215 * @param reply the reply the answer is in
216 * @param id the question's id
217 * @param options the option names the question offered
218 * @returns the picked option and the confidence, when one was reported
219 */
220export function choiceOf(
221  reply: Reply,
222  id: string,
223  options: readonly string[],
224): Picked {
225  const answer = reply.answers[id]
226
227  if (
228    !isRecord(answer) ||
229    typeof answer.choice !== 'string' ||
230    !options.includes(answer.choice)
231  ) {
232    throw new Error(`no offered option was answered for ${id}`)
233  }
234
235  const picked: Picked = { choice: answer.choice }
236  const confidence = numberOf(answer.confidence)
237
238  if (confidence !== undefined) {
239    picked.confidence = confidence
240  }
241
242  return picked
243}
244
hooks/openai.ts 201 lines
1import { idsOf, isRecord } from './systemone'
2import type { Question, Questions, Wire } from './systemone'
3
4/**
5 * One question as OpenAI's Decisions API takes it: a list entry that names
6 * itself, where System One keys a map by the id.
7 */
8type OpenAIQuestion =
9  | { name: string; type: 'predicate'; instructions: string }
10  | {
11      name: string
12      type: 'choice'
13      instructions: string
14      choices: { value: string; description?: string }[]
15    }
16
17/**
18 * The longest JSON of a value OpenAI sent that an error quotes.
19 */
20const QUOTED_CHARS = 64
21
22/**
23 * A value OpenAI sent, made fit to quote in an error: its JSON, which
24 * escapes every control character, when that is short, and only its length
25 * otherwise. It is never cut: the message is redacted where it is shown,
26 * and a credential cut short would no longer be found there.
27 */
28function quoted(value: unknown): string {
29  const json = JSON.stringify(value)
30
31  if (json === undefined) {
32    return 'missing'
33  }
34
35  return json.length <= QUOTED_CHARS
36    ? json
37    : `a value of ${json.length} characters`
38}
39
40/**
41 * A text made to end a sentence, so that two of them read as two sentences
42 * once appended to the instructions.
43 */
44function sentenceOf(text: string): string {
45  const trimmed = text.trim()
46
47  return /[.!?]$/.test(trimmed) ? trimmed : `${trimmed}.`
48}
49
50/**
51 * A question in OpenAI's form. A `predicate` has no field for what yes and
52 * no stand for, so a `noul` question's criteria are appended to its
53 * instructions, one sentence each; a `choice` option's description is
54 * omitted where System One gives `null`.
55 */
56function questionOf(name: string, question: Question): OpenAIQuestion {
57  if (question.type === 'choice') {
58    return {
59      name,
60      type: 'choice',
61      instructions: question.instructions,
62      choices: Object.entries(question.criteria).map(([value, description]) =>
63        description === null ? { value } : { value, description },
64      ),
65    }
66  }
67
68  const said = [question.instructions]
69  const criteria = question.criteria ?? {}
70
71  if (criteria.true !== undefined) {
72    said.push(`True means: ${sentenceOf(criteria.true)}`)
73  }
74
75  if (criteria.false !== undefined) {
76    said.push(`False means: ${sentenceOf(criteria.false)}`)
77  }
78
79  return { name, type: 'predicate', instructions: said.join(' ') }
80}
81
82/**
83 * Serialises one request for OpenAI: the state as a JSON string in `input`,
84 * which takes text and no object, and the questions as a list.
85 *
86 * @param model the model name
87 * @param state what every question is asked about
88 * @param questions the questions, by id
89 * @returns the JSON text of the request body
90 */
91function encodeOpenAI(
92  model: string,
93  state: unknown,
94  questions: Questions,
95): string {
96  return JSON.stringify({
97    model,
98    input: JSON.stringify(state),
99    questions: idsOf(questions).map(id =>
100      questionOf(id, questions[id] as Question),
101    ),
102  })
103}
104
105/**
106 * Reads OpenAI's list of answers as the System One map by id, which is what
107 * the readers of a reply take: a `predicate`'s probability as `noul`, a
108 * `choice`'s pick and confidence as they are. Only the fields read here are
109 * relied on, so a field OpenAI adds changes nothing; one it renames throws,
110 * naming the field. A refusal to answer throws, naming the question: every
111 * question was sent with a name, so every answer has one.
112 *
113 * An answer is taken only under the name of a question that was asked, and
114 * only once: any other name answers nothing the request asked, and a second
115 * answer to the same question would leave which one counts to chance. The
116 * map has no prototype, so a name such as `__proto__` is an ordinary key.
117 *
118 * @param payload the decoded response body
119 * @param asked the questions the request asked, by id
120 * @returns the body in the System One shape
121 */
122export function decodeOpenAI(payload: unknown, asked: Questions): unknown {
123  if (!isRecord(payload) || !Array.isArray(payload.answers)) {
124    throw new Error('the response holds no answers list')
125  }
126
127  const answers: Record<string, unknown> = Object.create(null)
128
129  payload.answers.forEach((answer: unknown, index) => {
130    const at = `answers[${index}]`
131
132    if (!isRecord(answer)) {
133      throw new Error(`${at} is not an object`)
134    }
135
136    if (typeof answer.name !== 'string' || answer.name === '') {
137      throw new Error(`${at}.name is missing`)
138    }
139
140    if (!Object.hasOwn(asked, answer.name)) {
141      throw new Error(
142        `${at}.name is ${quoted(answer.name)}, which was not asked`,
143      )
144    }
145
146    if (Object.hasOwn(answers, answer.name)) {
147      throw new Error(`${at} answers ${answer.name} a second time`)
148    }
149
150    if (answer.type === 'refusal') {
151      throw new Error(`refused to answer ${answer.name}`)
152    }
153
154    if (answer.type === 'predicate') {
155      answers[answer.name] = { noul: answer.probability }
156    } else if (answer.type === 'choice') {
157      // An option's value may be a boolean in OpenAI's schema. Read as text
158      // it can only match an option that was offered as that same text;
159      // any other pick is refused by the reader of the answer.
160      answers[answer.name] = {
161        choice:
162          typeof answer.choice === 'boolean'
163            ? String(answer.choice)
164            : answer.choice,
165        confidence: answer.confidence,
166      }
167    } else {
168      throw new Error(
169        `${at}.type is ${quoted(answer.type)}, neither predicate nor choice`,
170      )
171    }
172  })
173
174  const reported = isRecord(payload.usage) ? payload.usage : {}
175  const decoded: Record<string, unknown> = {
176    answers,
177    usage: { input_tokens: reported.input_tokens },
178  }
179
180  if (typeof payload.model === 'string') {
181    decoded.model = payload.model
182  }
183
184  return decoded
185}
186
187/**
188 * The body OpenAI's Decisions API takes. The state rides as a string, so its
189 * JSON is escaped once more and is measured in that form.
190 *
191 * OpenAI counts far more input tokens per question than its text: 129
192 * one-line predicates were counted as 19,139 input tokens, about 148 each,
193 * where Codiv counted about 26 each for the same questions.
194 */
195export const OPENAI: Wire = {
196  encode: encodeOpenAI,
197  questionJsonOf: (id, question) => JSON.stringify(questionOf(id, question)),
198  stateJsonOf: json => JSON.stringify(json).slice(1, -1),
199  tokensPerQuestion: 148,
200}
201
hooks/route.ts 428 lines
1import type { SessionMessage } from 'claude-code'
2
3import type { Ask } from './compact'
4import {
5  briefOf,
6  familyOf,
7  gatewayOf,
8  messageOf,
9  missingOf,
10  modelsOf,
11} from './providers'
12import type { Credentials, ProviderName, Route } from './providers'
13import type { Call } from './state'
14import { choiceOf } from './systemone'
15import type { ChoiceQuestion } from './systemone'
16
17/**
18 * A route the decision may pick: the name it is offered under and what is
19 * documented about the model behind it.
20 */
21type Candidate = {
22  key: string
23  route: Route
24  about: string
25}
26
27/**
28 * The job in a few numbers: what the routing question is asked about. It
29 * stands in for the full state, which would cost a second full-size request
30 * just to choose where to send the first.
31 */
32export type Profile = {
33  context: string
34  goal: string
35  messages: number
36  candidate_calls: number
37  state_tokens: number
38  tools: Record<string, number>
39  error_share: number
40  non_ascii_share: number
41}
42
43/**
44 * Says why the fitted state and its questions do not fit a route's limits,
45 * or nothing when they do.
46 */
47type SizeCheck = (route: Route) => string | undefined
48
49/**
50 * The route to use and one line saying how it was arrived at.
51 */
52type Routed = {
53  route: Route
54  why: string
55}
56
57/**
58 * The id of the one routing question.
59 */
60const ROUTE_QUESTION = 'route'
61
62/**
63 * How many tools the profile names; the rest are counted together.
64 */
65const TOOLS_SHOWN = 8
66
67/**
68 * What the providers' own documentation says each model is for. The
69 * decision is only as good as these lines, so each states published facts
70 * and nothing inferred.
71 */
72const ABOUT: Record<string, string> = {
73  clef:
74    'Clef, a 27B model. Cloudflare recommends it for highest-precision ' +
75    'decisions. It is slower and costs more per token than Clef-flash.',
76  'clef-flash':
77    'Clef-flash, a 9B model. Cloudflare recommends it for latency-critical ' +
78    'decisions on a hot path. It is the fastest of the models offered here.',
79  jev:
80    "Jev, TypeSafe's flagship System One model. It reads text only, and " +
81    'English is its primary training language: TypeSafe reports lower ' +
82    'accuracy on other languages, CJK scripts included.',
83  openjev:
84    'OpenJev, DiffusionGemma 26B-A4B made into a System One model; it ' +
85    'reads a distribution for every answer in one denoising step.',
86  'pplx-decider':
87    "Perplexity's 27B decision model. It reads text and images and " +
88    'returns typed answers with probabilities.',
89  'gpt-6-luna':
90    'GPT-6 Luna, which OpenAI describes as its most efficient model for ' +
91    'focused, high-volume tasks, served through its Decisions API.',
92}
93
94const PROFILE_CONTEXT =
95  'A coding assistant session is about to be compacted: a decision model ' +
96  'will be asked, for each old tool call, whether the call and its output ' +
97  'still have to stay in the transcript. This object describes that job: ' +
98  'how many messages and candidate tool calls there are, the estimated ' +
99  'size in tokens of the state every request will carry, which tools were ' +
100  'called how often, the share of calls that ended in an error, and the ' +
101  'share of the conversation text that is not ASCII. `goal` is what the ' +
102  'assistant is working on.'
103
104/**
105 * A route as a line of text shows it.
106 *
107 * @param route the route
108 * @returns `provider/model`
109 */
110export function labelOf(route: Route): string {
111  return `${route.provider}/${route.model}`
112}
113
114function keyOf(route: Route): string {
115  return `${route.provider}.${route.model.replace(/[^A-Za-z0-9_.-]/g, '-')}`
116}
117
118function aboutOf(route: Route): string {
119  const known = ABOUT[familyOf(route)]
120  const gateway = gatewayOf(route.provider)
121  const via = gateway === undefined ? '' : ` Reached through ${gateway}.`
122
123  return known === undefined
124    ? `The model ${route.model} at ${route.provider}; nothing is documented ` +
125        'here about what it is best at.'
126    : `${known}${via}`
127}
128
129/**
130 * Says why a route cannot take the job, or nothing when it can. Both checks
131 * are ones code can make with certainty, which is why they are made here
132 * and not put to a model.
133 */
134function refusalOf(
135  route: Route,
136  credentials: Credentials,
137  sizeRefusalOf: SizeCheck,
138): string | undefined {
139  const missing = missingOf(route.provider, credentials)
140
141  return missing.length > 0
142    ? `${missing.join(' and ')} unset`
143    : sizeRefusalOf(route)
144}
145
146/**
147 * Lists the routes that can take the job.
148 *
149 * Every eligible provider is considered with each model it offers; the
150 * configured provider with the configured model in place of its default. A
151 * route is left out when a credential it needs is unset or the fitted state
152 * exceeds its documented limit. A model family is offered once: over the
153 * route that reaches it directly whenever there is one, else over the
154 * first gateway to it in `PROVIDERS` order, and every other route to it is
155 * left out of the choice.
156 *
157 * @param configured the route the options name
158 * @param eligible the providers that may be sent anything, in `PROVIDERS`
159 * order
160 * @param credentials the resolved credentials
161 * @param sizeRefusalOf the check of the fitted state against a route's limits
162 * @returns the candidates, for every route not offered the reason, and
163 * whether the configured provider is among the routes that can take the job
164 */
165export function candidatesOf(
166  configured: Route,
167  eligible: readonly ProviderName[],
168  credentials: Credentials,
169  sizeRefusalOf: SizeCheck,
170): { candidates: Candidate[]; refused: string[]; isConfiguredUsable: boolean } {
171  const usable: Route[] = []
172  const refused: string[] = []
173
174  for (const provider of eligible) {
175    const offered = modelsOf(provider)
176    const models =
177      provider === configured.provider && !offered.includes(configured.model)
178        ? [configured.model]
179        : offered
180
181    for (const model of models) {
182      const route = { provider, model }
183      const refusal = refusalOf(route, credentials, sizeRefusalOf)
184
185      if (refusal === undefined) {
186        usable.push(route)
187      } else {
188        refused.push(`${labelOf(route)}: ${refusal}`)
189      }
190    }
191  }
192
193  const offered = new Map<string, Route>()
194
195  for (const route of [
196    ...usable.filter(route => gatewayOf(route.provider) === undefined),
197    ...usable.filter(route => gatewayOf(route.provider) !== undefined),
198  ]) {
199    if (!offered.has(familyOf(route))) {
200      offered.set(familyOf(route), route)
201    }
202  }
203
204  const candidates: Candidate[] = []
205
206  for (const route of usable) {
207    const same = offered.get(familyOf(route))
208
209    if (same !== undefined && same !== route) {
210      refused.push(`${labelOf(route)}: the same model as ${labelOf(same)}`)
211      continue
212    }
213
214    candidates.push({ key: keyOf(route), route, about: aboutOf(route) })
215  }
216
217  return {
218    candidates,
219    refused,
220    isConfiguredUsable: usable.some(
221      route => route.provider === configured.provider,
222    ),
223  }
224}
225
226function shareOf(part: number, whole: number): number {
227  return whole === 0 ? 0 : Math.round((part / whole) * 100) / 100
228}
229
230/**
231 * Describes the job for the routing question.
232 *
233 * @param messages the conversation
234 * @param calls the calls that will be asked about
235 * @param stateTokens the estimated size of the fitted state
236 * @param goal what the assistant is working on
237 * @returns the profile
238 */
239export function profileOf(
240  messages: readonly SessionMessage[],
241  calls: readonly Call[],
242  stateTokens: number,
243  goal: string,
244): Profile {
245  const counts = new Map<string, number>()
246
247  for (const call of calls) {
248    counts.set(call.tool, (counts.get(call.tool) ?? 0) + 1)
249  }
250
251  const ranked = [...counts].sort(([, a], [, b]) => b - a)
252  const tools = Object.fromEntries(ranked.slice(0, TOOLS_SHOWN))
253  const rest = ranked
254    .slice(TOOLS_SHOWN)
255    .reduce((sum, [, count]) => sum + count, 0)
256
257  if (rest > 0) {
258    tools.other = rest
259  }
260
261  let chars = 0
262  let nonAscii = 0
263
264  for (const { text } of messages) {
265    chars += text.length
266
267    for (let index = 0; index < text.length; index++) {
268      if (text.charCodeAt(index) > 127) {
269        nonAscii++
270      }
271    }
272  }
273
274  return {
275    context: PROFILE_CONTEXT,
276    goal,
277    messages: messages.length,
278    candidate_calls: calls.length,
279    state_tokens: stateTokens,
280    tools,
281    error_share: shareOf(
282      calls.filter(call => call.isError).length,
283      calls.length,
284    ),
285    non_ascii_share: shareOf(nonAscii, chars),
286  }
287}
288
289/**
290 * The one question the decision asks: which of the candidates should judge
291 * this job, each described by what is documented about it.
292 *
293 * @param candidates the routes on offer, two or more
294 * @returns the `choice` question
295 */
296function routeQuestionOf(
297  candidates: readonly Candidate[],
298): ChoiceQuestion {
299  return {
300    type: 'choice',
301    instructions:
302      'Which decision model should judge this compaction job? A wrong ' +
303      '"remove" loses output the assistant may still need, a wrong "keep" ' +
304      'only wastes context, and the answers are waited for while the ' +
305      'session is paused. Pick the model whose documented strengths fit ' +
306      'the job the state describes.',
307    criteria: Object.fromEntries(
308      candidates.map(candidate => [candidate.key, candidate.about]),
309    ),
310  }
311}
312
313/**
314 * Decides which route a compaction goes over.
315 *
316 * First by what code can check: `candidatesOf` leaves out every route that
317 * cannot take the job. Only when two or more remain is a model asked, with
318 * one `choice` question over the job's profile, on the fastest route there
319 * is: Cloudflare's Clef-flash when Cloudflare is eligible and configured,
320 * the configured route otherwise. The question carries a profile with the
321 * person's prompts in it, so it goes to no provider outside the eligible
322 * ones. With one candidate or a failed routing request the
323 * configured route is used, so the decision can never make a compaction
324 * fail that would have worked without it; the exception is a configured
325 * route that cannot take the job at all, where the one that can is used.
326 * The one candidate may be the direct route to the model a configured
327 * gateway also reaches: there is then one model and nothing to choose, and
328 * the configured gateway stays in use.
329 *
330 * With no candidate at all there is nothing to send the job to, the
331 * configured route included, so this throws, and the reason lists why each
332 * route was ruled out: the person has to see which credential or limit to
333 * change.
334 *
335 * @param configured the route the options name
336 * @param eligible the providers that may be sent anything, in `PROVIDERS`
337 * order
338 * @param credentials the resolved credentials
339 * @param sizeRefusalOf the check of the fitted state against a route's limits
340 * @param profile the job's profile
341 * @param askOn how a question is sent over a given route
342 * @returns the route and why it was picked; throws when no route can take
343 * the job
344 */
345export async function chooseRoute(
346  configured: Route,
347  eligible: readonly ProviderName[],
348  credentials: Credentials,
349  sizeRefusalOf: SizeCheck,
350  profile: Profile,
351  askOn: (route: Route) => Ask,
352): Promise<Routed> {
353  const { candidates, refused, isConfiguredUsable } = candidatesOf(
354    configured,
355    eligible,
356    credentials,
357    sizeRefusalOf,
358  )
359  const [first, second] = candidates
360  const left = refused.length === 0 ? '' : ` (${refused.join('; ')})`
361
362  if (first === undefined) {
363    throw new Error(`no route can take this job: ${refused.join('; ')}`)
364  }
365
366  if (second === undefined) {
367    if (!isConfiguredUsable) {
368      return {
369        route: first.route,
370        why: `only ${labelOf(first.route)} can take this job${left}`,
371      }
372    }
373
374    const found =
375      first.route.provider === configured.provider
376        ? 'one route is available'
377        : 'one model is available'
378
379    return {
380      route: configured,
381      why: `${found}${left}; using the configured ${labelOf(configured)}`,
382    }
383  }
384
385  const asked: Route =
386    eligible.includes('cloudflare') &&
387    missingOf('cloudflare', credentials).length === 0
388      ? { provider: 'cloudflare', model: 'clef-flash' }
389      : configured
390
391  try {
392    const reply = await askOn(asked)(profile, {
393      [ROUTE_QUESTION]: routeQuestionOf(candidates),
394    })
395    const picked = choiceOf(
396      reply,
397      ROUTE_QUESTION,
398      candidates.map(candidate => candidate.key),
399    )
400    const chosen = candidates.find(candidate => candidate.key === picked.choice)
401
402    if (chosen === undefined) {
403      throw new Error('the picked option is not a candidate')
404    }
405
406    const sure =
407      picked.confidence === undefined
408        ? ''
409        : `, confidence ${picked.confidence.toFixed(2)}`
410
411    return {
412      route: chosen.route,
413      why: `${labelOf(asked)} picked ${labelOf(chosen.route)}${sure}`,
414    }
415  } catch (error) {
416    // The reason may quote a provider's text, and it reaches the report
417    // line, so it is quoted like any other: redacted and on one short line.
418    const reason = briefOf(messageOf(error), credentials)
419
420    return {
421      route: configured,
422      why:
423        `the routing request failed (${reason}); using the configured ` +
424        labelOf(configured),
425    }
426  }
427}
428
hooks/ask.ts 280 lines
1import type { HttpInit, HttpResponse } from 'claude-code'
2
3import type { Ask } from './compact'
4import {
5  briefOf,
6  exchangeOf,
7  isRetryable,
8  messageOf,
9  replyFrom,
10} from './providers'
11import type { Credentials, Route } from './providers'
12
13/**
14 * What a request needs from the engine, handed in as functions because the
15 * engine interface itself cannot leave the hooks module's own file.
16 */
17export type Ports = {
18  /**
19   * `$.http.fetch`.
20   */
21  fetch: (url: string, init: HttpInit) => Promise<HttpResponse>
22  /**
23   * Waits before a retry. Resolves false, without waiting, when the hook has
24   * no time left to spend on it. `signal` aborts when the attempt the wait
25   * belongs to is over, and the wait must end with it then: a wait that
26   * outlived its attempt would hold the hook for nothing.
27   */
28  pause: (ms: number, signal: AbortSignal) => Promise<boolean>
29  /**
30   * `$.clock.after`: runs `fn` once after `ms` milliseconds unless the
31   * answer's `cancel` is called first.
32   */
33  after: (ms: number, fn: () => void) => { cancel: () => void }
34  /**
35   * `next.signal`: aborts when the compaction this hook serves is given up.
36   */
37  signal: AbortSignal
38}
39
40/**
41 * How long to wait before the second and the third attempt of one request.
42 * The hook's own code and its waits are charged to its ten seconds; the time
43 * a request is out is not. The two waits of one request stay under two
44 * seconds together, but that bound is per request, not per compaction: with
45 * `MAX_REQUESTS` batches going out `CONCURRENT_REQUESTS` at a time, as many
46 * rounds of batches as that makes (six, at sixteen and three) can each wait
47 * in turn. What bounds the waits of a compaction together is the `pause`
48 * port, which declines a wait when the hook has too little of its time left.
49 */
50const BACKOFF_MS = [400, 1200] as const
51
52/**
53 * How long one compaction may wait on its providers, from the moment its
54 * first request can go out: the routing question, every batch, every retry
55 * and every wait between retries together. A request has no timeout of its
56 * own, and the time it takes is not charged to the hook, so without a bound
57 * a provider that accepts a request and never answers would hold the
58 * compaction open for good, and the built-in summary would never get its
59 * turn. The bound is the longest the person waits on the providers, with the
60 * session paused, before that summary starts.
61 */
62const COMPACTION_DEADLINE_MS = 30_000
63
64/**
65 * The requests and waits of one compaction, under the one bound they share.
66 *
67 * The compaction is over, as far as its providers go, at the first of: its
68 * deadline passing, the engine giving the compaction up, `fail` being
69 * called, or `close` being called. From then on `fetch` and `pause` start
70 * nothing and reject with the reason. When it is over for any reason but
71 * `close`, whatever they had under way rejects with it at once; however it
72 * ends, a wait under way is aborted. A request that is out cannot be
73 * withdrawn, so its late answer is ignored.
74 */
75type Attempt = {
76  /**
77   * `Ports.fetch`, started only while the compaction is not over. A request
78   * the host refuses by throwing, instead of by rejecting, rejects all the
79   * same.
80   */
81  fetch: Ports['fetch']
82  /**
83   * `Ports.pause`, started only while the compaction is not over, with a
84   * signal that aborts when it is.
85   */
86  pause: (ms: number) => Promise<boolean>
87  /**
88   * Why the compaction is over, or undefined while it is not.
89   */
90  failure: () => Error | undefined
91  /**
92   * Ends the compaction with a reason. Only the first reason counts.
93   */
94  fail: (reason: Error) => void
95  /**
96   * Ends the compaction, when nothing ended it before, and releases the
97   * deadline's timer and the listener on the engine's signal. Called once
98   * the compaction has its outcome, whichever that is: a request still out
99   * then has nobody to answer, and must not retry or wait.
100   */
101  close: () => void
102}
103
104/**
105 * Puts one compaction's requests under one deadline, armed here and nowhere
106 * else, and under the engine's signal.
107 *
108 * @param ports the engine calls
109 * @returns the attempt; the caller closes it when the compaction is decided
110 */
111export function attemptOf(ports: Ports): Attempt {
112  let failure: Error | undefined
113  let disarm = () => {}
114  let reject: (reason: Error) => void = () => {}
115  // Aborted when the attempt is over, however it ends: the waits listen to
116  // it, so that none of them outlives the attempt.
117  const ended = new AbortController()
118
119  const cut = new Promise<never>((_resolve, rejectCut) => {
120    reject = rejectCut
121  })
122
123  // An end that comes while nothing is under way has nobody to reach, and
124  // must not surface as a rejection nobody handled.
125  cut.catch(() => {})
126
127  const fail = (reason: Error) => {
128    if (failure !== undefined) {
129      return
130    }
131
132    failure = reason
133    disarm()
134    reject(reason)
135    ended.abort(reason)
136  }
137
138  const interrupted = () =>
139    new Error(
140      'the compaction was interrupted before its requests were answered',
141    )
142
143  if (ports.signal.aborted) {
144    fail(interrupted())
145  } else {
146    const onAbort = () => fail(interrupted())
147
148    // Neither is left set up without the other: the listener goes on first
149    // and comes off again when the timer cannot be armed, so an attempt that
150    // could not be made leaves nothing on the engine's signal or its clock.
151    ports.signal.addEventListener('abort', onAbort, { once: true })
152
153    let timer: { cancel: () => void }
154
155    try {
156      timer = ports.after(COMPACTION_DEADLINE_MS, () =>
157        fail(
158          new Error(
159            'the requests of this compaction were not all answered within ' +
160              `${COMPACTION_DEADLINE_MS / 1000} seconds`,
161          ),
162        ),
163      )
164    } catch (error) {
165      ports.signal.removeEventListener('abort', onAbort)
166
167      throw error
168    }
169
170    disarm = () => {
171      disarm = () => {}
172      timer.cancel()
173      ports.signal.removeEventListener('abort', onAbort)
174    }
175  }
176
177  // The executor turns a `start` that throws into a rejection, so a call
178  // refused on the spot takes the same path as one refused later.
179  const within = <Value>(start: () => Promise<Value>): Promise<Value> =>
180    failure === undefined
181      ? Promise.race([new Promise<Value>(resolve => resolve(start())), cut])
182      : Promise.reject(failure)
183
184  return {
185    fetch: (url, init) => within(() => ports.fetch(url, init)),
186    pause: ms => within(() => ports.pause(ms, ended.signal)),
187    failure: () => failure,
188    fail,
189    // The shared cut is not rejected here: whatever is still out when the
190    // outcome is known is waited on by nobody and is left to settle on its
191    // own. Only what it would start from here on is refused.
192    close: () => {
193      failure ??= new Error('the compaction already has its outcome')
194      disarm()
195      ended.abort(failure)
196    },
197  }
198}
199
200/**
201 * Builds the function that sends one set of questions over a route.
202 *
203 * A response that may succeed on another attempt (rate limiting, an
204 * overloaded or failing server) is retried after a short wait, at most
205 * twice; any other failure is final at once. Only this one route is ever
206 * tried: there is no falling over to another provider.
207 *
208 * A failure of a decisive request ends the attempt on the spot, in the same
209 * step that found it, so that no other request of the compaction retries or
210 * waits after it. Once the attempt is over, every request fails with the
211 * attempt's reason, whatever its own would have been.
212 *
213 * What the host says when it refuses a call is quoted redacted and short:
214 * it may repeat the request's headers.
215 *
216 * @param route the provider and model
217 * @param credentials the resolved credentials
218 * @param attempt the compaction's requests and their shared bound
219 * @param isDecisive true for a request the compaction cannot do without (a
220 * batch of questions about calls), false for one it can (the routing
221 * question)
222 * @returns the function `compact` and the routing decision ask through
223 */
224export function askerOf(
225  route: Route,
226  credentials: Credentials,
227  attempt: Attempt,
228  isDecisive: boolean,
229): Ask {
230  const hosted = <Value>(what: string, call: Promise<Value>): Promise<Value> =>
231    call.catch(error => {
232      throw (
233        attempt.failure() ??
234        new Error(`${what} failed: ${briefOf(messageOf(error), credentials)}`)
235      )
236    })
237
238  return async (state, questions) => {
239    try {
240      const exchange = exchangeOf(route, credentials, state, questions)
241      const send = () =>
242        hosted(
243          `the request to ${route.provider}`,
244          attempt.fetch(exchange.url, exchange.init),
245        )
246
247      let response = await send()
248
249      for (const wait of BACKOFF_MS) {
250        if (response.ok || !isRetryable(response.status)) {
251          break
252        }
253
254        const waited = hosted(
255          `the wait before asking ${route.provider} again`,
256          attempt.pause(wait),
257        )
258
259        if (!(await waited)) {
260          break
261        }
262
263        response = await send()
264      }
265
266      return replyFrom(route, response, credentials, questions)
267    } catch (error) {
268      const reason =
269        attempt.failure() ??
270        (error instanceof Error ? error : new Error(String(error)))
271
272      if (isDecisive) {
273        attempt.fail(reason)
274      }
275
276      throw reason
277    }
278  }
279}
280