Continuous engineering loop for Claude Code: Jev-based size/clarity triage, a plan-clarify-build-test-revise harness for small tasks, and for large tasks…

<img src="docs/assets/forge-banner.jpg" alt="forge: Continuous Engineering Loop for Claude Code" width="100%">
<strong>A continuous engineering loop for Claude Code.</strong><br> Triage every request, run small tasks through a test-driven harness, and drive large tasks to a measurable goal with an orchestrator, parallel subagents, and a deterministic judge.
<img alt="version" src="https://img.shields.io/badge/version-0.4.2-orange"> <img alt="Claude Code" src="https://img.shields.io/badge/Claude%20Code-%E2%89%A5%202.1.280-black"> <img alt="Node" src="https://img.shields.io/badge/node-%E2%89%A5%2020-339933"> <img alt="dependencies" src="https://img.shields.io/badge/dependencies-none-lightgrey"> <img alt="license" src="https://img.shields.io/badge/license-MIT-blue">
forge takes the ratchet idea from karpathy/autoresearch (one change, one commit, one measurement, keep it or reset it) and wraps it in the process a real team follows: a written spec, a mapped scope, clarifying questions, tests beyond the happy path, and a human sign-off before anything runs unattended.
The worker is always Claude. Jev, TypeSafe's decision model, only makes cheap, fast judgements: triage, model routing, finding pre-sorting, failure labelling, and context compaction. Pass or fail always comes from test exit codes, a baseline, and a metric.
<table> <tr> <td width="33%" valign="top"> <h3>Triage</h3> Each engineering prompt is classified as a question, a vague request, a small task, or a large task, and routed to the matching flow. </td> <td width="33%" valign="top"> <h3>Orchestrated goal loop</h3> Parallel subagents in their own worktrees, an orchestrator that rules on findings by citing the spec, and a judge that keeps or resets each iteration. </td> <td width="33%" valign="top"> <h3>Self-learning</h3> Repo-scoped instincts distilled from run evidence and owner corrections, fed back into the next task. </td> </tr> </table>
<img src="docs/assets/forge-flow.jpg" alt="forge flow: triage, small harness, large pipeline, ratchet loop, report" width="100%">
Plan with happy, unhappy, and edge cases, clarify open decisions up front, then build. When Claude ends a turn, a Stop hook runs the cases and the test command: red keeps the session working with the failing cases attached, green is done, and three red revisions escalate to you. A plan that grows past a few files is proposed for promotion to a large task.
Every stage has a gate, and forge check tells you what is missing.
| Stage | Output |
|---|---|
| Intake | frozen copy of the source (Jira, Google Doc, PDF, Markdown, PRD, Figma) and the report target |
| Goal and spec | goal.json (test and regression commands, metric, target, budget, locked files) and spec.md, with a source for every acceptance criterion |
| Scope | scope.json: in scope, impacted, out of scope, and conflicts, each with file:line evidence |
| Worklist and matrix | vertical slices with owned files; happy, unhappy, and edge cases per criterion plus a regression case per impacted path |
| Lanes | subagent count computed from file collisions, dependencies, free RAM, and quota; shared files belong to the orchestrator |
| Baseline | matrix, suite, and metric measured twice at HEAD to catch flaky tests |
| Clarify | numbered questions A to J, each with a recommendation, including the subagent count |
| Spec lock | explicit approval; artifacts and evaluator files are hashed, and edits to them are denied |
Then the loop runs in the background. Each iteration:
claude -p in its own git worktree and may only write the files it owns;git reset.The loop stops when the target is reached, the budget is spent, progress stalls, or every remaining item is BLOCKED. report.md lists finished work, BLOCKED items with questions and options, ticket proposals, decisions, model usage, and the metric curve. forge never merges or pushes.
Requirements: Claude Code 2.1.280+ (logged in), Node.js 20+, and git. There are no npm dependencies.
claude plugin marketplace add sudikama/forge
claude plugin install forge@forge-local
Open a new Claude Code session inside a git repository and run /forge:doctor. Claude Code, auth, and git must be ok. Jev lanes without a key show FAIL and can be ignored as long as one lane works; without any, forge falls back to deterministic rules.
Then work as usual. Every prompt is triaged automatically:
| Triage | What happens |
|---|---|
none | a question or conversation; forge stays out of the way |
clarify | no checkable definition of done; Claude asks before writing code |
small | small harness |
large | large pipeline; Claude may not write product code directly |
| Command | Purpose | ||
|---|---|---|---|
| `/forge:start <request \ | ticket \ | doc link>` | triage manually and start |
/forge:status | state, failing gates, loop progress, BLOCKED items | ||
/forge:doctor | check prerequisites |
To update, run claude plugin marketplace update forge-local && claude plugin update forge@forge-local.
Set options with /plugin configure forge@forge-local inside Claude Code, or in ~/.forge/env as FORGE_<OPTION>=value (for example FORGE_MAX_WORKERS=2; API keys keep their own names). The process environment wins over the file. claude plugin configure forge@forge-local lists the current values.
| Option | Default | Description |
|---|---|---|
jevLanes | zen,commandcode | Jev backend failover order: zen (free, keyless), commandcode (200 requests/day), typesafe, openrouter |
COMMANDCODE_API_KEY, TYPESAFE_API_KEY, OPENROUTER_API_KEY | empty | Keys for the matching lanes; typesafe is the only lane with calibrated confidence |
autoTriage | true | Triage every engineering prompt |
maxWorkers | 4 | Hard cap on parallel subagents (also capped by free RAM, about 1.2 GB each) |
smallMaxRevisions | 3 | Red test runs before the small harness escalates |
fastModel, balancedModel, deepModel | haiku, claude-sonnet-5-5, claude-opus-5-5 | Model per tier |
ladderFailsFast, ladderFailsBalanced, ladderFailsDeep | 1, 2, 2 | Gate failures allowed on each tier before moving up, or before BLOCKED on the top tier |
maxEffort | high | Highest effort routing may pick |
jevDecide, jevRejectBar | true, 0.6 | Let Jev reject a finding that matches a spec exclusion line with at least this confidence |
jevCompaction, compactAtPercent | true, 60 | Jev-guided compaction and its trigger (function hooks only) |
JIRA_BASE_URL, JIRA_EMAIL, JIRA_TOKEN | empty | Jira REST fallback when no Jira MCP is registered |
reportCmd | empty | Shell command for the command report target |
extraInstincts | empty | Optional read-only instinct store (same <repo>/*.yaml layout) merged into prompts |
The full list with descriptions lives in .claude-plugin/plugin.json. Every Jev call fails open, so an outage never blocks a session.
Large tasks read their source through MCP and freeze a hashed copy under .forge/tasks/<KEY>/source/. Register the servers you use:
claude mcp add jira -- <any Jira MCP server exposing get_issue>
claude mcp add --transport http figma https://mcp.figma.com/mcp
claude mcp add gdrive -- <a Google Drive MCP server with OAuth>
Without MCP, forge source add accepts a Jira key (REST fallback), a PDF, a Markdown or PRD file, or any text piped on --stdin.
Each loop agent is a fresh claude -p, so switching models costs no prompt cache. Jev picks the starting tier and effort per lane: raising them needs 0.3 confidence, lowering them needs 0.7. L-sized items never start on Haiku, and the orchestrator never runs below Sonnet 5.5.
From then on, the gate moves each item up a model ladder:
| Tier | Model | Failures allowed | Then |
|---|---|---|---|
| fast | Haiku | 1 | Sonnet 5.5 |
| balanced | Sonnet 5.5 | 2 | Opus 5.5 |
| deep | Opus 5.5 | 2 | BLOCKED, with every attempt as evidence |
A failure is a red case owned by the item, a rejected or unmergeable lane, or a regression proven to come from that lane. When a shared check goes red, forge re-runs it in each lane's worktree on its own, so only the lane that breaks it is blamed. Passing never moves an item down.
| Point | What Jev does | Guard |
|---|---|---|
| Triage | size, parallelism, presence of a verifier | a low verifier score forces clarify |
| Findings | maps a finding to a Non-goals, MUST NOT, or out-of-scope line | rejects only above jevRejectBar with a valid citation |
| Failures | labels why tests failed, to steer the next revision | never affects the verdict |
| Compaction | scores which tool calls are still needed | files still referenced are kept; falls back to the built-in summary |
| Session subagents | picks the model of Agent-tool subagents | explicit models and forks are untouched |
Compaction and session subagent routing are function hooks (hooks/forge-fn.js). They load only with CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1, an early-access API that may change between releases.
Instincts are stored per repository in ~/.forge/learn/repos/<repo>/ and come only from run evidence: test commands proven green, recurring errors that were resolved, cases that caught a real regression, and your corrections (forge answer <ID> "<answer>" --correction "<lesson>"). Confidence starts at 0.5, rises 0.1 when confirmed, drops 0.2 when contradicted, and the instinct is dropped below 0.4. Changes to forge itself are only proposed in the report, never applied automatically.
Skills and hooks drive the CLI for you; these commands are useful directly. Run node bin/forge.mjs doctor --install-shim from a clone to put forge on your PATH.
forge doctor | triage "<request>"
forge new <KEY> --mode small|large --title ".." --report-to file|jira:KEY|webhook:URL|command
forge source add --kind jira|gdoc|pdf|markdown|prd|figma|file (--ref X | --file P | --stdin)
forge check | lanes | baseline | clarify | answer <ID> "<answer>" | lock --approve
forge run | status | pause | stop | resume
forge findings | report | learn list
Report targets: file always writes report.md; jira:KEY posts it as a comment; webhook:URL POSTs {subject, report, path} as JSON (Slack, Discord, n8n, anything); command runs the reportCmd option with FORGE_REPORT_PATH and FORGE_REPORT_SUBJECT set, for any notifier you already use.
git clone https://github.com/sudikama/forge.git && cd forge
claude --plugin-dir . # run Claude Code with the working copy
bash tests/e2e/run-large.sh # full large flow with a stub agent, deterministic, no cost
bash tests/e2e/run-small.sh # small harness, lane gate, live triage on Jev zen
node tests/unit/route.test.mjs # routing policy, model ladder, Jev finding decisions
node tests/unit/forge-fn.test.mjs # function hooks on a fake engine (LIVE=1 hits Jev zen)
node tests/unit/deliver.test.mjs # report delivery targets
bash tests/e2e/run-layouts.sh # submodules, worktrees, missing .git/info, non-git dirs
FORGE_E2E_WORKTREE=1 bash tests/e2e/run-large.sh # the large flow started from a linked worktree
| Path | Contents |
|---|---|
bin/forge.mjs, lib/ | CLI and core: triage, routing, lanes, judge, runner, clarify, learning, report |
hooks/ | classic hooks, function hooks, vendored fast-jev-compaction |
skills/, commands/, templates/ | instructions, slash commands, and artifact templates for Claude |
tests/ | unit tests, end-to-end fixture, stub agent, Jev mock |
In a target repository forge writes only to .forge/ (excluded from git automatically) and to forge/<KEY> branches.
| Symptom | Fix |
|---|---|
| Hooks do not run | start a new session after installing; check claude plugin list |
claude auth: not logged in | run claude then /login, or set ANTHROPIC_API_KEY |
Jev lane returns HTTP 403 | Cloudflare; retry or switch lanes |
cannot lock, gates still failing | run forge check and fix the listed items |
baseline ran at X but HEAD is Y | a commit landed after the baseline; run forge baseline again |
Loop stops with every remaining item is blocked | answer the BLOCKED items in report.md, then open a follow-up task |
| Only one worker | low free RAM or file collisions; see forge lanes |
forge new fails with ENOTDIR or ENOENT on .git/info/exclude | forge older than 0.4.1 in a worktree or submodule; update the plugin and start a new session (/forge:doctor shows the version) |
Logs are in ~/.forge/logs/ and .forge/tasks/<KEY>/runner.log.
forge triage to check a misclassified request.MIT. Built on ideas and code from karpathy/autoresearch, tamaratran/fast-jev-compaction (vendored under MIT), moelahmady/jev-model-router, ifoster01/jev-effort, and shitianfang/jev-use.
hooks/forge-fn.js 201 lines1// forge function-hook module (EARLY ACCESS API; only loaded when Claude Code runs with
2// CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1). Classic hooks in hooks.json keep working without it.
3//
4// session.compact : Jev-guided verbatim compaction (vendored fast-jev-compaction, MIT):
5// Jev scores every tool call/result as still-needed or not; unneeded results
6// are truncated, unneeded calls dropped, no summary. Falls back to the
7// built-in summary below minReductionRatio or on any error.
8// turn.complete : asks for a compaction at compactAtPercent of the context window.
9// agent.spawn : routes the model of Agent-tool subagents inside an interactive session.
10//
11// Runs in Claude Code's hook sandbox: no Node, only `$` (http.fetch, env, settings, ui, session).
12import { compact, reductionRatio, applyDecisions, messageChars } from './vendor/fast-jev/compact.js'
13import { collectToolCalls } from './vendor/fast-jev/state.js'
14
15// Reference guard (forge, deterministic): the zen lane is not calibrated and was seen scoring
16// the file under edit at 0.17. A call whose input names a file/identifier that the goal or any
17// LATER message text mentions is still in use, so it is kept whatever Jev says.
18function refsOf(input) {
19 const out = new Set()
20 const walk = (v) => {
21 if (typeof v === 'string') {
22 for (const m of v.matchAll(/[\w.\-/]*[\w-]+\.[a-z0-9]{1,6}\b|[\w-]+\/[\w.\-/]+/gi)) {
23 const p = m[0]
24 if (p.length < 4 || /^\d/.test(p)) continue
25 out.add(p)
26 const base = p.split('/').pop()
27 if (base && base.length >= 4 && base.includes('.')) out.add(base)
28 }
29 } else if (v && typeof v === 'object') Object.values(v).forEach(walk)
30 }
31 walk(input)
32 return [...out]
33}
34export function guardDecisions(messages, decisions, goal, preserveRecent) {
35 const calls = collectToolCalls(messages, preserveRecent)
36 const byId = new Map(calls.map((c) => [c.id, c]))
37 let rescued = 0
38 const out = decisions.map((d) => {
39 if (d.action === 'keep') return d
40 const call = byId.get(d.id)
41 if (!call) return d
42 const later = goal + '\n' + messages.slice(call.resultIndex + 1).map((m) => m.text || '').join('\n')
43 const hit = refsOf(call.input).find((r) => later.includes(r))
44 if (!hit) return d
45 rescued++
46 return { ...d, action: 'keep', reason: 'kept', guard: hit }
47 })
48 return { decisions: out, calls, rescued }
49}
50
51const LANES = {
52 zen: { url: 'https://opencode.ai/zen/v1/systemone', model: 'jev-1.13-free', key: null },
53 commandcode: { url: 'https://api.commandcode.ai/provider/v1/systemone', model: 'typesafe/jev', key: 'COMMANDCODE_API_KEY' },
54 typesafe: { url: 'https://api.typesafe.ai/v1/systemone', model: 'jev-latest', key: 'TYPESAFE_API_KEY' },
55 openrouter: { url: 'https://openrouter.ai/api/alpha/decisions', model: 'typesafe/jev-latest', key: 'OPENROUTER_API_KEY' },
56}
57const UA = 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126 Safari/537.36 forge'
58
59const num = (o, k, d) => (typeof o[k] === 'number' && Number.isFinite(o[k]) ? o[k] : d)
60const str = (o, k, d) => (typeof o[k] === 'string' && o[k] ? o[k] : d)
61const bool = (o, k, d) => (typeof o[k] === 'boolean' ? o[k] : d)
62
63async function keyOf($, options, name) {
64 if (!name) return ''
65 if (typeof options[name] === 'string' && options[name]) return options[name]
66 switch (name) {
67 case 'COMMANDCODE_API_KEY': return (await $.env.get('COMMANDCODE_API_KEY')) || ''
68 case 'TYPESAFE_API_KEY': return (await $.env.get('TYPESAFE_API_KEY')) || ''
69 case 'OPENROUTER_API_KEY': return (await $.env.get('OPENROUTER_API_KEY')) || ''
70 }
71 return ''
72}
73
74// A JevAsker over $.http.fetch with forge's lane failover.
75function asker($, options) {
76 const order = str(options, 'jevLanes', 'zen,commandcode').split(',').map((s) => s.trim()).filter((s) => LANES[s])
77 return {
78 async ask(state, questions) {
79 const errors = []
80 for (const name of order) {
81 const lane = LANES[name]
82 const key = await keyOf($, options, lane.key)
83 if (lane.key && !key) { errors.push(`${name}: no key`); continue }
84 const headers = { 'content-type': 'application/json', 'user-agent': UA }
85 if (key) headers.authorization = `Bearer ${key}`
86 try {
87 const res = await $.http.fetch(lane.url, { method: 'POST', headers, body: JSON.stringify({ model: lane.model, state, questions }) })
88 if (!res.ok) { errors.push(`${name}: HTTP ${res.status}`); continue }
89 const body = JSON.parse(res.text)
90 if (!body || typeof body.answers !== 'object' || body.answers === null) { errors.push(`${name}: no answers`); continue }
91 return body
92 } catch (e) { errors.push(`${name}: ${e instanceof Error ? e.message : String(e)}`) }
93 }
94 throw new Error(`all Jev lanes failed (${errors.join('; ')})`)
95 },
96 }
97}
98
99// Same message mapping as fast-jev-compaction: untouched objects keep the engine's handle.
100function toSessionMessages(input, output) {
101 const messages = new Map(); const uses = new Map(); const results = new Map()
102 for (const m of input) {
103 messages.set(m, m)
104 for (const u of m.toolUses) uses.set(u, u)
105 for (const r of m.toolResults ?? []) results.set(r, r)
106 }
107 return output.map((m) => {
108 const own = messages.get(m)
109 if (own) return own
110 const rebuilt = { role: m.role, text: m.text, toolUses: m.toolUses.map((u) => uses.get(u) ?? { tool_use_id: u.tool_use_id, tool: u.tool, input: u.input, ...(u.text !== undefined ? { text: u.text } : {}), ...(u.isError ? { isError: true } : {}) }) }
111 if (m.toolResults?.length) rebuilt.toolResults = m.toolResults.map((r) => results.get(r) ?? { tool_use_id: r.tool_use_id, text: r.text, isError: r.isError ?? false })
112 return rebuilt
113 })
114}
115
116
117/** @type {import('claude-code').Register} */
118export const register = (on, options) => {
119 const TIER = { fast: str(options, 'fastModel', 'haiku'), balanced: str(options, 'balancedModel', 'claude-sonnet-5-5'), deep: str(options, 'deepModel', 'claude-opus-5-5') }
120 const compactOn = bool(options, 'jevCompaction', true)
121 const minReduction = num(options, 'minReductionRatio', 0.25)
122 const compactAt = num(options, 'compactAtPercent', 60)
123 let compacting = false
124
125 if (compactOn) {
126 on('session.compact', async ($, e, next) => {
127 try {
128 const goal = typeof e.instructions === 'string' ? e.instructions : ''
129 const preserve = num(options, 'preserveRecentMessages', 6)
130 const head = num(options, 'truncateHeadChars', 300)
131 const raw = await compact(e.messages, asker($, options), {
132 goal, keepThreshold: num(options, 'keepThreshold', 0.5), preserveRecentMessages: preserve,
133 maxStateTokens: num(options, 'maxStateTokens', 25000), maxRequestTokens: num(options, 'maxRequestTokens', 30000), truncateHeadChars: head,
134 })
135 let result = raw
136 if (bool(options, 'compactionRefGuard', true)) {
137 const g = guardDecisions(e.messages, raw.decisions, goal, preserve)
138 if (g.rescued) {
139 const kept = applyDecisions(e.messages, g.decisions, g.calls, head)
140 result = { messages: kept, decisions: g.decisions, stats: { ...raw.stats, messagesAfter: kept.length, charsAfter: kept.reduce((s, m) => s + messageChars(m), 0), rescued: g.rescued,
141 resultsDropped: g.decisions.filter((d) => d.reason === 'result_dropped').length, callsDropped: g.decisions.filter((d) => d.reason === 'call_dropped').length } }
142 }
143 }
144 const ratio = reductionRatio(result)
145 if (ratio < minReduction) {
146 $.ui.log(`forge compaction: ${Math.round(ratio * 100)}% < ${Math.round(minReduction * 100)}%, using built-in summary`)
147 return next(e)
148 }
149 const messages = toSessionMessages(e.messages, result.messages)
150 $.ui.log(`forge compaction: kept ${messages.length}/${e.messages.length} messages verbatim, ${Math.round(ratio * 100)}% smaller, ${result.stats.resultsDropped} results truncated, ${result.stats.callsDropped} calls dropped, ${result.stats.rescued || 0} rescued by ref guard`)
151 return { messages }
152 } catch (err) {
153 $.ui.log(`forge compaction fallback to built-in summary (${err instanceof Error ? err.message : String(err)})`)
154 return next(e)
155 }
156 }).catch(($, e, next) => {
157 // Budget overrun or misreturn: the built-in summary, never a stuck compaction.
158 $.ui.log(`forge compaction ${next.error.kind}: built-in summary`)
159 return next(e)
160 })
161
162 on('turn.complete', async ($, e, next) => {
163 if (compacting) return next(e)
164 try {
165 const { context } = await $.session.usage()
166 if ((context.percent ?? 0) >= compactAt) { compacting = true; await $.session.compact() }
167 } catch (err) {
168 $.ui.log(`forge auto-compact skipped (${err instanceof Error ? err.message : String(err)})`)
169 } finally { compacting = false }
170 return next(e)
171 })
172 }
173
174 if (bool(options, 'routeSessionSubagents', true)) {
175 // Interactive Agent-tool subagents: pick the tier from the task text. Only when the caller
176 // left the model open; forks always inherit; spending less needs high confidence.
177 on('agent.spawn', async ($, e, next) => {
178 if (e.fork || e.model) return next(e)
179 try {
180 const body = await asker($, options).ask({ subagent: e.subagentType, task: String(e.prompt).slice(0, 4000) }, {
181 tier: { type: 'choice', instructions: 'Which model tier does this subagent task need?', criteria: {
182 fast: 'search, read, list, summarise, or a mechanical edit',
183 balanced: 'ordinary engineering: implement or fix with tests',
184 deep: 'hard or risky: design, concurrency, security, data integrity, or debugging an unclear failure',
185 } },
186 })
187 const a = body.answers?.tier
188 const choice = a?.choice ?? a?.option
189 const conf = Number(a?.confidence ?? a?.probabilities?.[choice] ?? 0)
190 if (!TIER[choice]) return next(e)
191 const down = choice === 'fast'
192 if ((down && conf < num(options, 'minDowngradeConfidence', 0.7)) || (!down && conf < num(options, 'minUpgradeConfidence', 0.3))) return next(e)
193 $.ui.log(`forge route: ${e.subagentType} subagent on ${TIER[choice]} (jev ${choice} ${conf.toFixed(2)})`)
194 return next({ ...e, model: TIER[choice] })
195 } catch {
196 return next(e)
197 }
198 }).catch(($, e, next) => next(e))
199 }
200}
201hooks/vendor/fast-jev/compact.js 237 lines1import { noulAnswer } from "./request.js";
2import { collectToolCalls, estimateTokens, fitState } from "./state.js";
3const DEFAULT_OPTIONS = {
4 goal: "",
5 keepThreshold: 0.5,
6 preserveRecentMessages: 6,
7 maxStateTokens: 25e3,
8 maxRequestTokens: 3e4,
9 truncateHeadChars: 300
10};
11const REQUEST_OVERHEAD_TOKENS = 20;
12function finite(value, fallback) {
13 return typeof value === "number" && Number.isFinite(value) ? value : fallback;
14}
15function resolveOptions(options = {}) {
16 return {
17 goal: options.goal ?? DEFAULT_OPTIONS.goal,
18 keepThreshold: finite(options.keepThreshold, DEFAULT_OPTIONS.keepThreshold),
19 preserveRecentMessages: Math.max(
20 0,
21 Math.floor(
22 finite(options.preserveRecentMessages, DEFAULT_OPTIONS.preserveRecentMessages)
23 )
24 ),
25 maxStateTokens: Math.max(1, finite(options.maxStateTokens, DEFAULT_OPTIONS.maxStateTokens)),
26 maxRequestTokens: Math.max(
27 1,
28 finite(options.maxRequestTokens, DEFAULT_OPTIONS.maxRequestTokens)
29 ),
30 truncateHeadChars: Math.max(
31 0,
32 Math.floor(finite(options.truncateHeadChars, DEFAULT_OPTIONS.truncateHeadChars))
33 )
34 };
35}
36function questionsFor(call) {
37 return {
38 [`call_${call.id}`]: {
39 type: "noul",
40 instructions: `Tool call ${call.id} (${call.tool}) should stay in the history: knowing this call was made, with its input, still matters for what the assistant does next`
41 },
42 [`result_${call.id}`]: {
43 type: "noul",
44 instructions: `The full output of tool call ${call.id} (${call.tool}, ${call.resultChars} chars) should stay in the history verbatim: the assistant still needs its contents and re-running the tool would not do`
45 }
46 };
47}
48function batchCalls(calls, stateTokens, options) {
49 const budget = options.maxRequestTokens - stateTokens - REQUEST_OVERHEAD_TOKENS;
50 const batches = [];
51 let current = [];
52 let currentTokens = 0;
53 for (const call of calls) {
54 const tokens = estimateTokens(JSON.stringify(questionsFor(call)));
55 if (current.length > 0 && currentTokens + tokens > budget) {
56 batches.push(current);
57 current = [];
58 currentTokens = 0;
59 }
60 if (current.length === 0 && tokens > budget) {
61 throw new Error(
62 `state leaves no room for questions (~${stateTokens} of ${options.maxRequestTokens} tokens)`
63 );
64 }
65 current.push(call);
66 currentTokens += tokens;
67 }
68 if (current.length > 0) batches.push(current);
69 return batches;
70}
71function decideCall(call, answer, options) {
72 const base = { id: call.id, tool: call.tool, ...answer };
73 if (call.pinned) return { ...base, action: "keep", reason: "pinned" };
74 if (answer.keepResult >= options.keepThreshold) {
75 return { ...base, action: "keep", reason: "kept" };
76 }
77 if (answer.keepCall >= options.keepThreshold) {
78 return { ...base, action: "drop_result", reason: "result_dropped" };
79 }
80 return { ...base, action: "drop_call", reason: "call_dropped" };
81}
82async function askBatch(asker, state, batch) {
83 const questions = Object.assign({}, ...batch.map(questionsFor));
84 const { answers } = await asker.ask(state, questions);
85 return new Map(
86 batch.map((call) => [
87 call.id,
88 {
89 keepCall: noulAnswer(answers, `call_${call.id}`),
90 keepResult: noulAnswer(answers, `result_${call.id}`)
91 }
92 ])
93 );
94}
95function truncatedResultText(text, isError, headChars) {
96 if (text.length <= headChars + 120) return text;
97 const head = headChars > 0 ? `${text.slice(0, headChars)}
98` : "";
99 return `${head}[fast-jev-compaction truncated ${text.length - headChars} chars of this tool result${isError ? " (error)" : ""}; re-run the tool if needed]`;
100}
101function applyDecisions(messages, decisions, calls, headChars) {
102 const byId = new Map(calls.map((call) => [call.id, call]));
103 const actions = /* @__PURE__ */ new Map();
104 for (const decision of decisions) {
105 const call = byId.get(decision.id);
106 if (call && decision.action !== "keep") actions.set(call.tool_use_id, decision.action);
107 }
108 const kept = [];
109 for (const message of messages) {
110 const touched = message.toolUses.some((tool) => actions.has(tool.tool_use_id)) || (message.toolResults ?? []).some((result) => actions.has(result.tool_use_id));
111 if (!touched) {
112 kept.push(message);
113 continue;
114 }
115 const toolUses = message.toolUses.filter((tool) => actions.get(tool.tool_use_id) !== "drop_call").map((tool) => {
116 if (actions.get(tool.tool_use_id) !== "drop_result") return tool;
117 const text = truncatedResultText(
118 tool.text ?? "",
119 tool.isError ?? false,
120 headChars
121 );
122 if ((tool.text ?? "") === text) return tool;
123 const copy = {
124 tool_use_id: tool.tool_use_id,
125 tool: tool.tool,
126 input: tool.input,
127 text
128 };
129 if (tool.isError) copy.isError = true;
130 return copy;
131 });
132 const toolResults = (message.toolResults ?? []).filter((result) => actions.get(result.tool_use_id) !== "drop_call").map((result) => {
133 if (actions.get(result.tool_use_id) !== "drop_result") return result;
134 const text = truncatedResultText(result.text, result.isError ?? false, headChars);
135 return text === result.text ? result : {
136 tool_use_id: result.tool_use_id,
137 text,
138 isError: result.isError
139 };
140 });
141 if (!message.toolUses.some(
142 (tool) => actions.get(tool.tool_use_id) === "drop_call"
143 ) && !(message.toolResults ?? []).some(
144 (result) => actions.get(result.tool_use_id) === "drop_call"
145 ) && toolUses.every((tool, index) => tool === message.toolUses[index]) && toolResults.every(
146 (result, index) => result === message.toolResults?.[index]
147 )) {
148 kept.push(message);
149 continue;
150 }
151 if (message.text.trim().length === 0 && toolUses.length === 0 && toolResults.length === 0) {
152 continue;
153 }
154 const rebuilt = { role: message.role, text: message.text, toolUses };
155 if (toolResults.length > 0) rebuilt.toolResults = toolResults;
156 kept.push(rebuilt);
157 }
158 return kept;
159}
160function messageChars(message) {
161 let total = message.text.length;
162 for (const tool of message.toolUses) {
163 try {
164 total += JSON.stringify(tool.input).length;
165 } catch {
166 total += 20;
167 }
168 }
169 for (const result of message.toolResults ?? []) total += result.text.length;
170 return total;
171}
172function reductionRatio(result) {
173 const { charsBefore, charsAfter } = result.stats;
174 return charsBefore === 0 ? 0 : (charsBefore - charsAfter) / charsBefore;
175}
176function count(decisions, reason) {
177 return decisions.filter((decision) => decision.reason === reason).length;
178}
179async function compact(messages, asker, options = {}) {
180 const started = Date.now();
181 const resolved = resolveOptions(options);
182 const calls = collectToolCalls(messages, resolved.preserveRecentMessages);
183 const candidates = calls.filter((call) => !call.pinned);
184 const charsBefore = messages.reduce((sum, message) => sum + messageChars(message), 0);
185 let fitted = { tokens: 0, stage: "" };
186 let batches = [];
187 const answers = /* @__PURE__ */ new Map();
188 if (candidates.length > 0) {
189 const state = fitState(messages, calls, resolved);
190 fitted = state;
191 batches = batchCalls(candidates, state.tokens, resolved);
192 const answered = await Promise.all(
193 batches.map((batch) => askBatch(asker, state.state, batch))
194 );
195 for (const map of answered) for (const [id, answer] of map) answers.set(id, answer);
196 }
197 const decisions = calls.map(
198 (call) => decideCall(call, answers.get(call.id) ?? { keepCall: 1, keepResult: 1 }, resolved)
199 );
200 const kept = applyDecisions(
201 messages,
202 decisions,
203 calls,
204 resolved.truncateHeadChars
205 );
206 return {
207 messages: kept,
208 decisions,
209 stats: {
210 messagesBefore: messages.length,
211 messagesAfter: kept.length,
212 charsBefore,
213 charsAfter: kept.reduce((sum, message) => sum + messageChars(message), 0),
214 calls: calls.length,
215 kept: count(decisions, "kept"),
216 resultsDropped: count(decisions, "result_dropped"),
217 callsDropped: count(decisions, "call_dropped"),
218 pinned: count(decisions, "pinned"),
219 stateTokens: fitted.tokens,
220 stateStage: fitted.stage,
221 requests: batches.length,
222 ms: Date.now() - started
223 }
224 };
225}
226export {
227 DEFAULT_OPTIONS,
228 applyDecisions,
229 batchCalls,
230 compact,
231 decideCall,
232 messageChars,
233 questionsFor,
234 reductionRatio,
235 resolveOptions
236};
237hooks/vendor/fast-jev/state.js 224 lines1const STATE_CONTEXT = "A coding assistant conversation is being compacted to free context. `history` is the whole conversation so far, oldest first; tool outputs are replaced by a short `result` note and long texts may be abridged. Each question asks whether one tool call, or the full output of that call, still needs to stay in the history verbatim. Whatever is not kept is deleted permanently, but the assistant can always re-run a tool or re-read a file.";
2const INPUT_CHARS = [1e3, 200, 60];
3const TEXT_HEAD = 400;
4const TEXT_TAIL = 150;
5const TOKEN_PIECES = /[A-Za-z]+|\d+|[^\sA-Za-z\d]/g;
6function estimateTokens(text) {
7 let tokens = 0;
8 for (const [piece] of text.matchAll(TOKEN_PIECES)) {
9 const first = piece.charCodeAt(0);
10 if (first >= 48 && first <= 57) tokens += piece.length / 2;
11 else if (first >= 65 && first <= 90 || first >= 97 && first <= 122) {
12 tokens += 1 + Math.floor((piece.length - 1) / 6);
13 } else tokens += 0.9;
14 }
15 return Math.ceil(tokens);
16}
17function truncate(text, limit) {
18 return text.length <= limit ? text : `${text.slice(0, Math.max(0, limit - 1))}\u2026`;
19}
20function abridge(text, head, tail) {
21 if (text.length <= head + tail + 40) return text;
22 const omitted = text.length - head - tail;
23 return `${text.slice(0, head)}
24[\u2026 ${omitted} chars omitted \u2026]
25${text.slice(-tail)}`;
26}
27function isPinned(index, total, preserveRecentMessages) {
28 return index === 0 || index >= total - preserveRecentMessages;
29}
30function collectToolCalls(messages, preserveRecentMessages) {
31 const results = /* @__PURE__ */ new Map();
32 messages.forEach((message, index) => {
33 for (const result of message.toolResults ?? []) {
34 results.set(result.tool_use_id, { index, result });
35 }
36 });
37 const calls = [];
38 messages.forEach((message, callIndex) => {
39 for (const tool of message.toolUses) {
40 const found = results.get(tool.tool_use_id);
41 if (!found) continue;
42 calls.push({
43 id: `t${calls.length + 1}`,
44 tool_use_id: tool.tool_use_id,
45 tool: tool.tool,
46 input: tool.input,
47 callIndex,
48 resultIndex: found.index,
49 resultChars: found.result.text.length,
50 isError: found.result.isError ?? false,
51 pinned: isPinned(callIndex, messages.length, preserveRecentMessages) || isPinned(found.index, messages.length, preserveRecentMessages)
52 });
53 }
54 });
55 return calls;
56}
57function inputText(input, limit) {
58 let json = "";
59 try {
60 json = JSON.stringify(input);
61 } catch {
62 json = "[unserializable input]";
63 }
64 return truncate(json, limit);
65}
66function resultNote(call) {
67 return `${call.isError ? "error" : "ok"}, ${call.resultChars} chars (omitted)`;
68}
69function compactCall(call) {
70 const input = Object.entries(call.input).map(([key, value]) => {
71 const text = typeof value === "string" ? value : inputText({ [key]: value }, 200);
72 return `${key}=${text.replace(/\s+/g, " ")}`;
73 }).join(" ");
74 return `${call.id} ${call.tool} ${truncate(input, INPUT_CHARS[2])} \u2192 ${call.isError ? "error" : "ok"} ${call.resultChars}ch`;
75}
76function mergeCallRuns(history, pinned) {
77 const merged = [];
78 for (const entry of history) {
79 const previous = merged[merged.length - 1];
80 const foldable = (e) => !pinned(e) && e.text.length === 0 && typeof e.tool_calls?.[0] === "string";
81 if (previous && foldable(previous) && foldable(entry) && previous.role === entry.role) {
82 previous.tool_calls = [...previous.tool_calls, ...entry.tool_calls];
83 continue;
84 }
85 merged.push({ ...entry });
86 }
87 return merged;
88}
89function callsByMessage(calls) {
90 const byMessage = /* @__PURE__ */ new Map();
91 for (const call of calls) {
92 const list = byMessage.get(call.callIndex) ?? [];
93 list.push(call);
94 byMessage.set(call.callIndex, list);
95 }
96 return byMessage;
97}
98function historyEntries(messages, calls, inputChars) {
99 const byMessage = callsByMessage(calls);
100 const entries = [];
101 messages.forEach((message, i) => {
102 const toolCalls = (byMessage.get(i) ?? []).map((call) => ({
103 id: call.id,
104 tool: call.tool,
105 input: inputText(call.input, inputChars),
106 result: resultNote(call)
107 }));
108 if (message.text.trim().length === 0 && toolCalls.length === 0) return;
109 const entry = { i, role: message.role, text: message.text };
110 if (toolCalls.length > 0) entry.tool_calls = toolCalls;
111 entries.push(entry);
112 });
113 return entries;
114}
115function goalFromMessages(messages) {
116 return messages.filter(
117 (message) => message.role === "user" && message.text.trim().length > 0 && (message.toolResults ?? []).length === 0
118 ).slice(-3).map((message) => truncate(message.text, 500)).join("\n");
119}
120function fitState(messages, calls, options) {
121 const goal = options.goal || goalFromMessages(messages);
122 const stateOf = (history2) => ({
123 context: STATE_CONTEXT,
124 goal,
125 history: history2
126 });
127 const entryTokens = (entry) => estimateTokens(JSON.stringify(entry)) + 1;
128 const baseTokens = estimateTokens(JSON.stringify(stateOf([])));
129 const fitted = (history2, tokens2, stage) => ({
130 state: stateOf(history2),
131 tokens: tokens2,
132 stage
133 });
134 let history = [];
135 let perEntry = [];
136 let tokens = 0;
137 const rebuild = (inputChars) => {
138 history = historyEntries(messages, calls, inputChars);
139 perEntry = history.map(entryTokens);
140 tokens = baseTokens + perEntry.reduce((sum, n) => sum + n, 0);
141 };
142 const fits = () => tokens <= options.maxStateTokens;
143 const shrink = (index, change) => {
144 const entry = history[index];
145 if (!entry) return;
146 change(entry);
147 const now = entryTokens(entry);
148 tokens += now - (perEntry[index] ?? 0);
149 perEntry[index] = now;
150 };
151 rebuild(INPUT_CHARS[0]);
152 if (fits()) return fitted(history, tokens, "full");
153 for (const limit of INPUT_CHARS.slice(1)) {
154 rebuild(limit);
155 if (fits()) return fitted(history, tokens, `inputs<=${limit}`);
156 }
157 const pinned = (entry) => isPinned(entry.i, messages.length, options.preserveRecentMessages);
158 const indices = history.map((_, index) => index);
159 const order = [
160 ...indices.filter((index) => !pinned(history[index])),
161 ...indices.filter((index) => pinned(history[index]))
162 ];
163 for (const index of order) {
164 const entry = history[index];
165 if (entry.text.length <= TEXT_HEAD + TEXT_TAIL + 40) continue;
166 shrink(index, (e) => {
167 e.text = abridge(e.text, TEXT_HEAD, TEXT_TAIL);
168 });
169 if (fits()) return fitted(history, tokens, "texts abridged");
170 }
171 for (const index of order) {
172 const entry = history[index];
173 if (pinned(entry) || entry.text.length === 0) continue;
174 const original = messages[entry.i]?.text.length ?? entry.text.length;
175 shrink(index, (e) => {
176 e.text = `[\u2026 ${original} chars omitted \u2026]`;
177 });
178 if (fits()) return fitted(history, tokens, "old messages collapsed");
179 }
180 const byMessage = callsByMessage(calls);
181 for (const index of order) {
182 const entry = history[index];
183 const own = byMessage.get(entry.i);
184 if (pinned(entry) || !own) continue;
185 shrink(index, (e) => {
186 e.tool_calls = own.map(compactCall);
187 });
188 if (fits()) return fitted(history, tokens, "old calls compacted");
189 }
190 const left = /* @__PURE__ */ new Set();
191 for (const index of order) {
192 const entry = history[index];
193 if (pinned(entry) || entry.tool_calls) continue;
194 left.add(index);
195 tokens -= perEntry[index] ?? 0;
196 if (fits()) {
197 return fitted(
198 history.filter((_, i) => !left.has(i)),
199 tokens,
200 "old messages left out"
201 );
202 }
203 }
204 history = mergeCallRuns(
205 history.filter((_, i) => !left.has(i)),
206 pinned
207 );
208 perEntry = history.map(entryTokens);
209 tokens = baseTokens + perEntry.reduce((sum, n) => sum + n, 0);
210 if (fits()) return fitted(history, tokens, "old calls merged");
211 throw new Error(
212 `history too large for Jev (~${tokens} tokens after truncation, limit ${options.maxStateTokens})`
213 );
214}
215export {
216 STATE_CONTEXT,
217 collectToolCalls,
218 estimateTokens,
219 fitState,
220 goalFromMessages,
221 isPinned,
222 truncate
223};
224hooks/vendor/fast-jev/request.js 47 lines1const SYSTEM_ONE_URL = "https://api.typesafe.ai/v1/systemone";
2const DEFAULT_MODEL = "jev-latest";
3function buildJevRequest(params, state, questions) {
4 return {
5 url: params.baseUrl ?? SYSTEM_ONE_URL,
6 method: "POST",
7 headers: {
8 authorization: `Bearer ${params.apiKey}`,
9 "content-type": "application/json"
10 },
11 body: JSON.stringify({
12 model: params.model ?? DEFAULT_MODEL,
13 state,
14 questions
15 })
16 };
17}
18function parseJevResponse(status, ok, text) {
19 if (!ok) {
20 throw new Error(`Jev request failed (${status}): ${text.slice(0, 200)}`);
21 }
22 let parsed;
23 try {
24 parsed = JSON.parse(text);
25 } catch {
26 throw new Error("Jev returned malformed JSON");
27 }
28 if (parsed === null || typeof parsed !== "object" || !("answers" in parsed) || parsed.answers === null || typeof parsed.answers !== "object") {
29 throw new Error("Jev response is missing answers");
30 }
31 return parsed;
32}
33function noulAnswer(answers, name) {
34 const answer = answers[name];
35 if (!answer || !("noul" in answer) || typeof answer.noul !== "number" || !Number.isFinite(answer.noul)) {
36 throw new Error(`Invalid Jev answer for ${name}`);
37 }
38 return answer.noul;
39}
40export {
41 DEFAULT_MODEL,
42 SYSTEM_ONE_URL,
43 buildJevRequest,
44 noulAnswer,
45 parseJevResponse
46};
47