SLOPSHOPPER

forge

Continuous engineering loop for Claude Code: Jev-based size/clarity triage, a plan-clarify-build-test-revise harness for small tasks, and for large tasks…

newnetworkagents
v0.4.2MITupdated 2026-10-09sudikama/forge
A shopper browsing a rack in a slop shop
README

<img src="docs/assets/forge-banner.jpg" alt="forge: Continuous Engineering Loop for Claude Code" width="100%">

<strong>A continuous engineering loop for Claude Code.</strong><br> Triage every request, run small tasks through a test-driven harness, and drive large tasks to a measurable goal with an orchestrator, parallel subagents, and a deterministic judge.

<img alt="version" src="https://img.shields.io/badge/version-0.4.2-orange"> <img alt="Claude Code" src="https://img.shields.io/badge/Claude%20Code-%E2%89%A5%202.1.280-black"> <img alt="Node" src="https://img.shields.io/badge/node-%E2%89%A5%2020-339933"> <img alt="dependencies" src="https://img.shields.io/badge/dependencies-none-lightgrey"> <img alt="license" src="https://img.shields.io/badge/license-MIT-blue">


forge takes the ratchet idea from karpathy/autoresearch (one change, one commit, one measurement, keep it or reset it) and wraps it in the process a real team follows: a written spec, a mapped scope, clarifying questions, tests beyond the happy path, and a human sign-off before anything runs unattended.

The worker is always Claude. Jev, TypeSafe's decision model, only makes cheap, fast judgements: triage, model routing, finding pre-sorting, failure labelling, and context compaction. Pass or fail always comes from test exit codes, a baseline, and a metric.

<table> <tr> <td width="33%" valign="top"> <h3>Triage</h3> Each engineering prompt is classified as a question, a vague request, a small task, or a large task, and routed to the matching flow. </td> <td width="33%" valign="top"> <h3>Orchestrated goal loop</h3> Parallel subagents in their own worktrees, an orchestrator that rules on findings by citing the spec, and a judge that keeps or resets each iteration. </td> <td width="33%" valign="top"> <h3>Self-learning</h3> Repo-scoped instincts distilled from run evidence and owner corrections, fed back into the next task. </td> </tr> </table>

Contents

How it works

<img src="docs/assets/forge-flow.jpg" alt="forge flow: triage, small harness, large pipeline, ratchet loop, report" width="100%">

Small tasks

Plan with happy, unhappy, and edge cases, clarify open decisions up front, then build. When Claude ends a turn, a Stop hook runs the cases and the test command: red keeps the session working with the failing cases attached, green is done, and three red revisions escalate to you. A plan that grows past a few files is proposed for promotion to a large task.

Large tasks

Every stage has a gate, and forge check tells you what is missing.

StageOutput
Intakefrozen copy of the source (Jira, Google Doc, PDF, Markdown, PRD, Figma) and the report target
Goal and specgoal.json (test and regression commands, metric, target, budget, locked files) and spec.md, with a source for every acceptance criterion
Scopescope.json: in scope, impacted, out of scope, and conflicts, each with file:line evidence
Worklist and matrixvertical slices with owned files; happy, unhappy, and edge cases per criterion plus a regression case per impacted path
Lanessubagent count computed from file collisions, dependencies, free RAM, and quota; shared files belong to the orchestrator
Baselinematrix, suite, and metric measured twice at HEAD to catch flaky tests
Clarifynumbered questions A to J, each with a recommendation, including the subagent count
Spec lockexplicit approval; artifacts and evaluator files are hashed, and edits to them are denied

Then the loop runs in the background. Each iteration:

  • every lane is a fresh claude -p in its own git worktree and may only write the files it owns;
  • findings go to the orchestrator, which must cite a real line of the spec, goal, clarify answers, or scope. Anything it cannot ground becomes BLOCKED, and the loop continues with the other items;
  • the judge keeps the iteration only if nothing regressed, acceptance did not drop, the pre-existing suite is green, and acceptance or the metric improved. Otherwise it runs git reset.

The loop stops when the target is reached, the budget is spent, progress stalls, or every remaining item is BLOCKED. report.md lists finished work, BLOCKED items with questions and options, ticket proposals, decisions, model usage, and the metric curve. forge never merges or pushes.

Quick start

Requirements: Claude Code 2.1.280+ (logged in), Node.js 20+, and git. There are no npm dependencies.

claude plugin marketplace add sudikama/forge
claude plugin install forge@forge-local

Open a new Claude Code session inside a git repository and run /forge:doctor. Claude Code, auth, and git must be ok. Jev lanes without a key show FAIL and can be ignored as long as one lane works; without any, forge falls back to deterministic rules.

Then work as usual. Every prompt is triaged automatically:

TriageWhat happens
nonea question or conversation; forge stays out of the way
clarifyno checkable definition of done; Claude asks before writing code
smallsmall harness
largelarge pipeline; Claude may not write product code directly
CommandPurpose
`/forge:start <request \ticket \doc link>`triage manually and start
/forge:statusstate, failing gates, loop progress, BLOCKED items
/forge:doctorcheck prerequisites

To update, run claude plugin marketplace update forge-local && claude plugin update forge@forge-local.

Configuration

Set options with /plugin configure forge@forge-local inside Claude Code, or in ~/.forge/env as FORGE_<OPTION>=value (for example FORGE_MAX_WORKERS=2; API keys keep their own names). The process environment wins over the file. claude plugin configure forge@forge-local lists the current values.

OptionDefaultDescription
jevLaneszen,commandcodeJev backend failover order: zen (free, keyless), commandcode (200 requests/day), typesafe, openrouter
COMMANDCODE_API_KEY, TYPESAFE_API_KEY, OPENROUTER_API_KEYemptyKeys for the matching lanes; typesafe is the only lane with calibrated confidence
autoTriagetrueTriage every engineering prompt
maxWorkers4Hard cap on parallel subagents (also capped by free RAM, about 1.2 GB each)
smallMaxRevisions3Red test runs before the small harness escalates
fastModel, balancedModel, deepModelhaiku, claude-sonnet-5-5, claude-opus-5-5Model per tier
ladderFailsFast, ladderFailsBalanced, ladderFailsDeep1, 2, 2Gate failures allowed on each tier before moving up, or before BLOCKED on the top tier
maxEfforthighHighest effort routing may pick
jevDecide, jevRejectBartrue, 0.6Let Jev reject a finding that matches a spec exclusion line with at least this confidence
jevCompaction, compactAtPercenttrue, 60Jev-guided compaction and its trigger (function hooks only)
JIRA_BASE_URL, JIRA_EMAIL, JIRA_TOKENemptyJira REST fallback when no Jira MCP is registered
reportCmdemptyShell command for the command report target
extraInstinctsemptyOptional read-only instinct store (same <repo>/*.yaml layout) merged into prompts

The full list with descriptions lives in .claude-plugin/plugin.json. Every Jev call fails open, so an outage never blocks a session.

Spec sources

Large tasks read their source through MCP and freeze a hashed copy under .forge/tasks/<KEY>/source/. Register the servers you use:

claude mcp add jira -- <any Jira MCP server exposing get_issue>
claude mcp add --transport http figma https://mcp.figma.com/mcp
claude mcp add gdrive -- <a Google Drive MCP server with OAuth>

Without MCP, forge source add accepts a Jira key (REST fallback), a PDF, a Markdown or PRD file, or any text piped on --stdin.

Model routing

Each loop agent is a fresh claude -p, so switching models costs no prompt cache. Jev picks the starting tier and effort per lane: raising them needs 0.3 confidence, lowering them needs 0.7. L-sized items never start on Haiku, and the orchestrator never runs below Sonnet 5.5.

From then on, the gate moves each item up a model ladder:

TierModelFailures allowedThen
fastHaiku1Sonnet 5.5
balancedSonnet 5.52Opus 5.5
deepOpus 5.52BLOCKED, with every attempt as evidence

A failure is a red case owned by the item, a rejected or unmergeable lane, or a regression proven to come from that lane. When a shared check goes red, forge re-runs it in each lane's worktree on its own, so only the lane that breaks it is blamed. Passing never moves an item down.

PointWhat Jev doesGuard
Triagesize, parallelism, presence of a verifiera low verifier score forces clarify
Findingsmaps a finding to a Non-goals, MUST NOT, or out-of-scope linerejects only above jevRejectBar with a valid citation
Failureslabels why tests failed, to steer the next revisionnever affects the verdict
Compactionscores which tool calls are still neededfiles still referenced are kept; falls back to the built-in summary
Session subagentspicks the model of Agent-tool subagentsexplicit models and forks are untouched

Compaction and session subagent routing are function hooks (hooks/forge-fn.js). They load only with CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1, an early-access API that may change between releases.

Self-learning

Instincts are stored per repository in ~/.forge/learn/repos/<repo>/ and come only from run evidence: test commands proven green, recurring errors that were resolved, cases that caught a real regression, and your corrections (forge answer <ID> "<answer>" --correction "<lesson>"). Confidence starts at 0.5, rises 0.1 when confirmed, drops 0.2 when contradicted, and the instinct is dropped below 0.4. Changes to forge itself are only proposed in the report, never applied automatically.

CLI reference

Skills and hooks drive the CLI for you; these commands are useful directly. Run node bin/forge.mjs doctor --install-shim from a clone to put forge on your PATH.

forge doctor | triage "<request>"
forge new <KEY> --mode small|large --title ".." --report-to file|jira:KEY|webhook:URL|command
forge source add --kind jira|gdoc|pdf|markdown|prd|figma|file (--ref X | --file P | --stdin)
forge check | lanes | baseline | clarify | answer <ID> "<answer>" | lock --approve
forge run | status | pause | stop | resume
forge findings | report | learn list

Report targets: file always writes report.md; jira:KEY posts it as a comment; webhook:URL POSTs {subject, report, path} as JSON (Slack, Discord, n8n, anything); command runs the reportCmd option with FORGE_REPORT_PATH and FORGE_REPORT_SUBJECT set, for any notifier you already use.

Development

git clone https://github.com/sudikama/forge.git && cd forge
claude --plugin-dir .                 # run Claude Code with the working copy

bash tests/e2e/run-large.sh           # full large flow with a stub agent, deterministic, no cost
bash tests/e2e/run-small.sh           # small harness, lane gate, live triage on Jev zen
node tests/unit/route.test.mjs        # routing policy, model ladder, Jev finding decisions
node tests/unit/forge-fn.test.mjs     # function hooks on a fake engine (LIVE=1 hits Jev zen)
node tests/unit/deliver.test.mjs      # report delivery targets
bash tests/e2e/run-layouts.sh         # submodules, worktrees, missing .git/info, non-git dirs
FORGE_E2E_WORKTREE=1 bash tests/e2e/run-large.sh   # the large flow started from a linked worktree
PathContents
bin/forge.mjs, lib/CLI and core: triage, routing, lanes, judge, runner, clarify, learning, report
hooks/classic hooks, function hooks, vendored fast-jev-compaction
skills/, commands/, templates/instructions, slash commands, and artifact templates for Claude
tests/unit tests, end-to-end fixture, stub agent, Jev mock

In a target repository forge writes only to .forge/ (excluded from git automatically) and to forge/<KEY> branches.

Troubleshooting

SymptomFix
Hooks do not runstart a new session after installing; check claude plugin list
claude auth: not logged inrun claude then /login, or set ANTHROPIC_API_KEY
Jev lane returns HTTP 403Cloudflare; retry or switch lanes
cannot lock, gates still failingrun forge check and fix the listed items
baseline ran at X but HEAD is Ya commit landed after the baseline; run forge baseline again
Loop stops with every remaining item is blockedanswer the BLOCKED items in report.md, then open a follow-up task
Only one workerlow free RAM or file collisions; see forge lanes
forge new fails with ENOTDIR or ENOENT on .git/info/excludeforge older than 0.4.1 in a worktree or submodule; update the plugin and start a new session (/forge:doctor shows the version)

Logs are in ~/.forge/logs/ and .forge/tasks/<KEY>/runner.log.

Known limitations

  • The large loop is verified end to end with a stub agent; real runs need a logged-in Claude Code.
  • Function hooks are tested on a fake engine plus the live Jev zen lane, not yet in a real session.
  • Jev zen confidence is uncalibrated, so most findings still go to the Claude orchestrator.
  • Triage thresholds are calibrated on a small sample; use forge triage to check a misclassified request.
  • Regression detection is only as strong as the existing tests plus the characterization tests written at baseline.

License and credits

MIT. Built on ideas and code from karpathy/autoresearch, tamaratran/fast-jev-compaction (vendored under MIT), moelahmady/jev-model-router, ifoster01/jev-effort, and shitianfang/jev-use.

Source 4 files
hooks/forge-fn.js 201 lines
1// forge function-hook module (EARLY ACCESS API; only loaded when Claude Code runs with
2// CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1). Classic hooks in hooks.json keep working without it.
3//
4//   session.compact : Jev-guided verbatim compaction (vendored fast-jev-compaction, MIT):
5//                     Jev scores every tool call/result as still-needed or not; unneeded results
6//                     are truncated, unneeded calls dropped, no summary. Falls back to the
7//                     built-in summary below minReductionRatio or on any error.
8//   turn.complete   : asks for a compaction at compactAtPercent of the context window.
9//   agent.spawn     : routes the model of Agent-tool subagents inside an interactive session.
10//
11// Runs in Claude Code's hook sandbox: no Node, only `$` (http.fetch, env, settings, ui, session).
12import { compact, reductionRatio, applyDecisions, messageChars } from './vendor/fast-jev/compact.js'
13import { collectToolCalls } from './vendor/fast-jev/state.js'
14
15// Reference guard (forge, deterministic): the zen lane is not calibrated and was seen scoring
16// the file under edit at 0.17. A call whose input names a file/identifier that the goal or any
17// LATER message text mentions is still in use, so it is kept whatever Jev says.
18function refsOf(input) {
19  const out = new Set()
20  const walk = (v) => {
21    if (typeof v === 'string') {
22      for (const m of v.matchAll(/[\w.\-/]*[\w-]+\.[a-z0-9]{1,6}\b|[\w-]+\/[\w.\-/]+/gi)) {
23        const p = m[0]
24        if (p.length < 4 || /^\d/.test(p)) continue
25        out.add(p)
26        const base = p.split('/').pop()
27        if (base && base.length >= 4 && base.includes('.')) out.add(base)
28      }
29    } else if (v && typeof v === 'object') Object.values(v).forEach(walk)
30  }
31  walk(input)
32  return [...out]
33}
34export function guardDecisions(messages, decisions, goal, preserveRecent) {
35  const calls = collectToolCalls(messages, preserveRecent)
36  const byId = new Map(calls.map((c) => [c.id, c]))
37  let rescued = 0
38  const out = decisions.map((d) => {
39    if (d.action === 'keep') return d
40    const call = byId.get(d.id)
41    if (!call) return d
42    const later = goal + '\n' + messages.slice(call.resultIndex + 1).map((m) => m.text || '').join('\n')
43    const hit = refsOf(call.input).find((r) => later.includes(r))
44    if (!hit) return d
45    rescued++
46    return { ...d, action: 'keep', reason: 'kept', guard: hit }
47  })
48  return { decisions: out, calls, rescued }
49}
50
51const LANES = {
52  zen: { url: 'https://opencode.ai/zen/v1/systemone', model: 'jev-1.13-free', key: null },
53  commandcode: { url: 'https://api.commandcode.ai/provider/v1/systemone', model: 'typesafe/jev', key: 'COMMANDCODE_API_KEY' },
54  typesafe: { url: 'https://api.typesafe.ai/v1/systemone', model: 'jev-latest', key: 'TYPESAFE_API_KEY' },
55  openrouter: { url: 'https://openrouter.ai/api/alpha/decisions', model: 'typesafe/jev-latest', key: 'OPENROUTER_API_KEY' },
56}
57const UA = 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126 Safari/537.36 forge'
58
59const num = (o, k, d) => (typeof o[k] === 'number' && Number.isFinite(o[k]) ? o[k] : d)
60const str = (o, k, d) => (typeof o[k] === 'string' && o[k] ? o[k] : d)
61const bool = (o, k, d) => (typeof o[k] === 'boolean' ? o[k] : d)
62
63async function keyOf($, options, name) {
64  if (!name) return ''
65  if (typeof options[name] === 'string' && options[name]) return options[name]
66  switch (name) {
67    case 'COMMANDCODE_API_KEY': return (await $.env.get('COMMANDCODE_API_KEY')) || ''
68    case 'TYPESAFE_API_KEY': return (await $.env.get('TYPESAFE_API_KEY')) || ''
69    case 'OPENROUTER_API_KEY': return (await $.env.get('OPENROUTER_API_KEY')) || ''
70  }
71  return ''
72}
73
74// A JevAsker over $.http.fetch with forge's lane failover.
75function asker($, options) {
76  const order = str(options, 'jevLanes', 'zen,commandcode').split(',').map((s) => s.trim()).filter((s) => LANES[s])
77  return {
78    async ask(state, questions) {
79      const errors = []
80      for (const name of order) {
81        const lane = LANES[name]
82        const key = await keyOf($, options, lane.key)
83        if (lane.key && !key) { errors.push(`${name}: no key`); continue }
84        const headers = { 'content-type': 'application/json', 'user-agent': UA }
85        if (key) headers.authorization = `Bearer ${key}`
86        try {
87          const res = await $.http.fetch(lane.url, { method: 'POST', headers, body: JSON.stringify({ model: lane.model, state, questions }) })
88          if (!res.ok) { errors.push(`${name}: HTTP ${res.status}`); continue }
89          const body = JSON.parse(res.text)
90          if (!body || typeof body.answers !== 'object' || body.answers === null) { errors.push(`${name}: no answers`); continue }
91          return body
92        } catch (e) { errors.push(`${name}: ${e instanceof Error ? e.message : String(e)}`) }
93      }
94      throw new Error(`all Jev lanes failed (${errors.join('; ')})`)
95    },
96  }
97}
98
99// Same message mapping as fast-jev-compaction: untouched objects keep the engine's handle.
100function toSessionMessages(input, output) {
101  const messages = new Map(); const uses = new Map(); const results = new Map()
102  for (const m of input) {
103    messages.set(m, m)
104    for (const u of m.toolUses) uses.set(u, u)
105    for (const r of m.toolResults ?? []) results.set(r, r)
106  }
107  return output.map((m) => {
108    const own = messages.get(m)
109    if (own) return own
110    const rebuilt = { role: m.role, text: m.text, toolUses: m.toolUses.map((u) => uses.get(u) ?? { tool_use_id: u.tool_use_id, tool: u.tool, input: u.input, ...(u.text !== undefined ? { text: u.text } : {}), ...(u.isError ? { isError: true } : {}) }) }
111    if (m.toolResults?.length) rebuilt.toolResults = m.toolResults.map((r) => results.get(r) ?? { tool_use_id: r.tool_use_id, text: r.text, isError: r.isError ?? false })
112    return rebuilt
113  })
114}
115
116
117/** @type {import('claude-code').Register} */
118export const register = (on, options) => {
119  const TIER = { fast: str(options, 'fastModel', 'haiku'), balanced: str(options, 'balancedModel', 'claude-sonnet-5-5'), deep: str(options, 'deepModel', 'claude-opus-5-5') }
120  const compactOn = bool(options, 'jevCompaction', true)
121  const minReduction = num(options, 'minReductionRatio', 0.25)
122  const compactAt = num(options, 'compactAtPercent', 60)
123  let compacting = false
124
125  if (compactOn) {
126    on('session.compact', async ($, e, next) => {
127      try {
128        const goal = typeof e.instructions === 'string' ? e.instructions : ''
129        const preserve = num(options, 'preserveRecentMessages', 6)
130        const head = num(options, 'truncateHeadChars', 300)
131        const raw = await compact(e.messages, asker($, options), {
132          goal, keepThreshold: num(options, 'keepThreshold', 0.5), preserveRecentMessages: preserve,
133          maxStateTokens: num(options, 'maxStateTokens', 25000), maxRequestTokens: num(options, 'maxRequestTokens', 30000), truncateHeadChars: head,
134        })
135        let result = raw
136        if (bool(options, 'compactionRefGuard', true)) {
137          const g = guardDecisions(e.messages, raw.decisions, goal, preserve)
138          if (g.rescued) {
139            const kept = applyDecisions(e.messages, g.decisions, g.calls, head)
140            result = { messages: kept, decisions: g.decisions, stats: { ...raw.stats, messagesAfter: kept.length, charsAfter: kept.reduce((s, m) => s + messageChars(m), 0), rescued: g.rescued,
141              resultsDropped: g.decisions.filter((d) => d.reason === 'result_dropped').length, callsDropped: g.decisions.filter((d) => d.reason === 'call_dropped').length } }
142          }
143        }
144        const ratio = reductionRatio(result)
145        if (ratio < minReduction) {
146          $.ui.log(`forge compaction: ${Math.round(ratio * 100)}% < ${Math.round(minReduction * 100)}%, using built-in summary`)
147          return next(e)
148        }
149        const messages = toSessionMessages(e.messages, result.messages)
150        $.ui.log(`forge compaction: kept ${messages.length}/${e.messages.length} messages verbatim, ${Math.round(ratio * 100)}% smaller, ${result.stats.resultsDropped} results truncated, ${result.stats.callsDropped} calls dropped, ${result.stats.rescued || 0} rescued by ref guard`)
151        return { messages }
152      } catch (err) {
153        $.ui.log(`forge compaction fallback to built-in summary (${err instanceof Error ? err.message : String(err)})`)
154        return next(e)
155      }
156    }).catch(($, e, next) => {
157      // Budget overrun or misreturn: the built-in summary, never a stuck compaction.
158      $.ui.log(`forge compaction ${next.error.kind}: built-in summary`)
159      return next(e)
160    })
161
162    on('turn.complete', async ($, e, next) => {
163      if (compacting) return next(e)
164      try {
165        const { context } = await $.session.usage()
166        if ((context.percent ?? 0) >= compactAt) { compacting = true; await $.session.compact() }
167      } catch (err) {
168        $.ui.log(`forge auto-compact skipped (${err instanceof Error ? err.message : String(err)})`)
169      } finally { compacting = false }
170      return next(e)
171    })
172  }
173
174  if (bool(options, 'routeSessionSubagents', true)) {
175    // Interactive Agent-tool subagents: pick the tier from the task text. Only when the caller
176    // left the model open; forks always inherit; spending less needs high confidence.
177    on('agent.spawn', async ($, e, next) => {
178      if (e.fork || e.model) return next(e)
179      try {
180        const body = await asker($, options).ask({ subagent: e.subagentType, task: String(e.prompt).slice(0, 4000) }, {
181          tier: { type: 'choice', instructions: 'Which model tier does this subagent task need?', criteria: {
182            fast: 'search, read, list, summarise, or a mechanical edit',
183            balanced: 'ordinary engineering: implement or fix with tests',
184            deep: 'hard or risky: design, concurrency, security, data integrity, or debugging an unclear failure',
185          } },
186        })
187        const a = body.answers?.tier
188        const choice = a?.choice ?? a?.option
189        const conf = Number(a?.confidence ?? a?.probabilities?.[choice] ?? 0)
190        if (!TIER[choice]) return next(e)
191        const down = choice === 'fast'
192        if ((down && conf < num(options, 'minDowngradeConfidence', 0.7)) || (!down && conf < num(options, 'minUpgradeConfidence', 0.3))) return next(e)
193        $.ui.log(`forge route: ${e.subagentType} subagent on ${TIER[choice]} (jev ${choice} ${conf.toFixed(2)})`)
194        return next({ ...e, model: TIER[choice] })
195      } catch {
196        return next(e)
197      }
198    }).catch(($, e, next) => next(e))
199  }
200}
201
hooks/vendor/fast-jev/compact.js 237 lines
1import { noulAnswer } from "./request.js";
2import { collectToolCalls, estimateTokens, fitState } from "./state.js";
3const DEFAULT_OPTIONS = {
4  goal: "",
5  keepThreshold: 0.5,
6  preserveRecentMessages: 6,
7  maxStateTokens: 25e3,
8  maxRequestTokens: 3e4,
9  truncateHeadChars: 300
10};
11const REQUEST_OVERHEAD_TOKENS = 20;
12function finite(value, fallback) {
13  return typeof value === "number" && Number.isFinite(value) ? value : fallback;
14}
15function resolveOptions(options = {}) {
16  return {
17    goal: options.goal ?? DEFAULT_OPTIONS.goal,
18    keepThreshold: finite(options.keepThreshold, DEFAULT_OPTIONS.keepThreshold),
19    preserveRecentMessages: Math.max(
20      0,
21      Math.floor(
22        finite(options.preserveRecentMessages, DEFAULT_OPTIONS.preserveRecentMessages)
23      )
24    ),
25    maxStateTokens: Math.max(1, finite(options.maxStateTokens, DEFAULT_OPTIONS.maxStateTokens)),
26    maxRequestTokens: Math.max(
27      1,
28      finite(options.maxRequestTokens, DEFAULT_OPTIONS.maxRequestTokens)
29    ),
30    truncateHeadChars: Math.max(
31      0,
32      Math.floor(finite(options.truncateHeadChars, DEFAULT_OPTIONS.truncateHeadChars))
33    )
34  };
35}
36function questionsFor(call) {
37  return {
38    [`call_${call.id}`]: {
39      type: "noul",
40      instructions: `Tool call ${call.id} (${call.tool}) should stay in the history: knowing this call was made, with its input, still matters for what the assistant does next`
41    },
42    [`result_${call.id}`]: {
43      type: "noul",
44      instructions: `The full output of tool call ${call.id} (${call.tool}, ${call.resultChars} chars) should stay in the history verbatim: the assistant still needs its contents and re-running the tool would not do`
45    }
46  };
47}
48function batchCalls(calls, stateTokens, options) {
49  const budget = options.maxRequestTokens - stateTokens - REQUEST_OVERHEAD_TOKENS;
50  const batches = [];
51  let current = [];
52  let currentTokens = 0;
53  for (const call of calls) {
54    const tokens = estimateTokens(JSON.stringify(questionsFor(call)));
55    if (current.length > 0 && currentTokens + tokens > budget) {
56      batches.push(current);
57      current = [];
58      currentTokens = 0;
59    }
60    if (current.length === 0 && tokens > budget) {
61      throw new Error(
62        `state leaves no room for questions (~${stateTokens} of ${options.maxRequestTokens} tokens)`
63      );
64    }
65    current.push(call);
66    currentTokens += tokens;
67  }
68  if (current.length > 0) batches.push(current);
69  return batches;
70}
71function decideCall(call, answer, options) {
72  const base = { id: call.id, tool: call.tool, ...answer };
73  if (call.pinned) return { ...base, action: "keep", reason: "pinned" };
74  if (answer.keepResult >= options.keepThreshold) {
75    return { ...base, action: "keep", reason: "kept" };
76  }
77  if (answer.keepCall >= options.keepThreshold) {
78    return { ...base, action: "drop_result", reason: "result_dropped" };
79  }
80  return { ...base, action: "drop_call", reason: "call_dropped" };
81}
82async function askBatch(asker, state, batch) {
83  const questions = Object.assign({}, ...batch.map(questionsFor));
84  const { answers } = await asker.ask(state, questions);
85  return new Map(
86    batch.map((call) => [
87      call.id,
88      {
89        keepCall: noulAnswer(answers, `call_${call.id}`),
90        keepResult: noulAnswer(answers, `result_${call.id}`)
91      }
92    ])
93  );
94}
95function truncatedResultText(text, isError, headChars) {
96  if (text.length <= headChars + 120) return text;
97  const head = headChars > 0 ? `${text.slice(0, headChars)}
98` : "";
99  return `${head}[fast-jev-compaction truncated ${text.length - headChars} chars of this tool result${isError ? " (error)" : ""}; re-run the tool if needed]`;
100}
101function applyDecisions(messages, decisions, calls, headChars) {
102  const byId = new Map(calls.map((call) => [call.id, call]));
103  const actions = /* @__PURE__ */ new Map();
104  for (const decision of decisions) {
105    const call = byId.get(decision.id);
106    if (call && decision.action !== "keep") actions.set(call.tool_use_id, decision.action);
107  }
108  const kept = [];
109  for (const message of messages) {
110    const touched = message.toolUses.some((tool) => actions.has(tool.tool_use_id)) || (message.toolResults ?? []).some((result) => actions.has(result.tool_use_id));
111    if (!touched) {
112      kept.push(message);
113      continue;
114    }
115    const toolUses = message.toolUses.filter((tool) => actions.get(tool.tool_use_id) !== "drop_call").map((tool) => {
116      if (actions.get(tool.tool_use_id) !== "drop_result") return tool;
117      const text = truncatedResultText(
118        tool.text ?? "",
119        tool.isError ?? false,
120        headChars
121      );
122      if ((tool.text ?? "") === text) return tool;
123      const copy = {
124        tool_use_id: tool.tool_use_id,
125        tool: tool.tool,
126        input: tool.input,
127        text
128      };
129      if (tool.isError) copy.isError = true;
130      return copy;
131    });
132    const toolResults = (message.toolResults ?? []).filter((result) => actions.get(result.tool_use_id) !== "drop_call").map((result) => {
133      if (actions.get(result.tool_use_id) !== "drop_result") return result;
134      const text = truncatedResultText(result.text, result.isError ?? false, headChars);
135      return text === result.text ? result : {
136        tool_use_id: result.tool_use_id,
137        text,
138        isError: result.isError
139      };
140    });
141    if (!message.toolUses.some(
142      (tool) => actions.get(tool.tool_use_id) === "drop_call"
143    ) && !(message.toolResults ?? []).some(
144      (result) => actions.get(result.tool_use_id) === "drop_call"
145    ) && toolUses.every((tool, index) => tool === message.toolUses[index]) && toolResults.every(
146      (result, index) => result === message.toolResults?.[index]
147    )) {
148      kept.push(message);
149      continue;
150    }
151    if (message.text.trim().length === 0 && toolUses.length === 0 && toolResults.length === 0) {
152      continue;
153    }
154    const rebuilt = { role: message.role, text: message.text, toolUses };
155    if (toolResults.length > 0) rebuilt.toolResults = toolResults;
156    kept.push(rebuilt);
157  }
158  return kept;
159}
160function messageChars(message) {
161  let total = message.text.length;
162  for (const tool of message.toolUses) {
163    try {
164      total += JSON.stringify(tool.input).length;
165    } catch {
166      total += 20;
167    }
168  }
169  for (const result of message.toolResults ?? []) total += result.text.length;
170  return total;
171}
172function reductionRatio(result) {
173  const { charsBefore, charsAfter } = result.stats;
174  return charsBefore === 0 ? 0 : (charsBefore - charsAfter) / charsBefore;
175}
176function count(decisions, reason) {
177  return decisions.filter((decision) => decision.reason === reason).length;
178}
179async function compact(messages, asker, options = {}) {
180  const started = Date.now();
181  const resolved = resolveOptions(options);
182  const calls = collectToolCalls(messages, resolved.preserveRecentMessages);
183  const candidates = calls.filter((call) => !call.pinned);
184  const charsBefore = messages.reduce((sum, message) => sum + messageChars(message), 0);
185  let fitted = { tokens: 0, stage: "" };
186  let batches = [];
187  const answers = /* @__PURE__ */ new Map();
188  if (candidates.length > 0) {
189    const state = fitState(messages, calls, resolved);
190    fitted = state;
191    batches = batchCalls(candidates, state.tokens, resolved);
192    const answered = await Promise.all(
193      batches.map((batch) => askBatch(asker, state.state, batch))
194    );
195    for (const map of answered) for (const [id, answer] of map) answers.set(id, answer);
196  }
197  const decisions = calls.map(
198    (call) => decideCall(call, answers.get(call.id) ?? { keepCall: 1, keepResult: 1 }, resolved)
199  );
200  const kept = applyDecisions(
201    messages,
202    decisions,
203    calls,
204    resolved.truncateHeadChars
205  );
206  return {
207    messages: kept,
208    decisions,
209    stats: {
210      messagesBefore: messages.length,
211      messagesAfter: kept.length,
212      charsBefore,
213      charsAfter: kept.reduce((sum, message) => sum + messageChars(message), 0),
214      calls: calls.length,
215      kept: count(decisions, "kept"),
216      resultsDropped: count(decisions, "result_dropped"),
217      callsDropped: count(decisions, "call_dropped"),
218      pinned: count(decisions, "pinned"),
219      stateTokens: fitted.tokens,
220      stateStage: fitted.stage,
221      requests: batches.length,
222      ms: Date.now() - started
223    }
224  };
225}
226export {
227  DEFAULT_OPTIONS,
228  applyDecisions,
229  batchCalls,
230  compact,
231  decideCall,
232  messageChars,
233  questionsFor,
234  reductionRatio,
235  resolveOptions
236};
237
hooks/vendor/fast-jev/state.js 224 lines
1const STATE_CONTEXT = "A coding assistant conversation is being compacted to free context. `history` is the whole conversation so far, oldest first; tool outputs are replaced by a short `result` note and long texts may be abridged. Each question asks whether one tool call, or the full output of that call, still needs to stay in the history verbatim. Whatever is not kept is deleted permanently, but the assistant can always re-run a tool or re-read a file.";
2const INPUT_CHARS = [1e3, 200, 60];
3const TEXT_HEAD = 400;
4const TEXT_TAIL = 150;
5const TOKEN_PIECES = /[A-Za-z]+|\d+|[^\sA-Za-z\d]/g;
6function estimateTokens(text) {
7  let tokens = 0;
8  for (const [piece] of text.matchAll(TOKEN_PIECES)) {
9    const first = piece.charCodeAt(0);
10    if (first >= 48 && first <= 57) tokens += piece.length / 2;
11    else if (first >= 65 && first <= 90 || first >= 97 && first <= 122) {
12      tokens += 1 + Math.floor((piece.length - 1) / 6);
13    } else tokens += 0.9;
14  }
15  return Math.ceil(tokens);
16}
17function truncate(text, limit) {
18  return text.length <= limit ? text : `${text.slice(0, Math.max(0, limit - 1))}\u2026`;
19}
20function abridge(text, head, tail) {
21  if (text.length <= head + tail + 40) return text;
22  const omitted = text.length - head - tail;
23  return `${text.slice(0, head)}
24[\u2026 ${omitted} chars omitted \u2026]
25${text.slice(-tail)}`;
26}
27function isPinned(index, total, preserveRecentMessages) {
28  return index === 0 || index >= total - preserveRecentMessages;
29}
30function collectToolCalls(messages, preserveRecentMessages) {
31  const results = /* @__PURE__ */ new Map();
32  messages.forEach((message, index) => {
33    for (const result of message.toolResults ?? []) {
34      results.set(result.tool_use_id, { index, result });
35    }
36  });
37  const calls = [];
38  messages.forEach((message, callIndex) => {
39    for (const tool of message.toolUses) {
40      const found = results.get(tool.tool_use_id);
41      if (!found) continue;
42      calls.push({
43        id: `t${calls.length + 1}`,
44        tool_use_id: tool.tool_use_id,
45        tool: tool.tool,
46        input: tool.input,
47        callIndex,
48        resultIndex: found.index,
49        resultChars: found.result.text.length,
50        isError: found.result.isError ?? false,
51        pinned: isPinned(callIndex, messages.length, preserveRecentMessages) || isPinned(found.index, messages.length, preserveRecentMessages)
52      });
53    }
54  });
55  return calls;
56}
57function inputText(input, limit) {
58  let json = "";
59  try {
60    json = JSON.stringify(input);
61  } catch {
62    json = "[unserializable input]";
63  }
64  return truncate(json, limit);
65}
66function resultNote(call) {
67  return `${call.isError ? "error" : "ok"}, ${call.resultChars} chars (omitted)`;
68}
69function compactCall(call) {
70  const input = Object.entries(call.input).map(([key, value]) => {
71    const text = typeof value === "string" ? value : inputText({ [key]: value }, 200);
72    return `${key}=${text.replace(/\s+/g, " ")}`;
73  }).join(" ");
74  return `${call.id} ${call.tool} ${truncate(input, INPUT_CHARS[2])} \u2192 ${call.isError ? "error" : "ok"} ${call.resultChars}ch`;
75}
76function mergeCallRuns(history, pinned) {
77  const merged = [];
78  for (const entry of history) {
79    const previous = merged[merged.length - 1];
80    const foldable = (e) => !pinned(e) && e.text.length === 0 && typeof e.tool_calls?.[0] === "string";
81    if (previous && foldable(previous) && foldable(entry) && previous.role === entry.role) {
82      previous.tool_calls = [...previous.tool_calls, ...entry.tool_calls];
83      continue;
84    }
85    merged.push({ ...entry });
86  }
87  return merged;
88}
89function callsByMessage(calls) {
90  const byMessage = /* @__PURE__ */ new Map();
91  for (const call of calls) {
92    const list = byMessage.get(call.callIndex) ?? [];
93    list.push(call);
94    byMessage.set(call.callIndex, list);
95  }
96  return byMessage;
97}
98function historyEntries(messages, calls, inputChars) {
99  const byMessage = callsByMessage(calls);
100  const entries = [];
101  messages.forEach((message, i) => {
102    const toolCalls = (byMessage.get(i) ?? []).map((call) => ({
103      id: call.id,
104      tool: call.tool,
105      input: inputText(call.input, inputChars),
106      result: resultNote(call)
107    }));
108    if (message.text.trim().length === 0 && toolCalls.length === 0) return;
109    const entry = { i, role: message.role, text: message.text };
110    if (toolCalls.length > 0) entry.tool_calls = toolCalls;
111    entries.push(entry);
112  });
113  return entries;
114}
115function goalFromMessages(messages) {
116  return messages.filter(
117    (message) => message.role === "user" && message.text.trim().length > 0 && (message.toolResults ?? []).length === 0
118  ).slice(-3).map((message) => truncate(message.text, 500)).join("\n");
119}
120function fitState(messages, calls, options) {
121  const goal = options.goal || goalFromMessages(messages);
122  const stateOf = (history2) => ({
123    context: STATE_CONTEXT,
124    goal,
125    history: history2
126  });
127  const entryTokens = (entry) => estimateTokens(JSON.stringify(entry)) + 1;
128  const baseTokens = estimateTokens(JSON.stringify(stateOf([])));
129  const fitted = (history2, tokens2, stage) => ({
130    state: stateOf(history2),
131    tokens: tokens2,
132    stage
133  });
134  let history = [];
135  let perEntry = [];
136  let tokens = 0;
137  const rebuild = (inputChars) => {
138    history = historyEntries(messages, calls, inputChars);
139    perEntry = history.map(entryTokens);
140    tokens = baseTokens + perEntry.reduce((sum, n) => sum + n, 0);
141  };
142  const fits = () => tokens <= options.maxStateTokens;
143  const shrink = (index, change) => {
144    const entry = history[index];
145    if (!entry) return;
146    change(entry);
147    const now = entryTokens(entry);
148    tokens += now - (perEntry[index] ?? 0);
149    perEntry[index] = now;
150  };
151  rebuild(INPUT_CHARS[0]);
152  if (fits()) return fitted(history, tokens, "full");
153  for (const limit of INPUT_CHARS.slice(1)) {
154    rebuild(limit);
155    if (fits()) return fitted(history, tokens, `inputs<=${limit}`);
156  }
157  const pinned = (entry) => isPinned(entry.i, messages.length, options.preserveRecentMessages);
158  const indices = history.map((_, index) => index);
159  const order = [
160    ...indices.filter((index) => !pinned(history[index])),
161    ...indices.filter((index) => pinned(history[index]))
162  ];
163  for (const index of order) {
164    const entry = history[index];
165    if (entry.text.length <= TEXT_HEAD + TEXT_TAIL + 40) continue;
166    shrink(index, (e) => {
167      e.text = abridge(e.text, TEXT_HEAD, TEXT_TAIL);
168    });
169    if (fits()) return fitted(history, tokens, "texts abridged");
170  }
171  for (const index of order) {
172    const entry = history[index];
173    if (pinned(entry) || entry.text.length === 0) continue;
174    const original = messages[entry.i]?.text.length ?? entry.text.length;
175    shrink(index, (e) => {
176      e.text = `[\u2026 ${original} chars omitted \u2026]`;
177    });
178    if (fits()) return fitted(history, tokens, "old messages collapsed");
179  }
180  const byMessage = callsByMessage(calls);
181  for (const index of order) {
182    const entry = history[index];
183    const own = byMessage.get(entry.i);
184    if (pinned(entry) || !own) continue;
185    shrink(index, (e) => {
186      e.tool_calls = own.map(compactCall);
187    });
188    if (fits()) return fitted(history, tokens, "old calls compacted");
189  }
190  const left = /* @__PURE__ */ new Set();
191  for (const index of order) {
192    const entry = history[index];
193    if (pinned(entry) || entry.tool_calls) continue;
194    left.add(index);
195    tokens -= perEntry[index] ?? 0;
196    if (fits()) {
197      return fitted(
198        history.filter((_, i) => !left.has(i)),
199        tokens,
200        "old messages left out"
201      );
202    }
203  }
204  history = mergeCallRuns(
205    history.filter((_, i) => !left.has(i)),
206    pinned
207  );
208  perEntry = history.map(entryTokens);
209  tokens = baseTokens + perEntry.reduce((sum, n) => sum + n, 0);
210  if (fits()) return fitted(history, tokens, "old calls merged");
211  throw new Error(
212    `history too large for Jev (~${tokens} tokens after truncation, limit ${options.maxStateTokens})`
213  );
214}
215export {
216  STATE_CONTEXT,
217  collectToolCalls,
218  estimateTokens,
219  fitState,
220  goalFromMessages,
221  isPinned,
222  truncate
223};
224
hooks/vendor/fast-jev/request.js 47 lines
1const SYSTEM_ONE_URL = "https://api.typesafe.ai/v1/systemone";
2const DEFAULT_MODEL = "jev-latest";
3function buildJevRequest(params, state, questions) {
4  return {
5    url: params.baseUrl ?? SYSTEM_ONE_URL,
6    method: "POST",
7    headers: {
8      authorization: `Bearer ${params.apiKey}`,
9      "content-type": "application/json"
10    },
11    body: JSON.stringify({
12      model: params.model ?? DEFAULT_MODEL,
13      state,
14      questions
15    })
16  };
17}
18function parseJevResponse(status, ok, text) {
19  if (!ok) {
20    throw new Error(`Jev request failed (${status}): ${text.slice(0, 200)}`);
21  }
22  let parsed;
23  try {
24    parsed = JSON.parse(text);
25  } catch {
26    throw new Error("Jev returned malformed JSON");
27  }
28  if (parsed === null || typeof parsed !== "object" || !("answers" in parsed) || parsed.answers === null || typeof parsed.answers !== "object") {
29    throw new Error("Jev response is missing answers");
30  }
31  return parsed;
32}
33function noulAnswer(answers, name) {
34  const answer = answers[name];
35  if (!answer || !("noul" in answer) || typeof answer.noul !== "number" || !Number.isFinite(answer.noul)) {
36    throw new Error(`Invalid Jev answer for ${name}`);
37  }
38  return answer.noul;
39}
40export {
41  DEFAULT_MODEL,
42  SYSTEM_ONE_URL,
43  buildJevRequest,
44  noulAnswer,
45  parseJevResponse
46};
47