Three-level decision layer on TypeSafe's Jev: picks model and effort per session and per subagent, compacts context without rewriting it, and gates code review…

A three-level decision layer on TypeSafe's Jev, a model that answers closed questions with calibrated probabilities. Jev never decides on its own: it answers yes/no, scores and choices, and the code turns those numbers into a decision through fixed thresholds.
| Level | What it decides | Claude Code | Pi |
|---|---|---|---|
| Router | model and effort, once per session (and per subagent on Claude Code) | hooks/router.ts | pi/extensions/jev-router/ |
| Compact | which tool calls to keep, truncate or drop when compacting, without rewriting anything | hooks/compact.ts | pi/extensions/jev-compact/ |
| Review | the verdict on a diff, computed in code from Jev's answers; the agent takes the handoff | skills/code-review/ | pi/skills/jev-code-review/ |
Thresholds and questions live once, in shared/ (router and compact) and review/ (review), and both runtimes import them: they cannot drift apart. With lexi, the review runs after the last GREEN of the bug and feature flows.
TYPESAFE_API_KEY: under env in the project's .claude/settings.local.json, one key per project. Claude Code merges that file's env itself; on Pi the router, compaction and review read the file directly. A shell export also works and wins. The file must be gitignored (git check-ignore .claude/settings.local.json): never put the key in a settings file that is committed. Without it every level stays off and says so once.Claude Code — jev is a dependency of lexi (since 0.4.0): installing lexi@lexi installs it. Older installs: /plugin install jev@lexi. Then run /lexi:init in the project.
init asks whether to enable Jev and, on yes, writes "jev": {} in .lexi.json (the review step) and, in the project's .claude/settings.json, CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 (the hooks module stays off without it) and CLAUDE_AUTOCOMPACT_PCT_OVERRIDE=50. It also writes two project shims, .claude/skills/code-review/SKILL.md and .claude/skills/review/SKILL.md, that hand off to jev:code-review: a project skill replaces the built-in skill of the same name, and /review needs its own file. Commit them so the team shares the review; delete them to get the built-in review back. The project folder must be trusted, or Claude Code reads neither the project env nor its skills.
The settings value alone is not enough yet. While Claude Code's plugin-hooks rollout is off server-side, the flag counts only in the environment of the claude process; the settings env is applied too late. Add export CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 to your shell profile (~/.zshrc, ~/.bashrc) and open a new session; every teammate needs it. The debug log shows hooks module jev@lexi not loaded when it is missing. review.py does not use hooks, so the review works either way.
Pi has no built-in code review, so there is nothing to replace: the skill is /skill:jev-code-review.
Pi — jev ships inside the lexi package. /skill:lexi-init asks whether to enable it, which model to use per tier, and writes a jev object in .lexi.json; without that object the extensions do nothing, even when the package is installed.
{ "jev": { "tiers": { "fast": "claude-sonnet-5-5", "balanced": "claude-opus-5-5", "deep": "claude-opus-5-5" } } }
tiers is optional (those are the defaults). A bare id resolves on the provider the session already runs on, to the newest model of that family there (claude-opus-5-5 → claude-bridge/claude-opus-5 on a bridge session); provider/id pins a tier to that provider exactly. It also sets compaction.modelOverrides[<model>].reserveTokens in .pi/settings.json to half each model's context window, so compaction runs at ~50%.
On the first prompt, Jev classifies the task: tier (trivial, fast, balanced, deep), effort (low … xhigh) and risk. The choice is applied once and then held for the whole session, so the prompt cache is never thrown away; a manual /model is never overridden. Going up needs confidence ≥ 0.3, going down ≥ 0.6; risk > 0.7 forces deep with effort at least xhigh; deep alone runs at least high, every other session at least medium (Sonnet 5.5's starting point for agentic coding: at low it can report a change done without checking it). If Jev does not answer (1.5 s timeout, error), the next prompt asks again.
On Claude Code every subagent spawn (forks excepted) is classified on its own: trivial → haiku, fast → sonnet, balanced → opus, deep → opus. Opus 5.5 beats Fable 5.1 at every cost point, so deep is Opus at a higher effort. A session is never put on haiku: it holds its model to the end, so trivial stays at fast. On Pi subagent models are pinned by the subagent extension and left alone.
You see [jev-router] session: sonnet · effort medium in the transcript and jev · … in the status line.
lexi flow. In a project with .lexi.json, every prompt that is not a slash command also gets a second question, bug, feature, open or none (shared/flow-policy.ts). At confidence ≥ 0.6 the answer rides along as hidden context ordering the matching skill (lexi:bug, lexi:feature, lexi:grill), with a way out to lexi:lexi when it is clearly wrong; below 0.6 it only points at lexi:lexi; none adds nothing. Unlike the model, the flow is asked on every prompt: each prompt can be a new task. On Pi the same answer arrives as a hidden message. A Jev failure leaves the prompt as typed; lexi's own one-line session hint still applies.
Each routed prompt shows one line, [jev-flow] bug (0.82) → lexi:bug (or → lexi:lexi below 0.6); none stays in the debug log.
Native compaction replaces the conversation with a summary, and paths, line numbers and exact errors can drift. Here Jev sees the structure of the conversation (without tool result contents) and, per tool call, decides whether to keep it, truncate its result to 300 characters, or drop it with its result. What stays is verbatim; user and assistant text is never touched. It falls back to the native summary when the key is missing, Jev does not answer within 8 s, the answer is invalid, there is no candidate call, or the reduction is under 30%: below that, rewriting the prompt cache does not pay off.
On Pi compaction can only return text, so the kept history is written back verbatim as text ([User]: …, [Assistant tool calls]: …, [Tool result <tool>]: …) under a header saying it is not a summary.
python3 jev/review/review.py --working # uncommitted changes, new files included
python3 jev/review/review.py --git $(git merge-base HEAD main) # the branch's commits
python3 jev/review/review.py --diff pr.diff --title "..." --description "..."
python3 jev/review/review.py ... --json # for CI and the skills
python3 jev/review/review.py ... --escalate # the handoff as a ready prompt
python3 jev/review/review.py ... --compare before.json # after a fix: per-check deltas
Exit codes: 0 MERGE, 1 NITS/CONVENTIONS/QUALITY ("fix before merge, no risk outside the codebase"), 2 SECURITY REVIEW, 3 BLOCK, 4 error. A diff over Jev's request ceiling (max_request_tokens in policy.json, ~40K tokens with the questions) is split into parts of whole files, one parallel call each (~2 s, ~$0.002 a call), and the answers merged per check: the worst part wins, except checks marked "aggregate": "min" (docs_only, outside_test_perimeter) that must hold for every part. Files are ordered by the most severe lane their critical checks feed (BLOCK first: tests, secret-looking files), then the rest; lockfiles, generated code, docs and agent tooling (drop_first_patterns) go last. Past max_parts (32) the rest is omitted, the JSON says so (omitted_files, suspended_checks) and the rules on higher_is_better checks (adds_tests, description_matches) are suspended: absence is not provable on a partial diff. Every merged number remembers the part that produced it: the handoff's files for that check come from that part only. From an agent: /jev:code-review (Claude Code; /code-review and /review too in a project init set up) or /skill:jev-code-review (Pi).
| File | Holds |
|---|---|
review/checks.json | the core questions: generic risk (secrets, injection, auth, weakened tests, API breaks, migrations…) and code quality (nesting, length, parameters, naming, duplication, over-engineering…) |
review/policy.json | the lanes BLOCK → SECURITY REVIEW → CONVENTIONS → QUALITY → NITS → MERGE, one threshold per check, the uncertainty band, state limits |
review/review.py | the CLI: builds the state, calls Jev, applies the policy. No question, threshold or check name lives here |
Jev says that there is a problem and how likely, not where. The JSON carries a handoff: every fired rule to locate (path:line and the minimal fix), every critical check in the 0.35–0.65 band to verify, every check with escalate_to to delegate. After the fix, --compare reruns Jev: the fix is judged by the numbers, not by opinion.
The state also carries .lexi.json → testable as testable_paths, so outside_test_perimeter knows where a unit test is expected.
.lexi/review.jsonThe CONVENTIONS lane ships empty. A project fills it:
{
"checks": {
"layer_bypass": {
"label": "Layer bypass",
"critical": false,
"higher_is_better": false,
"escalation_patterns": ["components/", "pages/"],
"type": "noul",
"instructions": { "question": "Does a UI file in `diff` call the HTTP client directly instead of going through a repository?" },
"criteria": { "true": "…boundary cases that count…", "false": "…boundary cases that do not…" }
}
},
"rules": { "CONVENTIONS": [{ "check": "layer_bypass", "op": "gte", "value": 0.7 }] },
"drop_first_patterns": ["^fixtures/"]
}
checks uses the checks.json format; a project check with a core id overrides it for that project. rules maps a lane name to rules appended to it. An unknown lane, a rule on a check that does not exist or an invalid regex in escalation_patterns or drop_first_patterns is an error (exit 4), never a silent no-op. escalate_to: "agent:<name>" on a critical check hands it to that agent.
A verdict is wrong when a probability is wrong. Fix it in the criteria of the question that answered badly: add the boundary case, with an example, to the true or false side that should absorb it. Never move a threshold in policy.json to let one case through, and never touch review.py: the threshold is the cost you accept for that error, not a knob to tune on the last diff.
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude plugin test jev # router, flow, compact and shared/flow-policy, Claude Code
claude plugin validate jev
CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 claude -p "/plugin-types jev/types" && npx -y -p typescript@5 tsc -p jev/tsconfig.json
node --test pi/extensions/jev-compact/transcript.test.ts
npx -y -p typescript@5 tsc -p pi/tsconfig.json # Pi extensions (paths assume a brew install of pi)
python3 jev/review/test_review.py
A Claude Code plugin gets exactly one hooks module, and register must call on("<event>", hook) directly: hooks/index.ts wires both levels, router.ts and compact.ts export only handlers.
Claude Code: /plugin uninstall jev@lexi, or drop CLAUDE_CODE_ENABLE_FUNCTION_HOOKS from the project settings and your shell. Pi: remove the jev object from .lexi.json. Either runtime, one machine only: remove TYPESAFE_API_KEY from .claude/settings.local.json.
hooks/index.ts 14 lines1import type { Register } from 'claude-code'
2import { compact } from './compact.ts'
3import { onAgentSpawn, onPromptSubmit, onSessionEnd, onSessionStart, onTurnStep } from './router.ts'
4
5// A plugin gets exactly one hooks module: this one wires both levels.
6export const register: Register = (on) => {
7 on('session.start', onSessionStart)
8 on('session.end', onSessionEnd)
9 on('prompt.submit', onPromptSubmit)
10 on('turn.step', onTurnStep)
11 on('agent.spawn', onAgentSpawn)
12 on('session.compact', compact)
13}
14hooks/compact.ts 190 lines1import type { EngineInterface, Hook, SessionCompactInput, SessionCompactResult, SessionMessage, ToolResultSummary, ToolUseSummary } from 'claude-code'
2import type { CallInfo, JevState as State, Question, StateCall as Call, StateEntry as Entry, Verdict } from '../shared/compact-policy.ts'
3import {
4 bar,
5 describeCall,
6 fitState,
7 isOpen,
8 JEV_MODEL,
9 JEV_TIMEOUT_MS,
10 JEV_URL,
11 kilo,
12 MAX_STATE_TOKENS,
13 MIN_REDUCTION,
14 parseAnswers,
15 percent,
16 toBatches,
17 truncated,
18 verdictOf,
19} from '../shared/compact-policy.ts'
20
21// Claude Code side only: thresholds and questions live in shared/compact-policy.ts, shared with Pi.
22const RESULT_VISIBLE_MS = 30_000 // how long the outcome stays in the status line
23
24type Next = Parameters<Hook<'session.compact'>>[2]
25type Candidate = { use: ToolUseSummary; result: ToolResultSummary }
26type Plan = { messages: SessionMessage[]; candidates: Candidate[]; verdicts: Map<string, Verdict>; before: number; after: number; count: number }
27
28export async function compact($: EngineInterface, e: SessionCompactInput, next: Next): Promise<SessionCompactResult> {
29 const startedAt = await $.clock.now()
30 const plan = await planCompaction($, e, next.signal).catch((error: unknown) => `internal error: ${String(error)}`)
31 if (typeof plan === 'string') return delegate($, e, next, plan)
32 report($, plan, (await $.clock.now()) - startedAt)
33 return { messages: plan.messages }
34}
35
36// Every reason not to compact becomes a string: the caller defers to the native summary.
37async function planCompaction($: EngineInterface, e: SessionCompactInput, signal: AbortSignal): Promise<Plan | string> {
38 const key = await $.env.get('TYPESAFE_API_KEY')
39 if (!key) return 'TYPESAFE_API_KEY missing'
40 progress($, 10, 'finding candidates')
41 const candidates = findCandidates(e.messages)
42 if (!candidates.length) return 'no candidate tool call'
43 const state = fitState(buildState(e), e.messages.length)
44 if (!state) return `state over ${MAX_STATE_TOKENS} tokens`
45 const answers = await askJev($, { key, state, calls: candidates.map(infoOf) }, signal)
46 if (typeof answers === 'string') return answers
47 progress($, 90, 'rebuilding')
48 const verdicts = new Map(candidates.map((candidate) => [candidate.use.tool_use_id, verdictOf(answers, infoOf(candidate))]))
49 const messages = rebuild(e.messages, verdicts)
50 const before = charsOf(e.messages)
51 const after = charsOf(messages)
52 if (1 - after / before < MIN_REDUCTION) return `reduction ${percent(1 - after / before)} below ${percent(MIN_REDUCTION)}`
53 return { messages, candidates, verdicts, before, after, count: e.messages.length }
54}
55
56const infoOf = ({ use, result }: Candidate): CallInfo => ({ id: use.tool_use_id, tool: use.tool, resultChars: result.text.length })
57
58// Candidate: a call with its result, both outside the immune messages.
59const findCandidates = (messages: readonly SessionMessage[]): Candidate[] => {
60 const open = messages.filter((_, i) => isOpen(i, messages.length))
61 const results = new Map(open.flatMap((m) => m.toolResults ?? []).map((r) => [r.tool_use_id, r]))
62 return open.flatMap((m) => (m.role === 'assistant' ? m.toolUses : [])).flatMap((use) => {
63 const result = results.get(use.tool_use_id)
64 return result ? [{ use, result }] : []
65 })
66}
67
68// Structure and history without result contents: only a note with outcome and length.
69const buildState = ({ messages, instructions }: SessionCompactInput): State => {
70 const results = new Map(messages.flatMap((m) => m.toolResults ?? []).map((r) => [r.tool_use_id, r]))
71 const note = (id: string) => {
72 const r = results.get(id)
73 return r ? `${r.isError ? 'error' : 'ok'}, ${r.text.length} chars` : 'no result'
74 }
75 const conversation = messages.flatMap((m, i): Entry[] => {
76 const calls = m.toolUses.map((u) => ({ id: u.tool_use_id, tool: u.tool, input: u.input, result: note(u.tool_use_id) }))
77 if (!m.text && !calls.length) return []
78 return [{ i, role: m.role, text: m.text, ...(calls.length ? { tool_calls: calls } : {}) }]
79 })
80 const goal = messages.filter((m) => m.role === 'user' && m.text).slice(-3).map((m) => m.text)
81 return { goal, ...(instructions ? { instructions } : {}), conversation }
82}
83
84// Batches run in parallel under one timeout: an in-flight fetch does not consume the hook budget,
85// waiting on the clock does, so the timeout stays under 10 s.
86async function askJev($: EngineInterface, ask: { key: string; state: State; calls: CallInfo[] }, signal: AbortSignal): Promise<Record<string, number> | string> {
87 const batches = toBatches(ask.state, ask.calls)
88 const timer = new AbortController()
89 signal.addEventListener('abort', () => timer.abort())
90 const timeout = $.clock.sleep(JEV_TIMEOUT_MS, { signal: timer.signal }).then(() => 'timeout', () => 'timeout')
91 const status = { done: 0, total: batches.length, calls: ask.calls.length }
92 showBatches($, status)
93 const requests = Promise.all(batches.map((questions) => askBatch($, { ...ask, questions }, status)))
94 const outcome = await Promise.race([requests, timeout])
95 timer.abort()
96 if (typeof outcome === 'string') return outcome
97 return outcome.find((answer): answer is string => typeof answer === 'string') ?? Object.assign({}, ...outcome)
98}
99
100async function askBatch($: EngineInterface, ask: { key: string; state: State; questions: Record<string, Question> }, status: { done: number; total: number; calls: number }): Promise<Record<string, number> | string> {
101 const response = await $.http
102 .fetch(JEV_URL, {
103 method: 'POST',
104 headers: { authorization: `Bearer ${ask.key}`, 'content-type': 'application/json' },
105 body: JSON.stringify({ model: JEV_MODEL, state: ask.state, questions: ask.questions }),
106 })
107 .catch((error: unknown) => `network error: ${String(error)}`)
108 status.done++
109 showBatches($, status)
110 if (typeof response === 'string') return response
111 return response.ok ? parseAnswers(response.text, Object.keys(ask.questions)) : `HTTP ${response.status}`
112}
113
114// Call and result share one verdict: never a tool_use without its tool_result.
115const reshape = <T extends { tool_use_id: string; text?: string; result?: unknown }>(blocks: readonly T[], verdicts: Map<string, Verdict>): T[] =>
116 blocks.flatMap((block) => {
117 const verdict = verdicts.get(block.tool_use_id)
118 if (verdict === 'drop') return []
119 if (verdict !== 'truncate' || block.text === undefined) return [block]
120 const { result: _stored, ...rest } = block
121 return [{ ...rest, text: truncated(block.text) } as T]
122 })
123
124// Without a handle the engine rebuilds each message from role, text and blocks; with the handle it
125// would ignore the edits and reload the old chain on --resume.
126const rebuild = (messages: readonly SessionMessage[], verdicts: Map<string, Verdict>): SessionMessage[] =>
127 joinParallelCalls(
128 messages.flatMap(({ handle: _handle, ...message }) => {
129 const toolUses = reshape(message.toolUses, verdicts)
130 const toolResults = message.toolResults && reshape(message.toolResults, verdicts)
131 const rebuilt: SessionMessage = { ...message, toolUses, ...(toolResults ? { toolResults } : {}) }
132 return rebuilt.text || toolUses.length || toolResults?.length ? [rebuilt] : []
133 }),
134 )
135
136// Parallel calls of one response arrive as consecutive assistant messages, one per block. Rebuilt,
137// they would get new ids and on --resume the loader would no longer rejoin them
138// (ensureToolResultPairing): they go back to one message, like the original response.
139const joinParallelCalls = (messages: SessionMessage[]): SessionMessage[] => {
140 const joined: SessionMessage[] = []
141 for (const message of messages) {
142 const previous = joined[joined.length - 1]
143 if (message.role !== 'assistant' || previous?.role !== 'assistant' || !previous.toolUses.length) joined.push(message)
144 else joined[joined.length - 1] = { ...previous, text: [previous.text, message.text].filter(Boolean).join('\n\n'), toolUses: [...previous.toolUses, ...message.toolUses] }
145 }
146 return joined
147}
148
149const charsOf = (messages: readonly SessionMessage[]) =>
150 messages.reduce((sum, m) => sum + m.text.length + m.toolUses.reduce((s, u) => s + JSON.stringify(u.input).length, 0) + (m.toolResults ?? []).reduce((s, r) => s + r.text.length, 0), 0)
151
152function progress($: EngineInterface, pct: number, phase: string): void {
153 $.ui.status(`jev-compact ${bar(pct)} ${pct}% · ${phase}`)
154}
155
156function showBatches($: EngineInterface, { done, total, calls }: { done: number; total: number; calls: number }): void {
157 progress($, 20 + Math.round((60 * done) / total), `asking Jev (${done}/${total} batches, ${calls} calls)`)
158}
159
160// The outcome stays readable for RESULT_VISIBLE_MS: Jev answers in a second, the bar alone would vanish.
161function settle($: EngineInterface, outcome: string): void {
162 $.ui.status(`jev-compact ${outcome}`)
163 $.ui.toast(`jev-compact ${outcome}`, { timeoutMs: 8_000 })
164 $.clock.after(RESULT_VISIBLE_MS, () => $.ui.status(undefined))
165}
166
167function report($: EngineInterface, plan: Plan, ms: number): void {
168 const counts = { keep: 0, truncate: 0, drop: 0 }
169 for (const verdict of plan.verdicts.values()) counts[verdict]++
170 const saved = plan.before - plan.after
171 const reduction = percent(saved / plan.before)
172 settle($, `${bar(100)} 100% · -${reduction} (${kilo(saved)} chars) · ${counts.truncate} truncated, ${counts.drop} dropped, ${counts.keep} kept`)
173 $.ui.log(`Jev compacted the context: -${reduction} (${kilo(saved)} of ${kilo(plan.before)} chars) in ${ms}ms`)
174 $.ui.log(`${plan.verdicts.size} tool calls judged → ${counts.keep} kept, ${counts.truncate} truncated, ${counts.drop} dropped · messages ${plan.count} → ${plan.messages.length}`)
175 for (const candidate of plan.candidates) {
176 const verdict = plan.verdicts.get(candidate.use.tool_use_id) ?? 'keep'
177 if (verdict !== 'keep') $.ui.log(describeCall(infoOf(candidate), verdict, candidate.use.input))
178 }
179}
180
181async function delegate($: EngineInterface, e: SessionCompactInput, next: Next, reason: string): Promise<SessionCompactResult> {
182 $.ui.log(`Jev did not compact (${reason}): native Claude Code summary`)
183 progress($, 90, 'native summary')
184 try {
185 return await next(e)
186 } finally {
187 settle($, `Jev not used (${reason}) → native summary`)
188 }
189}
190hooks/router.ts 170 lines1import type { EngineInterface, EventName, Hook } from 'claude-code'
2import { flowContext, flowSkill, FLOW_QUESTIONS, isRoutable, parseFlow } from '../shared/flow-policy.ts'
3import type { Answer, Effort } from '../shared/router-policy.ts'
4import {
5 aliasOfRank,
6 causeOf,
7 EFFORTS,
8 effortRankOf,
9 isMoveAllowed,
10 isRisky,
11 JEV_MODEL,
12 JEV_URL,
13 parseAnswer,
14 QUESTIONS,
15 sessionTierRank,
16 shortModel,
17 targetEffortRank,
18 targetTierRank,
19 TIER_ALIAS,
20 tierRankOfModel,
21 TIMEOUT_MS,
22} from '../shared/router-policy.ts'
23
24// Claude Code side only: thresholds and questions live in shared/router-policy.ts, shared with Pi.
25const SHOW_STATUS = true
26const MODEL_ID: Record<string, string> = { sonnet: 'claude-sonnet-5-5', opus: 'claude-opus-5-5' }
27const DEBUG = { to: 'debug' } as const
28
29type HookArgs<N extends EventName> = Parameters<Hook<N>>
30type Plan = { fromModel: string; fromEffort: unknown; model?: string; effort?: Effort }
31
32let apiKey: string | undefined | null = null // null: not read yet
33let sessionAnswer: Answer | undefined
34let sessionPlan: Plan | undefined
35let hasWarnedUnavailable = false
36
37// session.start does not fire on /clear (session.end with reason 'clear' does): reset on both.
38export async function onSessionStart($: EngineInterface, e: HookArgs<'session.start'>[1], next: HookArgs<'session.start'>[2]) {
39 resetSession($)
40 return next(e)
41}
42
43export async function onSessionEnd($: EngineInterface, e: HookArgs<'session.end'>[1], next: HookArgs<'session.end'>[2]) {
44 resetSession($)
45 return next(e)
46}
47
48function resetSession($: EngineInterface) {
49 sessionAnswer = undefined
50 sessionPlan = undefined
51 if (SHOW_STATUS) $.ui.status(undefined)
52}
53
54// One hook per event: the model is chosen once per session, the lexi flow on every prompt.
55export async function onPromptSubmit($: EngineInterface, e: HookArgs<'prompt.submit'>[1], next: HookArgs<'prompt.submit'>[2]) {
56 if (sessionAnswer) $.ui.log('[jev-router] reusing the first prompt\'s choice', DEBUG)
57 // ponytail: sequential, so the first prompt can wait up to two Jev timeouts; run both at once if it shows.
58 else sessionAnswer = await classify($, 'prompt', { prompt: e.text })
59 const context = await routeFlow($, e.text)
60 return next(context ? { ...e, context: [...(e.context ?? []), context] } : e)
61}
62
63// In a lexi project (`.lexi.json` is the opt-in), Jev names the flow and its order rides along as context.
64async function routeFlow($: EngineInterface, prompt: string): Promise<string | undefined> {
65 if (!isRoutable(prompt) || !(await $.fs.read('.lexi.json').then(() => true, () => false))) return undefined
66 const key = await readApiKey($)
67 if (!key) return undefined
68 const outcome = await postJev($, { model: JEV_MODEL, state: { prompt }, questions: FLOW_QUESTIONS }, key)
69 const answer = typeof outcome === 'string' ? outcome : parseFlow(outcome.text)
70 if (typeof answer === 'string') {
71 $.ui.log(`[jev-flow] jev unavailable: ${answer}, no routing for this prompt`, DEBUG)
72 return undefined
73 }
74 // Shown like the router's line, so the person sees which skill Jev sent the prompt to; `none` stays in debug.
75 const skill = flowSkill(answer.flow, answer.confidence)
76 $.ui.log(`[jev-flow] ${answer.flow} (${answer.confidence.toFixed(2)})${skill ? ` → ${skill}` : ''}`, skill ? {} : DEBUG)
77 return flowContext(answer.flow, answer.confidence)
78}
79
80export async function* onTurnStep($: EngineInterface, e: HookArgs<'turn.step'>[1], next: HookArgs<'turn.step'>[2]) {
81 if (e.agentId !== undefined || !sessionAnswer) return yield* next(e)
82 sessionPlan ??= startSessionPlan($, sessionAnswer, e.model, e.effort)
83 const isModelOurs = sessionPlan.model && e.model === sessionPlan.fromModel
84 const isEffortOurs = sessionPlan.effort && e.effort === sessionPlan.fromEffort
85 return yield* next({
86 ...e,
87 ...(isModelOurs ? { model: sessionPlan.model } : {}),
88 ...(isEffortOurs ? { effort: sessionPlan.effort } : {}),
89 })
90}
91
92export async function onAgentSpawn($: EngineInterface, e: HookArgs<'agent.spawn'>[1], next: HookArgs<'agent.spawn'>[2]) {
93 if (e.fork) return next(e)
94 const label = `subagent ${e.subagentType}`
95 const answer = await classify($, label, { prompt: e.prompt, description: e.description, agentType: e.subagentType })
96 if (!answer) return next(e)
97 const current = e.model ?? e.parentModel
98 const rank = targetTierRank(answer)
99 const alias = isMoveAllowed(tierRankOfModel(current), rank, answer.tierConfidence, isRisky(answer)) ? aliasOfRank(rank) : undefined
100 $.ui.log(`[jev-router] ${label}: ${alias ? `model ${shortModel(current)} → ${alias}` : 'unchanged'} (${causeOf(answer)})`, DEBUG)
101 const started = await next(alias ? { ...e, model: alias } : e)
102 if (started.model) $.ui.log(`[jev-router] ${label}: ${shortModel(started.model)}`)
103 return started
104}
105
106function startSessionPlan($: EngineInterface, answer: Answer, model: string, effort: unknown): Plan {
107 const tierRank = sessionTierRank(answer)
108 const effortRank = targetEffortRank(answer)
109 const moveModel = isMoveAllowed(tierRankOfModel(model), tierRank, answer.tierConfidence, isRisky(answer))
110 const moveEffort = isMoveAllowed(effortRankOf(effort), effortRank, answer.effortConfidence, isRisky(answer))
111 const plan: Plan = { fromModel: model, fromEffort: effort }
112 if (moveModel) plan.model = MODEL_ID[aliasOfRank(tierRank)] ?? aliasOfRank(tierRank)
113 if (moveEffort) plan.effort = EFFORTS[effortRank]
114 const changes = [
115 plan.model && `model ${shortModel(model)} → ${shortModel(plan.model)}`,
116 plan.effort && `effort ${String(effort ?? 'default')} → ${plan.effort}`,
117 ].filter(Boolean)
118 const summary = `${shortModel(plan.model ?? model)} · effort ${String(plan.effort ?? effort ?? 'default')}`
119 $.ui.log(`[jev-router] ${changes.length ? changes.join(', ') : `unchanged — staying on ${summary}`} (${causeOf(answer)})`, DEBUG)
120 $.ui.log(`[jev-router] session: ${summary}`)
121 if (SHOW_STATUS) $.ui.status(`jev · ${summary}`)
122 return plan
123}
124
125async function classify($: EngineInterface, label: string, state: Record<string, string>): Promise<Answer | undefined> {
126 const key = await readApiKey($)
127 if (!key) return undefined
128 const result = await askJev($, state, key)
129 if (typeof result === 'string') {
130 // First failure goes to the transcript (bad key, blocked network), later ones to debug only.
131 $.ui.log(`[jev-router] jev unavailable (${label}): ${result}, retrying next time`, hasWarnedUnavailable ? DEBUG : {})
132 hasWarnedUnavailable = true
133 return undefined
134 }
135 const effortName = EFFORTS[result.effort] ?? result.effort
136 $.ui.log(`[jev-router] classification (${label}): ${result.tier} → ${TIER_ALIAS[result.tier]} (${result.tierConfidence.toFixed(2)}) · effort ${effortName} (${result.effortConfidence.toFixed(2)}) · risk ${result.risk.toFixed(2)} · ${result.ms}ms`, DEBUG)
137 return result
138}
139
140async function readApiKey($: EngineInterface): Promise<string | undefined> {
141 if (apiKey !== null) return apiKey
142 apiKey = await $.env.get('TYPESAFE_API_KEY')
143 if (apiKey) $.ui.log(`[jev-router] ready — ${JEV_URL}, timeout ${TIMEOUT_MS}ms`, DEBUG)
144 else $.ui.log('[jev-router] TYPESAFE_API_KEY missing: add it under "env" in .claude/settings.local.json or export it in the shell. Router off.')
145 return apiKey
146}
147
148async function askJev($: EngineInterface, state: Record<string, string>, key: string): Promise<Answer | string> {
149 const outcome = await postJev($, { model: JEV_MODEL, state, questions: QUESTIONS }, key)
150 return typeof outcome === 'string' ? outcome : parseAnswer(outcome.text, outcome.ms)
151}
152
153// Any failure (timeout, non-2xx, exception) becomes a reason string, never a throw.
154async function postJev($: EngineInterface, body: object, key: string): Promise<{ text: string; ms: number } | string> {
155 const startedAt = await $.clock.now()
156 const timer = new AbortController()
157 const timeout = $.clock.sleep(TIMEOUT_MS, { signal: timer.signal }).then(() => 'timeout', () => 'timeout')
158 const request = $.http
159 .fetch(JEV_URL, {
160 method: 'POST',
161 headers: { authorization: `Bearer ${key}`, 'content-type': 'application/json' },
162 body: JSON.stringify(body),
163 })
164 .then((response) => (response.ok ? response : `HTTP ${response.status}`), (error: unknown) => `error ${String(error)}`)
165 const outcome = await Promise.race([request, timeout])
166 timer.abort()
167 if (typeof outcome === 'string') return outcome
168 return { text: outcome.text, ms: (await $.clock.now()) - startedAt }
169}
170shared/compact-policy.ts 110 lines1// Compaction policy: questions to Jev, thresholds, verdict and formatting. No runtime dependency, so
2// the Claude Code hook (`jev/hooks/compact.ts`) and the Pi extension (`pi/extensions/jev-compact/`)
3// share the same numbers.
4
5export const KEEP_THRESHOLD = 0.5 // minimum probability to keep
6export const PRESERVE_RECENT = 6 // last messages never touched (the first is immune too)
7export const TRUNCATE_HEAD = 300 // what is left of a truncated result
8export const MAX_STATE_TOKENS = 25_000 // cap on the state sent to Jev
9export const MAX_REQUEST_TOKENS = 30_000 // cap on one request (state + questions)
10export const MIN_REDUCTION = 0.3 // below this, defer to native (rewriting the cache does not pay off)
11export const JEV_TIMEOUT_MS = 8_000
12export const JEV_MODEL = 'jev-latest'
13export const JEV_URL = 'https://api.typesafe.ai/v1/systemone'
14export const ARG_KEYS = ['file_path', 'command', 'pattern', 'url', 'query', 'path']
15
16export type Verdict = 'keep' | 'truncate' | 'drop'
17export type Question = { type: 'noul'; instructions: string; criteria: { true: string; false: string } }
18// The minimum needed to ask the questions and read the verdict: each runtime maps its own tool
19// call representation onto it.
20export type CallInfo = { id: string; tool: string; resultChars: number }
21
22export const isOpen = (i: number, total: number) => i > 0 && i < total - PRESERVE_RECENT
23
24// The state Jev reads: structure and history without tool result contents. Each runtime builds it
25// from its own messages; the JSON shape is the same.
26export type StateCall = { id: string; tool: string; input?: unknown; result: string }
27export type StateEntry = { i: number; role: string; text: string; tool_calls?: StateCall[] }
28export type JevState = { goal: string[]; instructions?: string; conversation: StateEntry[] }
29
30// Shrinks in steps until it fits: call inputs, then long texts, then whole messages, always from
31// the oldest unprotected ones. undefined if that is not enough.
32export const fitState = (state: JevState, total: number): JevState | undefined => {
33 const open = state.conversation.filter((entry) => isOpen(entry.i, total))
34 const lengths = new Map(open.map((entry) => [entry, entry.text.length]))
35 const clipInputs = () => state.conversation.forEach((entry) => entry.tool_calls?.forEach((c) => (c.input = clip(JSON.stringify(c.input), 200))))
36 const shorten = (entry: StateEntry) => () => (entry.text = headTail(entry.text))
37 const collapse = (entry: StateEntry) => () => {
38 entry.text = `[… ${lengths.get(entry)} chars]`
39 entry.tool_calls?.forEach((c) => (c.input = undefined))
40 }
41 for (const step of [() => {}, clipInputs, ...open.map(shorten), ...open.map(collapse)]) {
42 step()
43 if (tokensOf(state) <= MAX_STATE_TOKENS) return state
44 }
45 return undefined
46}
47
48export const questionsFor = ({ id, tool, resultChars }: CallInfo): Record<string, Question> => ({
49 [`keep_call_${id}`]: {
50 type: 'noul',
51 instructions: `Tool call \`${id}\` (\`${tool}\`) must stay in the history: knowing it was made, with its input, still matters for what the assistant does next.`,
52 criteria: { true: 'The assistant will still refer to this call or its input.', false: 'Step superseded: exploration finished, empty search, file later rewritten, nothing that steers the work.' },
53 },
54 [`keep_result_${id}`]: {
55 type: 'noul',
56 instructions: `The full output of tool call \`${id}\` (\`${tool}\`, ${resultChars} characters) must stay verbatim: the assistant still needs its content and re-running the tool would not do.`,
57 criteria: { true: 'Needed word for word and not recoverable by re-running the tool (exact error, non-repeatable output, changed state).', false: 'Superseded or recoverable by re-running the tool: file later edited or re-readable, listing, search, output already summarised.' },
58 },
59})
60
61// Batches under MAX_REQUEST_TOKENS: each one resends the full state.
62export const toBatches = (state: unknown, calls: readonly CallInfo[]): Record<string, Question>[] => {
63 const room = MAX_REQUEST_TOKENS - tokensOf(state)
64 const batches: Record<string, Question>[] = [{}]
65 let used = 0
66 for (const call of calls) {
67 const questions = questionsFor(call)
68 if (used > 0 && used + tokensOf(questions) > room) (batches.push({}), (used = 0))
69 Object.assign(batches[batches.length - 1] ?? {}, questions)
70 used += tokensOf(questions)
71 }
72 return batches
73}
74
75export const parseAnswers = (text: string, ids: string[]): Record<string, number> | string => {
76 let answers: Record<string, { noul?: unknown } | undefined> | undefined
77 try {
78 answers = (JSON.parse(text) as { answers?: typeof answers } | null)?.answers
79 } catch {
80 return 'answer is not JSON'
81 }
82 const entries = ids.flatMap((id) => {
83 const noul = answers?.[id]?.noul
84 return typeof noul === 'number' && Number.isFinite(noul) ? [[id, noul] as const] : []
85 })
86 return entries.length === ids.length ? Object.fromEntries(entries) : 'answer without valid noul'
87}
88
89export const verdictOf = (answers: Record<string, number>, { id, resultChars }: CallInfo): Verdict => {
90 if ((answers[`keep_result_${id}`] ?? 1) >= KEEP_THRESHOLD) return 'keep'
91 if ((answers[`keep_call_${id}`] ?? 1) < KEEP_THRESHOLD) return 'drop'
92 return resultChars > TRUNCATE_HEAD ? 'truncate' : 'keep'
93}
94
95export const truncated = (text: string) => `${text.slice(0, TRUNCATE_HEAD)}\n[… ${text.length - TRUNCATE_HEAD} chars removed]`
96
97export const tokensOf = (value: unknown) => JSON.stringify(value).length / 4
98export const clip = (text: string, max: number) => (text.length > max ? `${text.slice(0, max - 1)}…` : text)
99export const headTail = (text: string) => (text.length > 800 ? `${text.slice(0, 400)} […] ${text.slice(-400)}` : text)
100export const percent = (ratio: number) => `${Math.round(ratio * 100)}%`
101export const kilo = (chars: number) => (chars >= 1000 ? `${(chars / 1000).toFixed(1)}k` : `${chars}`)
102export const bar = (pct: number) => '█'.repeat(Math.round(pct / 10)) + '░'.repeat(10 - Math.round(pct / 10))
103export const argumentOf = (input: Record<string, unknown>) =>
104 clip((ARG_KEYS.map((key) => input[key]).find((v): v is string => typeof v === 'string') ?? JSON.stringify(input)).replace(/\s+/g, ' ').trim(), 70)
105
106export const describeCall = ({ tool, resultChars }: CallInfo, verdict: Verdict, input: Record<string, unknown>) =>
107 verdict === 'drop'
108 ? `✗ dropped ${tool} · ${argumentOf(input)} (-${kilo(resultChars + JSON.stringify(input).length)} chars)`
109 : `✂ truncated ${tool} · ${argumentOf(input)} (${kilo(resultChars)} → ${TRUNCATE_HEAD} chars)`
110shared/flow-policy.ts 57 lines1// Flow policy: which lexi skill a prompt goes to. No runtime dependency, so the Claude Code hook
2// (`jev/hooks/flow.ts`) and the Pi extension (`pi/extensions/jev-flow.ts`) share it.
3
4export const FLOWS = ['bug', 'feature', 'open', 'none'] as const
5export type Flow = (typeof FLOWS)[number]
6export type FlowAnswer = { flow: Flow; confidence: number }
7
8export const FLOW_QUESTIONS = {
9 flow: {
10 type: 'choice',
11 instructions: 'In a project that writes a failing test before any code change, which flow fits this prompt?',
12 criteria: {
13 bug: 'Existing behaviour is broken: an error, a crash, a wrong result, a failing test, a regression the user reports.',
14 feature: 'A new or changed behaviour whose design the prompt already settles: what to build is stated, only how remains.',
15 open: 'A new or changed behaviour with several defensible designs or unstated product intent: decisions come before tests can be named.',
16 none: 'No change to production behaviour: a question, an explanation, docs, config, a review, a pure refactor, or a reply that continues the current task ("ok", "go on").',
17 },
18 },
19}
20type JevResponse = { answers?: { flow?: { choice?: string; confidence?: number } } }
21
22export const parseFlow = (text: string): FlowAnswer | string => {
23 let body: JevResponse
24 try {
25 body = JSON.parse(text) as JevResponse
26 } catch {
27 return 'unreadable JSON'
28 }
29 const answer = body.answers?.flow
30 const flow = FLOWS.find((name) => name === answer?.choice)
31 if (!flow) return 'answer without a valid flow'
32 return { flow, confidence: answer?.confidence ?? 0 }
33}
34
35export const MIN_CONFIDENCE = 0.6
36const SKILL_OF: Record<Exclude<Flow, 'none'>, string> = { bug: 'bug', feature: 'feature', open: 'grill' }
37
38// `skillOf` names a skill for the runtime: `lexi:bug` on Claude Code, `lexi-bug` on Pi.
39type SkillOf = (name: string) => string
40const LEXI_SKILL: SkillOf = (name) => `lexi:${name}`
41
42// The skill the flow sends a prompt to: its own when confident, lexi's router when not, none for `none`.
43export const flowSkill = (flow: Flow, confidence: number, skillOf = LEXI_SKILL): string | undefined => {
44 if (flow === 'none') return undefined
45 return skillOf(confidence < MIN_CONFIDENCE ? 'lexi' : SKILL_OF[flow])
46}
47
48export const flowContext = (flow: Flow, confidence: number, skillOf = LEXI_SKILL): string | undefined => {
49 const skill = flowSkill(flow, confidence, skillOf)
50 if (!skill) return undefined
51 if (confidence < MIN_CONFIDENCE) return `lexi flow (Jev): unsure (${flow}, ${confidence.toFixed(2)}). If this prompt asks for a code change, load the skill "${skill}" first: it routes.`
52 return `lexi flow (Jev): ${flow}, confidence ${confidence.toFixed(2)}. Load the skill "${skill}" before anything else. If that is clearly wrong for this prompt, say why in one line and load "${skillOf('lexi')}" instead.`
53}
54
55// A slash command already names what runs: classifying it would only second-guess the user.
56export const isRoutable = (prompt: string): boolean => !prompt.trimStart().startsWith('/')
57shared/router-policy.ts 84 lines1// Router policy: questions to Jev, thresholds and decision. No runtime dependency, so the Claude
2// Code hook (`jev/hooks/router.ts`) and the Pi extension (`pi/extensions/jev-router/`) share it.
3
4export const JEV_URL = 'https://api.typesafe.ai/v1/systemone'
5export const JEV_MODEL = 'jev-latest'
6export const TIMEOUT_MS = 1500
7export const UP_CONFIDENCE = 0.3
8export const DOWN_CONFIDENCE = 0.6
9export const RISK_THRESHOLD = 0.7
10export const TIERS = ['trivial', 'fast', 'balanced', 'deep'] as const
11// Opus 5.5 beats Fable 5.1 at every cost point, so `deep` is Opus at a higher effort.
12export const TIER_ALIAS: Record<Tier, string> = { trivial: 'haiku', fast: 'sonnet', balanced: 'opus', deep: 'opus' }
13export const EFFORTS = ['low', 'medium', 'high', 'xhigh', 'max'] as const
14export const MIN_EFFORT = 1 // medium: Sonnet 5.5's start for agentic coding; at low it can call a change done unchecked
15export const DEEP_EFFORT = 2 // high
16export const RISKY_EFFORT = 3 // xhigh: max scores lower and costs more on Opus 5.5
17
18export const QUESTIONS = {
19 tier: {
20 type: 'choice',
21 instructions: 'Which is the cheapest tier that can complete this coding task well?',
22 criteria: {
23 trivial: 'Read-only or single shell command: read/search/summarise files, run one command and report its output, answer from context. No edits.',
24 fast: 'Small or well-specified edit: rename a symbol, fix a typo, apply an exact change already described, fix a bug whose cause is already known, add tests for a change in one or two files.',
25 balanced: 'Ordinary engineering: implement a well-specified change across several files, fix a bug with clear symptoms that still has to be located in the code, review a diff.',
26 deep: 'Hard or high-stakes: architecture and design, debugging a failure whose cause is unknown, security, data migrations, concurrency, anything touching production or money.',
27 },
28 },
29 effort: {
30 type: 'score',
31 instructions: 'How much step-by-step reasoning does this task need?',
32 criteria: ['almost none', 'some', 'a lot', 'as much as possible'],
33 },
34 risky: { type: 'noul', instructions: 'Carrying out this task would itself change production, move real money, or alter data that cannot be restored. Writing or testing code that deals with such things, without running it against the real system, does not count.' },
35}
36
37export type Tier = (typeof TIERS)[number]
38export type SessionTier = Exclude<Tier, 'trivial'>
39export type Effort = (typeof EFFORTS)[number]
40export type Answer = { tier: Tier; tierConfidence: number; effort: number; effortConfidence: number; risk: number; ms: number }
41type JevAnswers = { tier?: { choice?: string; confidence?: number }; effort?: { score?: number; confidence?: number }; risky?: { noul?: number } }
42type JevResponse = { answers?: JevAnswers }
43
44export const parseAnswer = (text: string, ms: number): Answer | string => {
45 let body: JevResponse
46 try {
47 body = JSON.parse(text) as JevResponse
48 } catch {
49 return 'unreadable JSON'
50 }
51 const answers = body.answers
52 const tier = TIERS.find((name) => name === answers?.tier?.choice)
53 const score = answers?.effort?.score
54 // Jev's score is continuous (0.78 = between "almost none" and "some"): take the nearest level.
55 if (!tier || typeof score !== 'number' || !(score >= 0 && score <= 3)) return 'answer without a valid tier/effort'
56 const effort = Math.round(score)
57 const confidenceOf = (value: unknown) => (typeof value === 'number' && Number.isFinite(value) ? value : 0)
58 return { tier, tierConfidence: confidenceOf(answers?.tier?.confidence), effort, effortConfidence: confidenceOf(answers?.effort?.confidence), risk: confidenceOf(answers?.risky?.noul), ms }
59}
60
61// An unknown current value (-1: haiku, a numeric effort, an unrecognised id) counts as below
62// every level, so any move from it is a raise: it can only go up.
63export const isMoveAllowed = (from: number, to: number, confidence: number, isForced: boolean): boolean => {
64 if (to === from) return false
65 if (isForced) return to > from
66 return confidence >= (to > from ? UP_CONFIDENCE : DOWN_CONFIDENCE)
67}
68
69export const isRisky = (answer: Answer) => answer.risk > RISK_THRESHOLD
70// An off-scale level ('off', 'minimal', a number) ranks -1: below all, so it can only go up.
71export const effortRankOf = (level: unknown) => (EFFORTS as readonly unknown[]).indexOf(level)
72export const targetTierRank = (answer: Answer) => (isRisky(answer) ? TIERS.length - 1 : TIERS.indexOf(answer.tier))
73export const targetEffortRank = (answer: Answer) => {
74 if (isRisky(answer)) return Math.max(answer.effort, RISKY_EFFORT)
75 return Math.max(answer.effort, answer.tier === 'deep' ? DEEP_EFFORT : MIN_EFFORT)
76}
77// The session holds its model to the end: haiku is for one-shot subagents only.
78export const sessionTierRank = (answer: Answer) => Math.max(targetTierRank(answer), TIERS.indexOf('fast'))
79export const aliasOfRank = (rank: number) => TIER_ALIAS[TIERS[rank] ?? 'deep']
80export const tierRankOfModel = (model: string) => TIERS.findIndex((tier) => model.includes(TIER_ALIAS[tier]))
81export const shortModel = (model: string) => Object.values(TIER_ALIAS).find((alias) => model.includes(alias)) ?? model
82export const causeOf = (answer: Answer) =>
83 isRisky(answer) ? `risk ${answer.risk.toFixed(2)}` : `${answer.tier}, confidence ${answer.tierConfidence.toFixed(2)}`
84