SLOPSHOPPER

quality-router

Quality-first routing: picks main-loop effort (medium to max) per request, assigns model and effort to subagents by kind and weight, checks Workflow agent()…

newguardcommandtoaststatusprompt
★ 2v0.3.1MITupdated 2026-10-07nextscape/ns-mods/mods/quality-router
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · quality-router
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /quality-router ⎿ quality-router: quality-router v? の設定 ⎿ quality-router: - 振り分け(本体の Effort・本体モデルの守り・子): on ⎿ quality-router: - Workflow の点検: on ⎿ quality-router: - 本体モデルの守り: off(上位モデル: opus) ⎿ quality-router: - 本体の段階: floor medium / ceiling max ⎿ quality-router: - スキル連動の下限: なし ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts ⚠ quality-router: main waiting · sub 0 · gate on
README

quality-router

Claude Code の作業に合わせて、Effort(推論の深さ)とモデルを自動で選ぶ mod です。 方針は「品質を外さない」ことです。難しい依頼や議論の最中に浅い設定で動くのを防ぎ、軽い作業には軽い設定を割り当てます。

機能内容既定
本体の Effort依頼ごとに、Effort を medium / high / xhigh / max から選びますon
スキル連動の下限壁打ち・計画・原因調査・レビューのスキルが呼ばれたら、Effort の下限を引き上げますon
途中の格上げツールのエラーが2回続いたら、Effort を1段上げますon
サブエージェントの振り分けAgent ツールで起動するサブエージェントに、作業の種類と重さでモデルと Effort を割り当てますon
Workflow の点検Workflow の各 agent() が、作業の種類ごとの下限・上限に収まっているかを確かめ、外れていれば差し戻しますon
判定の記録判定とその結果を手元に記録し、/qr stats で集計します(本文は記録しません)on
本体モデルの守り会話の相手のモデルが Sonnet などになったら、上位モデル(Opus など)に戻しますoff(/qr guard on で有効)

導入

ns-mods の README の手順で入れます。

/plugin install quality-router --marketplace nextscape/ns-mods
  • Claude Code 2.1.287 以降が必要です。2.1.292 で動作を確かめています。
  • 新しく開いたセッションから効きます。
  • 入ったことは、画面下の状態表示(v0.3.1 · main … · sub … · gate on)か、/qr version で確かめられます。

使い方

入れるだけで動きます。設定を変えたいときは、次のコマンドを使います。/qr と /quality-router は同じ働きです(/qr がほかのコマンドと重なるときは /quality-router を使ってください)。

コマンド動作
/qr現在の設定と直近の判定を表示します
/qr on / /qr off本体の Effort、本体モデルの守り、サブエージェントの振り分けをまとめて有効・無効にします
/qr gate on / /qr gate offWorkflow の点検を有効・無効にします
/qr guard on / /qr guard off本体モデルの守りを有効・無効にします(既定は off)
/qr floor <段階> / /qr ceiling <段階>本体の Effort の下限・上限を変えます(medium / high / xhigh / max。既定は medium / max)
/qr top opus / /qr top fable上位モデルを変えます(既定は opus)
/qr skill clearスキル連動の下限を解除します
/qr up / /qr down(/qr up sub で直前のサブエージェント)直前の判定が浅すぎた・重すぎたと記録します
/qr stats / /qr stats today / /qr stats 30d判定の記録を集計して表示します(既定は直近7日)
/qr log / /qr log main / /qr log sub / /qr log gate判定・振り分け・点検の直近の記録を表示します
/qr eval同梱の評価用の例文を実際に判定させ、判定の精度を採点します
/qr version動いている版と読み込み元のフォルダを表示します

設定は手元に保存され、次のセッションにも引き継がれます。

状態表示

v0.3.1 · main xhigh (judged medium, skill: brainstorming) · sub 2 · gate on
表示意味
v0.3.1動いている版
main xhigh会話の相手がいま使っている Effort の段階
judged medium判定は medium だったが、下限などで引き上げたこと
skill: brainstormingスキル連動の下限が効いていること
+1ツールのエラーが続いたため、途中で1段上げたこと
main sonnet! high本体モデルの守りが on で、会話の相手を上位モデルに戻せなかったこと
main paused (effort-router)effort-router に本体の Effort を任せていること
sub 2このセッションで振り分けたサブエージェントの数
gate onWorkflow の点検が有効なこと

effort-router と一緒に使うとき

effort-router(同じマーケットプレイスの mod)も、本体の Effort を選びます。両方を入れたときは、本体の Effort は effort-router に任せ、quality-router は次を続けます。

  • サブエージェントの振り分け
  • Workflow の点検
  • 本体モデルの守り(on にしているとき)
  • 判定の記録(本体のターンの判定は effort-router が行うので、本体ターンの記録は残しません)

effort-router が入っていることは、ユーザー設定と、読み込まれているコマンドの一覧から、どのスコープで入れても検知します。検知したときは、通知で一度だけ知らせ、状態表示に main paused (effort-router) と出します。本体の Effort も quality-router で選びたいときは、effort-router を外してください。

組み込みの /effort との関係

本体の Effort の振り分けが有効な間は、組み込みの /effort で選んだ段階も、次の依頼で上書きします。段階を固定したいときは /qr off にしてから /effort を使ってください。

Workflow を書く人へ

Workflow の各 agent() には、ラベルの頭に作業の種類を書き、その種類の下限と上限の範囲で、モデルと Effort(重さ)を明示してください。規則に合わないスクリプトは、直し方を添えて差し戻します。規則はシステムプロンプトにも入っているので、Claude が自分で書き直します。

ラベルの頭対象の作業下限上限例(軽い/重い)
chore:一覧・収集・抽出・整形・コミットhaiku・lowsonnet・medium件数を数える=haiku・low/コミット=sonnet・medium
impl:修正・機能追加sonnet・mediumなし名前の変更=sonnet・high/設計を伴う実装=上位・xhigh
investigate:調べもの・原因調査・報告sonnet・mediumなしAPI の調べもの=sonnet・high/原因調査=上位・xhigh
verify:テストの実行・事実の照合など、根拠が機械的に得られる確認sonnet・highなしテストを流す=sonnet・high/報告を API で照合=上位・high
review:差分や設計の良し悪しの評価上位・highなし小さな差分=上位・high/設計レビュー=上位・xhigh 以上
refute:結論を崩しにいく上位・highなし—
decide:方針や合否を決める上位・highなし—
fix:指摘を受けたやり直し上位・high、または LADDER の展開なし段を1つ以上上げる(...LADDER[tier + 1])
  • 「上位」は Opus 系・Fable 系・Mythos 系のモデルです。Fable は正式ID claude-fable-5-1 で書きます。
  • 力量は、モデルの格を先に比べ、同じなら Effort を比べます。model を省くと会話の相手のモデル(上位)を引き継ぎます。たとえば上位・low は sonnet・medium より強く、sonnet・xhigh は上位・high より弱い扱いです。
  • fix: 以外は、effort を文字列で書くか、添字が数字の ...LADDER[n] を展開してください(力量を読めないと差し戻します)。
  • chore: で model を省くと上位とみなされ、上限(sonnet・medium)を超えて差し戻されます。...LADDER[0] か ...LADDER[1] を使ってください。
const LADDER = [
  { model: 'haiku', effort: 'low' },
  { model: 'sonnet', effort: 'medium' },
  { model: 'opus', effort: 'high' },
  { model: 'opus', effort: 'xhigh' },
  { model: 'opus', effort: 'max' },
]
await agent(prompt, { label: 'chore:scan', ...LADDER[1] })
await agent(prompt, { label: 'impl:impl', effort: 'high' })
await agent(prompt, { label: 'review:diff', model: 'opus', effort: 'xhigh' })
await agent(prompt, { label: `fix:${n}`, ...LADDER[tier + 1] })
  • オプションは agent() の中に直接書いてください。変数で渡すと点検できないため、差し戻します。
  • 展開(...)は ...LADDER[n] だけが使えます。ほかの展開は点検できないため、差し戻します。
  • 保存済みの Workflow(.claude/workflows/ に置いたもの)は、規則に合わなくても通知だけして通します。

費用・待ち時間・送るデータ

項目内容
Haiku の呼び出し依頼ごと(20文字以上)に1回。振り分けの対象のサブエージェントを起動するたびに2回(種類と重さを同時に)
待ち時間1回あたり約1秒。本体の判定は最大5秒待ち、間に合わなければ前の段階を引き継いで進みます
送るデータ本体:依頼の冒頭と末尾、前の回答の末尾。サブエージェント:説明と指示の冒頭と末尾。送り先は Claude Code 自身の接続(Anthropic)だけで、第三者には送りません
使用量判定の分も、利用者のプランまたは API キーの使用量に数えられます
手元に残るデータ下の「判定の記録」と、/qr log 用の直近200件(依頼は冒頭40文字だけ)

判定の記録

判定の改善に使うため、判定とその結果を手元に記録します。

  • 場所:~/.claude/quality-router/log/YYYY-MM/<セッションID>.jsonl(1行が1件)
  • 記録するもの:
  • 本体の1ターンごとの判定(段階・決まり方・効いた下限)と結果(ステップ数・ツールエラー数・所要時間・トークン量)
  • サブエージェントと Workflow の子の割り当てと結果
  • Workflow の点検の結果
  • 好みの手がかり:組み込みの /effort・/model を手で変えたこと、/qr up・/qr down、同じ作業のやり直しが3回目になったこと
  • 依頼や応答の本文は書きません。 文字数と、会話ログの行を指す ID だけを残します。
  • 記録は手元にだけ残り、どこにも送りません。
  • /qr off の間は、本体ターンの記録と手動の /effort・/model の記録は止まります。サブエージェントや Workflow の記録、/qr up・/qr down は続きます。
  • 記録のフォルダは消してかまいません。動いているセッションは次の記録で自分のファイルを書き直すので、消すのはセッションを閉じてからにしてください。

仕組み

本体の Effort

  • 判定は Claude Haiku が行い、前のやり取りの様子(前の段階、ステップ数、ツール数、エラー数、所要時間、前の依頼と回答の抜粋)も材料にします。
  • 上げるのは即座、下げるのは1回の依頼につき1段までです。20文字未満の短い返事(「続けて」など)は、前の段階を引き継ぎます。
  • 1回の依頼の中では段階を下げません。
  • Effort を変えても、プロンプトのキャッシュは無効になりません(実測で確かめています)。

スキル連動の下限

スキル下限
superpowers:brainstorming、superpowers:writing-plans、superpowers:systematic-debuggingxhigh
code-review、superpowers:requesting-code-review、superpowers:receiving-code-reviewxhigh
superpowers:executing-plans、superpowers:subagent-driven-development、superpowers:test-driven-developmenthigh

表にないスキルを呼んでも、下限は変わりません。

サブエージェントの振り分け

作業の種類と重さ(軽い・重い)を Claude Haiku が同時に判定し、次の表で割り当てます。

種類軽い(light)重い(heavy)
chore 機械的な作業(一覧・収集・抽出・整形・コミット)Sonnet・mediumSonnet・medium
impl 実装Sonnet・high会話の相手と同じ・high
investigate 調査・報告Sonnet・high会話の相手と同じ・xhigh
verify 検証(テストの実行・事実の照合)Sonnet・high上位モデル・high
review レビュー上位モデル・high上位モデル・xhigh(会話の相手が max なら max)
refute 反証上位モデル・xhigh上位モデル・xhigh(会話の相手が max なら max)
decide 判断上位モデル・high上位モデル・xhigh(会話の相手が max なら max)
種類が判定できない・判定に失敗した触らない触らない
  • 種類は判定できたが重さが判定できなかったときは、重い(heavy)として扱います。
  • 同じ説明のサブエージェントが起動されたらやり直しとみなし、前回より1段上げます。3回目は通知で知らせます。
  • 独自に定義したサブエージェント(.claude/agents/ などで作ったもの)や、モデルを明示して起動したものには触りません。

本体モデルの守り(既定 off)

/qr guard on にすると、会話の相手のモデルが Opus 系・Fable 系以外(Sonnet など)になったとき、上位モデルの正式IDに直して送ります。直せないとき(プランや設定で使えないなど)は、元のモデルのまま続け、通知で知らせます。上位モデルを使う分、使用量は増えます。

止める・外す

  • 一時的に止める:/qr off(Workflow の点検も止めるなら /qr gate off)
  • 外す:claude plugin uninstall quality-router。新しいセッションから外れます。手元の記録と設定は残るので、要らなければ ~/.claude/quality-router/ を消してください。

制約と既知の課題

  • 試験版です。判定がおかしいと感じたら、/qr log main の表示を添えて GitHub Issues で知らせてください。
  • 本体の Effort の判定は、Opus 系・Sonnet 系・Fable 系など、Effort を受け付けるモデルで効きます。
  • mod の API は新しく、Claude Code の版によって動かなくなることがあります。

変更履歴

CHANGELOG.md を見てください。

出典

本体の Effort の判定の仕組みと、評価用の例文の一部は、このマーケットプレイスの effort-router(0.3.0)を元にしています。

ライセンス

MIT

Source 16 files
hooks/register.ts 1219 lines
1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, FsEntry, Register, ToolCallResult, TurnCompleteInput, TurnStepChunk, TurnStepResult } from 'claude-code'
3
4import type { ChildRecord, Effort, Level, TurnSignals } from '../types'
5import { MAIN_CASES, SUB_CASES } from './eval-cases'
6import { scanAgentCalls } from './gate-lexer'
7import type { AgentCall } from './gate-lexer'
8import { checkCalls, formatDeny, isSessionScriptPath } from './gate-rules'
9import { guidanceText } from './guidance'
10import { TOP_IDS, applyBounds, asLevel, boundOf, compareStrength, isLevel, isTop, isTopFamily, minLevel, rank, up } from './levels'
11import type { Bound, Bounds } from './levels'
12import { LABELS, MID_TURN_BUMP, MIN_CHARS, compose, decide, head, levelOf, shortDecision, tail } from './main-effort'
13import type { Decision } from './main-effort'
14import { PART_LIMIT, common, gateCounts, recordPath, tokensOf } from './record'
15import type { QrRecord, RecordBody, SubWhy } from './record'
16import { effortOverride, modelOverride, parseFeedback } from './signals'
17import { parseRecords, periodOf, sinceOf, summarize } from './stats'
18import { DEFAULTS, SETTING_KEYS, mentionsEffortRouter, parseSettings } from './settings'
19import type { Settings } from './settings'
20import { skillFloorFor } from './skills'
21import {
22  SUB_KINDS,
23  UNSURE_TIER,
24  WEIGHTS,
25  assign,
26  composeKind,
27  composeWeight,
28  escalate,
29  retryKey,
30  shouldRoute,
31  strengthOfAssignment,
32  subKindOf,
33  weightOf,
34} from './sub-route'
35import type { Assignment, SpawnFacts, SubKind, Weight } from './sub-route'
36import { TEXT, describeBuild, describeSettings, formatGateLog, formatMainLog, formatSubLog, statusLine } from './ui'
37import type { Build, GateLog, MainLog, MainView, SubLog } from './ui'
38
39// Quality-first routing of model and effort (see README.md).
40// A: the main loop's effort, skill floors, the mid-turn bump and the model guard.
41// B: the Workflow gate. C: the Workflow guidance. D: Agent-tool subagents.
42// E: status, toasts, commands, logs and eval. Decisions live in the pure modules.
43//
44// The loader follows `$` only into functions declared at the top of this file,
45// so every function that takes `$` lives here, not inside `register`. The
46// module's own state for one load is the `Ctx` that `register` creates.
47
48const COMMANDS = ['quality-router', 'qr'] as const
49const GUIDANCE_ID = 'quality-router:guidance'
50const LOG_LIMIT = 200
51const RECENT = 3
52// How long a request waits for its turn's judgment. The wait is the step
53// hook's own time (its budget is 10 s), so it stays well inside it.
54const JUDGE_WAIT_MS = 5_000
55const OPEN_TURNS = 8
56
57const prevTurn = atom({ plugin: 'quality-router', key: 'prev' } as const, null)
58const recentLevels = atom({ plugin: 'quality-router', key: 'recent' } as const, [])
59const skillFloorAtom = atom({ plugin: 'quality-router', key: 'skillFloor' } as const, null)
60const childrenAtom = atom({ plugin: 'quality-router', key: 'children' } as const, {})
61const retriesAtom = atom({ plugin: 'quality-router', key: 'retries' } as const, {})
62
63type Turn = Decision & {
64  request: string
65  // The judgment in flight, set by turn.start; `decided` once it settled.
66  judging: Promise<void> | null
67  decided: boolean
68  steps: number
69  tools: number
70  toolErrors: number
71  streak: number
72  bump: number
73  // The highest level this turn reached before the bump, and the bound that
74  // set it: the turn never drops below it, and the bump adds to it.
75  held: { level: Level; bound: Bound | null } | null
76  // The level last written into a request, and the bound that set it.
77  applied: Level | null
78  bound: Bound | null
79  answeredBy: string | null
80  guarded: boolean
81  // The transcript rows this turn's prompt and last answer were stored as
82  // (the transcript rows); null for a turn with no typed prompt.
83  promptRow: PromptRow | null
84  answerUuid: string | null
85  judgeMs: number | null
86}
87
88// A main prompt row as stored: its uuid at once, its time once the clock answers.
89type PromptRow = { uuid: string; at: number | null }
90
91// The session's record file, written whole on each record (fs.write has no
92// append). `resolved` once the path and any lines already there were read.
93type Sink = {
94  lines: string[]
95  size: number
96  path: string | null
97  home: string | null
98  session: string | null
99  startedAt: number | null
100  part: number
101  resolved: boolean
102  warned: boolean
103  // Set after /clear: no session.start follows, so the first record of the
104  // new file is preceded by a session record.
105  announce: boolean
106  chain: Promise<void>
107}
108
109const newSink = (): Sink => ({
110  lines: [],
111  size: 0,
112  path: null,
113  home: null,
114  session: null,
115  startedAt: null,
116  part: 1,
117  resolved: false,
118  warned: false,
119  announce: false,
120  chain: Promise.resolve(),
121})
122
123// What agent.spawn knew of an Agent-tool child, for its record when it ends.
124type SubInfo = {
125  turn: string | null
126  type: string
127  descChars: number
128  kindJudged: SubKind | null
129  weightJudged: Weight | null
130  model: string | null
131  effort: Effort | null
132  why: SubWhy
133  retries: number
134}
135
136// A child's requests so far: the main turn it started under, steps, tool errors, and what it ran on.
137type KidStats = { turn: string | null; steps: number; toolErrors: number; model: string | null; effort: string | null }
138
139// What one load of the module keeps between events.
140type Ctx = {
141  settings: Settings
142  paused: boolean
143  // plugin.json's version and the folder it was loaded from, read with the settings.
144  build: Build
145  // The settings and the effort-router check, read once; every hook awaits it.
146  loading: Promise<void> | null
147  current: string | undefined
148  subCount: number
149  mainView: MainView
150  turns: Map<string, Turn>
151  // Children's assignments, ahead of childrenAtom: a child's first request
152  // can come before the agent.spawn hook has written the atom.
153  children: Map<string, ChildRecord>
154  // Whether the command list was looked at for effort-router (once, at the first turn).
155  routerChecked: boolean
156  // The turn whose request the top model failed: its later steps skip the guard.
157  guardFailed: string | undefined
158  // Log writes, one after another (parallel spawns append at once).
159  logChain: Promise<void>
160  // Records.
161  sink: Sink
162  // The main prompt row last stored, taken by the next turn.start that has text.
163  pendingPrompt: PromptRow | null
164  // The last main turn and child that ended, for /qr up and down.
165  lastMain: string | null
166  lastSub: string | null
167  // The model that answered the main loop last, for a manual /model.
168  lastModel: string | null
169  subs: Map<string, SubInfo>
170  kids: Map<string, KidStats>
171}
172
173type WorkflowInput = { script?: string; scriptPath?: string; name?: string }
174type Source = { text: string; source: GateLog['source']; fromSession: boolean }
175
176async function loadSettings($: EngineInterface): Promise<Settings> {
177  const values: Partial<Record<keyof Settings, unknown>> = {}
178  for (const key of SETTING_KEYS) {
179    try {
180      values[key] = await $.store.get(key)
181    } catch {
182      values[key] = undefined
183    }
184  }
185  return parseSettings(values)
186}
187
188async function saveSetting<K extends keyof Settings>($: EngineInterface, key: K, value: Settings[K]): Promise<void> {
189  await $.store.set(key, value)
190}
191
192// effort-router loaded as a plugin from ~/.claude/settings.json pauses A.
193// effort-router loaded from any scope (user, project, local or a plugin
194// folder): its own /effort-router command is listed.
195async function effortRouterLoaded($: EngineInterface): Promise<boolean> {
196  try {
197    return (await $.command.list()).some(c => c.plugin === 'effort-router' || c.name === 'effort-router')
198  } catch {
199    return false
200  }
201}
202
203// effort-router may register its command after this mod loads: look again
204// once, as the first turn is judged, and say so if it is there.
205async function checkEffortRouter($: EngineInterface, ctx: Ctx): Promise<void> {
206  if (ctx.paused || ctx.routerChecked) return
207  ctx.routerChecked = true
208  if (!(await effortRouterLoaded($))) return
209  ctx.paused = true
210  $.ui.toast(TEXT.effortRouterDetected)
211  refresh($, ctx)
212}
213
214async function effortRouterInstalled($: EngineInterface): Promise<boolean> {
215  try {
216    return mentionsEffortRouter(await $.settings.read({ source: 'user' }))
217  } catch {
218    return false
219  }
220}
221
222async function classify($: EngineInterface, text: string, labels: readonly string[]): Promise<string | undefined | Error> {
223  try {
224    return await $.model.classify(text, labels, { model: 'haiku' })
225  } catch (error) {
226    return error instanceof Error ? error : new Error(String(error))
227  }
228}
229
230// A subagent's kind and weight, judged at once. A failed judgment
231// reads as unknown.
232async function classifySub(
233  $: EngineInterface,
234  e: Pick<SpawnFacts, 'subagentType' | 'description' | 'prompt'>,
235): Promise<{ kind: SubKind | undefined; weight: Weight | undefined }> {
236  const [kind, weight] = await Promise.all([classify($, composeKind(e), SUB_KINDS), classify($, composeWeight(e), WEIGHTS)])
237  return {
238    kind: kind instanceof Error ? undefined : subKindOf(kind),
239    weight: weight instanceof Error ? undefined : weightOf(weight),
240  }
241}
242
243// Never throws: a log that cannot be written must not undo the work it records.
244async function appendLog($: EngineInterface, ctx: Ctx, key: string, entry: unknown): Promise<void> {
245  const write = ctx.logChain.then(async () => {
246    const list = ((await $.store.get(key)) as unknown[] | undefined) ?? []
247    await $.store.set(key, [...list, entry].slice(-LOG_LIMIT))
248  })
249  const done = write.catch(() => undefined)
250  ctx.logChain = done
251  await done
252}
253
254// The file's place, found once: the home folder, the session and its first
255// time, the last part already written, and that part's lines (a hot reload
256// or a resume keeps writing the same file without losing them).
257// A lookup that throws leaves the sink unresolved, so the next record tries again.
258async function resolveSink($: EngineInterface, sink: Sink): Promise<void> {
259  if (sink.resolved) return
260  const home = await homeDir($)
261  if (home === null) {
262    sink.resolved = true
263    return
264  }
265  const session = await $.session.id()
266  const at = await $.clock.now()
267  let part = 1
268  while (await $.fs.exists(recordPath(home, at, session, part + 1)).catch(() => false)) part += 1
269  let path = recordPath(home, at, session, part)
270  let kept: string[] = []
271  if (await $.fs.exists(path).catch(() => false)) {
272    try {
273      const text = await $.fs.read(path)
274      if (typeof text !== 'string') throw new Error('not text')
275      kept = text.split('\n').filter(line => line !== '')
276    } catch {
277      // There but unreadable (busy, or past 4 MiB): never overwrite it, go on in the next part.
278      part += 1
279      path = recordPath(home, at, session, part)
280    }
281  }
282  Object.assign(sink, { home, session, startedAt: at, part, path, lines: [...kept, ...sink.lines], resolved: true })
283  sink.size = sink.lines.reduce((n, line) => n + line.length + 1, 0)
284}
285
286// Appends one record and writes the session's file. Never throws: a record
287// that cannot be written must not undo the work it records. The first
288// failure is toasted, later ones are not.
289async function record($: EngineInterface, ctx: Ctx, turn: string | null, body: RecordBody): Promise<void> {
290  const sink = ctx.sink
291  const write = sink.chain.then(async () => {
292    await resolveSink($, sink)
293    if (sink.path === null || sink.home === null || sink.session === null || sink.startedAt === null) return
294    const at = await $.clock.now()
295    const bodies: RecordBody[] = []
296    if (sink.announce && sink.lines.length === 0 && body.kind !== 'session') bodies.push(await sessionBody($, ctx, 'start'))
297    bodies.push(body)
298    sink.announce = false
299    for (const b of bodies) {
300      const line = JSON.stringify({ ...common(at, sink.session, b.kind === 'session' ? null : turn, ctx.build.version), ...b } as QrRecord)
301      // One read or write is at most 4 MiB: a long session goes on in a new part.
302      if (sink.lines.length > 0 && sink.size + line.length + 1 > PART_LIMIT) {
303        sink.part += 1
304        sink.path = recordPath(sink.home, sink.startedAt, sink.session, sink.part)
305        sink.lines = []
306        sink.size = 0
307      }
308      sink.lines.push(line)
309      sink.size += line.length + 1
310    }
311    await $.fs.write(sink.path!, `${sink.lines.join('\n')}\n`)
312  })
313  const done = write.catch(() => {
314    if (sink.warned) return
315    sink.warned = true
316    $.ui.toast(TEXT.recordFailed)
317  })
318  sink.chain = done
319  await done
320}
321
322async function sessionBody($: EngineInterface, ctx: Ctx, event: 'start' | 'settings'): Promise<RecordBody> {
323  const cwd = await $.session.cwd().catch(() => null)
324  return { kind: 'session', event, root: ctx.build.root, cwd, settings: ctx.settings, paused: ctx.paused }
325}
326
327async function recordSession($: EngineInterface, ctx: Ctx, event: 'start' | 'settings'): Promise<void> {
328  await record($, ctx, null, await sessionBody($, ctx, event))
329}
330
331// A child's run ended: an Agent-tool child (agent.spawn saw it) is a sub
332// record, any other (a Workflow agent) a wf child record. Its counts start
333// over for its next run.
334async function recordChild($: EngineInterface, ctx: Ctx, e: TurnCompleteInput, agentId: string): Promise<void> {
335  const stats = ctx.kids.get(agentId)
336  ctx.kids.delete(agentId)
337  const info = ctx.subs.get(agentId)
338  const result = { steps: stats?.steps ?? 0, toolErrors: stats?.toolErrors ?? 0, tokens: tokensOf(e.usage), durationMs: e.durationMs, reason: e.reason }
339  if (info) {
340    ctx.lastSub = agentId
341    const { turn, ...rest } = info
342    await record($, ctx, turn, { kind: 'sub', agentId, ...rest, ...result })
343  } else {
344    await record($, ctx, stats?.turn ?? null, { kind: 'wf', phase: 'child', agentId, model: stats?.model ?? null, effort: stats?.effort ?? null, ...result })
345  }
346}
347
348// The records since `since`, from every session's file: month folders from
349// the one `since` falls in, files changed since then, lines at or after it.
350async function readRecords($: EngineInterface, since: number): Promise<QrRecord[]> {
351  const home = await homeDir($)
352  if (home === null) return []
353  const base = `${home}/.claude/quality-router/log`
354  // A session's file stays in the month it began: look one month further back.
355  const d = new Date(since)
356  const fromMonth = new Date(Date.UTC(d.getUTCFullYear(), d.getUTCMonth() - 1, 1)).toISOString().slice(0, 7)
357  let months: readonly FsEntry[]
358  try {
359    months = await $.fs.list(base)
360  } catch {
361    return []
362  }
363  const out: QrRecord[] = []
364  for (const month of months) {
365    if (month.kind !== 'dir' || month.name < fromMonth) continue
366    let files: readonly FsEntry[]
367    try {
368      files = await $.fs.list(`${base}/${month.name}`)
369    } catch {
370      continue
371    }
372    for (const file of files) {
373      if (file.kind !== 'file' || !file.name.endsWith('.jsonl') || file.mtimeMs < since) continue
374      try {
375        const text = await $.fs.read(`${base}/${month.name}/${file.name}`)
376        if (typeof text === 'string') out.push(...parseRecords(text).filter(r => Date.parse(r.at) >= since))
377      } catch {
378        // A file being written or removed: skip it.
379      }
380    }
381  }
382  return out
383}
384
385// The user's home with forward slashes, or null when neither variable answers.
386async function homeDir($: EngineInterface): Promise<string | null> {
387  const home = (await $.env.get('USERPROFILE').catch(() => undefined)) || (await $.env.get('HOME').catch(() => undefined))
388  return home ? home.replace(/\\/g, '/') : null
389}
390
391// A saved workflow `name` as a file in `dir`, or null when there is none.
392async function savedWorkflow($: EngineInterface, dir: string, name: string): Promise<string | null> {
393  for (const ext of ['js', 'mjs']) {
394    const path = `${dir}/.claude/workflows/${name}.${ext}`
395    if (!(await $.fs.exists(path))) continue
396    const text = await $.fs.read(path)
397    if (typeof text === 'string') return text
398  }
399  return null
400}
401
402// scriptPath takes precedence over script and name, as the Workflow tool reads them.
403// A name is looked up in the project, then in the home folder, read only as far as needed.
404async function workflowSource($: EngineInterface, e: WorkflowInput): Promise<Source | null> {
405  if (e.scriptPath) {
406    const text = await $.fs.read(e.scriptPath)
407    if (typeof text !== 'string') return null
408    return { text, source: 'scriptPath', fromSession: isSessionScriptPath(e.scriptPath, await $.session.id()) }
409  }
410  if (e.script) return { text: e.script, source: 'script', fromSession: true }
411  if (e.name) {
412    const cwd = (await $.session.cwd()).replace(/\\/g, '/')
413    let text = await savedWorkflow($, cwd, e.name)
414    if (text === null) {
415      const home = await homeDir($)
416      if (home !== null) text = await savedWorkflow($, home, e.name)
417    }
418    if (text !== null) return { text, source: 'name', fromSession: false }
419  }
420  return null
421}
422
423// The manifest sits in .claude-plugin/ of the plugin's folder; a folder that
424// holds plugin.json itself is read too. Null when neither can be read.
425async function readVersion($: EngineInterface, root: string): Promise<string | null> {
426  for (const path of [`${root}/.claude-plugin/plugin.json`, `${root}/plugin.json`]) {
427    try {
428      const text = await $.fs.read(path)
429      if (typeof text !== 'string') continue
430      const version: unknown = JSON.parse(text).version
431      if (typeof version === 'string' && version !== '') return version
432    } catch {
433      // Not there or not JSON: try the next one.
434    }
435  }
436  return null
437}
438
439async function loadBuild($: EngineInterface): Promise<Build> {
440  let root: string | null = null
441  try {
442    root = $.plugin.root || null
443  } catch {
444    root = null
445  }
446  return { version: root ? await readVersion($, root) : null, root }
447}
448
449async function load($: EngineInterface, ctx: Ctx): Promise<void> {
450  ctx.build = await loadBuild($)
451  ctx.settings = await loadSettings($)
452  ctx.paused = (await effortRouterInstalled($)) || (await effortRouterLoaded($))
453  if (ctx.paused) $.ui.toast(TEXT.effortRouterDetected)
454}
455
456// Every hook reads the settings through here. The first call loads them; the
457// calls that come while it loads wait for the same load, not the defaults.
458async function cfg($: EngineInterface, ctx: Ctx): Promise<Settings> {
459  ctx.loading ??= load($, ctx)
460  await ctx.loading
461  return ctx.settings
462}
463
464function refresh($: EngineInterface, ctx: Ctx): void {
465  $.ui.status(statusLine({ settings: ctx.settings, paused: ctx.paused, main: ctx.mainView, sub: ctx.subCount, version: ctx.build.version }))
466}
467
468// A main-loop turn, opened by turn.start or, when the engine sends the turn's
469// first request before turn.start's hook has run, by that request's step.
470// turn.complete closes it; the few kept beyond that bound the ones it never closed.
471function openTurn(ctx: Ctx, turnId: string): Turn {
472  for (const id of ctx.turns.keys()) {
473    if (ctx.turns.size < OPEN_TURNS) break
474    ctx.turns.delete(id)
475  }
476  const turn: Turn = {
477    base: null,
478    judged: null,
479    why: 'late',
480    request: '',
481    judging: null,
482    decided: false,
483    steps: 0,
484    tools: 0,
485    toolErrors: 0,
486    streak: 0,
487    bump: 0,
488    held: null,
489    applied: null,
490    bound: null,
491    answeredBy: null,
492    guarded: false,
493    promptRow: null,
494    answerUuid: null,
495    judgeMs: null,
496  }
497  ctx.turns.set(turnId, turn)
498  ctx.current = turnId
499  return turn
500}
501
502// A: judge the turn's request. A request shorter than
503// MIN_CHARS ("", a continuation) inherits; a failed or unmatched judgment
504// keeps the previous level. While A is off the turn stays undecided ('late').
505async function judgeTurn($: EngineInterface, ctx: Ctx, turn: Turn): Promise<void> {
506  try {
507    const s = await cfg($, ctx)
508    if (s.mode === 'off') return
509    await checkEffortRouter($, ctx)
510    if (ctx.paused) return
511    const prev = await read($, prevTurn)
512    let decision: Decision
513    if (turn.request.length < MIN_CHARS) {
514      decision = shortDecision(prev)
515    } else {
516      const started = await $.clock.now()
517      decision = decide(await classify($, compose(turn.request, prev, await read($, recentLevels)), LABELS), prev)
518      turn.judgeMs = (await $.clock.now()) - started
519    }
520    turn.base = decision.base
521    turn.judged = decision.judged
522    turn.why = decision.why
523  } catch {
524    // The steps inherit the previous level, as for a late judgment.
525  } finally {
526    turn.decided = true
527  }
528}
529
530// Waits for `promise` at most `ms`, or until `signal` aborts; never throws.
531async function within($: EngineInterface, promise: Promise<void>, ms: number, signal: AbortSignal): Promise<void> {
532  if (signal.aborted) return
533  const stop = new AbortController()
534  const quit = () => stop.abort()
535  signal.addEventListener('abort', quit)
536  // Aborted (the wait ended first, or the dispatch did): go on.
537  const timer = $.clock.sleep(ms, { signal: stop.signal }).catch(() => undefined)
538  try {
539    await Promise.race([promise, timer])
540  } finally {
541    signal.removeEventListener('abort', quit)
542    stop.abort()
543  }
544}
545
546// A level written into a main request this turn: the log, the status and /qr read it.
547function markApplied(ctx: Ctx, turn: Turn, level: Level, skill: string | null, bound: Bound | null): void {
548  turn.applied = level
549  turn.bound = bound
550  ctx.mainView = { ...ctx.mainView, level, judged: turn.judged, skill, bump: turn.bump > 0 }
551}
552
553// A: the main loop's calls and failures for the mid-turn bump.
554// A deny is neither an error nor a success: the tool never ran.
555function countTool(turn: Turn, ran: ToolCallResult): void {
556  turn.tools += 1
557  if (ran.deny !== undefined) return
558  if (ran.isError === true) {
559    turn.toolErrors += 1
560    turn.streak += 1
561    if (MID_TURN_BUMP && turn.streak >= 2 && turn.bump === 0) turn.bump = 1
562  } else {
563    turn.streak = 0
564  }
565}
566
567// The response's own content: once one of these has gone up the chain, it is
568// shown and recorded, so the request can no longer be sent again unseen.
569function isContent(c: TurnStepChunk): boolean {
570  return c.kind === 'text' || c.kind === 'thinking' || c.kind === 'tool' || c.kind === 'input'
571}
572
573// B: the deny text for a session script that breaks a rule, or null to let
574// the call run (a script that cannot be read runs).
575async function gate($: EngineInterface, ctx: Ctx, input: WorkflowInput): Promise<string | null> {
576  const s = await cfg($, ctx)
577  if (s.gate === 'off') return null
578  const at = await $.clock.now()
579  const fallbackSource: GateLog['source'] = input.scriptPath ? 'scriptPath' : input.script ? 'script' : 'name'
580  let src: Source | null = null
581  let calls = 0
582  let found: AgentCall[] = []
583  let violations: ReturnType<typeof checkCalls> | null = null
584  try {
585    src = await workflowSource($, input)
586    if (src) {
587      found = scanAgentCalls(src.text)
588      calls = found.length
589      violations = checkCalls(found)
590    }
591  } catch {
592    violations = null
593  }
594  if (src === null || violations === null) {
595    // A name not found under .claude/workflows is a built-in one
596    // (review-changes): logged, not toasted.
597    const unresolvedName = input.scriptPath === undefined && input.script === undefined && input.name !== undefined
598    if (!unresolvedName || (violations === null && src !== null)) $.ui.toast(TEXT.gateFailed)
599    await appendLog($, ctx, 'log.gate', { at, source: src?.source ?? fallbackSource, verdict: 'unchecked', count: 0, calls } satisfies GateLog)
600    await record($, ctx, ctx.current ?? null, { kind: 'wf', phase: 'gate', verdict: 'unchecked', source: src?.source ?? fallbackSource, calls, kinds: {}, rules: {}, fixes: 0 })
601    return null
602  }
603  const verdict: GateLog['verdict'] = violations.length === 0 ? 'pass' : src.fromSession ? 'deny' : 'warn'
604  await appendLog($, ctx, 'log.gate', { at, source: src.source, verdict, count: violations.length, calls } satisfies GateLog)
605  await record($, ctx, ctx.current ?? null, { kind: 'wf', phase: 'gate', verdict, source: src.source, calls, ...gateCounts(found, violations) })
606  if (verdict === 'pass') return null
607  if (verdict === 'warn') {
608    $.ui.toast(TEXT.warned(violations.length))
609    return null
610  }
611  $.ui.toast(TEXT.denied(violations.length))
612  return formatDeny(violations)
613}
614
615async function set<K extends keyof Settings>($: EngineInterface, ctx: Ctx, key: K, value: Settings[K]): Promise<void> {
616  ctx.settings = { ...ctx.settings, [key]: value }
617  await saveSetting($, key, value)
618}
619
620// /clear and an in-process /resume end the conversation but not the process,
621// and no session.start follows: what this session kept would leak into the
622// next one.
623async function resetSession($: EngineInterface, ctx: Ctx): Promise<void> {
624  // A /clear starts a new session id: its records go to a new file.
625  ctx.sink = { ...newSink(), announce: true }
626  ctx.pendingPrompt = null
627  ctx.lastMain = null
628  ctx.lastSub = null
629  ctx.lastModel = null
630  ctx.subs.clear()
631  ctx.kids.clear()
632  ctx.turns.clear()
633  ctx.children.clear()
634  ctx.current = undefined
635  ctx.guardFailed = undefined
636  ctx.mainView = {}
637  ctx.subCount = 0
638  try {
639    await update($, prevTurn, () => null)
640    await update($, recentLevels, () => [])
641    await update($, skillFloorAtom, () => null)
642    await update($, childrenAtom, () => ({}))
643    await update($, retriesAtom, () => ({}))
644  } catch {
645    // The host may already have dropped the session's state.
646  }
647  refresh($, ctx)
648}
649
650async function formatLogs($: EngineInterface, which: string): Promise<string> {
651  const parts: string[] = []
652  if (which === '' || which === 'main') parts.push(formatMainLog(((await $.store.get('log.main')) as MainLog[] | undefined) ?? []))
653  if (which === '' || which === 'sub') parts.push(formatSubLog(((await $.store.get('log.sub')) as SubLog[] | undefined) ?? []))
654  if (which === '' || which === 'gate') parts.push(formatGateLog(((await $.store.get('log.gate')) as GateLog[] | undefined) ?? []))
655  return parts.length > 0 ? parts.join('\n\n') : TEXT.unknown(`log ${which}`)
656}
657
658async function evalMain($: EngineInterface): Promise<string> {
659  let hits = 0
660  let low = 0
661  const lines: string[] = []
662  for (const one of MAIN_CASES) {
663    const answer = await classify($, compose(one.request, one.prev, one.recent), LABELS)
664    const judged = answer instanceof Error ? undefined : levelOf(answer)
665    const isLow = judged !== undefined && rank(judged) < rank(one.expect)
666    if (judged === one.expect) hits += 1
667    if (isLow) low += 1
668    const mark = judged === one.expect ? 'ok ' : isLow ? 'LOW' : 'off'
669    lines.push(`  ${mark} ${one.name}: expect ${one.expect}, judged ${judged ?? '-'}`)
670  }
671  const passed = low === 0 && MAIN_CASES.length > 0 && hits / MAIN_CASES.length >= 0.8
672  return [
673    `eval main: ${hits}/${MAIN_CASES.length} match, ${low} judged too low`,
674    passed ? '判定: 合格' : '判定: 不合格(低すぎる判定が0件、かつ一致が80%以上が条件)',
675    ...lines,
676  ].join('\n')
677}
678
679async function evalSub($: EngineInterface): Promise<string> {
680  let hits = 0
681  let low = 0
682  const lines: string[] = []
683  for (const one of SUB_CASES) {
684    const { kind, weight } = await classifySub($, { subagentType: one.type, description: one.description, prompt: one.prompt })
685    const judged = assign(kind, weight, null, 'opus')
686    const expected = assign(one.expect, one.weight, null, 'opus')
687    const isLow = judged !== null && expected !== null && compareStrength(strengthOfAssignment(judged), strengthOfAssignment(expected)) < 0
688    const hit = kind === one.expect
689    if (hit) hits += 1
690    if (isLow) low += 1
691    const mark = isLow ? 'LOW' : hit ? 'ok ' : 'off'
692    lines.push(`  ${mark} ${one.name}: expect ${one.expect}/${one.weight}, judged ${kind ?? '-'}/${weight ?? '-'}`)
693  }
694  const passed = low === 0 && SUB_CASES.length > 0 && hits / SUB_CASES.length >= 0.8
695  return [
696    `eval sub: ${hits}/${SUB_CASES.length} match, ${low} judged too low`,
697    passed ? '判定: 合格' : '判定: 不合格(割り当てが正解より弱くなる判定が0件、かつ種類の一致が80%以上が条件)',
698    ...lines,
699  ].join('\n')
700}
701
702async function runCommand($: EngineInterface, ctx: Ctx, raw: string): Promise<string> {
703  await cfg($, ctx)
704  const trimmed = raw.trim()
705  const [verb = '', arg = ''] = trimmed.toLowerCase().split(/\s+/)
706  switch (verb) {
707    case '':
708      break
709    case 'on':
710    case 'off':
711      await set($, ctx, 'mode', verb)
712      break
713    case 'gate':
714    case 'guard':
715      if (arg !== 'on' && arg !== 'off') return TEXT.unknown(trimmed)
716      await set($, ctx, verb, arg)
717      break
718    case 'floor':
719    case 'ceiling':
720      if (!isLevel(arg)) return TEXT.unknown(trimmed)
721      await set($, ctx, verb, arg)
722      break
723    case 'top':
724      if (!isTop(arg)) return TEXT.unknown(trimmed)
725      await set($, ctx, 'top', arg)
726      break
727    case 'skill':
728      if (arg !== 'clear') return TEXT.unknown(trimmed)
729      await update($, skillFloorAtom, () => null)
730      ctx.mainView = { ...ctx.mainView, skill: null }
731      break
732    case 'up':
733    case 'down':
734    case '↑':
735    case '↓': {
736      const fb = parseFeedback(verb, arg)
737      if (fb === null) return TEXT.unknown(trimmed)
738      const targetTurn = fb.target === 'main' ? ctx.lastMain : null
739      const targetAgent = fb.target === 'sub' ? ctx.lastSub : null
740      if (targetTurn === null && targetAgent === null) return TEXT.feedbackNoTarget(fb.target)
741      await record($, ctx, ctx.current ?? null, { kind: 'signal', type: 'feedback', dir: fb.dir, target: fb.target, targetTurn, targetAgent })
742      return TEXT.feedbackSaved(fb.dir, fb.target)
743    }
744    case 'stats': {
745      const period = periodOf(arg)
746      if (period === null) return TEXT.unknown(trimmed)
747      return summarize(await readRecords($, sinceOf(period, await $.clock.now())), period)
748    }
749    case 'log':
750      return formatLogs($, arg)
751    case 'version':
752      return describeBuild(ctx.build)
753    case 'eval': {
754      const parts: string[] = []
755      if (arg === '' || arg === 'main') parts.push(await evalMain($))
756      if (arg === '' || arg === 'sub') parts.push(await evalSub($))
757      return parts.length > 0 ? parts.join('\n\n') : TEXT.unknown(trimmed)
758    }
759    default:
760      return TEXT.unknown(trimmed)
761  }
762  if (['on', 'off', 'gate', 'guard', 'floor', 'ceiling', 'top'].includes(verb)) await recordSession($, ctx, 'settings')
763  refresh($, ctx)
764  const floor = await read($, skillFloorAtom)
765  return describeSettings(ctx.settings, ctx.paused, floor?.skill ?? null, ctx.mainView, ctx.build)
766}
767
768export const register: Register = on => {
769  const ctx: Ctx = {
770    settings: DEFAULTS,
771    paused: false,
772    build: { version: null, root: null },
773    loading: null,
774    current: undefined,
775    subCount: 0,
776    mainView: {},
777    turns: new Map<string, Turn>(),
778    children: new Map<string, ChildRecord>(),
779    guardFailed: undefined,
780    routerChecked: false,
781    logChain: Promise.resolve(),
782    sink: newSink(),
783    pendingPrompt: null,
784    lastMain: null,
785    lastSub: null,
786    lastModel: null,
787    subs: new Map<string, SubInfo>(),
788    kids: new Map<string, KidStats>(),
789  }
790
791  on('session.start', async ($, e, next) => {
792    await cfg($, ctx)
793    // One name failing to register leaves the other and the status line.
794    for (const name of COMMANDS) {
795      try {
796        await $.command.register({
797          name,
798          description: 'Quality-first model and effort routing: on|off, gate, guard, floor, ceiling, top, skill, up|down, stats, log, eval, version',
799          argumentHint: '[on|off|gate|guard|floor|ceiling|top|skill|up|down|stats|log|eval|version]',
800          immediate: true,
801        })
802      } catch {
803        $.ui.toast(TEXT.commandFailed(name))
804      }
805    }
806    await recordSession($, ctx, 'start')
807    refresh($, ctx)
808    return next(e)
809  })
810
811  // The other reasons end the process: nothing is left to leak into.
812  on('session.end', async ($, e, next) => {
813    if (e.reason === 'clear' || e.reason === 'resume') await resetSession($, ctx)
814    return next(e)
815  })
816
817  // Literal matchers, one per name in COMMANDS, so the host can read them statically.
818  on('command.run', { command: 'quality-router' }, async ($, e) => ({ text: await runCommand($, ctx, e.args) }))
819  on('command.run', { command: 'qr' }, async ($, e) => ({ text: await runCommand($, ctx, e.args) }))
820
821  // Records: a manual /effort or /model, read against what A sent and what answered.
822  // While A is off or paused, a manual choice is the person's way to pin a
823  // level, not a verdict on a judgment: it is not recorded.
824  on('command.run', { command: 'effort' }, async ($, e, next) => {
825    const s = await cfg($, ctx)
826    if (s.mode === 'on' && !ctx.paused) await record($, ctx, ctx.current ?? null, { kind: 'signal', ...effortOverride(e.args, ctx.mainView.level ?? null) })
827    return next(e)
828  })
829  on('command.run', { command: 'model' }, async ($, e, next) => {
830    const s = await cfg($, ctx)
831    if (s.mode === 'on' && !ctx.paused) await record($, ctx, ctx.current ?? null, { kind: 'signal', ...modelOverride(e.args, ctx.lastModel) })
832    return next(e)
833  })
834
835  // A: judge once when the request arrives. The engine sends the turn's first
836  // request after waiting at most 3 s for this hook (and while the previous
837  // turn's events are still queued, before it runs), so the turn is opened
838  // before any await, and turn.step waits for the judgment, bounded.
839  on('turn.start', async ($, e, next) => {
840    const turn = ctx.turns.get(e.turnId) ?? openTurn(ctx, e.turnId)
841    ctx.current = e.turnId
842    turn.request = e.text
843    // The prompt row comes just before; a turn with no typed text (a
844    // continuation) takes none, and a row left over is dropped either way.
845    if (e.text !== '' && ctx.pendingPrompt !== null) turn.promptRow = ctx.pendingPrompt
846    ctx.pendingPrompt = null
847    turn.judging = judgeTurn($, ctx, turn)
848    await turn.judging
849    return next(e)
850  })
851
852  // Records: the transcript row ids that lead back to the prompt and the answer.
853  on('session.append', { door: 'prompt' }, async ($, e, next) => {
854    if (e.agentId === undefined) {
855      // Taken before any await: turn.start can run while this hook waits on the clock.
856      const row: PromptRow = { uuid: e.uuid, at: null }
857      ctx.pendingPrompt = row
858      row.at = await $.clock.now()
859    }
860    return next(e)
861  })
862  on('session.append', { door: 'response' }, async ($, e, next) => {
863    const turn = e.agentId === undefined && ctx.current !== undefined ? ctx.turns.get(ctx.current) : undefined
864    if (turn) turn.answerUuid = e.uuid
865    return next(e)
866  })
867
868  on('turn.step', async function* ($, e, next) {
869    const s = await cfg($, ctx)
870
871    // D: a subagent the Agent tool started runs at its assigned effort.
872    if (e.agentId !== undefined) {
873      const kid = s.mode === 'on' ? (ctx.children.get(e.agentId) ?? (await read($, childrenAtom))[e.agentId]) : undefined
874      const sent = kid?.effort !== undefined && e.effort !== undefined ? { ...e, effort: kid.effort } : e
875      // Records: the child's first request names the main turn it started under.
876      const stats = ctx.kids.get(e.agentId) ?? { turn: ctx.current ?? null, steps: 0, toolErrors: 0, model: null, effort: null }
877      stats.steps += 1
878      stats.model = sent.model
879      stats.effort = typeof sent.effort === 'string' ? sent.effort : null
880      ctx.kids.set(e.agentId, stats)
881      return yield* next(sent)
882    }
883
884    // A: the main loop's effort for this request.
885    const active = s.mode === 'on' && !ctx.paused
886    const turn = ctx.turns.get(e.turnId) ?? (active ? openTurn(ctx, e.turnId) : undefined)
887    let sent = e
888    let level: Level | null = null
889    let skill: string | null = null
890    let bound: Bound | null = null
891    if (turn && active) {
892      if (!turn.decided && turn.judging) await within($, turn.judging, JUDGE_WAIT_MS, next.signal)
893      turn.steps += 1
894      // No judgment in time: inherit the previous level, as a short prompt does.
895      // The judgment, when it comes, takes over from the next request; the
896      // never-lower rule below keeps the turn from dropping meanwhile.
897      const decision = turn.why === 'late' ? shortDecision(await read($, prevTurn)) : turn
898      const base = decision.base ?? asLevel(e.effort)
899      if (base !== null) {
900        const floor = await read($, skillFloorAtom)
901        skill = floor?.skill ?? null
902        const bounds: Bounds = { floor: s.floor, ceiling: s.ceiling, skillFloor: floor?.level ?? null }
903        level = applyBounds(base, bounds)
904        bound = boundOf(base, bounds)
905        // Never lower within a turn: a lower skill floor, `/qr
906        // skill clear` or a late lower judgment takes effect at the next turn.
907        // The ceiling still bounds it.
908        const held = turn.held
909        if (held !== null && rank(held.level) > rank(level)) {
910          const kept = minLevel(held.level, s.ceiling)
911          if (rank(kept) > rank(level)) {
912            level = kept
913            bound = kept === held.level ? held.bound : 'ceiling'
914          }
915        }
916        if (held === null || rank(level) > rank(held.level)) turn.held = { level, bound }
917        if (turn.bump > 0) level = minLevel(up(level, turn.bump), s.ceiling)
918        if (e.effort !== undefined) {
919          sent = { ...sent, effort: level }
920          markApplied(ctx, turn, level, skill, bound)
921        }
922      }
923    }
924
925    // A: the main model guard, part of A, so `/qr off` stops it too.
926    // effort-router only sets effort, so its pause leaves the guard on.
927    if (s.mode === 'on' && s.guard === 'on' && !isTopFamily(e.model)) {
928      if (ctx.guardFailed !== e.turnId) {
929        const target = TOP_IDS[s.top]
930        // A model without effort sends none; the top model takes this turn's level.
931        const added = sent.effort === undefined ? level : null
932        const attempt = next(added !== null ? { ...sent, model: target, effort: added } : { ...sent, model: target })
933        // The engine's items ahead of the first content are held back: a failed
934        // attempt's (its API error item among them) are dropped for the resend.
935        const held: TurnStepChunk[] = []
936        let shown = false
937        let r: TurnStepResult | undefined
938        let failure: { error: unknown } | undefined
939        try {
940          for await (const c of attempt) {
941            if (!shown && !isContent(c)) {
942              held.push(c)
943              continue
944            }
945            if (!shown) {
946              shown = true
947              yield* held.splice(0)
948            }
949            yield c
950          }
951          r = await attempt.result
952        } catch (error) {
953          failure = { error }
954        }
955        const answered = r !== undefined && (r.stopReason !== null || r.usage !== null)
956        const by = r?.usage?.model
957        if (r !== undefined && answered && (by === undefined || isTopFamily(by))) {
958          yield* held
959          if (turn) {
960            turn.guarded = true
961            turn.answeredBy = by ?? target
962            if (added !== null) markApplied(ctx, turn, added, skill, bound)
963          }
964          ctx.mainView = { ...ctx.mainView, offModel: null }
965          refresh($, ctx)
966          return r
967        }
968        if (r !== undefined && answered && by !== undefined) {
969          // The engine kept another model (a policy that does not allow the
970          // rewrite): its answer stands, and the rest of the turn stops trying.
971          yield* held
972          ctx.guardFailed = e.turnId
973          $.ui.toast(TEXT.guardFallback(by))
974          if (turn) turn.answeredBy = by
975          ctx.mainView = { ...ctx.mainView, offModel: by }
976          refresh($, ctx)
977          return r
978        }
979        // No response because the person interrupted (Esc): the top model did
980        // not fail, so no toast, no off-model mark and no resend. The test
981        // kit cannot abort a dispatch, so this branch has no test.
982        if (next.signal.aborted) {
983          if (failure) throw failure.error
984          yield* held
985          return r!
986        }
987        // Told once a turn: its later steps go straight to the original model.
988        ctx.guardFailed = e.turnId
989        $.ui.toast(TEXT.guardFallback(e.model))
990        ctx.mainView = { ...ctx.mainView, offModel: e.model }
991        if (shown) {
992          // Part of the failed answer is already shown and recorded: a resend
993          // would follow it. The engine's own error handling takes the step,
994          // and the turn's next request goes out on the original model.
995          refresh($, ctx)
996          if (failure) throw failure.error
997          return r!
998        }
999        // Nothing shown: resend on the original model.
1000      }
1001      ctx.mainView = { ...ctx.mainView, offModel: e.model }
1002      const r = yield* next(sent)
1003      if (turn) turn.answeredBy = r.usage?.model ?? e.model
1004      refresh($, ctx)
1005      return r
1006    }
1007
1008    const r = yield* next(sent)
1009    if (turn) turn.answeredBy = r.usage?.model ?? e.model
1010    ctx.mainView = { ...ctx.mainView, offModel: null }
1011    refresh($, ctx)
1012    return r
1013  })
1014
1015  on('tool.call', async ($, e, next) => {
1016    // The main loop's turn, read before the call: a long tool can outlast it.
1017    const turn = e.agentId === undefined && ctx.current !== undefined ? ctx.turns.get(ctx.current) : undefined
1018
1019    // B: the Workflow gate.
1020    const deny = e.tool === 'Workflow' ? await gate($, ctx, { script: e.script, scriptPath: e.scriptPath, name: e.name }) : null
1021    const ran: ToolCallResult = deny !== null ? { deny } : await next(e)
1022
1023    // A: every main-loop call counts toward the mid-turn bump, Workflow included.
1024    if (turn) countTool(turn, ran)
1025
1026    // Records: a child's failed calls.
1027    if (e.agentId !== undefined && ran.isError === true) {
1028      const stats = ctx.kids.get(e.agentId)
1029      if (stats) stats.toolErrors += 1
1030    }
1031
1032    // A: a skill the main loop called sets the skill floor from its next
1033    // request; a denied or failed call sets nothing.
1034    if (e.tool === 'Skill' && e.agentId === undefined && ran.deny === undefined && ran.isError !== true) {
1035      const level = skillFloorFor(e.skill)
1036      if (level) {
1037        const skill = e.skill.trim().replace(/^\//, '')
1038        await update($, skillFloorAtom, () => ({ level, skill }))
1039        ctx.mainView = { ...ctx.mainView, skill }
1040      }
1041    }
1042    return ran
1043  })
1044
1045  // C: teach the staffing rules while the gate is on, to a model that is
1046  // offered the Workflow tool (not a subagent's or a Workflow agent's render
1047  // without it). `--bare` asked for a one-line prompt, so it gets none.
1048  on('prompt.compose', async ($, e, next) => {
1049    const result = await next(e)
1050    const s = await cfg($, ctx)
1051    if (s.gate === 'off' || e.traits.includes('bare') || !e.tools.includes('Workflow')) return result
1052    return { ...result, sections: [...result.sections, { id: GUIDANCE_ID, text: guidanceText(s.top), scope: 'session' as const }] }
1053  })
1054
1055  // D: route a subagent the Agent tool starts.
1056  on('agent.spawn', async ($, e, next) => {
1057    const s = await cfg($, ctx)
1058    if (s.mode === 'off' || !shouldRoute(e)) {
1059      const result = await next(e)
1060      if (result.deny === undefined && result.agentId !== undefined) {
1061        ctx.subs.set(result.agentId, {
1062          turn: ctx.current ?? null,
1063          type: e.subagentType,
1064          descChars: e.description.length,
1065          kindJudged: null,
1066          weightJudged: null,
1067          model: result.model ?? null,
1068          effort: null,
1069          why: s.mode === 'off' ? 'off' : 'explicit',
1070          retries: 0,
1071        })
1072      }
1073      return result
1074    }
1075    const key = retryKey(e.description)
1076    const prior = (await read($, retriesAtom))[key]
1077    let a: Assignment | null
1078    let why: SubLog['why']
1079    let role: string
1080    let kindJudged: SubKind | null = null
1081    let weightJudged: Weight | null = null
1082    if (prior) {
1083      a = escalate(prior.tier, s.top)
1084      why = 'retry'
1085      role = 'retry'
1086    } else {
1087      const { kind, weight } = await classifySub($, e)
1088      kindJudged = kind ?? null
1089      weightJudged = weight ?? null
1090      a = assign(kind, weight, ctx.mainView.level ?? null, s.top)
1091      why = a ? 'auto' : 'unsure'
1092      role = kind !== undefined ? `${kind}/${weight ?? 'heavy'}` : 'unsure'
1093    }
1094    const result = await next(a?.model !== undefined ? { ...e, model: a.model } : e)
1095    const agentId = result.deny === undefined ? result.agentId : undefined
1096    if (agentId !== undefined && a?.effort !== undefined) {
1097      const child: ChildRecord = { tier: a.tier, effort: a.effort }
1098      // In memory before any await: the child's first request can come before
1099      // this hook ends. The atom keeps it across a hot reload.
1100      ctx.children.set(agentId, child)
1101      try {
1102        await update($, childrenAtom, map => ({ ...map, [agentId]: child }))
1103      } catch {
1104        // The in-memory record serves this load.
1105      }
1106    }
1107    // A denied spawn started nothing: no retry to count, no child routed.
1108    if (result.deny === undefined) {
1109      const count = (prior?.count ?? 0) + 1
1110      const tier = a?.tier ?? UNSURE_TIER
1111      try {
1112        await update($, retriesAtom, map => ({ ...map, [key]: { count, tier } }))
1113      } catch {
1114        // A lost retry record only means the next start is judged afresh.
1115      }
1116      if (count === 3) {
1117        $.ui.toast(TEXT.retryThird(e.description))
1118        await record($, ctx, ctx.current ?? null, { kind: 'signal', type: 'retry3', agentId: agentId ?? null })
1119      }
1120      ctx.subCount += 1
1121      if (agentId !== undefined) {
1122        ctx.subs.set(agentId, {
1123          turn: ctx.current ?? null,
1124          type: e.subagentType,
1125          descChars: e.description.length,
1126          kindJudged,
1127          weightJudged,
1128          model: result.model ?? null,
1129          effort: a?.effort ?? null,
1130          why,
1131          retries: prior?.count ?? 0,
1132        })
1133      }
1134    }
1135    await appendLog($, ctx, 'log.sub', {
1136      at: await $.clock.now(),
1137      description: head(e.description, 40),
1138      type: e.subagentType,
1139      role,
1140      why,
1141      model: result.deny === undefined ? result.model : null,
1142      effort: a?.effort ?? null,
1143      agentId: agentId ?? null,
1144    } satisfies SubLog)
1145    refresh($, ctx)
1146    return result
1147  })
1148
1149  on('turn.complete', async ($, e, next) => {
1150    if (e.agentId !== undefined) {
1151      await recordChild($, ctx, e, e.agentId)
1152      return next(e)
1153    }
1154    const turn = e.agentId === undefined ? ctx.turns.get(e.turnId) : undefined
1155    if (turn) {
1156      ctx.turns.delete(e.turnId)
1157      if (ctx.current === e.turnId) ctx.current = undefined
1158    }
1159    // A turn A never stepped (off, paused, or aborted before a request) leaves no record.
1160    if (turn && turn.steps > 0) {
1161      const prev = await read($, prevTurn)
1162      const floor = await read($, skillFloorAtom)
1163      const signals: TurnSignals = {
1164        level: turn.applied,
1165        judged: turn.judged,
1166        steps: turn.steps,
1167        tools: turn.tools,
1168        toolErrors: turn.toolErrors,
1169        durationMs: e.durationMs,
1170        // A continuation ("") keeps the request it continues for the next judgment.
1171        request: turn.request !== '' ? head(turn.request) : (prev?.request ?? ''),
1172        answerTail: tail(e.answer),
1173      }
1174      await update($, prevTurn, () => signals)
1175      await update($, recentLevels, list => [...list, turn.applied].slice(-RECENT))
1176      await appendLog($, ctx, 'log.main', {
1177        at: await $.clock.now(),
1178        head: head(turn.request, 40),
1179        prev: prev?.level ?? null,
1180        judged: turn.judged,
1181        level: turn.applied,
1182        why: turn.why,
1183        bound: turn.bound,
1184        skill: floor?.skill ?? null,
1185        bump: turn.bump > 0,
1186        model: turn.answeredBy,
1187        guarded: turn.guarded,
1188        durationMs: e.durationMs,
1189      } satisfies MainLog)
1190      ctx.lastMain = e.turnId
1191      ctx.lastModel = turn.answeredBy
1192      await record($, ctx, e.turnId, {
1193        kind: 'main',
1194        promptUuid: turn.promptRow?.uuid ?? null,
1195        answerUuid: turn.answerUuid,
1196        promptAt: turn.promptRow?.at != null ? new Date(turn.promptRow.at).toISOString() : null,
1197        promptChars: turn.request.length,
1198        prev: prev?.level ?? null,
1199        judged: turn.judged,
1200        level: turn.applied,
hooks/eval-cases.ts 91 lines
1import type { Level, TurnSignals } from '../types'
2import type { SubKind, Weight } from './sub-route'
3
4// Cases for /qr eval. The first 16 main cases are from effort-router v0.3.0
5// (in this marketplace); the max cases and the sub cases are new. Add real requests
6// from /qr log main that were judged wrong, rewritten so they name no
7// customer, project or local path (use placeholders such as 案件A): the
8// cases may be sent outside Anthropic to compare another judge.
9
10export type MainCase = { name: string; request: string; prev: TurnSignals | null; recent: Array<Level | null>; expect: Level }
11export type SubCase = { name: string; type: string; description: string; prompt: string; expect: SubKind; weight: Weight }
12
13const heavy = (request: string, answerTail: string, toolErrors = 0): TurnSignals => ({
14  level: 'xhigh',
15  judged: 'xhigh',
16  steps: 9,
17  tools: 14,
18  toolErrors,
19  durationMs: 210_000,
20  request,
21  answerTail,
22})
23
24const light = (request: string, answerTail: string): TurnSignals => ({
25  level: 'medium',
26  judged: 'medium',
27  steps: 1,
28  tools: 0,
29  toolErrors: 0,
30  durationMs: 6_000,
31  request,
32  answerTail,
33})
34
35export const MAIN_CASES: MainCase[] = [
36  // From effort-router: no context, the request alone decides.
37  { name: 'fresh-simple-question', request: 'TypeScript の satisfies 演算子って何をするものか一言で教えて', prev: null, recent: [], expect: 'medium' },
38  { name: 'fresh-rename', request: 'utils.ts の関数 fmtDate を formatDate にリネームして', prev: null, recent: [], expect: 'high' },
39  { name: 'fresh-explain-file', request: 'src/pipeline/split.ts が何をしているか説明してほしい', prev: null, recent: [], expect: 'high' },
40  { name: 'fresh-design', request: 'スライド生成パイプラインにキャッシュ層を入れたい。設計案を複数出してトレードオフを比較して', prev: null, recent: [], expect: 'xhigh' },
41  { name: 'fresh-debug', request: 'npm test で split.test.ts だけ落ちる。原因を調べて直して', prev: null, recent: [], expect: 'xhigh' },
42  { name: 'fresh-stack-trace', request: "これが出て起動しない\n```\nTypeError: Cannot read properties of undefined (reading 'body')\n    at splitHtml (split.ts:42:13)\n    at run (pipeline.ts:88:5)\n    at main (index.ts:12:3)\n```", prev: null, recent: [], expect: 'xhigh' },
43  { name: 'fresh-commit-msg', request: 'いまの差分にコミットメッセージを付けてコミットして', prev: null, recent: [], expect: 'medium' },
44  // From effort-router: context makes a plain-looking request heavy.
45  { name: 'ctx-pick-option-after-design', request: '案2の方向で、さっきの設計どおりに実装を進めてください', prev: heavy('キャッシュ層の設計案を比較して', '…案1はシンプルだが無効化が難しい。案2は複雑だが整合性を保てる。どちらで進めますか?'), recent: ['xhigh', 'xhigh'], expect: 'xhigh' },
46  { name: 'ctx-still-broken', request: '直してもらったけど、まだ同じエラーが出ます。確認してください', prev: heavy('split.test.ts が落ちる原因を調べて', '…body 側の CSS が落ちていたのが原因でした。修正してテストが通ることを確認しました。', 2), recent: ['xhigh'], expect: 'xhigh' },
47  { name: 'ctx-continue-steps', request: 'その調子で残りのステップ3から5もお願いします', prev: heavy('移行計画の実装をステップ1から', '…ステップ1と2を実装し、テストも通りました。残りはステップ3〜5です。'), recent: ['xhigh', 'xhigh', 'xhigh'], expect: 'xhigh' },
48  // From effort-router: context makes a request light even after heavy work.
49  { name: 'ctx-thanks-after-heavy', request: 'ありがとうございます、助かりました。今日はここまでにします', prev: heavy('原因調査', '…修正してテストが通りました。'), recent: ['xhigh'], expect: 'medium' },
50  { name: 'ctx-typo-after-heavy', request: 'README の「インストール」の誤字だけ直しておいて', prev: heavy('設計レビュー', '…以上がレビュー結果です。'), recent: ['xhigh'], expect: 'medium' },
51  { name: 'ctx-yesno-confirm', request: 'さっきの変更はもう push 済みという認識で合ってる?', prev: light('main に push して', '…push しました。'), recent: ['medium'], expect: 'medium' },
52  // From effort-router: context keeps an ordinary task ordinary.
53  { name: 'ctx-add-test-after-light', request: 'formatDate にうるう年のテストケースを1つ追加して', prev: light('fmtDate をリネームして', '…リネームしました。'), recent: ['high', 'medium'], expect: 'high' },
54  { name: 'ctx-review-after-light', request: '今の差分をレビューして、問題があれば指摘して', prev: light('コミットして', '…コミットしました。'), recent: ['medium'], expect: 'high' },
55  { name: 'ctx-new-topic-heavy-after-light', request: '話は変わるけど、認証まわりを JWT からセッション方式に移行したい。影響範囲を洗い出して計画を立てて', prev: light('typo を直して', '…直しました。'), recent: ['medium'], expect: 'xhigh' },
56  // New: max.
57  { name: 'max-fix-failed-again', request: '修正してもらったのに、まだ同じテストが落ちます。これで3回目です。根本原因から洗い直してください', prev: heavy('split.test.ts を直して', '…修正しました。テストが通ることを確認しました。', 2), recent: ['xhigh', 'xhigh'], expect: 'max' },
58  { name: 'max-irreversible-migration', request: '本番DBのスキーマ移行の方針を決めたい。ロールバックできない変更なので、リスクを徹底的に洗い出して', prev: null, recent: [], expect: 'max' },
59  { name: 'max-explicit-deepest', request: 'この設計判断は後戻りできません。考えうる最も深いレベルで分析してください。時間はかかって構いません', prev: null, recent: [], expect: 'max' },
60  { name: 'max-prod-incident', request: '本番で決済が二重に計上されている。直前の修正では直らなかった。原因を特定して', prev: heavy('決済の重複を直して', '…冪等キーを追加しました。', 1), recent: ['xhigh'], expect: 'max' },
61  // New: discussion, planning and quick status.
62  { name: 'xhigh-pick-with-condition', request: '案Bでいきましょう。ただしロールバック手順も設計に含めてください', prev: heavy('移行方式の案を比べて', '…案Aは早いが戻せない。案Bは段階的に戻せる。どちらにしますか?'), recent: ['xhigh'], expect: 'xhigh' },
63  { name: 'xhigh-review-plan', request: 'この実装計画をレビューして、抜け漏れや危ない作業順がないか指摘して', prev: null, recent: [], expect: 'xhigh' },
64  { name: 'high-explain-config', request: 'tsconfig の strict と noUncheckedIndexedAccess の違いを説明して', prev: null, recent: [], expect: 'high' },
65  { name: 'medium-status-one-line', request: 'いまどこまで進んだか、一言で教えてもらえますか', prev: heavy('実装を進めて', '…タスク3まで終わりました。'), recent: ['xhigh'], expect: 'medium' },
66]
67
68export const SUB_CASES: SubCase[] = [
69  { name: 'chore-list', type: 'general-purpose', description: 'list hook files', prompt: 'プロジェクトの .claude/hooks フォルダのファイル名を一覧にして返して', expect: 'chore', weight: 'light' },
70  { name: 'chore-count', type: 'general-purpose', description: 'count commands', prompt: '.claude/commands のファイル数を数えて数字だけ返して', expect: 'chore', weight: 'light' },
71  { name: 'chore-commit', type: 'general-purpose', description: 'commit typo fix', prompt: 'いまの差分をコミットして。メッセージは「fix: 誤字を修正」で', expect: 'chore', weight: 'light' },
72  { name: 'impl-rename', type: 'general-purpose', description: 'rename function', prompt: 'utils.ts の関数 fmtDate を formatDate にリネームして、呼び出し元も直して', expect: 'impl', weight: 'light' },
73  { name: 'impl-lexer', type: 'general-purpose', description: 'implement parser', prompt: 'gate-lexer.ts に正規表現リテラルの読み飛ばしを実装して、テストも足して', expect: 'impl', weight: 'heavy' },
74  { name: 'impl-auth', type: 'general-purpose', description: 'implement session auth', prompt: '認証をセッション方式に切り替える実装を、設計メモに沿って複数のファイルにわたって進めて', expect: 'impl', weight: 'heavy' },
75  { name: 'investigate-api', type: 'general-purpose', description: 'research API', prompt: 'Azure DevOps の PR スレッド API でコメントを解決済みにする方法を調べて、公式ドキュメントの根拠つきで報告して', expect: 'investigate', weight: 'light' },
76  { name: 'investigate-failing', type: 'general-purpose', description: 'investigate failing test', prompt: 'npm test で register.test.ts だけ落ちる原因を調べて、原因と根拠を報告して', expect: 'investigate', weight: 'heavy' },
77  { name: 'investigate-status', type: 'general-purpose', description: 'sprint status', prompt: '案件A の今週の Sprint の進捗を調べて、遅れているチケットと理由をまとめて報告して', expect: 'investigate', weight: 'light' },
78  { name: 'verify-tests', type: 'general-purpose', description: 'run tests', prompt: 'テストを実行して、全部通ったかどうかと失敗したテスト名だけを返して', expect: 'verify', weight: 'light' },
79  { name: 'verify-build', type: 'general-purpose', description: 'check build', prompt: 'ビルドが通るか確かめて、エラーがあれば最初の5行を返して', expect: 'verify', weight: 'light' },
80  { name: 'verify-report', type: 'general-purpose', description: 'verify report', prompt: '次の調査報告に書かれた固有名詞とタイムスタンプを実 API で照合し、誤りがあれば指摘して', expect: 'verify', weight: 'heavy' },
81  { name: 'review-small', type: 'general-purpose', description: 'review small diff', prompt: 'この10行の差分をレビューして、明らかな誤りがあれば指摘して', expect: 'review', weight: 'light' },
82  { name: 'review-diff', type: 'general-purpose', description: 'review diff', prompt: 'この差分をレビューして、正しさのバグと抜けているテストを指摘して', expect: 'review', weight: 'heavy' },
83  { name: 'review-design', type: 'general-purpose', description: 'review design coverage', prompt: '設計書の決定事項がすべて実装に反映されているかレビューし、抜けと矛盾を列挙して', expect: 'review', weight: 'heavy' },
84  { name: 'refute-cache', type: 'general-purpose', description: 'refute finding', prompt: 'この指摘「キャッシュが無効になる」を反証して。反証できなければ refuted=false と答えて', expect: 'refute', weight: 'heavy' },
85  { name: 'refute-cause', type: 'general-purpose', description: 'refute root cause', prompt: '「この不具合の原因は設定ファイルの誤り」という結論の反例を探して崩してみて', expect: 'refute', weight: 'heavy' },
86  { name: 'refute-claim', type: 'general-purpose', description: 'refute simple claim', prompt: '「この関数は常に正の数を返す」という主張が明らかに誤っていないかだけ確かめて', expect: 'refute', weight: 'light' },
87  { name: 'decide-merge', type: 'general-purpose', description: 'merge decision', prompt: 'レビュー結果を読んで、マージしてよいかを yes か no で答えて', expect: 'decide', weight: 'light' },
88  { name: 'decide-migration', type: 'general-purpose', description: 'choose migration', prompt: '移行方式の案A と案B を比べ、どちらで進めるべきか理由とともに決めて', expect: 'decide', weight: 'heavy' },
89  { name: 'decide-audit', type: 'general-purpose', description: 'audit verdict', prompt: '実装計画が仕様のすべての要件を満たしているか監査して、合格か不合格かを決めて', expect: 'decide', weight: 'heavy' },
90]
91
hooks/gate-lexer.ts 283 lines
1// Reads a Workflow script token by token and finds its agent( ... ) calls.
2// Strings, template literals, comments and regex literals are skipped, so a
3// description that mentions "agent()" is not a call (the spike's false hit).
4
5export type Token = {
6  kind: 'ident' | 'punct' | 'string' | 'template' | 'regex' | 'number'
7  text: string
8  start: number
9  end: number
10}
11
12export type OptionsInfo =
13  | { kind: 'none' }
14  | { kind: 'dynamic'; text: string }
15  | { kind: 'object'; entries: Record<string, Token[]>; spreads: string[] }
16
17export type AgentCall = { ordinal: number; options: OptionsInfo }
18
19const REGEX_AFTER_PUNCT = new Set([
20  '(', ',', '=', ':', '[', '!', '&', '|', '?', '{', '}', ';', '+', '-', '*', '%', '<', '>', '~', '^', '=>',
21])
22const REGEX_AFTER_WORD = new Set([
23  'return', 'typeof', 'case', 'do', 'else', 'in', 'of', 'new', 'delete', 'void', 'throw', 'await', 'yield',
24])
25// A ')' that closes `if (`, `while (`, `for (` or `with (` ends a statement head, so a '/' after it starts a regex.
26const STATEMENT_HEAD = new Set(['if', 'while', 'for', 'with'])
27const OPEN = new Set(['(', '[', '{'])
28const CLOSE = new Set([')', ']', '}'])
29
30function skipQuoted(src: string, i: number, quote: string): number {
31  i += 1
32  while (i < src.length) {
33    const c = src[i]
34    if (c === '\\') {
35      i += 2
36      continue
37    }
38    i += 1
39    if (c === quote || c === '\n') break
40  }
41  return i
42}
43
44// From just after `${` to just after its closing `}`.
45function skipExpression(src: string, i: number): number {
46  let depth = 1
47  while (i < src.length && depth > 0) {
48    const c = src[i]
49    if (c === '"' || c === "'") {
50      i = skipQuoted(src, i, c)
51      continue
52    }
53    if (c === '`') {
54      i = skipTemplate(src, i)
55      continue
56    }
57    if (c === '{') depth += 1
58    else if (c === '}') depth -= 1
59    i += 1
60  }
61  return i
62}
63
64function skipTemplate(src: string, i: number): number {
65  i += 1
66  while (i < src.length) {
67    const c = src[i]
68    if (c === '\\') {
69      i += 2
70      continue
71    }
72    if (c === '`') return i + 1
73    if (c === '$' && src[i + 1] === '{') {
74      i = skipExpression(src, i + 2)
75      continue
76    }
77    i += 1
78  }
79  return i
80}
81
82function skipRegex(src: string, i: number): number {
83  i += 1
84  let inClass = false
85  while (i < src.length) {
86    const c = src[i]
87    if (c === '\\') {
88      i += 2
89      continue
90    }
91    if (c === '\n') break
92    i += 1
93    if (inClass) {
94      if (c === ']') inClass = false
95    } else if (c === '[') {
96      inClass = true
97    } else if (c === '/') {
98      while (i < src.length && /[a-z]/i.test(src[i] ?? '')) i += 1
99      break
100    }
101  }
102  return i
103}
104
105export function tokenize(src: string): Token[] {
106  const out: Token[] = []
107  const push = (kind: Token['kind'], start: number, end: number) =>
108    out.push({ kind, text: src.slice(start, end), start, end })
109  // One entry per open '(': whether it follows a STATEMENT_HEAD word.
110  const parens: boolean[] = []
111  // Indexes in `out` of the ')' tokens that close a statement head.
112  const headCloses = new Set<number>()
113  let i = 0
114  while (i < src.length) {
115    const c = src[i] ?? ''
116    if (/\s/.test(c)) {
117      i += 1
118      continue
119    }
120    if (c === '/' && src[i + 1] === '/') {
121      const j = src.indexOf('\n', i)
122      i = j < 0 ? src.length : j
123      continue
124    }
125    if (c === '/' && src[i + 1] === '*') {
126      const j = src.indexOf('*/', i + 2)
127      i = j < 0 ? src.length : j + 2
128      continue
129    }
130    const start = i
131    if (c === '"' || c === "'") {
132      i = skipQuoted(src, i, c)
133      push('string', start, i)
134      continue
135    }
136    if (c === '`') {
137      i = skipTemplate(src, i)
138      push('template', start, i)
139      continue
140    }
141    if (c === '/') {
142      const prev = out[out.length - 1]
143      const isRegex =
144        prev === undefined ||
145        (prev.kind === 'punct' && REGEX_AFTER_PUNCT.has(prev.text)) ||
146        (prev.kind === 'punct' && prev.text === ')' && headCloses.has(out.length - 1)) ||
147        (prev.kind === 'ident' && REGEX_AFTER_WORD.has(prev.text))
148      if (isRegex) {
149        i = skipRegex(src, i)
150        push('regex', start, i)
151        continue
152      }
153    }
154    if (/[A-Za-z_$]/.test(c)) {
155      while (i < src.length && /[\w$]/.test(src[i] ?? '')) i += 1
156      push('ident', start, i)
157      continue
158    }
159    if (/[0-9]/.test(c)) {
160      while (i < src.length && /[\w.]/.test(src[i] ?? '')) i += 1
161      push('number', start, i)
162      continue
163    }
164    const three = src.slice(i, i + 3)
165    const two = src.slice(i, i + 2)
166    // `++` and `--` are one token and not in REGEX_AFTER_PUNCT, so `i++ / 2` reads as a division.
167    const text = three === '...' ? three : two === '=>' || two === '?.' || two === '++' || two === '--' ? two : c
168    i += text.length
169    if (text === '(') {
170      const prev = out[out.length - 1]
171      parens.push(prev?.kind === 'ident' && STATEMENT_HEAD.has(prev.text))
172    } else if (text === ')' && parens.pop()) {
173      headCloses.add(out.length)
174    }
175    push('punct', start, i)
176  }
177  return out
178}
179
180// From the '(' at `open`: the index of its matching ')' (toks.length when it is never closed).
181function closeOf(toks: Token[], open: number): number {
182  let depth = 0
183  for (let j = open; j < toks.length; j += 1) {
184    const t = toks[j]!
185    if (t.kind !== 'punct') continue
186    if (OPEN.has(t.text)) depth += 1
187    else if (CLOSE.has(t.text)) {
188      depth -= 1
189      if (depth === 0) return j
190    }
191  }
192  return toks.length
193}
194
195// From the '(' at `open`: each top-level argument's tokens.
196function splitArgs(toks: Token[], open: number): Token[][] {
197  const args: Token[][] = [[]]
198  let depth = 0
199  for (let j = open; j < toks.length; j += 1) {
200    const t = toks[j]!
201    if (t.kind === 'punct' && OPEN.has(t.text)) {
202      depth += 1
203      if (depth === 1) continue
204    } else if (t.kind === 'punct' && CLOSE.has(t.text)) {
205      depth -= 1
206      if (depth === 0) break
207    } else if (depth === 1 && t.kind === 'punct' && t.text === ',') {
208      args.push([])
209      continue
210    }
211    args[args.length - 1]!.push(t)
212  }
213  return args.filter(arg => arg.length > 0)
214}
215
216function keyName(t: Token): string | null {
217  if (t.kind === 'ident') return t.text
218  if (t.kind === 'string') return t.text.slice(1, -1)
219  return null
220}
221
222function optionsOf(src: string, arg: Token[] | undefined): OptionsInfo {
223  if (!arg || arg.length === 0) return { kind: 'none' }
224  const first = arg[0]!
225  const last = arg[arg.length - 1]!
226  const isObject = first.kind === 'punct' && first.text === '{' && last.kind === 'punct' && last.text === '}'
227  if (!isObject) return { kind: 'dynamic', text: src.slice(first.start, last.end) }
228
229  const entries: Record<string, Token[]> = {}
230  const spreads: string[] = []
231  let depth = 0
232  let entry: Token[] = []
233  const flush = () => {
234    const lead = entry[0]
235    if (lead === undefined) return
236    if (lead.kind === 'punct' && lead.text === '...') {
237      const rest = entry.slice(1)
238      const a = rest[0]
239      const b = rest[rest.length - 1]
240      if (a && b) spreads.push(src.slice(a.start, b.end))
241    } else {
242      const key = keyName(lead)
243      const colon = entry[1]
244      if (key !== null && colon?.kind === 'punct' && colon.text === ':') entries[key] = entry.slice(2)
245      else if (key !== null && entry.length === 1) entries[key] = [lead]
246    }
247    entry = []
248  }
249  for (const t of arg.slice(1, -1)) {
250    if (t.kind === 'punct' && OPEN.has(t.text)) depth += 1
251    else if (t.kind === 'punct' && CLOSE.has(t.text)) depth -= 1
252    if (depth === 0 && t.kind === 'punct' && t.text === ',') {
253      flush()
254      continue
255    }
256    entry.push(t)
257  }
258  flush()
259  return { kind: 'object', entries, spreads }
260}
261
262export function scanAgentCalls(src: string): AgentCall[] {
263  const toks = tokenize(src)
264  const calls: AgentCall[] = []
265  for (let i = 0; i < toks.length; i += 1) {
266    const t = toks[i]!
267    const before = toks[i - 1]
268    const after = toks[i + 1]
269    if (t.kind !== 'ident' || t.text !== 'agent') continue
270    if (after?.kind !== 'punct' || after.text !== '(') continue
271    if (before?.kind === 'punct' && (before.text === '.' || before.text === '?.')) continue
272    if (before?.kind === 'ident' && before.text === 'function') continue
273    // `agent(p) { ... }` is a method named agent (object or class shorthand, `function* agent`), not a call.
274    // A real call is followed by '{' only across ASI (`agent(x)` then a block on the next line), which
275    // Workflow scripts do not write; skipping it errs on the side of letting the script run.
276    const next = toks[closeOf(toks, i + 1) + 1]
277    if (next?.kind === 'punct' && next.text === '{') continue
278    const args = splitArgs(toks, i + 1)
279    calls.push({ ordinal: calls.length + 1, options: optionsOf(src, args[1]) })
280  }
281  return calls
282}
283
hooks/gate-rules.ts 216 lines
1import type { Effort } from '../types'
2import type { AgentCall, Token } from './gate-lexer'
3import { KIND_RULES, LEGACY_KINDS, kindOf } from './kinds'
4import { EFFORTS, RUNG_STRENGTH, compareStrength, formatStrength, isEffort, modelClass } from './levels'
5import type { ModelClass, Strength } from './levels'
6
7// The Workflow gate's rules R0-R5 and the Japanese
8// text that sends a script back.
9
10export type Value = { kind: 'literal'; value: string } | { kind: 'template'; head: string } | { kind: 'expr' }
11export type Violation = {
12  ordinal: number
13  label: string | null
14  rule: 'R0' | 'R1' | 'R2' | 'R3' | 'R4' | 'R5'
15  message: string
16}
17
18const OPTIONS_INLINE = 'agent() の第2引数はオブジェクトリテラルで直接書いてください(変数や関数の戻り値では静的に確かめられないため)'
19const LABEL_LITERAL = 'label は文字列かテンプレート文字列で直接書き、頭に種類(chore: など)を付けてください(静的に確かめるため)'
20const KIND_PREFIX = 'ラベルの頭に種類(chore: / impl: / investigate: / verify: / review: / refute: / decide: / fix:)のいずれかを付けてください'
21const MODEL_LITERAL = 'model は文字列で直接書いてください(静的に確かめるため)'
22const EFFORT_LITERAL = 'effort は文字列で直接書いてください(静的に確かめるため)'
23const EFFORT_MISSING = 'effort を文字列で書くか、...LADDER[n](n は数字)を使ってください(重さを静的に確かめるため)'
24const INDEX_LITERAL = 'LADDER の添字は数字で書いてください(添字に式を使えるのは fix: だけです)'
25const INDEX_RANGE = `LADDER の添字は 0〜${RUNG_STRENGTH.length - 1} で書いてください`
26const FIX_UP = 'fix は段を1つ以上上げてください(...LADDER[tier + 1] の展開か、上位モデル(opus か claude-fable-5-1)で effort high 以上)'
27const INHERITS = '。model を省くと本体のモデルを引き継ぎます'
28
29const LADDER_SPREAD = /^LADDER\s*\[/
30// A rung whose index is a literal number, so its strength can be read.
31const NUMERIC_RUNG = /^LADDER\s*\[\s*(\d+)\s*\]$/
32
33// Only fix may index the ladder with an expression.
34function spreadOnlyLadder(kind: string): string {
35  return `展開(...)は ...LADDER[n]${kind === 'fix' ? ' ' : '(n は数字)'}だけにしてください(静的に確かめるため)`
36}
37
38// A string or template literal, or one followed by `+` ('impl:' + name), whose
39// literal head is what the rules read; anything else is an expression.
40export function valueOf(tokens: Token[] | undefined): Value | undefined {
41  if (!tokens || tokens.length === 0) return undefined
42  const first = tokens[0]!
43  const joined = tokens.length > 1 && tokens[1]?.kind === 'punct' && tokens[1].text === '+'
44  if (tokens.length > 1 && !joined) return { kind: 'expr' }
45  let value: Value
46  if (first.kind === 'string') value = { kind: 'literal', value: first.text.slice(1, -1) }
47  else if (first.kind === 'template') {
48    const body = first.text.slice(1, -1)
49    const at = body.indexOf('${')
50    value = at < 0 ? { kind: 'literal', value: body } : { kind: 'template', head: body.slice(0, at) }
51  } else return { kind: 'expr' }
52  if (!joined) return value
53  return { kind: 'template', head: value.kind === 'literal' ? value.value : value.head }
54}
55
56function headOf(value: Value | undefined): string {
57  return value?.kind === 'literal' ? value.value : value?.kind === 'template' ? value.head : ''
58}
59
60function isModelName(value: string): boolean {
61  return ['haiku', 'sonnet', 'opus'].includes(value) || /^claude-[a-z0-9-]+$/.test(value)
62}
63
64// Explicit keys and a rung spread may both be written, and the lexer keeps no
65// order, so each key may come from either of them (in JavaScript the one
66// written last wins, key by key). The floor reads the weakest pairing, the
67// lower tier with the lower effort; the ceiling the strongest. An omitted
68// model inherits the session model, which counts as the top tier.
69// undefined: no effort was written and no rung gives one.
70function bounds(
71  cls: ModelClass | undefined,
72  effort: Effort | undefined,
73  rung: Strength | undefined,
74): { weaker: Strength; stronger: Strength } | undefined {
75  const classes = [cls, rung?.cls].filter((c): c is ModelClass => c !== undefined)
76  const efforts = [effort, rung?.effort].filter((e): e is Effort => e !== undefined)
77  if (efforts.length === 0) return undefined
78  if (classes.length === 0) classes.push(3)
79  efforts.sort((x, y) => EFFORTS.indexOf(x) - EFFORTS.indexOf(y))
80  return {
81    weaker: { cls: Math.min(...classes) as ModelClass, effort: efforts[0]! },
82    stronger: { cls: Math.max(...classes) as ModelClass, effort: efforts[efforts.length - 1]! },
83  }
84}
85
86export function checkCalls(calls: AgentCall[]): Violation[] {
87  const out: Violation[] = []
88  for (const call of calls) {
89    const o = call.options
90    if (o.kind === 'dynamic') {
91      // Options in a variable or built by a call cannot be read statically,
92      // and reading past them would let a script step around the rules.
93      out.push({ ordinal: call.ordinal, label: null, rule: 'R0', message: OPTIONS_INLINE })
94      continue
95    }
96    const entries = o.kind === 'object' ? o.entries : {}
97    const spreads = o.kind === 'object' ? o.spreads : []
98    const label = valueOf(entries.label)
99    const labelText = label?.kind === 'literal' ? label.value : label?.kind === 'template' ? `${label.head}…` : null
100    const add = (rule: Violation['rule'], message: string) =>
101      out.push({ ordinal: call.ordinal, label: labelText, rule, message })
102
103    // R4 applies to every call, whatever its kind.
104    const model = valueOf(entries.model)
105    const effort = valueOf(entries.effort)
106    const badModel = model?.kind === 'literal' && !isModelName(model.value)
107    // A claude- id outside the known families has no tier, so neither its
108    // floor nor its ceiling could be checked.
109    const unknownTier = model?.kind === 'literal' && !badModel && modelClass(model.value) === null
110    const badEffort = effort?.kind === 'literal' && !isEffort(effort.value)
111    if (model?.kind === 'literal' && badModel) {
112      add('R4', `model「${model.value}」は使えません。haiku / sonnet / opus か、claude- で始まる正式IDにしてください`)
113    }
114    if (model?.kind === 'literal' && unknownTier) {
115      add(
116        'R4',
117        `model「${model.value}」は格(haiku / sonnet / 上位)が分かりません。haiku / sonnet / opus か、claude-haiku- / claude-sonnet- / claude-opus- / claude-fable- / claude-mythos- で始まる正式IDにしてください`,
118      )
119    }
120    if (effort?.kind === 'literal' && badEffort) {
121      add('R4', `effort「${effort.value}」は使えません。low / medium / high / xhigh / max のいずれかにしてください`)
122    }
123
124    // R0: the label starts with one of the eight kinds.
125    const head = headOf(label)
126    const kind = kindOf(head)
127    if (kind === null) {
128      const legacy = /^(light|work|judge):/.exec(head)?.[1]
129      if (legacy !== undefined) add('R0', `${legacy}: は旧ラベルです。${LEGACY_KINDS[legacy]} に置き換えてください`)
130      else add('R0', label?.kind === 'expr' ? LABEL_LITERAL : KIND_PREFIX)
131      continue
132    }
133    // A value R4 rejected leaves no strength to read.
134    if (badModel || unknownTier || badEffort) continue
135
136    // R1: any spread but a ladder rung could set model or effort to anything
137    // at run time, so the strength could not be read. This holds for
138    // fix too: a rung with an expression index is the only spread it may use.
139    if (spreads.some(spread => !LADDER_SPREAD.test(spread))) {
140      add('R1', spreadOnlyLadder(kind))
141      continue
142    }
143    const rungIndexes = spreads.flatMap(spread => {
144      const match = NUMERIC_RUNG.exec(spread)
145      return match ? [Number(match[1])] : []
146    })
147    if (rungIndexes.some(index => index >= RUNG_STRENGTH.length)) {
148      add('R1', INDEX_RANGE)
149      continue
150    }
151    // R5: fix only has to use the ladder; its index may be an expression,
152    // since "one rung above the last" cannot be read statically.
153    if (kind === 'fix' && spreads.length > 0) continue
154
155    // Every rung sets both model and effort, so the last one written wins.
156    const rungIndex = rungIndexes.at(-1)
157    const rung = rungIndex !== undefined ? RUNG_STRENGTH[rungIndex] : undefined
158    const explicitCls = model?.kind === 'literal' ? (modelClass(model.value) ?? undefined) : undefined
159    const explicitEffort = effort?.kind === 'literal' && isEffort(effort.value) ? effort.value : undefined
160    const unreadableModel = model !== undefined && model.kind !== 'literal'
161
162    if (kind === 'fix') {
163      const read = bounds(explicitCls, explicitEffort, rung)
164      if (unreadableModel || read === undefined || compareStrength(read.weaker, KIND_RULES.fix.floor) < 0) add('R5', FIX_UP)
165      continue
166    }
167
168    // R1: the weight must be readable.
169    if (unreadableModel) {
170      add('R1', MODEL_LITERAL)
171      continue
172    }
173    // An effort written as an expression could override a rung at run time.
174    if (effort !== undefined && effort.kind !== 'literal') {
175      add('R1', EFFORT_LITERAL)
176      continue
177    }
178    // An expression index could override the explicit keys at run time.
179    if (spreads.some(spread => !NUMERIC_RUNG.test(spread))) {
180      add('R1', INDEX_LITERAL)
181      continue
182    }
183    const read = bounds(explicitCls, explicitEffort, rung)
184    if (read === undefined) {
185      add('R1', EFFORT_MISSING)
186      continue
187    }
188
189    // R2 reads the weakest pairing against the floor, R3 the strongest
190    // against the ceiling (only chore has one).
191    const { weaker, stronger } = read
192    const { floor, ceiling } = KIND_RULES[kind]
193    if (compareStrength(weaker, floor) < 0) {
194      add('R2', `${kind} の下限は ${formatStrength(floor)} です(現在 ${formatStrength(weaker)})`)
195    }
196    if (ceiling !== undefined && compareStrength(stronger, ceiling) > 0) {
197      const inherits = model === undefined && rung === undefined ? INHERITS : ''
198      add('R3', `${kind} の上限は ${formatStrength(ceiling)} です(現在 ${formatStrength(stronger)}${inherits})`)
199    }
200  }
201  return out
202}
203
204export function formatDeny(violations: Violation[]): string {
205  return [
206    `quality-router: Workflow を差し戻しました(${violations.length}件)。`,
207    ...violations.map(v => `- ${v.ordinal}番目の agent(${v.label ?? 'ラベルなし'}): ${v.message}`),
208    '直したら Workflow をもう一度呼んでください。種類と段(LADDER)の定義はシステムプロンプトの quality-router 節にあります。',
209  ].join('\n')
210}
211
212// Session-written scripts live in ~/.claude/projects/<project>/<session>/workflows/scripts/.
213export function isSessionScriptPath(path: string, sessionId: string): boolean {
214  return path.replace(/\\/g, '/').includes(`/${sessionId}/workflows/scripts/`)
215}
216
hooks/guidance.ts 31 lines
1import type { Top } from '../types'
2import { ladder } from './levels'
3
4// The section prompt.compose adds while the gate is on. English,
5// because the main loop is its reader.
6export function guidanceText(top: Top): string {
7  const rungs = ladder(top)
8    .map((rung, i) => `  { model: '${rung.model}', effort: '${rung.effort}' }, // ${i}`)
9    .join('\n')
10  return [
11    '# quality-router: staffing Workflow agents',
12    'Every agent() call in a Workflow script carries a label that starts with its kind, and a weight the gate can read: effort as a string, or ...LADDER[n] with a numeric n. quality-router reads the script when Workflow is called and sends it back when a rule is broken.',
13    "Write the options of each agent() call inline as an object literal, with the label as a string or template literal whose kind prefix comes before any ${…}: options held in a variable cannot be checked and are sent back, and so is any spread but ...LADDER[n]. Write either model and effort or a rung: when both are written, the gate checks whichever could win.",
14    'Strength compares the model tier first (haiku < sonnet < top: opus or fable), then effort. Omitting model inherits the session model, which counts as the top tier.',
15    "- chore: mechanical work (listing, counting, collecting, formatting, committing). From haiku·low up to sonnet·medium, e.g. ...LADDER[1]. Leaving model out inherits the top tier and is sent back.",
16    '- impl: implementing and fixing. At least sonnet·medium; heavier work on the top model.',
17    '- investigate: research, root-cause analysis, reports someone relies on. At least sonnet·medium.',
18    '- verify: checks with mechanical evidence (running tests or builds, matching facts against a source). At least sonnet·high.',
19    '- review: judging whether a diff or design is good. At least the top model at high.',
20    '- refute: trying to break a conclusion. At least the top model at high.',
21    '- decide: choosing a direction or a pass/fail verdict. At least the top model at high.',
22    "- fix: a redo after a quality failure. Move at least one rung up with ...LADDER[tier + 1] (only fix: may use an expression index) and pass the reviewer's findings verbatim.",
23    'Mechanical failures (timeout, API error, malformed output) retry on the same rung.',
24    'After two escalations without a pass, end the workflow and hand the decision back to the user.',
25    'Declare the ladder in the script exactly as:',
26    'const LADDER = [',
27    rungs,
28    ']',
29  ].join('\n')
30}
31
hooks/levels.ts 129 lines
1import type { Effort, Level, Top } from '../types'
2
3// The four levels the main loop runs at, lowest first.
4export const LEVELS: readonly Level[] = ['medium', 'high', 'xhigh', 'max']
5export const EFFORTS: readonly Effort[] = ['low', 'medium', 'high', 'xhigh', 'max']
6
7// turn.step needs full ids. agent.spawn and Workflow agent() take aliases,
8// except Fable, whose alias is unverified.
9export const TOP_IDS: Readonly<Record<Top, string>> = {
10  opus: 'claude-opus-5-5',
11  fable: 'claude-fable-5-1',
12}
13export const TOP_SPAWN: Readonly<Record<Top, string>> = {
14  opus: 'opus',
15  fable: 'claude-fable-5-1',
16}
17
18export function isLevel(value: unknown): value is Level {
19  return typeof value === 'string' && (LEVELS as readonly string[]).includes(value)
20}
21
22export function isEffort(value: unknown): value is Effort {
23  return typeof value === 'string' && (EFFORTS as readonly string[]).includes(value)
24}
25
26export function isTop(value: unknown): value is Top {
27  return value === 'opus' || value === 'fable'
28}
29
30export function rank(level: Level): number {
31  return LEVELS.indexOf(level)
32}
33
34// The session's effort as a step names it, read as one of our levels.
35export function asLevel(effort: string | number | undefined): Level | null {
36  if (effort === 'low') return 'medium'
37  return isLevel(effort) ? effort : null
38}
39
40export function maxLevel(...levels: Array<Level | null | undefined>): Level | null {
41  let best: Level | null = null
42  for (const level of levels) {
43    if (level && (best === null || rank(level) > rank(best))) best = level
44  }
45  return best
46}
47
48export function minLevel(a: Level, b: Level): Level {
49  return rank(a) <= rank(b) ? a : b
50}
51
52export function up(level: Level, steps = 1): Level {
53  return LEVELS[Math.min(LEVELS.length - 1, rank(level) + steps)] ?? 'max'
54}
55
56export function down(level: Level): Level {
57  return LEVELS[Math.max(0, rank(level) - 1)] ?? 'medium'
58}
59
60export type Bounds = { floor: Level; ceiling: Level; skillFloor: Level | null }
61
62// level = min(ceiling, max(floor, skillFloor, base)).
63export function applyBounds(base: Level, bounds: Bounds): Level {
64  return minLevel(bounds.ceiling, maxLevel(bounds.floor, bounds.skillFloor, base) ?? base)
65}
66
67// Which bound set the level applyBounds returns (the log records
68// it): the ceiling when it cut, else the floor that lifted the base, the skill
69// floor when both lift it as far. null: the base went out as it was.
70export type Bound = 'floor' | 'skill' | 'ceiling'
71
72export function boundOf(base: Level, bounds: Bounds): Bound | null {
73  const level = applyBounds(base, bounds)
74  const wanted = maxLevel(bounds.floor, bounds.skillFloor, base) ?? base
75  if (rank(level) < rank(wanted)) return 'ceiling'
76  if (level === base) return null
77  return bounds.skillFloor === level ? 'skill' : 'floor'
78}
79
80export function isTopFamily(model: string): boolean {
81  return /^claude-(opus|fable|mythos)-/.test(model) || model === 'opus' || model === 'fable'
82}
83
84export type Rung = { model: string; effort: Effort }
85
86// The ladder: model first, then effort; the last rung is the top model.
87export function ladder(top: Top): readonly Rung[] {
88  return [
89    { model: 'haiku', effort: 'low' },
90    { model: 'sonnet', effort: 'medium' },
91    { model: 'opus', effort: 'high' },
92    { model: 'opus', effort: 'xhigh' },
93    { model: TOP_SPAWN[top], effort: 'max' },
94  ]
95}
96
97// Strength: the model tier first (haiku < sonnet < top), then effort.
98// An omitted model inherits the main model, which the guard keeps top.
99export type ModelClass = 1 | 2 | 3
100export type Strength = { cls: ModelClass; effort: Effort }
101
102export function modelClass(model: string | undefined): ModelClass | null {
103  if (model === undefined) return 3
104  if (model === 'haiku' || /^claude-haiku-/.test(model)) return 1
105  if (model === 'sonnet' || /^claude-sonnet-/.test(model)) return 2
106  if (isTopFamily(model)) return 3
107  return null
108}
109
110export function compareStrength(a: Strength, b: Strength): number {
111  if (a.cls !== b.cls) return a.cls - b.cls
112  return EFFORTS.indexOf(a.effort) - EFFORTS.indexOf(b.effort)
113}
114
115// The ladder's rungs as strengths; the top rung is the top model whatever it is.
116export const RUNG_STRENGTH: readonly Strength[] = [
117  { cls: 1, effort: 'low' },
118  { cls: 2, effort: 'medium' },
119  { cls: 3, effort: 'high' },
120  { cls: 3, effort: 'xhigh' },
121  { cls: 3, effort: 'max' },
122]
123
124const CLASS_NAMES: Readonly<Record<ModelClass, string>> = { 1: 'haiku', 2: 'sonnet', 3: '上位' }
125
126export function formatStrength(s: Strength): string {
127  return `${CLASS_NAMES[s.cls]}・${s.effort}`
128}
129
hooks/main-effort.ts 130 lines
1import type { Level, TurnSignals } from '../types'
2import { LEVELS, down, maxLevel, rank } from './levels'
3
4// Pure judgment of the main loop's level. Adapted from effort-router v0.3.0
5// (hooks/route.ts, in this marketplace): compose, compact, the one-word labels and the
6// one-level-a-turn drop; max, the second-judgment confirm and the judged
7// value are new here.
8
9// The classifier must answer a label verbatim, so labels are single words
10// and their definitions lead the text it reads.
11export const LABELS: readonly string[] = LEVELS
12
13// Set from docs/decisions/effort-cache.md (Task 1).
14export const DOWN_CONFIRM_TURNS: 1 | 2 = 1
15export const MID_TURN_BUMP = true
16
17const RUBRIC = [
18  '[rubric] Judge how much reasoning the [request] needs, reading it in the context of [prev]:',
19  'medium = no code change or a one-spot text fix: a quick answer, a confirmation, a typo, a commit, thanks or closing;',
20  'high = an ordinary code change or explanation: a rename, adding a test, explaining or reviewing code;',
21  'xhigh = anything that must find a cause (a failing test, an error, a stack trace, "it does not work"), design or planning, a multi-file change, deep analysis, or a follow-up that continues such work (implementing an option just chosen from a design, a fix that still fails);',
22  'max = a cause still not found after a fix already failed ([prev] tool_errors above 0 and the same problem goes on), an irreversible design, migration or data decision, or an explicit request for the deepest possible analysis.',
23].join(' ')
24
25// Prompts this short ("続けて", "OK") skip the classifier and inherit.
26export const MIN_CHARS = 20
27
28const PREV_REQUEST_CHARS = 100
29const ANSWER_TAIL_CHARS = 300
30const REQUEST_HEAD_CHARS = 600
31const REQUEST_TAIL_CHARS = 200
32
33export function levelOf(label: string | undefined): Level | undefined {
34  const first = label?.split(':')[0]?.trim()
35  return LEVELS.find(level => level === first)
36}
37
38// Code and stack traces become markers: the signal stays, the tokens go.
39export function compact(text: string): string {
40  return text
41    .replace(/```(\w*)\n([\s\S]*?)```/g, (_, lang: string, body: string) => {
42      const lines = body.split('\n').filter(Boolean).length
43      return `[code${lang ? `: ${lang}` : ''} ${lines} lines]`
44    })
45    // A frame: a JavaScript `at …` line, or a Python `File "…", line N` line
46    // with the source line(s) Python prints under it, indented deeper.
47    .replace(/(?:^[ \t]*(?:at .+|File ".+", line \d+.*(?:\n[ \t]{4,}\S.*){0,2})\n?){3,}/gm, run => {
48      const lines = run.split('\n').filter(Boolean).length
49      return `[stack trace: ${lines} lines]\n`
50    })
51    .replace(/[ \t]+/g, ' ')
52    .replace(/\n{3,}/g, '\n\n')
53    .trim()
54}
55
56function oneLine(text: string): string {
57  return text.replace(/\s+/g, ' ').trim()
58}
59
60function clip(text: string): string {
61  if (text.length <= REQUEST_HEAD_CHARS + REQUEST_TAIL_CHARS) return text
62  return `${text.slice(0, REQUEST_HEAD_CHARS)} … ${text.slice(-REQUEST_TAIL_CHARS)}`
63}
64
65export function tail(text: string, chars = ANSWER_TAIL_CHARS): string {
66  const line = oneLine(compact(text))
67  return line.length <= chars ? line : `…${line.slice(-chars)}`
68}
69
70export function head(text: string, chars = PREV_REQUEST_CHARS): string {
71  const line = oneLine(compact(text))
72  return line.length <= chars ? line : `${line.slice(0, chars)}…`
73}
74
75// What the classifier reads: the previous turn as signals and two short
76// excerpts, then this request, compacted and clipped.
77export function compose(
78  request: string,
79  prev: TurnSignals | null,
80  recent: ReadonlyArray<Level | null>,
81): string {
82  const lines: string[] = [RUBRIC]
83  if (prev) {
84    const max3 = maxLevel(...recent) ?? 'default'
85    lines.push(
86      `[prev] effort=${prev.level ?? 'default'} steps=${prev.steps} tools=${prev.tools} tool_errors=${prev.toolErrors} dur=${Math.round(prev.durationMs / 1000)}s max3=${max3}`,
87      `[prev_request] ${prev.request}`,
88      `[prev_answer_tail] ${prev.answerTail}`,
89    )
90  }
91  lines.push(`[request chars=${request.length}] ${clip(compact(request))}`)
92  return lines.join('\n')
93}
94
95// Up at once; down one level a turn at most. With confirm 2, down only when
96// the previous turn was also judged lower than the level it ran at.
97export function limitDown(
98  judged: Level,
99  prev: TurnSignals | null,
100  confirm: 1 | 2 = DOWN_CONFIRM_TURNS,
101): Level {
102  const prevLevel = prev?.level ?? null
103  if (prevLevel === null || rank(judged) >= rank(prevLevel)) return judged
104  if (confirm === 2 && !(prev?.judged && rank(prev.judged) < rank(prevLevel))) return prevLevel
105  return maxLevel(judged, down(prevLevel)) ?? judged
106}
107
108// 'late': no judgment had arrived (turn.step stopped waiting at its bound, or
109// A was off when the turn began); the steps inherited the previous level.
110export type Why = 'auto' | 'short' | 'unsure' | 'error' | 'late'
111
112// base null: the session's own effort, resolved at the step.
113export type Decision = { base: Level | null; judged: Level | null; why: Why }
114
115export function decide(
116  answer: string | undefined | Error,
117  prev: TurnSignals | null,
118  confirm: 1 | 2 = DOWN_CONFIRM_TURNS,
119): Decision {
120  const kept = prev?.level ?? null
121  if (answer instanceof Error) return { base: kept, judged: null, why: 'error' }
122  const judged = levelOf(answer)
123  if (judged === undefined) return { base: kept, judged: null, why: 'unsure' }
124  return { base: limitDown(judged, prev, confirm), judged, why: 'auto' }
125}
126
127export function shortDecision(prev: TurnSignals | null): Decision {
128  return { base: prev?.level ?? null, judged: null, why: 'short' }
129}
130
hooks/record.ts 127 lines
1import type { Effort, Level } from '../types'
2import type { AgentCall } from './gate-lexer'
3import { valueOf } from './gate-rules'
4import type { Violation } from './gate-rules'
5import { kindOf } from './kinds'
6import type { Bound } from './levels'
7import type { Why } from './main-effort'
8import type { Settings } from './settings'
9import type { SubKind, Weight } from './sub-route'
10
11// Records: one JSON line per record. No request, description,
12// label or answer text: the transcript's row uuids lead back to them.
13
14export const PART_LIMIT = 3_500_000
15
16export type Common = { v: 1; at: string; session: string; turn: string | null; qr: string | null; tune: string | null }
17export type Tokens = { input: number; output: number; cacheRead: number }
18export type EndReason = 'answer' | 'aborted' | 'refusal' | 'error'
19
20export type MainRecord = Common & {
21  kind: 'main'
22  promptUuid: string | null
23  answerUuid: string | null
24  promptAt: string | null
25  promptChars: number
26  prev: Level | null
27  judged: Level | null
28  level: Level | null
29  why: Why
30  bound: Bound | null
31  skill: string | null
32  guarded: boolean
33  judgeMs: number | null
34  steps: number
35  tools: number
36  toolErrors: number
37  bump: boolean
38  durationMs: number
39  tokens: Tokens | null
40  answeredBy: string | null
41  reason: EndReason
42}
43
44export type SubWhy = 'auto' | 'retry' | 'unsure' | 'explicit' | 'off'
45type ChildResult = { steps: number; toolErrors: number; tokens: Tokens | null; durationMs: number; reason: EndReason }
46
47export type SubRecord = Common &
48  ChildResult & {
49    kind: 'sub'
50    agentId: string
51    type: string
52    descChars: number
53    kindJudged: SubKind | null
54    weightJudged: Weight | null
55    model: string | null
56    effort: Effort | null
57    why: SubWhy
58    retries: number
59  }
60
61export type GateCounts = { kinds: Record<string, number>; rules: Record<string, number>; fixes: number }
62export type WfGateRecord = Common &
63  GateCounts & {
64    kind: 'wf'
65    phase: 'gate'
66    verdict: 'pass' | 'deny' | 'warn' | 'unchecked'
67    source: 'script' | 'scriptPath' | 'name'
68    calls: number
69  }
70export type WfChildRecord = Common & ChildResult & { kind: 'wf'; phase: 'child'; agentId: string; model: string | null; effort: string | null }
71
72export type Dir = 'up' | 'down' | 'same' | 'unknown'
73export type OverrideSignal = { type: 'override'; what: 'effort' | 'model'; from: string | null; to: string | null; dir: Dir }
74export type FeedbackSignal = { type: 'feedback'; dir: 'up' | 'down'; target: 'main' | 'sub'; targetTurn: string | null; targetAgent: string | null }
75export type Retry3Signal = { type: 'retry3'; agentId: string | null }
76// One intersection per signal, so RecordBody distributes over them.
77export type SignalRecord =
78  | (Common & { kind: 'signal' } & OverrideSignal)
79  | (Common & { kind: 'signal' } & FeedbackSignal)
80  | (Common & { kind: 'signal' } & Retry3Signal)
81
82export type SessionRecord = Common & {
83  kind: 'session'
84  event: 'start' | 'settings'
85  root: string | null
86  cwd: string | null
87  settings: Settings
88  paused: boolean
89}
90
91export type QrRecord = MainRecord | SubRecord | WfGateRecord | WfChildRecord | SignalRecord | SessionRecord
92// A record without the shared fields, which the writer adds.
93export type RecordBody<R = QrRecord> = R extends QrRecord ? Omit<R, keyof Common> : never
94
95type UsageLike = { input_tokens: number; output_tokens: number; cache_read_input_tokens: number }
96
97export function common(at: number, session: string, turn: string | null, qr: string | null): Common {
98  return { v: 1, at: new Date(at).toISOString(), session, turn, qr, tune: null }
99}
100
101export function tokensOf(usage: UsageLike | undefined): Tokens | null {
102  if (!usage) return null
103  return { input: usage.input_tokens, output: usage.output_tokens, cacheRead: usage.cache_read_input_tokens }
104}
105
106// The month is the UTC one of `at` (the session's first record), so a session
107// stays in one folder; a part past the first carries its number.
108export function recordPath(home: string, at: number, session: string, part: number): string {
109  const month = new Date(at).toISOString().slice(0, 7)
110  return `${home}/.claude/quality-router/log/${month}/${session}${part > 1 ? `.${part}` : ''}.jsonl`
111}
112
113// A label's kind from its literal head ('unknown' when it has none), the
114// violations by rule, and the fix: calls (each one a quality failure redone).
115export function gateCounts(calls: AgentCall[], violations: Violation[]): GateCounts {
116  const kinds: Record<string, number> = {}
117  for (const call of calls) {
118    const label = call.options.kind === 'object' ? valueOf(call.options.entries['label']) : undefined
119    const text = label?.kind === 'literal' ? label.value : label?.kind === 'template' ? label.head : ''
120    const kind = kindOf(text) ?? 'unknown'
121    kinds[kind] = (kinds[kind] ?? 0) + 1
122  }
123  const rules: Record<string, number> = {}
124  for (const v of violations) rules[v.rule] = (rules[v.rule] ?? 0) + 1
125  return { kinds, rules, fixes: kinds['fix'] ?? 0 }
126}
127
hooks/signals.ts 39 lines
1import { EFFORTS, isEffort, modelClass } from './levels'
2import type { Dir, OverrideSignal } from './record'
3
4// What a manual /effort or /model, and /qr up or down, say.
5
6const compare = (a: number, b: number): Dir => (a > b ? 'up' : a < b ? 'down' : 'same')
7
8// `from` is the level quality-router last sent: the change is read against it.
9export function effortOverride(args: string, from: string | null): OverrideSignal {
10  const to = args.trim().toLowerCase() || null
11  const dir = to !== null && from !== null && isEffort(to) && isEffort(from) ? compare(EFFORTS.indexOf(to), EFFORTS.indexOf(from)) : 'unknown'
12  return { type: 'override', what: 'effort', from, to, dir }
13}
14
15// An alias (`opus[1m]` keeps its bracket off) or a full id; `default` and
16// anything else unknown is null.
17export function modelClassOf(name: string | null): 1 | 2 | 3 | null {
18  if (!name) return null
19  const n = name.trim().toLowerCase().replace(/\[.*\]$/, '')
20  if (n === 'haiku') return 1
21  if (n === 'sonnet') return 2
22  if (n === 'opus' || n === 'fable') return 3
23  return n.startsWith('claude-') ? modelClass(n) : null
24}
25
26// `from` is the model that answered the main loop last.
27export function modelOverride(args: string, from: string | null): OverrideSignal {
28  const to = args.trim() || null
29  const a = modelClassOf(to)
30  const b = modelClassOf(from)
31  return { type: 'override', what: 'model', from, to, dir: a !== null && b !== null ? compare(a, b) : 'unknown' }
32}
33
34export function parseFeedback(verb: string, arg: string): { dir: 'up' | 'down'; target: 'main' | 'sub' } | null {
35  const dir = verb === 'up' || verb === '↑' ? 'up' : verb === 'down' || verb === '↓' ? 'down' : null
36  if (dir === null || (arg !== '' && arg !== 'sub')) return null
37  return { dir, target: arg === 'sub' ? 'sub' : 'main' }
38}
39
hooks/stats.ts 88 lines
1import { LEVELS } from './levels'
2import type { MainRecord, QrRecord, SignalRecord, SubRecord, WfGateRecord } from './record'
3import { SUB_KINDS } from './sub-route'
4
5// /qr stats, a few English lines over the period's records.
6
7export type Period = 'today' | '7d' | '30d'
8
9const DAY = 86_400_000
10const WAYS = ['auto', 'short', 'late', 'unsure', 'error'] as const
11
12export function periodOf(arg: string): Period | null {
13  if (arg === '') return '7d'
14  return arg === 'today' || arg === '7d' || arg === '30d' ? arg : null
15}
16
17// today: since local midnight; the others: rolling days.
18export function sinceOf(period: Period, now: number): number {
19  if (period === 'today') {
20    const d = new Date(now)
21    d.setHours(0, 0, 0, 0)
22    return d.getTime()
23  }
24  return now - (period === '7d' ? 7 : 30) * DAY
25}
26
27// Lines that are not version-1 records are skipped.
28export function parseRecords(text: string): QrRecord[] {
29  const out: QrRecord[] = []
30  for (const line of text.split('\n')) {
31    if (line.trim() === '') continue
32    try {
33      const o = JSON.parse(line) as { v?: unknown; kind?: unknown; at?: unknown }
34      if (o && o.v === 1 && typeof o.kind === 'string' && typeof o.at === 'string') out.push(o as QrRecord)
35    } catch {
36      // A torn or foreign line.
37    }
38  }
39  return out
40}
41
42const pct = (n: number, total: number) => `${Math.round((n * 100) / total)}%`
43const amount = (n: number) => (n >= 1_000_000 ? `${(n / 1_000_000).toFixed(1)}M` : `${Math.round(n / 1000)}k`)
44
45export function summarize(records: QrRecord[], period: Period): string {
46  const main = records.filter((r): r is MainRecord => r.kind === 'main')
47  const subs = records.filter((r): r is SubRecord => r.kind === 'sub')
48  const gates = records.filter((r): r is WfGateRecord => r.kind === 'wf' && r.phase === 'gate')
49  const signals = records.filter((r): r is SignalRecord => r.kind === 'signal')
50  if (main.length === 0 && subs.length === 0 && gates.length === 0) return `stats ${period}: まだ記録がありません。`
51
52  const lines = [`stats ${period}: main ${main.length} turns, sub ${subs.length}, workflow ${gates.length}`]
53  if (main.length > 0) {
54    const levels = LEVELS.map(l => `${l} ${pct(main.filter(r => r.level === l).length, main.length)}`).join(' · ')
55    const ways = WAYS.map(w => [w, main.filter(r => r.why === w).length] as const)
56      .filter(([, n]) => n > 0)
57      .map(([w, n]) => `${w} ${pct(n, main.length)}`)
58      .join(', ')
59    lines.push(`main: ${levels}  (${ways})`)
60  }
61
62  const count = (pred: (s: SignalRecord) => boolean) => signals.filter(pred).length
63  const overUp = count(s => s.type === 'override' && s.dir === 'up')
64  const qrUp = count(s => s.type === 'feedback' && s.dir === 'up')
65  const bumps = main.filter(r => r.bump).length
66  const errors = main.filter(r => !r.bump && r.toolErrors >= 2).length
67  const overDown = count(s => s.type === 'override' && s.dir === 'down')
68  const qrDown = count(s => s.type === 'feedback' && s.dir === 'down')
69  lines.push(
70    `shallow signals ${overUp + qrUp + bumps + errors} (override up ${overUp}, /qr up ${qrUp}, bump ${bumps}, tool errors ${errors})` +
71      ` · heavy signals ${overDown + qrDown} (override down ${overDown}, /qr down ${qrDown})`,
72  )
73
74  const byLevel = LEVELS.map(l => [l, main.filter(r => r.level === l).reduce((n, r) => n + (r.tokens?.output ?? 0), 0)] as const).filter(([, n]) => n > 0)
75  if (byLevel.length > 0) lines.push(`tokens by level (output): ${byLevel.map(([l, n]) => `${l} ${amount(n)}`).join(' · ')}`)
76
77  if (subs.length > 0) {
78    const kinds = SUB_KINDS.map(k => [k, subs.filter(r => r.kindJudged === k).length] as const).filter(([, n]) => n > 0)
79    const explicit = subs.filter(r => r.why === 'explicit' || r.why === 'off').length
80    const parts = [...kinds.map(([k, n]) => `${k} ${n}`), ...(explicit > 0 ? [`explicit ${explicit}`] : [])]
81    lines.push(`sub: ${parts.join(' · ')} (retry ${subs.filter(r => r.why === 'retry').length}, unsure ${subs.filter(r => r.why === 'unsure').length})`)
82  }
83  if (gates.length > 0) {
84    lines.push(`workflow: deny ${gates.filter(r => r.verdict === 'deny').length} · fix ${gates.reduce((n, r) => n + r.fixes, 0)}`)
85  }
86  return lines.join('\n')
87}
88
hooks/settings.ts 69 lines
1import type { Level, Top } from '../types'
2import { isLevel, isTop } from './levels'
3
4// Pure: no `$`. The loader follows `$` only into functions declared in the
5// file that holds the hooks, never across an import, so the reads and writes
6// of `$.store` and `$.settings` (loadSettings, saveSetting,
7// effortRouterInstalled) live at the top of register.ts, built on
8// SETTING_KEYS, parseSettings and mentionsEffortRouter here.
9
10export type Toggle = 'on' | 'off'
11export type Settings = { mode: Toggle; gate: Toggle; guard: Toggle; floor: Level; ceiling: Level; top: Top }
12
13// The guard is off by default: it moves the main loop to a top model, which
14// not every plan or budget wants. `/qr guard on` turns it on.
15export const DEFAULTS: Settings = { mode: 'on', gate: 'on', guard: 'off', floor: 'medium', ceiling: 'max', top: 'opus' }
16export const SETTING_KEYS: ReadonlyArray<keyof Settings> = ['mode', 'gate', 'guard', 'floor', 'ceiling', 'top']
17
18const isToggle = (value: unknown): value is Toggle => value === 'on' || value === 'off'
19
20export function parseSettings(values: Partial<Record<keyof Settings, unknown>>): Settings {
21  return {
22    mode: isToggle(values.mode) ? values.mode : DEFAULTS.mode,
23    gate: isToggle(values.gate) ? values.gate : DEFAULTS.gate,
24    guard: isToggle(values.guard) ? values.guard : DEFAULTS.guard,
25    floor: isLevel(values.floor) ? values.floor : DEFAULTS.floor,
26    ceiling: isLevel(values.ceiling) ? values.ceiling : DEFAULTS.ceiling,
27    top: isTop(values.top) ? values.top : DEFAULTS.top,
28  }
29}
30
31const record = (value: unknown): Record<string, unknown> =>
32  typeof value === 'object' && value !== null && !Array.isArray(value) ? (value as Record<string, unknown>) : {}
33
34// A CLAUDE_CODE_PLUGIN_DIRS list: `;` between entries on Windows, `:` on
35// POSIX. A drive letter's colon (`C:/...`, `C:\...`) belongs to its path.
36function pluginDirs(value: string): string[] {
37  const dirs: string[] = []
38  for (const raw of value.split(/[;:]/)) {
39    const part = raw.trim()
40    const prev = dirs.pop()
41    if (prev === undefined) dirs.push(part)
42    else if (/^[A-Za-z]$/.test(prev) && /^[\\/]/.test(part)) dirs.push(`${prev}:${part}`)
43    else dirs.push(prev, part)
44  }
45  return dirs.filter(dir => dir !== '')
46}
47
48// Installed means loaded as a plugin: a CLAUDE_CODE_PLUGIN_DIRS entry, or an
49// enabledPlugins key set to true. A path elsewhere (permissions) does not
50// count, nor does a plugin turned off. Takes ~/.claude/settings.json as text,
51// or as the object `$.settings.read({ source: 'user' })` answers.
52export function mentionsEffortRouter(settings: unknown): boolean {
53  let json: unknown = settings
54  if (typeof settings === 'string') {
55    try {
56      json = JSON.parse(settings)
57    } catch {
58      return false
59    }
60  }
61  const o = record(json)
62  const dirs = record(o.env).CLAUDE_CODE_PLUGIN_DIRS
63  const inDirs = typeof dirs === 'string' && pluginDirs(dirs).some(dir => /(^|[\\/])effort-router[\\/]?$/.test(dir))
64  const inPlugins = Object.entries(record(o.enabledPlugins)).some(
65    ([key, enabled]) => enabled === true && key.split('@')[0] === 'effort-router',
66  )
67  return inDirs || inPlugins
68}
69
hooks/skills.ts 20 lines
1import type { Level } from '../types'
2
3// A skill the main loop calls raises its floor until another
4// listed skill is called or the session ends.
5export const SKILL_FLOORS: Readonly<Record<string, Level>> = {
6  'superpowers:brainstorming': 'xhigh',
7  'superpowers:writing-plans': 'xhigh',
8  'superpowers:systematic-debugging': 'xhigh',
9  'code-review': 'xhigh',
10  'superpowers:requesting-code-review': 'xhigh',
11  'superpowers:receiving-code-review': 'xhigh',
12  'superpowers:executing-plans': 'high',
13  'superpowers:subagent-driven-development': 'high',
14  'superpowers:test-driven-development': 'high',
15}
16
17export function skillFloorFor(skill: string): Level | undefined {
18  return SKILL_FLOORS[skill.trim().replace(/^\//, '')]
19}
20