SLOPSHOPPER

effort-router

Picks low/medium/high/xhigh effort per prompt for Opus/Sonnet/Haiku 5.5+ and Fable/Mythos 5.1+ by how uncertain and costly the work is, raising it mid-turn…

newguardcommandstatusmodel
★ 2v0.10.0MITupdated 2026-10-08nextscape/ns-mods/mods/effort-router
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · effort-router
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /effort-router ⎿ effort-router: on (picks low/medium/high/xhigh each prompt for Opus/Sonnet 5.5+ and Fable/Mythos 5.1+; a typed /effort runs a ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts ⚠ effort-router: on, waiting for the first prompt
README

effort-router — プロンプトごとに effort を自動で選ぶ

English summary: effort-router is a Claude Code mod that picks the main thread's effort (low / medium / high / xhigh) for each prompt, judged by Claude Haiku from how uncertain and costly the work is, and raises it one level mid-turn when a turn runs long. It never changes the model. It applies only to Claude Opus / Sonnet / Haiku 5.5+ and Fable / Mythos 5.1+. Each judgment is one Haiku call on your plan or API key (about 1 second before the turn starts). Install: /plugin install effort-router --marketplace nextscape/ns-mods.

Claude Code の mod(関数フックで書くプラグイン)です。プロンプトごとに、メインスレッドの effort(推論の深さ)を low / medium / high / xhigh から選びます。簡単な依頼は速く安く、難しい依頼は深く考えて進めるのが目的です。

  • 判定は Claude Haiku が行います。軸は作業の量ではなく、正しい答えがどれだけ不確かか、見落としたときの損がどれだけ大きいかです(判定基準)。
  • 判定が低すぎたときは、ターンの途中で1段上げて補います(往復が10回を超えたとき、またはツールのエラーが2回出たとき)。
  • モデルは変えません。 モデルの振り分けは、サブエージェントの定義(~/.claude/agents/)に任せます。
  • 対象は Claude Opus / Sonnet / Haiku 5.5 以降と Fable / Mythos 5.1 以降です。それ以外のモデル、サブエージェントのリクエストには何もしません。

導入する前に、費用・待ち時間・送るデータを確認してください。判定のたびに Haiku を1回呼びます。

導入

Claude Code(ターミナル)で次を実行します。

/plugin install effort-router --marketplace nextscape/ns-mods

新しいセッションを開くと読み込まれます。

項目内容
Claude Code2.1.287 以降(mod が既定で有効になる版)。2.1.291 で動作を確認
モデルメインのモデルが Opus / Sonnet / Haiku 5.5 以降、または Fable / Mythos 5.1 以降。ほかのモデルでは effort を書き換えず、判定もしません
OS問いません

使い方

入れれば自動で動きます。操作が要るのは、止めたいときと結果を確かめたいときだけです。

コマンド動作
/effort-router現在の状態(on / off)を表示する
/effort-router on自動判定を有効にする(既定)
/effort-router off自動判定をやめる。組み込みの /effort の設定がそのまま効く
/effort-router log段階ごとのターン数と平均所要時間、直近10件の判定を表示する
/effort-router log clear判定ログを消す
/effort-router eval評価セット36件を実際の Haiku で判定し、一致数と低すぎる判定の数を表示する(Haiku を36回呼ぶ)

on / off は次のセッションにも引き継がれます。

組み込みの /effort との関係

  • /effort で変えると、そのターン(ターンの合間なら次のターン)だけ設定どおりの effort で動きます。その次のプロンプトからは、また自動で判定します。
  • ずっと固定したいときは、/effort-router off にしてから /effort を使ってください。
  • /effort medium のように段階名を付けると、その段階が以後のセッションの既定として保存されます。このセッションだけ変えたいときは、/effort のスライダーで s を押します(公式ドキュメント)。
  • /effort auto は自動判定ではありません。保存した段階を消して、モデル既定の段階(Opus 5.5 なら medium)に戻すだけです。

ステータスライン

先頭に effort-router: が付いて表示されます。

表示意味
on, waiting for the first prompt有効だが、まだ判定していない(/effort-router on の直後にも出る)
high判定の結果、high で実行している
low (choice 3) / low (choice 1,3)前の回答の選択肢が番号で選ばれたので、その本文で判定した
high (reply to question)前の回答の質問への同意の返事として、提案された作業の重さで判定した
xhigh → medium (asked: 進め方)ターンの途中で AskUserQuestion(見出し「進め方」)に回答したので、判定し直して xhigh から medium に変えた
medium → high (raised after 10 requests) / (raised after 2 tool errors)ターンが長引いたか、エラーが続いたので、途中で1段上げた
high (task notice, kept)バックグラウンド作業の完了通知なので、前ターンの段階を引き継いでいる
high (short prompt, kept)「OK」「続けて」のような同意だけの短い入力で、前の回答が質問や選択肢で終わっていないので、前ターンの段階を引き継いでいる
high (unsure, kept) / high (error, kept)判定がどのラベルにも当てはまらなかったか、失敗したため、前ターンの段階を引き継いでいる
xhigh (session effort, …)mod は書き換えず、セッション本来の effort で実行している
high (set by /effort) / max (set by /effort)/effort で変えた。そのターン(ターンの合間なら次のターン)は設定どおりに動く
changed (set by /effort)スライダーで変えたため、値がまだ分からない。次のリクエストで実際の値に置き換わる(スライダーを変えずに閉じたときも出る)
not routed (claude-haiku-…)対象外のモデルなので、判定も書き換えもしていない
off (/effort applies)無効

費用・待ち時間・送るデータ

項目内容
Haiku を呼ぶとき20文字以上の依頼、20文字未満でも同意の語だけではない入力(「全体を監査して」「やめて」など)、前の回答の選択肢を番号で選んだとき、前の回答の質問への同意の返事、メインスレッドの AskUserQuestion への回答。1回につき1回呼ぶ
呼ばないとき20文字未満で同意の語だけの入力(「OK」「はい、進めて」「続けて」など)のうち、前の回答が質問か選択肢で終わっていないもの、バックグラウンド作業の完了通知、/effort で変えたターン、対象外のモデル、off のとき
費用mod のモデル呼び出しは、利用者のプランまたは API キーを使います(公式ドキュメント)。1回の入力は最大でも1,000トークン程度の見込みです(日本語のトークン数を経験則で見積もった値で、実測ではありません)
待ち時間判定が終わるまでターンは始まりません。作者の環境で評価セットを判定したときの実測は、1回あたり約0.7〜1.1秒でした。AskUserQuestion への回答のあとは、約1〜3秒遅れて次の処理に進みます。判定には制限時間を付けられないため、Haiku の応答が遅いとそのぶん待ちます
Haiku に送るもの判定基準(固定)、今回の依頼(冒頭600文字と末尾200文字。コードブロックとスタックトレースは行数だけの印に置き換える)、前のターンの依頼の冒頭100文字、前の回答の末尾300文字、前のターンの数値(往復回数、ツール数、エラー数、所要時間)、番号で選んだ選択肢の本文。AskUserQuestion では質問文と選んだ選択肢
送らないもの会話履歴、ファイルの中身、ツールの結果
手元に残るもの$.state(このセッションの間だけ): 前のターンの信号(依頼の冒頭100文字、回答の末尾300文字、選択肢)。$.store(同じ PC のすべてのセッションで共有): on / off と、直近200ターンの判定ログ(依頼の冒頭40文字を含む)
消し方/effort-router log clear で判定ログを消せます。アンインストールで $.store が消えるかは公式ドキュメントに記載がないため、気になる場合はアンインストールの前に消してください

判定基準

段階対象
low見落としの余地がない作業。雑談、確認、お礼、翻訳、コミット、手順がすべて決まっている機械的な修正(誤字、リネーム、移動)
medium方針が決まっている普通の作業。決めた方針や設計の実装、テストの追加、コードの説明、調べもの
high正しい答えが不確かな作業。不具合の原因調査、設計案の比較、差分のレビュー、あいまいな要件の整理
xhigh難しく、見落としの損が大きい作業。コードベース全体のレビューや監査、プロジェクト全体からの該当箇所の洗い出し、全体にわたる移行、抜本的な設計の見直し、直したのに直らない不具合、並行処理・セキュリティ・データ消失に関わる作業、長時間の自律作業

max は選びません。基準は Opus 5.5 に合わせて作り、対象のほかのモデルにも同じ基準を使います。Opus 6 や Sonnet 5.6 のような後の版も自動で対象になりますが、その版で基準が妥当かは確かめていません。

基準の考え方(2026-10-07 時点):

  • 公式の説明では、Opus 5.5 の既定は medium です。medium から始め、xhigh と max は効果を測った作業にだけ使うよう勧めています。xhigh は長時間のエージェント作業やコーディング作業向けとされています(公式ドキュメント)。
  • 上の段階ほど、品質の伸びは小さく、費用は増えます。そのため既定を medium 寄りにし、上げるのは不確かさか損の大きさがはっきりしている作業に限っています。
  • 「effort を上げると抜け漏れが減る」ことを直接測ったデータは見つかりませんでした。xhigh の「広さ」の項目(全体レビュー、洗い出し)は、実測のない仮説です。

対象を Opus / Sonnet / Haiku 5.5 以降と Fable / Mythos 5.1 以降に絞っているのは、公式ドキュメントで「effort を変えてもプロンプトキャッシュが保たれる」とされているモデルだからです。それ以外のモデルでは、切り替えのたびにキャッシュが作り直されて費用が増えます。

仕組み

判定の流れ

turn.start
 ├ 対象外のモデル → 判定しない(最初のターンは $.session.model() でモデルを確かめる)
 ├ ターンの合間に /effort を実行していた → 判定せず、セッションの effort(設定どおり)で動かす
 ├ バックグラウンド作業の完了通知 → 前ターンの段階を引き継ぐ(判定は呼ばない)
 ├ 前の回答の選択肢を番号で選んだ("2" "3で" "1と3でお願いします")
 │            → classify([rubric] + [choice] + … + [selected] 選択肢の本文)
 ├ 入力が20文字未満で同意の語だけ("OK" "はい、進めて" "続けて")で、前の回答が質問か選択肢で終わっている
 │            → classify([rubric] + [reply] + 前ターンの信号と抜粋 + 「前の回答の提案を進める」という依頼)
 ├ 入力が20文字未満で同意の語だけで、前の回答が質問か選択肢で終わっていない
 │            → 前ターンの段階を引き継ぐ(判定は呼ばない。最初のターンなら、セッション本来の effort のまま)
 └ それ以外 → classify([rubric] + 前ターンの信号と抜粋 + 今回の依頼)
   判定結果をそのまま使う。判定に失敗したか、どのラベルにも当てはまらなければ前ターンの段階
turn.step     対象モデルのメインスレッドのリクエストの effort を書き換え、往復回数を数える
              判定から往復が10回を超えたか、ツールのエラーが2回出たら1段上げる(1回の判定につき1段まで。上限 xhigh)
tool.call     メインスレッドのツール呼び出し数とエラー数を数える
              AskUserQuestion の回答(メインスレッドのみ)
              → classify([rubric] + [choice] か [reply] + 今のターンの信号と質問文 + 回答 + [selected] label — description)
              → 次の turn.step から、ターンの残りをその段階で送る(途中の引き上げはここから数え直す。判定に失敗したら変えない)
turn.complete ターンの最後の回答から選択肢と「質問で終わったか」を抜き出し(最後の回答に選択肢がなければ、ターンの途中で選択肢を出したテキストから読む)、次のターン用の信号として $.state に保存する。ログを $.store に追記する

判定に渡す入力

会話履歴は読みません。mod が自分で集めた値だけを渡すので、会話が長くなっても入力量はほぼ一定です。

[rubric] 4段階の判定基準(固定)
[prev] effort=xhigh steps=9 tools=14 tool_errors=2 dur=210s max3=xhigh
[prev_request] 前ターンの依頼の冒頭100文字
[prev_answer_tail] 前ターンの回答の末尾300文字
[request chars=1234] 今回の依頼の冒頭600文字 … 末尾200文字
[selected] 3: 選んだ選択肢の本文(番号で選んだときだけ)
  • コードブロックは [code: ts 40 lines] に、スタックトレースは [stack trace: 25 lines] に置き換えます。トークン数を減らしつつ、コードやエラーを含むという信号は残します。
  • 番号で選んだときは [choice]、質問への同意の返事のときは [reply] の判定規則を [rubric] の直後に足します。この2つを [rubric] に常に入れると、普通の依頼まで低めに判定される傾向が評価セットで出たため、該当するときだけ足しています。

選択肢の読み取り

  • ターンの最後の回答から、行頭の 1. 1) (1) ① 1. A. 案1: option 1: 1、(箇条書きの - * が前に付いてもよい)と、表の行 | 1 | … | を拾います。最後のひと続き(1 か a で始まり直す)を選択肢として保存します。最後の回答に選択肢がなければ、ターンの途中で選択肢を含んでいた最後のテキストから読みます。選択肢を出したあとにツールを呼び、短い文で終えたターンに備えるためです。1件150文字、最大10件までです。コードブロックの中の行は拾いません。
  • 入力から番号と「で」「番」「お願いします」などの言葉を取り除いて何も残らなければ、選択とみなします。保存した選択肢にない番号が含まれていれば、選択とはみなしません。
  • 2で、テストも追加して のように番号以外の指示が付くと、選択とはみなしません。普通の依頼として判定します。

同意だけの入力の見分け方

  • 句読点と空白を除いた入力が、「OK」「はい」「了解」「お願いします」「進めて」「続けて」「それで」「いい」「LGTM」「go ahead」などの語と「です」「ください」「で」「よ」「ね」だけでできていれば、同意とみなします。
  • 「いいえ」「やめて」「もういい」のような否定や、作業を名指す語(「監査して」「コミットして」)が含まれれば、同意とはみなさず普通の依頼として判定します。否定の返事を「前の回答の提案を進める依頼」として渡さないためです。
  • 「いいです」「大丈夫です」は断りの意味にもなります。「いいです」は同意とみなすので、断りのつもりでも提案された作業の重さで判定されることがあります(費用が増えるだけで、品質は下がりません)。「大丈夫」は同意の語に含めていません。

classify のラベルは1語にする

ラベルは low / medium / high / xhigh の1語にし、判定基準は入力の冒頭の [rubric] に書いています。ラベルに定義文を入れると、Haiku が長いラベル文を言い換えて返すため完全一致の照合に失敗し、評価セットでほぼ全件が外れました。

状態とログ

保存先キー内容有効範囲
$.stateprev / recent前ターンの信号(選択肢と「質問で終わったか」を含む)、直近3ターンの段階セッション内
$.storemodeon / offセッションをまたいで残る。同じ PC のすべてのセッションで共有
$.storelog直近200ターンの判定ログ(依頼の冒頭40文字を含む)同上。/effort-router log clear で消える

制約と既知の課題

  • 判定の一致率は、0.8.0 時点の評価セット31件で約90%(5回の実行で 27〜29/31)です。外れはすべて1段違い(xhigh を high、medium を low と判定するなど)で、2段以上の外れはありませんでした。低すぎる判定は、途中の引き上げで補います。
  • 評価セットは自作で、判定基準もそれを見ながら調整しました。実際の使い方での精度は /effort-router log で確かめる必要があります。
  • 途中の引き上げは、往復が多いだけの作業(量は多いが単純な作業)でも1段上がります。閾値(10回、エラー2回)は仮の値です。
  • サブエージェントのツール呼び出しと往復は数えません。サブエージェントの中の AskUserQuestion も対象外です(この除外は、テストからサブエージェントとして呼べないため結合テストでは確かめていません)。
  • AskUserQuestion で作業内容と関係の薄い質問(ファイル名の選択など)に答えても判定し直すので、effort が不要に上下することがあります。
  • 番号付きの箇条書きが選択肢ではなく手順や変更点の一覧でも、番号で返せば選択肢として扱います(本文を見て判定するので、大きく外れることはない見込みです)。
  • 見出し(### 1. … ### 案1)、隅付き括弧(【案1】)、区切り記号のない形(案1 … 案1 …)、A案 の形で書いた選択肢は拾いません。番号で選んでも選択とはみなさず、本文で判定します。見出しの番号は章立てにも使われるため、選択肢と区別しにくく、拾う対象に加えていません。
  • 回答の末尾に ? や「ますか」があれば、質問で終わったとみなします。疑問文を引用しただけの回答でも、次の同意の返事は「質問への返事」として判定されます。
  • 逆に、提案が疑問文で終わっていない回答(「必要なら〜します。」「次は〜しましょう。」)のあとに「OK」と返すと、質問とはみなさず前ターンの段階を引き継ぎます。前のターンが軽く、提案が重い作業だと、低い段階のまま始まります(途中の引き上げは1段まで)。
  • 判定ログは $.store に読み書きするため、複数のセッションがほぼ同時にターンを終えると、どちらかのログが欠けることがあります。
  • プロンプトキャッシュ: 公式ドキュメントでは、対象のモデルは(API キーかサブスクリプションなら)effort を変えてもキャッシュが保たれるとされています。作者の手元の記録でも、effort が変わったリクエストでキャッシュの読み込みが続く例を1件確認しました(網羅的な検証ではありません)。Bedrock、Google Cloud、ゲートウェイ経由では、変えるたびにキャッシュが作り直されます。
  • mod の API は新しく、Claude Code の版によって仕様が変わることがあります。

開発

ファイル構成

effort-router/
├─ .claude-plugin/
│  └─ plugin.json        # マニフェスト(名前、版、types の場所)
├─ hooks/
│  ├─ hooks.json         # フックのモジュールを指定
│  ├─ register.ts        # フック本体(イベント、コマンド、ステータス表示、ログ、評価)
│  ├─ route.ts           # 純粋なロジック(判定基準、入力の組み立てと圧縮、ラベルの解釈、途中の引き上げ)
│  └─ eval-cases.ts      # /effort-router eval で使う評価セット
├─ types/
│  └─ index.d.ts         # $.state の型(PluginState)
├─ tests/
│  ├─ register.test.ts   # フックの結合テスト(エンジンの下層をテストが代行)
│  └─ route.test.ts      # route.ts の単体テスト
├─ CHANGELOG.md
└─ README.md

tsconfig.json と .claude-plugin/types/ は、mod を読み込んだときにエンジンが生成します。リポジトリには含めず、手で編集もしません。

確認に使うコマンド

リポジトリの直下で実行します。

claude plugin validate mods/effort-router   # マニフェストとフックの静的検査
claude plugin test mods/effort-router       # tests/*.test.ts を実行
cd mods/effort-router && npx -p typescript@5 tsc -p .   # 型検査(mod を一度読み込んで型が生成されたあと)
claude --plugin-dir mods/effort-router      # 手元の版を読み込んで試す

判定基準を調整するときは

  1. hooks/route.ts の RUBRIC を編集する。実際に誤判定した入力は hooks/eval-cases.ts に足す。
  2. claude plugin test と tsc を通す。
  3. 新しいセッションで /effort-router eval を実行し(Haiku を36回呼ぶので、少額の費用がかかります)、一致数と LOW(低すぎる判定)の数を確かめる。Git Bash から claude -p "/effort-router eval" を実行するときは、MSYS_NO_PATHCONV=1 を付けないと / 始まりの引数がパスに書き換えられる。

変更履歴

CHANGELOG.md を見てください。

ライセンス

MIT(リポジトリの LICENSE)

Source 4 files
hooks/register.ts 460 lines
1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, Register } from 'claude-code'
3
4import type { Level, TurnSignals } from '../types'
5import { CASES } from './eval-cases'
6import type { EvalCase } from './eval-cases'
7import {
8  LABELS,
9  LEVELS,
10  asLevel,
11  asksOf,
12  compose,
13  decide,
14  effortArg,
15  escalate,
16  fromAsked,
17  head,
18  isNotice,
19  isRouted,
20  levelOf,
21  optionsOf,
22  plan,
23  rank,
24  tail,
25} from './route'
26import type { Asked, Decision, Route } from './route'
27
28// Picks the main thread's effort per prompt: low / medium / high / xhigh,
29// judged from the request and the previous turn's signals. A reply that
30// picks an offered option ("2") is judged by that option's text; a short
31// reply to a question is judged against the question; an AskUserQuestion
32// answer is judged mid-turn the same way. A turn that runs long or hits
33// errors goes up one level, so a judgment that was too low corrects itself.
34// A background task's notice keeps the level before it. Only the routed
35// models (isRouted: Opus, Sonnet and Haiku 5.5+, Fable and Mythos 5.1+) are touched;
36// on them Claude Code keeps the prompt cache across effort changes. The model
37// is never changed. A typed /effort runs as set for one turn; to pin a level,
38// turn routing off and use /effort.
39
40const LOG_KEY = 'log'
41const MODE_KEY = 'mode'
42const LOG_LIMIT = 200
43const RECENT = 3
44
45const prevTurn = atom({ plugin: 'effort-router', key: 'prev' } as const, null)
46const recentLevels = atom({ plugin: 'effort-router', key: 'recent' } as const, [])
47
48type Turn = Decision & {
49  request: string
50  picked: string | null
51  steps: number
52  tools: number
53  toolErrors: number
54  // `steps` and `toolErrors` when the level was last judged; escalation
55  // counts from here.
56  since: number
57  sinceErrors: number
58  // Levels escalation now adds over the judged level, and in all this turn.
59  bumps: number
60  raised: number
61  applied: Level | null
62  // Set by /effort: the session's effort goes as is, unjudged, unraised;
63  // `shown` is the effort the status line last showed for it.
64  manual: boolean
65  shown?: string
66  // The last step text that offered options: a closing line after a tool
67  // call ends the turn's answer without them.
68  offer?: string
69}
70
71const fresh = {
72  steps: 0,
73  tools: 0,
74  toolErrors: 0,
75  since: 0,
76  sinceErrors: 0,
77  bumps: 0,
78  raised: 0,
79  applied: null,
80  manual: false,
81}
82
83type LogEntry = {
84  at: number
85  head: string
86  chars: number
87  prev: Level | null
88  judged: Level | null
89  level: Level | null
90  why: Decision['why']
91  picked?: string | null
92  raised?: number
93  steps: number
94  toolErrors: number
95  durationMs: number
96  outputTokens: number | null
97}
98
99// The step's response as is; its text kept on the turn when it offers options.
100async function* answered<C, R extends { answer: string }>(turn: Turn, stream: AsyncGenerator<C, R>): AsyncGenerator<C, R> {
101  const result = yield* stream
102  if (optionsOf(result.answer).length > 0) turn.offer = result.answer
103  return result
104}
105
106export const register: Register = on => {
107  let isOn = true
108  let current: string | undefined
109  // The main thread's model as its last request named it.
110  let mainModel: string | undefined
111  // A /effort typed between turns: the next turn runs as set.
112  let manualNext = false
113  const turns = new Map<string, Turn>()
114
115  on('session.start', async ($, e, next) => {
116    isOn = (await $.store.get(MODE_KEY)) !== 'off'
117    await $.command.register({
118      name: 'effort-router',
119      description: 'Per-prompt effort routing (a typed /effort runs as set for one turn): on | off | log | log clear | eval',
120      argumentHint: '[on|off|log|log clear|eval]',
121      immediate: true,
122    })
123    show($, isOn)
124    return next(e)
125  })
126
127  on('command.run', { command: 'effort-router' }, async ($, e) => {
128    const arg = e.args.trim().toLowerCase()
129    if (arg === 'log') return { text: await summarize($) }
130    if (arg === 'log clear') {
131      await $.store.delete(LOG_KEY)
132      return { text: 'log cleared.' }
133    }
134    if (arg === 'eval') return { text: await evaluate($) }
135    if (arg === 'on' || arg === 'off') {
136      isOn = arg === 'on'
137      await $.store.set(MODE_KEY, arg)
138      show($, isOn)
139    } else if (arg !== '') {
140      return { text: `unknown "${arg}" (on | off | log | log clear | eval)` }
141    }
142    return {
143      text: isOn
144        ? 'on (picks low/medium/high/xhigh each prompt for Opus/Sonnet 5.5+ and Fable/Mythos 5.1+; a typed /effort runs as set for that turn)'
145        : 'off (/effort applies)',
146    }
147  })
148
149  // A typed /effort applies at once and shows at once, without pinning: the
150  // rest of a running turn, or else the next turn, keeps the session's effort
151  // as set; the turn after that is judged again.
152  on('command.run', { command: 'effort' }, async ($, e, next) => {
153    const ran = await next(e)
154    if (!isOn) return ran
155    const turn = current ? turns.get(current) : undefined
156    if (turn) turn.manual = true
157    else manualNext = true
158    const set = effortArg(e.args) ?? (await savedEffort($, mainModel))
159    if (turn && set) turn.shown = set
160    show($, isOn, { level: undefined, judged: null, why: 'manual', isSession: false, setTo: set })
161    return ran
162  }).catch(($, e, next) => next(e))
163
164  on('turn.start', async ($, e, next) => {
165    if (!isOn || e.text === '') return next(e)
166    // Before the first request names the model, ask the session for it.
167    if (mainModel === undefined) mainModel = await sessionModel($)
168    // Another model is in use: no judgment, and no classifier call, for it.
169    if (mainModel !== undefined && !isRouted(mainModel)) {
170      $.ui.status(`not routed (${mainModel})`)
171      return next(e)
172    }
173
174    const prev = await read($, prevTurn)
175    const prevLevel = prev?.level ?? null
176    if (manualNext) {
177      manualNext = false
178      current = e.turnId
179      turns.set(e.turnId, { level: undefined, judged: null, why: 'manual', ...fresh, manual: true, request: e.text, picked: null })
180      return next(e)
181    }
182    const route = isNotice(e.text) ? null : plan(e.text, prev)
183    let decision: Decision
184    let picked: string | null = null
185    if (route === null) {
186      decision = { level: prevLevel ?? undefined, judged: null, why: 'notice' }
187    } else if (route.kind === 'keep') {
188      decision = { level: prevLevel ?? undefined, judged: null, why: 'short' }
189    } else {
190      const input = compose(e.text, prev, await read($, recentLevels), route)
191      decision = decide(await judge($, input), prevLevel, route.why)
192      if (route.selected.length > 0) picked = route.selected.map(one => one.key).join(',')
193    }
194    current = e.turnId
195    turns.set(e.turnId, { ...decision, ...fresh, request: e.text, picked })
196    if (decision.level !== undefined) show($, isOn, { ...decision, picked, isSession: false })
197    return next(e)
198  })
199
200  on('tool.call', async ($, e, next) => {
201    const turn = e.agentId === undefined && current ? turns.get(current) : undefined
202    const ran = await next(e)
203    if (turn) {
204      turn.tools += 1
205      if (ran.deny === undefined && ran.isError === true) turn.toolErrors += 1
206    }
207    return ran
208    // Counting only watches: if it throws, the call's own answer stands.
209  }).catch(($, e, next) => next(e))
210
211  // An AskUserQuestion answer arrives mid-turn: judge what was picked and
212  // run the rest of the turn at that level, from the next request on.
213  on('tool.call', { tool: 'AskUserQuestion' }, async ($, e, next) => {
214    const ran = await next(e)
215    const turn = isOn && e.agentId === undefined && current ? turns.get(current) : undefined
216    // A turn the person set with /effort stays as set.
217    if (turn === undefined || turn.manual || ran.deny !== undefined || ran.isError === true) return ran
218    const asked = fromAsked(ran.result as Asked)
219    if (asked === null) return ran
220
221    const from = turn.applied ?? turn.level ?? null
222    const context: TurnSignals = {
223      level: from,
224      steps: turn.steps,
225      tools: turn.tools,
226      toolErrors: turn.toolErrors,
227      durationMs: 0,
228      request: head(turn.request),
229      answerTail: tail(asked.questions),
230    }
231    const input = compose(asked.request, context, await read($, recentLevels), asked)
232    const decision = decide(await judge($, input), from, asked.why)
233    if (decision.why !== 'choice' && decision.why !== 'reply') return ran
234    Object.assign(turn, {
235      level: decision.level,
236      judged: decision.judged,
237      why: 'asked',
238      picked: asked.picked,
239      since: turn.steps,
240      sinceErrors: turn.toolErrors,
241      bumps: 0,
242    })
243    show($, isOn, { ...decision, why: 'asked', picked: asked.picked, from, isSession: false })
244    return ran
245    // Re-judging only watches: if it throws, the answer still reaches the model.
246  }).catch(($, e, next) => next(e))
247
248  on('turn.step', async function* ($, e, next) {
249    if (e.agentId === undefined) mainModel = e.model
250    const turn = e.agentId === undefined ? turns.get(e.turnId) : undefined
251    if (turn === undefined) return yield* next(e)
252    turn.steps += 1
253    // Set by /effort: the session's effort goes as is; the status line shows
254    // it once it is known, as the request carries it.
255    if (turn.manual) {
256      turn.applied = asLevel(e.effort)
257      const set = e.effort === undefined ? undefined : String(e.effort)
258      if (set !== undefined && set !== turn.shown) {
259        turn.shown = set
260        show($, isOn, { level: undefined, judged: null, why: 'manual', isSession: false, setTo: set })
261      }
262      return yield* answered(turn, next(e))
263    }
264    // A model without effort leaves `effort` absent; it stays absent. A model
265    // not routed (isRouted) keeps its own effort.
266    const base = e.effort === undefined || !isRouted(e.model) ? null : (turn.level ?? asLevel(e.effort))
267    if (base === null) {
268      turn.applied = asLevel(e.effort)
269      return yield* answered(turn, next(e))
270    }
271    const level = escalate(base, turn.steps - turn.since, turn.toolErrors - turn.sinceErrors)
272    const bumps = rank(level) - rank(base)
273    if (bumps > turn.bumps) {
274      turn.raised += bumps - turn.bumps
275      const errors = turn.toolErrors - turn.sinceErrors
276      show($, isOn, { ...turn, level, from: turn.applied ?? base, isSession: false, raisedBy: errors > 1 ? `${errors} tool errors` : `${turn.steps - turn.since - 1} requests` })
277    }
278    turn.bumps = bumps
279    turn.applied = level
280    // No judgment and nothing raised: the session's own effort goes as is.
281    if (turn.level === undefined && level === base) {
282      if (turn.steps === 1) show($, isOn, { ...turn, level, isSession: true })
283      return yield* answered(turn, next(e))
284    }
285    return yield* answered(turn, next({ ...e, effort: level }))
286  })
287
288  on('turn.complete', async ($, e, next) => {
289    const turn = e.agentId === undefined ? turns.get(e.turnId) : undefined
290    if (turn) {
291      turns.delete(e.turnId)
292      if (current === e.turnId) current = undefined
293      const prev = await read($, prevTurn)
294      const final = optionsOf(e.answer)
295      const options = final.length > 0 || turn.offer === undefined ? final : optionsOf(turn.offer)
296      const signals: TurnSignals = {
297        level: turn.applied,
298        steps: turn.steps,
299        tools: turn.tools,
300        toolErrors: turn.toolErrors,
301        durationMs: e.durationMs,
302        request: head(turn.request),
303        answerTail: tail(e.answer),
304        options,
305        asks: asksOf(e.answer, options),
306      }
307      await update($, prevTurn, () => signals)
308      await update($, recentLevels, list => [...list, turn.applied].slice(-RECENT))
309      const entry: LogEntry = {
310        at: await $.clock.now(),
311        head: head(turn.request, 40),
312        chars: turn.request.length,
313        prev: prev?.level ?? null,
314        judged: turn.judged,
315        level: turn.applied,
316        why: turn.why,
317        picked: turn.picked,
318        raised: turn.raised,
319        steps: turn.steps,
320        toolErrors: turn.toolErrors,
321        durationMs: e.durationMs,
322        outputTokens: e.usage?.output_tokens ?? null,
323      }
324      const log = ((await $.store.get(LOG_KEY)) as LogEntry[] | undefined) ?? []
325      await $.store.set(LOG_KEY, [...log, entry].slice(-LOG_LIMIT))
326    }
327    return next(e)
328  })
329}
330
331// The classifier's answer, or the Error it rejected with.
332async function judge($: EngineInterface, input: string): Promise<string | undefined | Error> {
333  try {
334    return await $.model.classify(input, LABELS, { model: 'haiku' })
335  } catch (error) {
336    return error instanceof Error ? error : new Error(String(error))
337  }
338}
339
340// The main loop's model id, or undefined when the session does not name it
341// as an id (an alias such as "opus" is left to the first request to settle).
342async function sessionModel($: EngineInterface): Promise<string | undefined> {
343  try {
344    const model = await $.session.model()
345    return /^claude-/.test(model) ? model : undefined
346  } catch {
347    return undefined
348  }
349}
350
351// The effort /effort saved for the model, as the settings hold it, or null
352// when they do not name one (a change for this session only).
353async function savedEffort($: EngineInterface, model: string | undefined): Promise<string | null> {
354  if (model === undefined) return null
355  try {
356    const settings = await $.settings.read({})
357    const perModel = settings['modelSettings'] as Record<string, { effortLevel?: unknown }> | undefined
358    const set = perModel?.[model]?.effortLevel
359    return typeof set === 'string' ? set : null
360  } catch {
361    return null
362  }
363}
364
365type View = Pick<Decision, 'level' | 'judged' | 'why'> & {
366  picked?: string | null
367  isSession: boolean
368  from?: Level | null
369  // Set when the turn was raised: what ran long ("10 requests").
370  raisedBy?: string
371  // Set by /effort: the effort it set, null until the next request shows it.
372  setTo?: string | null
373}
374
375// The status line, after the engine's "effort-router:" prefix.
376function show($: EngineInterface, isOn: boolean, view?: View) {
377  if (!isOn) return $.ui.status('off (/effort applies)')
378  if (view === undefined) return $.ui.status('on, waiting for the first prompt')
379  if (view.why === 'manual') return $.ui.status(`${view.setTo ?? 'changed'} (set by /effort)`)
380  const notes: string[] = []
381  if (view.raisedBy !== undefined) notes.push(`raised after ${view.raisedBy}`)
382  else {
383    if (view.isSession) notes.push('session effort')
384    if (view.why === 'choice') notes.push(`choice ${view.picked ?? ''}`.trim())
385    if (view.why === 'reply') notes.push('reply to question')
386    if (view.why === 'asked') notes.push(`asked: ${view.picked ?? ''}`.trim())
387    if (view.why === 'short') notes.push('short prompt, kept')
388    if (view.why === 'notice') notes.push('task notice, kept')
389    if (view.why === 'unsure' || view.why === 'error') notes.push(`${view.why}, kept`)
390  }
391  const from = view.from && view.from !== view.level ? `${view.from} → ` : ''
392  $.ui.status(`${from}${view.level ?? 'n/a'}${notes.length ? ` (${notes.join(', ')})` : ''}`)
393}
394
395async function summarize($: EngineInterface): Promise<string> {
396  const log = ((await $.store.get(LOG_KEY)) as LogEntry[] | undefined) ?? []
397  if (log.length === 0) return 'no turns logged yet.'
398
399  const rows = new Map<string, { n: number; ms: number }>()
400  for (const one of log) {
401    const key = one.level ?? `default (${one.why})`
402    const row = rows.get(key) ?? { n: 0, ms: 0 }
403    rows.set(key, { n: row.n + 1, ms: row.ms + one.durationMs })
404  }
405  const table = [...rows].map(
406    ([key, row]) => `  ${key}: ${row.n} turns, avg ${Math.round(row.ms / row.n / 1000)}s`,
407  )
408  const recent = log.slice(-10).map(one => {
409    const judged = one.judged && one.judged !== one.level ? ` (judged ${one.judged})` : ''
410    const why = `${one.picked ? `${one.why} ${one.picked}` : one.why}${one.raised ? ` raised ${one.raised}` : ''}`
411    return `  ${one.prev ?? '-'} -> ${one.level ?? '-'}${judged}\t${why}\t${one.steps} steps\t${Math.round(one.durationMs / 1000)}s\t${one.head}`
412  })
413  return [`${log.length} turns logged`, ...table, 'recent (prev -> level):', ...recent].join('\n')
414}
415
416// Runs every eval case through the live classifier and scores the raw
417// judgment (before decide() falls back to the previous level) against the
418// expected level.
419async function evaluate($: EngineInterface): Promise<string> {
420  const lines: string[] = []
421  let hits = 0
422  let low = 0
423  let slowest = 0
424  for (const one of CASES) {
425    const started = await $.clock.now()
426    const { input, why } = caseInput(one)
427    const answer = await judge($, input)
428    const ms = (await $.clock.now()) - started
429    slowest = Math.max(slowest, ms)
430    const judged = answer instanceof Error ? undefined : levelOf(answer)
431    const final = decide(answer, one.prev?.level ?? null, why).level
432    const isLow = judged !== undefined && LEVELS.indexOf(judged) < LEVELS.indexOf(one.expect)
433    if (judged === one.expect) hits += 1
434    if (isLow) low += 1
435    const mark = judged === one.expect ? 'ok ' : isLow ? 'LOW' : 'off'
436    const error = answer instanceof Error ? ` error: ${answer.message}` : ''
437    lines.push(
438      `  ${mark} ${one.name}: expect ${one.expect}, judged ${judged ?? '-'}, final ${final ?? '-'} (${ms}ms)${error}`,
439    )
440  }
441  return [
442    `eval: ${hits}/${CASES.length} match, ${low} judged too low, slowest ${slowest}ms`,
443    ...lines,
444  ].join('\n')
445}
446
447// A case's classifier input, built the way the live hooks build it: an
448// AskUserQuestion answer mid-turn (`prev` standing for the turn so far), or
449// a prompt.
450function caseInput(one: EvalCase): { input: string; why: Route['why'] } {
451  const asked = one.asked ? fromAsked(one.asked) : null
452  if (asked) {
453    const prev = one.prev ? { ...one.prev, answerTail: tail(asked.questions) } : null
454    return { input: compose(asked.request, prev, one.recent, asked), why: asked.why }
455  }
456  const planned = plan(one.request, one.prev)
457  const route: Route = planned.kind === 'judge' ? planned : { why: 'auto', selected: [] }
458  return { input: compose(one.request, one.prev, one.recent, route), why: route.why }
459}
460
hooks/eval-cases.ts 116 lines
1import type { Level, Option, TurnSignals } from '../types'
2import type { Asked } from './route'
3
4// Cases for `/effort-router eval`: the classifier's judgment against the
5// level a person would pick for Claude Opus 5.5. Edit freely.
6
7export type EvalCase = {
8  name: string
9  request: string
10  prev: TurnSignals | null
11  recent: Array<Level | null>
12  expect: Level
13  // An AskUserQuestion answer mid-turn: `prev` is the turn so far and
14  // `request` only names the case.
15  asked?: Asked
16}
17
18const ask = (question: string, header: string, options: Array<[string, string]>, answer: string): Asked => ({
19  questions: [{ question, header, options: options.map(([label, description]) => ({ label, description })), multiSelect: false }],
20  answers: { [question]: answer },
21})
22
23const NEXT_STEP: Array<[string, string]> = [
24  ['実装する', 'キャッシュ層を追加し、pipeline と split と index の3ファイルを変更する'],
25  ['コミットして終了', 'ここまでの差分をコミットして今日は終わる'],
26]
27
28const heavy = (request: string, answerTail: string, toolErrors = 0, options: Option[] = []): TurnSignals => ({
29  level: 'xhigh',
30  steps: 9,
31  tools: 14,
32  toolErrors,
33  durationMs: 210_000,
34  request,
35  answerTail,
36  options,
37  asks: options.length > 0 || /[??]/.test(answerTail),
38})
39
40const light = (request: string, answerTail: string, options: Option[] = []): TurnSignals => ({
41  level: 'medium',
42  steps: 1,
43  tools: 0,
44  toolErrors: 0,
45  durationMs: 6_000,
46  request,
47  answerTail,
48  options,
49  asks: options.length > 0 || /[??]/.test(answerTail),
50})
51
52const choices = (...texts: string[]): Option[] => texts.map((text, i) => ({ key: String(i + 1), text }))
53
54const AFTER_DESIGN = choices(
55  '案2で実装を進める(キャッシュ層を追加し、pipeline と split と index の3ファイルを変更)',
56  '設計の前提をもう一度洗い出して比較し直す',
57  'ここまでの設計メモをコミットして今日は終了',
58)
59const AFTER_TYPO = choices(
60  'README の誤字だけ直して終わる',
61  'formatDate にうるう年のテストを1件追加する',
62  '認証まわりを JWT からセッション方式へ移す影響範囲を洗い出して移行計画を立てる',
63)
64
65export const CASES: EvalCase[] = [
66  // No context: the request alone decides.
67  { name: 'fresh-simple-question', request: 'TypeScript の satisfies 演算子って何をするものか一言で教えて', prev: null, recent: [], expect: 'low' },
68  { name: 'fresh-rename', request: 'utils.ts の関数 fmtDate を formatDate にリネームして', prev: null, recent: [], expect: 'low' },
69  { name: 'fresh-commit-msg', request: 'いまの差分にコミットメッセージを付けてコミットして', prev: null, recent: [], expect: 'low' },
70  { name: 'fresh-wording-everywhere', request: '全ページのヘッダーにある「問い合わせ」の表記を「お問い合わせ」にそろえて', prev: null, recent: [], expect: 'low' },
71  { name: 'fresh-explain-file', request: 'src/pipeline/split.ts が何をしているか説明してほしい', prev: null, recent: [], expect: 'medium' },
72  { name: 'fresh-implement-settled', request: 'さっき決めた仕様どおり、CSV エクスポートのボタンと API を追加して。テストも書いて', prev: null, recent: [], expect: 'medium' },
73  { name: 'fresh-design', request: 'スライド生成パイプラインにキャッシュ層を入れたい。設計案を複数出してトレードオフを比較して', prev: null, recent: [], expect: 'high' },
74  { name: 'fresh-debug', request: 'npm test で split.test.ts だけ落ちる。原因を調べて直して', prev: null, recent: [], expect: 'high' },
75  { name: 'fresh-stack-trace', request: 'これが出て起動しない\n```\nTypeError: Cannot read properties of undefined (reading \'body\')\n    at splitHtml (split.ts:42:13)\n    at run (pipeline.ts:88:5)\n    at main (index.ts:12:3)\n```', prev: null, recent: [], expect: 'high' },
76  { name: 'fresh-codebase-review', request: 'このリポジトリ全体をレビューして、バグや考慮漏れを洗い出して', prev: null, recent: [], expect: 'xhigh' },
77  { name: 'fresh-find-all', request: 'プロジェクト全体で旧 API の fetchUser を呼んでいる箇所を全部洗い出して、置き換えの方針を出して', prev: null, recent: [], expect: 'xhigh' },
78  { name: 'fresh-fundamental', request: '場当たり的な修正が続いて同じ種類の不具合が何度も出ている。根本的な解決策を検討して', prev: null, recent: [], expect: 'xhigh' },
79  { name: 'fresh-data-migration', request: '本番 DB の users テーブルを分割するマイグレーションを書いて。データは消せない', prev: null, recent: [], expect: 'xhigh' },
80
81  // Context makes a plain-looking request heavy, or keeps it ordinary.
82  { name: 'ctx-pick-option-after-design', request: '案2の方向で、さっきの設計どおりに実装を進めてください', prev: heavy('キャッシュ層の設計案を比較して', '…案1はシンプルだが無効化が難しい。案2は複雑だが整合性を保てる。どちらで進めますか?'), recent: ['xhigh', 'xhigh'], expect: 'medium' },
83  { name: 'ctx-still-broken', request: '直してもらったけど、まだ同じエラーが出ます。確認してください', prev: heavy('split.test.ts が落ちる原因を調べて', '…body 側の CSS が落ちていたのが原因でした。修正してテストが通ることを確認しました。', 2), recent: ['xhigh'], expect: 'xhigh' },
84  { name: 'ctx-continue-steps', request: 'その調子で残りのステップ3から5もお願いします', prev: heavy('移行計画の実装をステップ1から', '…ステップ1と2を実装し、テストも通りました。残りはステップ3〜5です。'), recent: ['xhigh', 'xhigh', 'xhigh'], expect: 'medium' },
85  { name: 'ctx-new-topic-heavy-after-light', request: '話は変わるけど、認証まわりを JWT からセッション方式に移行したい。影響範囲を洗い出して計画を立てて', prev: light('typo を直して', '…直しました。'), recent: ['medium'], expect: 'xhigh' },
86
87  // Context makes a request light even after heavy work.
88  { name: 'ctx-thanks-after-heavy', request: 'ありがとうございます、助かりました。今日はここまでにします', prev: heavy('原因調査', '…修正してテストが通りました。'), recent: ['xhigh'], expect: 'low' },
89  { name: 'ctx-typo-after-heavy', request: 'README の「インストール」の誤字だけ直しておいて', prev: heavy('設計レビュー', '…以上がレビュー結果です。'), recent: ['xhigh'], expect: 'low' },
90  { name: 'ctx-yesno-confirm', request: 'さっきの変更はもう push 済みという認識で合ってる?', prev: light('main に push して', '…push しました。'), recent: ['medium'], expect: 'low' },
91  { name: 'ctx-add-test-after-light', request: 'formatDate にうるう年のテストケースを1つ追加して', prev: light('fmtDate をリネームして', '…リネームしました。'), recent: ['high', 'medium'], expect: 'medium' },
92  { name: 'ctx-review-after-light', request: '今の差分をレビューして、問題があれば指摘して', prev: light('コミットして', '…コミットしました。'), recent: ['medium'], expect: 'high' },
93
94  // A number picks an offered option: its text decides, not the turn before.
95  { name: 'choice-low-after-design', request: '3', prev: heavy('キャッシュ層の設計案を比較して', '…1〜3から選んでください。', 0, AFTER_DESIGN), recent: ['xhigh', 'xhigh'], expect: 'low' },
96  { name: 'choice-medium-after-design', request: '1でお願いします', prev: heavy('キャッシュ層の設計案を比較して', '…1〜3から選んでください。', 0, AFTER_DESIGN), recent: ['xhigh'], expect: 'medium' },
97  { name: 'choice-xhigh-after-light', request: '3で', prev: light('typo を直して', '…直しました。次はどうしますか?', AFTER_TYPO), recent: ['medium'], expect: 'xhigh' },
98  { name: 'choice-medium-after-light', request: '2番', prev: light('typo を直して', '…直しました。次はどうしますか?', AFTER_TYPO), recent: ['medium'], expect: 'medium' },
99
100  // A short reply agrees to what the previous answer proposed.
101  { name: 'reply-yes-to-investigation', request: 'はい', prev: light('テストを流して', '…split.test.ts だけが落ちています。原因を調査して修正まで進めますか?'), recent: ['medium'], expect: 'high' },
102  { name: 'reply-ok-to-design', request: 'OK', prev: light('macOS と Linux にも対応させたい', '…かなり大きな作り直しになります。進めるなら、まず仕様の詰めから始め、設計文書にしてから実装に入りたいと思います。それでよいですか?'), recent: ['medium'], expect: 'high' },
103  { name: 'reply-ok-to-commit', request: 'OK', prev: heavy('原因を調べて直して', '…修正してテストが通りました。この内容でコミットしますか?'), recent: ['xhigh'], expect: 'low' },
104  { name: 'reply-ok-beside-list', request: 'OK', prev: light('fetchUser の置き換え方針をまとめて', '…1. 型定義の更新 2. API 層の差し替え 3. プロジェクト全体の呼び出し元の置き換え。この順で全体の移行作業に入りますか?', choices('型定義の更新', 'API 層の差し替え', 'プロジェクト全体の呼び出し元の置き換え')), recent: ['medium'], expect: 'xhigh' },
105
106  // A short prompt that names its own work, or declines, is judged as a request.
107  { name: 'short-audit-after-light', request: 'リポジトリ全体を監査して', prev: light('typo を直して', '…直しました。'), recent: ['medium'], expect: 'xhigh' },
108  { name: 'short-cause-after-question', request: '落ちる原因を調べて', prev: light('テストを流して', '…split.test.ts だけが落ちています。ほかは通ったのでコミットしますか?'), recent: ['medium'], expect: 'high' },
109  { name: 'short-decline-after-proposal', request: 'いや、やめておいて', prev: heavy('原因を調べて直して', '…根本的に直すなら設計の見直しが必要です。進めますか?'), recent: ['xhigh'], expect: 'low' },
110
111  // An AskUserQuestion answer mid-turn: the picked option's meaning decides.
112  { name: 'asked-commit-in-heavy-turn', request: '(AskUserQuestion)', asked: ask('次にどう進めますか?', '進め方', NEXT_STEP, 'コミットして終了'), prev: heavy('キャッシュ層の設計案を比較して', ''), recent: ['xhigh'], expect: 'low' },
113  { name: 'asked-implement-in-light-turn', request: '(AskUserQuestion)', asked: ask('次にどう進めますか?', '進め方', NEXT_STEP, '実装する'), prev: light('キャッシュ層の案を一言で', ''), recent: ['medium'], expect: 'medium' },
114  { name: 'asked-typed-investigate', request: '(AskUserQuestion)', asked: ask('次にどう進めますか?', '進め方', NEXT_STEP, 'その前に split.test.ts が落ちる原因を調べて'), prev: light('テストを流して', ''), recent: ['medium'], expect: 'high' },
115]
116
hooks/route.ts 343 lines
1import type { Level, Option, TurnSignals } from '../types'
2
3// Pure routing logic: what the classifier reads, and how its answer becomes
4// the turn's effort. No `$` here, so tests and the eval share it as is.
5
6// Tuned for Claude Opus 5.5 (and used for every routed model), whose default is medium and whose medium
7// already does ordinary multistep coding well; max is never picked.
8export const LEVELS: readonly Level[] = ['low', 'medium', 'high', 'xhigh']
9
10// The classifier must answer a label verbatim, so labels are single words
11// and their definitions lead the text it reads.
12export const LABELS: readonly string[] = LEVELS
13
14const RUBRIC = [
15  '[rubric] Judge how much reasoning the [request] needs, reading it in the context of [prev]. Judge by how uncertain the right answer is and how costly a miss would be, not by the amount of work:',
16  'low = nothing could be overlooked: chat, a confirmation, thanks or closing, a translation, a commit, a fully specified mechanical edit (a typo, a rename, a move);',
17  'medium = ordinary work with a clear approach: implementing a settled approach or design, adding a test, explaining code, looking something up;',
18  'high = the right answer is uncertain: finding the cause of a bug, comparing design options, reviewing a diff, sorting out vague requirements;',
19  'xhigh = a miss is costly and the work is hard: a review or audit of a whole codebase, finding every occurrence of something across a project, a project-wide migration, a fundamental redesign or root-cause fix, a fix that still fails, concurrency, security or data-loss risks, or long autonomous work.',
20].join(' ')
21
22// Added only when they apply, so an ordinary request reads the rubric alone.
23const CHOICE_RULE =
24  '[choice] The [request] picks the [selected] options the previous answer offered: judge the work those options name on the same scale (a commit or closing is low, implementing a settled approach is medium), however heavy the previous turn was.'
25const REPLY_RULE =
26  '[reply] The [request] is a short reply agreeing to what [prev_answer_tail] proposed or asked. Read it as a request for that proposed work, stated in full, and judge that work on the same scale; the agreement itself weighs nothing (agreeing to a commit is low, to drafting a spec or design is high, to investigating a cause is high).'
27
28// A turn still running after this many model requests since it was judged,
29// or after this many tool errors, goes up one level (once per judgment).
30export const ESCALATE_STEPS = 10
31export const ESCALATE_ERRORS = 2
32
33// Prompts this short made only of agreement ("続けて", "OK") skip the
34// classifier and inherit, unless they pick an offered option or answer a
35// question. Any other short prompt ("全体を監査して", "やめて") is judged.
36export const MIN_CHARS = 20
37
38const MAX_OPTIONS = 10
39const OPTION_CHARS = 150
40const ASK_TAIL_CHARS = 200
41
42const PREV_REQUEST_CHARS = 100
43const ANSWER_TAIL_CHARS = 300
44const REQUEST_HEAD_CHARS = 600
45const REQUEST_TAIL_CHARS = 200
46
47export function levelOf(label: string | undefined): Level | undefined {
48  const head = label?.split(':')[0]?.trim()
49  return LEVELS.find(level => level === head)
50}
51
52// The session's own effort as a step names it, read as one of ours.
53export function asLevel(effort: string | number | undefined): Level | null {
54  if (effort === 'max' || effort === 'xhigh') return 'xhigh'
55  if (effort === 'high' || effort === 'medium' || effort === 'low') return effort
56  return null
57}
58
59// Routed: Opus, Sonnet and Haiku from 5.5 on, Fable and Mythos from 5.1 on, the
60// models that keep the prompt cache across effort changes. A later version
61// (Opus 6, Sonnet 5.6) is routed too. The rubric is tuned for Opus 5.5.
62const ROUTED_FROM: Readonly<Record<string, number>> = { opus: 5.5, sonnet: 5.5, haiku: 5.5, fable: 5.1, mythos: 5.1 }
63
64export function isRouted(model: string): boolean {
65  const found = /claude-(opus|sonnet|haiku|fable|mythos)-(\d+)(?:-(\d{1,2})(?!\d))?/.exec(model)
66  if (!found) return false
67  const version = Number(found[2]) + Number(found[3] ?? 0) / 10
68  return version >= ROUTED_FROM[found[1]!]!
69}
70
71// What `/effort <args>` sets, when the args name a level.
72export function effortArg(args: string): string | null {
73  const word = args.trim().toLowerCase()
74  return ['low', 'medium', 'high', 'xhigh', 'max'].includes(word) ? word : null
75}
76
77// A background task's completion notice is not the person's request.
78export function isNotice(text: string): boolean {
79  return /^\s*<task-notification>/.test(text)
80}
81
82export function rank(level: Level): number {
83  return LEVELS.indexOf(level)
84}
85
86export function maxLevel(levels: ReadonlyArray<Level | null>): Level | null {
87  let best: Level | null = null
88  for (const level of levels) {
89    if (level && (best === null || rank(level) > rank(best))) best = level
90  }
91  return best
92}
93
94// The level for the next request: one up once the turn has run long or hit
95// errors since it was judged, never above xhigh.
96export function escalate(level: Level, steps: number, errors: number): Level {
97  const bump = steps > ESCALATE_STEPS || errors >= ESCALATE_ERRORS ? 1 : 0
98  return LEVELS[Math.min(rank(level) + bump, LEVELS.length - 1)]!
99}
100
101// Code and stack traces become markers: the signal stays, the tokens go.
102export function compact(text: string): string {
103  return text
104    .replace(/\r\n?/g, '\n')
105    .replace(/```(\w*)\n([\s\S]*?)```/g, (_, lang: string, body: string) => {
106      const lines = body.split('\n').filter(Boolean).length
107      return `[code${lang ? `: ${lang}` : ''} ${lines} lines]`
108    })
109    .replace(/(?:^[ \t]*(?:at .+|File ".+", line \d+.*)\n?){3,}/gm, run => {
110      const lines = run.split('\n').filter(Boolean).length
111      return `[stack trace: ${lines} lines]\n`
112    })
113    .replace(/[ \t]+/g, ' ')
114    .replace(/\n{3,}/g, '\n\n')
115    .trim()
116}
117
118function oneLine(text: string): string {
119  return text.replace(/\s+/g, ' ').trim()
120}
121
122function clip(text: string): string {
123  if (text.length <= REQUEST_HEAD_CHARS + REQUEST_TAIL_CHARS) return text
124  return `${text.slice(0, REQUEST_HEAD_CHARS)} … ${text.slice(-REQUEST_TAIL_CHARS)}`
125}
126
127export function tail(text: string, chars = ANSWER_TAIL_CHARS): string {
128  const line = oneLine(compact(text))
129  return line.length <= chars ? line : `…${line.slice(-chars)}`
130}
131
132export function head(text: string, chars = PREV_REQUEST_CHARS): string {
133  const line = oneLine(compact(text))
134  return line.length <= chars ? line : `${line.slice(0, chars)}…`
135}
136
137// What the classifier reads: the previous turn as signals and two short
138// excerpts, then this request, compacted and clipped, and the options it picked.
139export function compose(
140  request: string,
141  prev: TurnSignals | null,
142  recent: ReadonlyArray<Level | null>,
143  route: Route = { why: 'auto', selected: [] },
144): string {
145  const { why, selected } = route
146  const lines: string[] = [RUBRIC]
147  if (why === 'choice') lines.push(CHOICE_RULE)
148  if (why === 'reply') lines.push(REPLY_RULE)
149  if (prev) {
150    const max3 = maxLevel(recent) ?? 'default'
151    lines.push(
152      `[prev] effort=${prev.level ?? 'default'} steps=${prev.steps} tools=${prev.tools} tool_errors=${prev.toolErrors} dur=${Math.round(prev.durationMs / 1000)}s max3=${max3}`,
153      `[prev_request] ${prev.request}`,
154      `[prev_answer_tail] ${prev.answerTail}`,
155    )
156  }
157  // A bare "OK" pulls the judgment to low whatever it agrees to, so a reply
158  // to a question (only ever agreement, see plan) is stated as the request
159  // for the proposed work it is, offered options or not. Not so a typed
160  // AskUserQuestion answer: that says what to do itself.
161  const agrees = why === 'reply' && !('questions' in route)
162  lines.push(
163    agrees
164      ? `[request] go ahead with what [prev_answer_tail] proposed (the reply was: ${clip(compact(request))})`
165      : `[request chars=${request.length}] ${clip(compact(request))}`,
166  )
167  if (selected.length > 0) {
168    lines.push(`[selected] ${selected.map(one => `${one.key}: ${one.text}`).join(' | ')}`)
169  }
170  return lines.join('\n')
171}
172
173// `asked`: re-judged mid-turn from an AskUserQuestion answer. `notice`: a
174// background task's notice, which keeps the level before it. `manual`: the
175// person set the effort with /effort, which runs as set.
176export type Why = 'auto' | 'choice' | 'reply' | 'asked' | 'short' | 'notice' | 'manual' | 'unsure' | 'error'
177export type Decision = { level: Level | undefined; judged: Level | null; why: Why }
178
179// The turn's effort from the classifier's answer (undefined: none matched;
180// an Error: the call failed): the judgment as is, or the previous turn's
181// level when there is none. A judgment that was too low is raised mid-turn.
182export function decide(
183  answer: string | undefined | Error,
184  prev: Level | null,
185  why: 'auto' | 'choice' | 'reply' = 'auto',
186): Decision {
187  if (answer instanceof Error) return { level: prev ?? undefined, judged: null, why: 'error' }
188  const judged = levelOf(answer)
189  if (judged === undefined) return { level: prev ?? undefined, judged: null, why: 'unsure' }
190  return { level: judged, judged, why }
191}
192
193// CRLF reads as LF; full-width digits and letters as ASCII; a circled
194// number as "1.".
195function normalize(text: string): string {
196  return text
197    .replace(/\r\n?/g, '\n')
198    .replace(/[0-9A-Za-z]/g, c => String.fromCharCode(c.charCodeAt(0) - 0xfee0))
199    .replace(/[①-⑳]/g, c => `${c.charCodeAt(0) - 0x2460 + 1}.`)
200}
201
202function cleanOption(text: string): string {
203  const line = oneLine(text.replace(/\*\*|`/g, '').replace(/\|\s*$/, '').replace(/\s*\|\s*/g, ' / '))
204  return line.length <= OPTION_CHARS ? line : `${line.slice(0, OPTION_CHARS)}…`
205}
206
207const OPTION_LINE = /^\s*(?:[-*+]\s+)?(?:\*\*)?(?:案|option\s*)?[((]?(\d{1,2}|[A-Za-z])(?:[..))::、])(?:\*\*)?\s*(?!\d)(.+)$/i
208const OPTION_ROW = /^\s*\|\s*(?:\*\*)?(\d{1,2}|[A-Za-z])(?:\*\*)?\s*\|(.+)$/
209
210// The last run of numbered or lettered lines in the answer (a list, or the
211// rows of a table): a new run starts at 1 or a. Lines inside a code block
212// are code, not options.
213export function optionsOf(answer: string): Option[] {
214  let run: Option[] = []
215  let last: Option[] = []
216  let inCode = false
217  for (const raw of normalize(answer).split('\n')) {
218    if (/^\s*(```|~~~)/.test(raw)) {
219      inCode = !inCode
220      continue
221    }
222    if (inCode) continue
223    const found = OPTION_LINE.exec(raw) ?? OPTION_ROW.exec(raw)
224    if (!found) continue
225    const key = found[1]!.toLowerCase()
226    const text = cleanOption(found[2]!)
227    if (text === '') continue
228    if (key === '1' || key === 'a' || run.length === 0) {
229      run = []
230      last = run
231    }
232    if (run.length < MAX_OPTIONS && !run.some(one => one.key === key)) run.push({ key, text })
233  }
234  return last
235}
236
237// Whether the answer ends by offering choices or asking something.
238export function asksOf(answer: string, options: readonly Option[]): boolean {
239  if (options.length > 0) return true
240  const end = tail(answer, ASK_TAIL_CHARS)
241  return /[??]/.test(end) || /(ますか|でしょうか|しましょうか|いかがですか|ませんか)/.test(end)
242}
243
244const KEY = /\b(\d{1,2}|[A-Za-z])\b/g
245const FILLER =
246  /選択肢|オプション|option|and|please|お願い|おねがい|いたします|致します|します|ください|下さい|進めて|すすめて|して|やって|実行|でいい|いい|です|案|番|目|で|を|に|は|と|も|の|よ|ね|[\s、,。..!!&+・//~〜]/gi
247
248// The offered options a reply picks ("2", "2で", "1と3でお願いします"), or
249// null when it is anything else or names a key that was not offered.
250export function picked(text: string, options: readonly Option[]): Option[] | null {
251  if (options.length === 0) return null
252  const line = normalize(text).trim()
253  const keys = [...line.matchAll(KEY)].map(found => found[1]!.toLowerCase())
254  if (keys.length === 0) return null
255  if (line.replace(KEY, '').replace(FILLER, '') !== '') return null
256  const chosen = keys.map(key => options.find(one => one.key === key))
257  if (chosen.some(one => one === undefined)) return null
258  return chosen.filter((one, i): one is Option => one !== undefined && chosen.indexOf(one) === i)
259}
260
261const AGREE =
262  /^(?:ok|okay|おk|おけ|おっけー|オッケー|はい|うん|ええ|了解|りょうかい|承知|yes|yep|sure|please|go|ahead|lgtm|お願い|おねがい|よろしく|頼む|頼みます|どうぞ|ぜひ|それで|これで|そう|いい|良い|進めて|すすめて|やって|続けて|つづけて|続き|続行|して|します|いたします|致します|ください|下さい|です|じゃあ|では|で|よ|ね|も|お)+$/
263
264// Whether the prompt only agrees or says go on ("OK", "はい、進めてください",
265// "続けて"), naming no work of its own. "いいえ", "やめて", "もういい" do not.
266export function isAgreement(text: string): boolean {
267  return AGREE.test(normalize(text).toLowerCase().replace(/[\s、,。..!!~〜…・]/g, ''))
268}
269
270export type Route = { why: 'auto' | 'choice' | 'reply'; selected: readonly Option[] }
271export type Plan = { kind: 'keep' } | ({ kind: 'judge' } & Route)
272
273// How to read this prompt: a pick of offered options, a short agreement to a
274// question, a short agreement with nothing to read it against, or a request
275// (a short one naming its own work, or declining, included).
276export function plan(text: string, prev: TurnSignals | null): Plan {
277  const selected = picked(text, prev?.options ?? [])
278  if (selected) return { kind: 'judge', why: 'choice', selected }
279  if (text.length >= MIN_CHARS || !isAgreement(text)) return { kind: 'judge', why: 'auto', selected: [] }
280  if (prev?.asks === true) return { kind: 'judge', why: 'reply', selected: [] }
281  return { kind: 'keep' }
282}
283
284// What AskUserQuestion hands back: the questions with their options, the
285// answers keyed by question text (multi-select comma-separated), and the
286// text typed instead of picking.
287export type Asked = {
288  questions?: ReadonlyArray<{
289    question: string
290    header: string
291    options?: ReadonlyArray<{ label: string; description?: string }>
292    multiSelect?: boolean
293  }>
294  answers?: Readonly<Record<string, string>>
295  response?: string
296}
297
298export type AskedRoute = Route & {
299  why: 'choice' | 'reply'
300  request: string
301  questions: string
302  picked: string
303}
304
305// The answers as a request to judge: picked options carry their
306// descriptions; anything typed makes it a reply to the questions.
307export function fromAsked(asked: Asked): AskedRoute | null {
308  const selected: Option[] = []
309  const lines: string[] = []
310  const headers: string[] = []
311  let typed = false
312  for (const one of asked.questions ?? []) {
313    const answer = asked.answers?.[one.question]?.trim()
314    if (!answer) continue
315    lines.push(`${one.header}: ${answer}`)
316    headers.push(one.header)
317    const find = (label: string) => one.options?.find(option => option.label === label.trim())
318    const whole = find(answer)
319    const parts = whole ? [whole] : one.multiSelect ? answer.split(/,\s*/).map(find) : [undefined]
320    if (parts.some(part => part === undefined)) {
321      typed = true
322      continue
323    }
324    for (const part of parts) {
325      const text = part!.description ? `${part!.label} — ${part!.description}` : part!.label
326      selected.push({ key: one.header, text: cleanOption(text) })
327    }
328  }
329  const response = asked.response?.trim()
330  if (response) {
331    lines.push(response)
332    typed = true
333  }
334  if (lines.length === 0) return null
335  return {
336    why: typed ? 'reply' : 'choice',
337    selected,
338    request: lines.join('\n'),
339    questions: (asked.questions ?? []).map(one => one.question).join(' / '),
340    picked: [...new Set(headers)].join(',') || 'text',
341  }
342}
343
types/index.d.ts 25 lines
1export type Level = 'low' | 'medium' | 'high' | 'xhigh'
2
3// One numbered or lettered choice the answer offered ("2. テストを追加する").
4export type Option = { key: string; text: string }
5
6// What one main-thread turn left behind, for judging the next one.
7// `options` and `asks` may be absent in state an older version saved.
8export type TurnSignals = {
9  level: Level | null
10  steps: number
11  tools: number
12  toolErrors: number
13  durationMs: number
14  request: string
15  answerTail: string
16  options?: Option[]
17  asks?: boolean
18}
19
20declare module 'claude-code' {
21  interface PluginState {
22    'effort-router': { prev: TurnSignals | null; recent: Array<Level | null> }
23  }
24}
25