Dispatch Pilot:一个在 Claude 之外的决策模型(默认 Perplexity 的 pplx,没有 Perplexity key 时退回 TypeSafe 的 Jev,也可以固定用 Jev)在你每次发消息时决定主 agent 这一轮的…

Dispatch Pilot 是 alex-mods marketplace 里的一个 Claude Code mod。它在 Claude 之外调用一个决策模型(默认是 Perplexity 的 pplx;没有 Perplexity 的 key、但有 TypeSafe 的 key 时退回 TypeSafe 的 Jev,也可以指定用 Jev),替你决定 Claude Code 怎么干活:主 agent 每一轮用哪档 effort,派出的 agent 用哪个模型和哪档 effort,主 agent 该看哪几个 skill。
主 agent 的模型从不改变,所以 prompt cache 不受影响。这只在 Claude Code 订阅下成立,见「要求」。
决策模型只负责给出判断,不参与回答你。每个判断的结果写在看板上,依据记进决策日志,在依据面板里看(/dp),也可以用 /dp log N 在对话里列出。
max 只在它自己的概率达到 thetaMax 时才用。这一轮的每个模型请求都按这一档发出。一轮进行中,每隔几步(rejudgeEvery),以及主 agent 派出 agent、启动 Workflow 或加载 skill 时,会再判断一次(中途重判);升档和降档都有防抖:升档要置信度达到 thetaUp(0.3),降档要达到 thetaDown(0.75)而且一次只降一档,升档之后 holdSteps(5)步之内不降。派出 agent 交回结果、后台任务结束,会话空闲时它们会开始主 agent 的新一轮,这样的一轮也走同样的判断:决策模型读到的是那条报告的文字,不推荐 skill,也不在这一轮中途重判(会话自己的 xhigh 用在只是看一眼结果上是浪费);报告送进一轮正在进行的对话时不另外判断。其他会话的消息、插件自己发的消息和空消息都不判断。你输入的 skill 或 markdown 命令(例如 /implement #19)开始的一轮也照常判断:决策模型读你输入的命令和参数,以及这个命令是做什么的,不读命令展开后的正文;参数里的点名照样算;这样的一轮不推荐 skill。/dp、/clear 这类本地命令不开始一轮,也就不判断。agentFable 后加入 fable)和 effort;选了 haiku 就不设 effort,haiku 不支持。三个模型的分工按 Artificial Analysis 的基准定(见下面的「模型和 effort 的依据」):haiku 只做一两步就能跑完、结果只需要收集起来并按要求排版的只读查找,sonnet 承担大多数执行类工作(终端、需求明确的代码修改、自动化、调研汇报),opus 做需要审慎判断或细微错误代价高的工作(安全、并发、涉及钱、迁移、生产)、难推理、设计、原因未知的 bug、科学或算法类代码,以及结论取决于记忆中的事实、而且无法在仓库或文档里查证的调研;fable 在 AA 上没有领先 Opus 的地方,价格是 2.5 倍,所以仍默认关闭。决策出来的 effort 不低于所选模型的下限:sonnet 和 opus 至少 medium,haiku 不带 effort。你在消息里点名的模型或 effort(「用 opus」「effort 开 low」)一定照办,下限和往上取的一档都不会改你点名的值,你排除的模型(「别用 opus」)不会用(决策请求整体失败时除外,见「局限和待评测」)。主 agent 自己为这个 agent 指定了模型时,只有决策模型选了别的、而且置信度达到 agentOverride 才推翻它。agent() 调用点做同样的判断,把模型和 effort 写进脚本再运行,并告诉主 agent 写了什么;workflowMode 选 return 时改为退回脚本,附上逐个调用的推荐,让主 agent 自己写进去。脚本写不进去的(用 scriptPath 或 name 提交、恢复的运行、读不了的脚本),在每个 agent 启动时按它的 label 设置,这叫兜底。escalateAfter 次时,再问决策模型一次,把它的 effort 升一档(escalateMode 选 max 则直接升到 max);haiku 没有 effort 可升,改用 sonnet 接着做(escalateHaikuTo)。这些失败本来就在意料之中的,例如先写下、要看它红的测试,或者没找到东西而以非零退出的搜索,不升档;是不是预期内失败由决策模型判断,不靠关键词。你自己拒绝的调用从不算失败。/clear 和新会话清零,/compact 和热重载保留;agent 交回结果、后台任务通知开始的一轮不问。依据卡片显示次数和这一次的结论。次数和下面的摘要一起作为决策模型判断 effort 的依据,见「强提示」。这项功能(unresolved 开关,次数、摘要、强提示共用)默认关闭:评测里它对默认的决策模型没有提升,所以不问这道题、不数次数、不写摘要、不给强提示。/dp unresolved on 打开,对 pplx 和 Jev 都生效;/dp unresolved off 再关。开关存在 $.store,你手动设过的状态升级后保持原样;0.3.1 里它是默认开着的,开着时没有保存过任何设置,所以从 0.3.1 升级上来、一直开着的,现在变成了关闭,要用请自己打开。summaryModel,默认 haiku,用你的 Claude 登录和用量)在后台续写一份简短的摘要:问题是什么、试过哪些做法、现在到了哪一步,不超过 500 token,只记试过什么,不判断成没成功;你的下一条消息说「仍未解决」时,最后一次尝试标上「未解决」。决策模型发消息时读它,借此看到超出上下文预算的更早几轮;它从不等摘要写完,来不及时用上一份。写失败(出错、超时)保留旧摘要并在决策日志里记一条。问题解决、换了问题、/clear 和新会话时和次数一起清空,/compact 和热重载保留。依据面板里能看摘要全文。交给写摘要模型的内容先脱敏;contextMessages 为 0 时不写(摘要是对话的转述)。默认关闭,随 unresolved 开关一起打开和关闭。unresolvedMaxAfter(默认 3,0 到 10,0 表示不给)时,发消息时和中途重判时 effort 题的说明里多一条:这项工作属于多次尝试都没解决的故障。它只描述情境:不写档位的名字或序号,pickEffort 的规则和 thetaMax 都不变,用不用最高一档仍由决策模型结合摘要和对话自己决定。次数和摘要也进这两种请求的 state(unresolved_count、problem_summary);派出 agent 的决定读摘要和次数当背景,但从不加强提示(「查一个文件」这类子任务不该因此判到最高一档)。发消息时带的是这条消息之前的次数(这条消息自己的结论要同一个请求的回答才知道),所以默认设置下第四次说「仍未解决」的那条消息的中途重判先读到强提示,下一条消息的请求再带上。每次给了强提示,决策日志记一条(次数、已给强提示、决策模型判出的档位),依据卡片显示「已给强提示」;看板不新增事件。unresolved 开关关着(默认)时这些都不带。find_skill 工具按几个词查 skill。skill 本身和 Skill 工具都不变,主 agent 仍然可以按名字加载任何 skill。timeoutMs)、出错、回答无法解析,或者没有配密钥时,消息照常进入,不额外等待,这一轮用会话自己的 effort,看板上写明原因,并弹一个 toast。本文提到的配置项(例如 escalateAfter、timeoutMs),默认值都在「配置」的表里。每项功能都可以用 /dp 单独关掉,也可以整个 mod 一起关(见「控制:/dp」)。完整的行为规则(各种优先级、各种失败情形、看板和日志的写法)见 DEVELOPMENT.md 的「它做什么」。
模型的分工、往上取一档、中途的门槛和各模型的 effort 下限,都按 Artificial Analysis(AA)Intelligence Index v4.3.2 的十个分项校正过。几个要点:Sonnet 5.5 的分数随 effort 掉得很快(Terminal-Bench:low 20.7%、medium 29.8%、high 43.9%、max 63.6%;指数 36、41、47、56),所以低估 sonnet 的代价大;Opus 5.5 在 medium 就有 51;Sonnet 与 Opus 在终端、自动化、知识工作上持平或略高,拉开的是事实知识(Omniscience 32 对 46)、难推理(HLE 差 6.4)和科学代码(SciCode 差 5.9);Haiku 4.5 在 Terminal-Bench 上是 0%。数字、来源和已存评测回答按新规则离线重算的结果见 docs/research/aa-benchmarks-2026-10.md。
看板是 Dispatch Pilot 常驻的界面,只显示读数和结论,依据放在依据面板里。给人看的文字都是中文,模型名和 effort 档名照 /model、/effort 的写法。
find_skill 的查询,以及请求失败、回答迟到。一轮结束后 band 折成一行:主 agent 的模型和 effort、这一档是怎么来的、「可试 /x」。引擎给 band 的行数不到 4 行时,退成一行摘要。+N 个运行中的 agent;放不下时 effort 先写短(xhi),再省掉模型。/dp <功能> off)的决定、事件和计数不出现在 band 和脚部;/dp off 时 band 上没有 Dispatch Pilot 的内容,脚部写 ○ dp 已关。claude -p 没有界面,内容写进 debug log。/dp 打开依据面板,再输一次或按 Esc 关掉;/dp log 也是打开它。全屏且终端够宽时它停靠在对话右边(约 75 列),否则在 prompt 上方。面板窄时文字换行缩进,不截断。
hook-block-failures 也在内)、/dp lock 的锁定,以及本会话 skill 画像的情况(skill 画像:保留 52 · 新写 3 · 失败 1;写的过程中是「生成中 2/5」;停写时写原因;有失败时按 f 列出失败的 skill 和原因)。a 起,跳过 p、n、f)折叠或展开,最新两轮默认展开。日志保留最近 20 轮、最多 300 条。/dp 打开时会拿键盘;在 band 上按数字键打开的不抢键盘,可以接着按别的数字)。面板没被放出来时(比如接着的端不放面板)会立即关掉,并说明原因(/dp 在回复里说,按数字键打开的弹一个 toast),这时用 /dp log 10 在对话里看。contextMessages 条,总长度不超过 contextTokens)。主 agent 的 effort 题单独发一个请求,能读到的最近对话比和 skill 推荐一起发时长得多(Jev 默认 24000 token,对 6000);skill 推荐的请求和它并行发出,不增加等待。每条只有文字和调用过的工具名,不包含文件内容和工具输出。一轮中途重判时发的是这一轮最近几步(rejudgeSteps)的摘要:主 agent 写的文字、调用的工具和一句话结果,同样不含文件内容、写入的内容和工具输出。summaryModel 对你和主 agent 这个问题的转述,输入给它的内容先脱敏;它占 state 的预算,最近的对话让给它。prompt)、描述、agent 类型,加上你这一轮说的话。skillsProfileModel),用你的用量,不经过决策模型的提供方。[REDACTED]:各家的 API key 和 token、password=... 这类赋值、URL 里的密码、私钥和 JWT。api.typesafe.ai)。claude --debug-file <路径>),不进入会话。skillsProfileModel),算在你的用量里:每份约 2k 输入和 200 输出 token,每个 SKILL.md 版本只写一次,每次会话开始最多写 skillsProfilesPerSession 份。summaryModel),算在你的用量里:你本人的消息开始的每一轮结束后一次,输入是上一份摘要加这一轮(你的话、主 agent 的最终回复、工具汇总,各自截断),输出不超过 500 token。claude plugin details dispatch-pilot@alex-mods(2.1.289)显示 0 个组件、常驻开销约 0 token:它看不到 mod 在运行时附加和替换的内容。eval/ 和 scripts/ 里的 Node 脚本用的是 Node 26.5,用 mod 本身不需要 Node。docs/adr/0001-main-agent-effort-only-no-model-switch.md)。Dispatch Pilot 每轮、每一步都改主 agent 的 effort,但从不改它的模型:换模型必然让 prompt cache 失效,而在订阅下,同一个模型内切换 effort 保留缓存(2.1.289 上 Opus 5.5 加订阅实测)。Bedrock、Vertex 和各种网关上切换 effort 会让缓存失效,不在支持范围内。Claude Code 的文档只点名了 Opus 5.5、Sonnet 5.5 和 Fable 5.1 保留缓存,其他大多数模型上每档 effort 各有一份缓存,切换会重算整段请求(见 docs/research/decision-models-and-caching.md 的 3.4)。claude plugin marketplace add alexcz-a11y/claude-mods
claude plugin install dispatch-pilot@alex-mods --scope user
想改 mod 或者试未发布的改动时,也可以从本地克隆添加:claude plugin marketplace add ./claude-mods(相对路径要以 ./ 或 ../ 开头,否则会被当成 GitHub 仓库)。从本地目录添加的 marketplace,Claude Code 直接从那个目录加载 mod,克隆里检出的是哪个版本,用的就是哪个版本。
在 shell 里装好的 mod,下次启动 Claude Code 时才加载:装好后重启 Claude Code,或者在已经开着的会话里运行 /reload-plugins。在会话里运行 /plugin,看到 1 mod active · dispatch-pilot 说明 mod 已经加载;运行 /dp 会打开依据面板,/dp status 列出各项功能的开关。接着给它一个决策模型的密钥。
pplx(默认)。 给它 Perplexity 的 API key,有两种给法,同时给了就用第 1 种:
/plugin configure dispatch-pilot@alex-mods,在弹出的配置对话框里填 perplexityApiKey;输入会被遮住。在会话里用 /plugin 安装时也会弹出这个对话框;在 shell 里用 claude plugin install 安装不会弹,装好后用这条命令,或者用 stdin 传进去(需要 jq): read -rs PERPLEXITY_API_KEY && export PERPLEXITY_API_KEY # 粘贴密钥后回车,不回显,值也不会留在 shell 历史里
jq -n '{perplexityApiKey: env.PERPLEXITY_API_KEY}' | claude plugin configure dispatch-pilot@alex-mods --values-stdin
PERPLEXITY_API_KEY(启动 Claude Code 的环境里要有)。Claude Code Desktop 读不到敏感的 userConfig(#35),用环境变量就绕开了这个问题。不要用 claude plugin install --config perplexityApiKey=... 传密钥,也不要用 jq --arg:它们的值会出现在进程参数里,同一台机器上的 ps 看得到。
--values-stdin 读一个 JSON 对象,值都是单行字符串,没写到的选项保持原值。保存之后要重启 Claude Code 才生效,命令也会提示 Configuration saved. Restart Claude Code to apply it.。不带参数运行 claude plugin configure dispatch-pilot@alex-mods 会列出所有选项,并标出哪些还没有设置。
key 只放在请求头里,不会出现在 debug log、看板、依据面板和 $.state 里;会话开始时 debug log 只写 key 从哪里来(选项还是环境变量),不写 key 本身。
Jev。 想用 Jev,就把 decisionModel 设成 jev(在 /config 里选,或写进 pluginConfigs),再填 TypeSafe 的 API key typesafeApiKey,填法同上(配置对话框,或 stdin 的 JSON 里写 typesafeApiKey;环境变量 TYPESAFE_API_KEY 只用来放进 stdin,mod 本身不读它)。设成 jev 就始终用 Jev,哪怕也有 Perplexity 的 key。
只有 TypeSafe 的 key 的老用户不用改任何配置:decisionModel 没设(或设的是 pplx、已移除的 clef、拼错的值)又没有 Perplexity 的 key 时,用 Jev,每个会话的决策日志(/dp 打开的依据面板)里有一条「改用 Jev」写明原因。以后想换成 pplx,只要填上 Perplexity 的 key。
decisionModel 在 /config 里是下拉选择,有 pplx(默认)和 jev 两项;填了别的值(包括已移除的 clef),按没设处理,也就是默认的 pplx(没有 Perplexity 的 key 时按上面的规则退回 Jev)。
从 GitHub 添加 marketplace 的,在 shell 里运行:
claude plugin marketplace update alex-mods
claude plugin update dispatch-pilot@alex-mods
第一条刷新 marketplace 的列表,第二条更新 mod。更新之后重启 Claude Code,或者在开着的会话里运行 /reload-plugins。
claude plugin update 看的是版本号(.claude-plugin/plugin.json 的 version):版本号和你装的一样,它就回答已经是最新版本(is already at the latest version),不换掉本机的副本,哪怕分支上已经有新的提交。新的版本会升版本号。这个 marketplace 默认不自动更新;要打开,在会话里运行 /plugin,到 Marketplaces 里选 alex-mods,再选 Enable auto-update。自动更新同样只在版本号变了时才换。
从本地克隆安装的,不用 claude plugin update:在克隆里取新的提交(例如 git -C claude-mods pull),然后重启 Claude Code 或运行 /reload-plugins。Claude Code 直接从克隆加载 mod,不看版本号。
Dispatch Pilot 和 jev-pilot 不能共存:两者都在 turn.step 上改主 agent 的 effort,会互相覆盖。已经装了 jev-pilot 的话,按这个顺序切换:
/jev-pilot:setup restore。 它把 /jev-pilot:setup 改过的 skill 设置(settings 里的 skillOverrides)恢复成原样,并删掉备份。这条命令是 jev-pilot 自己的,只有它加载着才能用,所以要在停用它之前运行。不恢复的话,那些 skill 会一直对主 agent 隐藏,Dispatch Pilot 也推荐不了它们。没有运行过 /jev-pilot:setup 的话,这一步可以跳过。claude plugin disable jev-pilot@jev-pilot。claude 启动,不要用 claude-jev。 jev-pilot 的 marketplace 安装停用之后,claude-jev 启动器会改用 --plugin-dir 加载它在 ~/.claude/plugins/marketplaces/jev-pilot 里的本地副本,jev-pilot 又被加载进来(还会起它自己的 router),两者就又撞在一起了。选项的值存在两个地方:
typesafeApiKey、perplexityApiKey)存进平台的安全凭据存储(Claude Code 文档的说法),输入时被遮住,不在 /config 里出现。用配置对话框(/plugin configure dispatch-pilot@alex-mods),或者 claude plugin configure ... --values-stdin 填,见「安装」。pluginConfigs 下,在 /config 面板里一项一行,可以直接改(需要 Claude Code 2.1.269 及以上)。两个列表项 skillsAlwaysListed 和 skillsNeverSuggested 不在 /config 里出现,要在 ~/.claude/settings.json 里写成字符串数组。项目和本地 settings 里的 pluginConfigs 会被 Claude Code 忽略。{
"pluginConfigs": {
"dispatch-pilot@alex-mods": {
"options": { "timeoutMs": 2500, "skillsAlwaysListed": ["anthropic-skills:pdf"] }
}
}
}
选项留空就取决策模型的默认值:timeoutMs、contextMessages、contextTokens、thetaMax、thetaUp、thetaDown、rejudgeSteps、rejudgeWaitMs、thetaExpected、agentOverride、skillsMinRelevance、findSkillMinRelevance 这 12 项在 manifest 里没有默认值,所以 /config 里显示为空,你不设时 Dispatch Pilot 取表里你选的决策模型的那一列。你自己设了某一项,就用你设的值。会话开始时 debug log 会写一行哪些选项用了默认值。
另有几个值也随决策模型而定,但不是选项,/config 里改不了,也不在下面的表里:往上取一档的门槛(高一档的概率达到它,就往上取一档;Jev 是 0.3,pplx 是 0.45),问题怎么问(语言和题型:Jev 发消息时判 effort 的那一题用中文,其余所有问题用英文;pplx 所有问题都用英文;都是 Score),以及 contextMessages 的上限(Jev 是 32,pplx 是 2000)。它们和上面那些默认值一起写在 core/setup.ts 的 BACKEND_DEFAULTS,评测和 mod 读的是同一张表。依据卡片的规则推演里,「上取一档」一步写的就是这里的门槛。
pplx 的默认值(B′,ADR 0006,docs/adr/0006-dispatch-pilot-pplx-default-decision-model.md):pplx 的窗口大得多,不受 Jev 那 32k 的限制,所以发消息时的 effort 题、中途重判、派出 agent 和 Workflow 的 state 各取 48000 token、最近 2000 条消息(由 token 预算截断,不是条数),skill 的两段排序(它们带着每个 skill 的画像)仍是 6000,因为 B′ 没有评测过 skill 题。它回答要几秒(eval v2 里 48000 token 的请求约 5 秒),所以发消息等 8000 毫秒(hook 自己的上限是 10 秒),中途重判和 find_skill 各等 6000 毫秒。所有问题都用英文问。thetaMax 0.47、thetaUp 0、thetaDown 0.55 和往上取一档的 0.45 是按已存的 pplx 回答离线定的(DEVELOPMENT.md 的「eval v2 的结果」):0.45 让日常消息判高从 18.5% 降到 13.0%,max 的召回不变。thetaExpected、agentOverride 和两个相关度门槛没有为 pplx 校准过,沿用 Jev 的值。用户自己设的值总是优先。
Jev 的默认值是按 Jev 的上下文上限配置的。 Jev 一个请求最多收 64k token,其中 state 加上最长的那一道题不能超过 32k。最长的题是发消息时 skill 推荐的第一段(111 个 skill 带画像时 Jev 计约 2.19 万 token),所以带 skill 题的请求(发消息时打开了 skill 推荐,以及 find_skill 的第一段)的 contextTokens 取 6000:state 最多约 6.7k(Jev 的计数),加上那一题约 28.6k,在 32k 的 90% 以内,整个请求在 64k 之内,而且这个值正好让每个 skill 的画像不被裁短。其余的请求没有这么长的题(最长的不到 700 token),所以按种类各取一个更大的值:关着 skill 推荐的发消息、中途重判(包括卡住时的)、派出 agent、一批 Workflow 调用,state 最多 24000(Jev 的计数约 26.7k,加上最长的题仍在 32k 的 90% 以内)。你自己设了 contextTokens,所有种类都用它,但每个种类各取它和自己的上限里较小的一个(设 16000 的话,带 skill 题的请求仍是 6000,其余是 16000)。contextMessages 和 rejudgeSteps 取 manifest 允许的最大值(32 条、16 步),让 token 预算而不是条数决定发多少,旧的消息、步骤放不下就整条丢掉。仍然只发文字和工具名,不发工具的输入和输出。这样配置是为了让 Jev 看到尽量多的信息;它没有评测过:评测用的是 4 条、2000 个 token、4 步,更多上下文是否真的判得更准没有数据,延迟见「局限和待评测」。计算过程见 DEVELOPMENT.md 的「Jev 的上下文默认值怎么算」。
下表的 Jev 和 pplx 两列是各自的默认值。起点:这个默认值是暂定的起点,还没有评测数据支持。按 Jev 的上限:按上面的算法取 Jev 能接受的最大值,没有评测。按 AA 基准:按 Artificial Analysis 的基准校正过(见「模型和 effort 的依据」),没有在这套评测集上量过。没有标注的,是按测到的数据定的,或者本来就不需要校准。
| 选项 | 作用 | Jev | pplx |
|---|---|---|---|
decisionModel | 决策模型,在 /config 里是下拉选择:pplx(Perplexity 的 pplx-decider-v1.1-27b)或 jev(TypeSafe 的 Jev)。默认 pplx;填了别的值(包括已移除的 clef),按没设处理。设成 jev 就始终用 Jev;其他情况有 Perplexity 的 key 用 pplx,没有但有 TypeSafe 的 key 就退回 Jev(决策日志里记一条原因),两种都没有就不发请求 | pplx | pplx |
typesafeApiKey | TypeSafe 的 API key,选 jev 时用;没有 Perplexity 的 key 时默认的 pplx 退回 Jev,用的也是它。敏感项 | 空 | 空 |
perplexityApiKey | Perplexity 的 API key,默认的决策模型 pplx 用它。敏感项。为空时读环境变量 PERPLEXITY_API_KEY,两处都有就用这里的;Perplexity 和 TypeSafe 的 key 都没有时不发任何请求。key 只放在请求头里,不进日志、看板和 $.state | 空 | 空 |
pplxQps | 每秒最多向 Perplexity 发几个请求(1–50),只对 pplx 有用,Jev 不受它限制。Tier 0 账户限 1 QPS,一条消息又发 2 到 3 个请求,所以超出的在 mod 里排队,发消息时的 effort 题最先发,排队的时间算进该请求自己的等待。遇到 429 时,剩余等待时间不少于 Retry-After 加约 5 秒就按 Retry-After 重试一次,不够就不经路由地放行,看板写「被限速」。升级了账户就调高 | 1 仅 pplx 用 | 1 |
| 选项 | 作用 | Jev | pplx |
|---|---|---|---|
timeoutMs | 一条消息等决策模型的最长时间(200–8000 毫秒),超过就不经路由地放行 | 1500 | 8000 |
contextMessages | 随你的消息一起发的最近消息条数,从 0 到决策模型能接受的条数(Jev 是 32,pplx 是 2000)。两个模型都默认取上限,由 contextTokens 决定实际发多少 | 32 按 Jev 的上限 | 2000 |
contextTokens | 发给决策模型的 state 的 token 预算(100–16000,pplx 到 48000):你的消息加上最近的消息,按发出去的样子数(连同字段名和转义),中英文按同一个尺度数,旧消息先丢。按请求的种类取:默认值带 skill 题的请求 6000,其余 Jev 24000、pplx 48000,包括发消息时单独发的 effort 请求(见上面);你设了就每个种类取它和自己上限里较小的 | 6000 按 Jev 的上限,其余种类 24000 | 6000,其余种类 48000 |
| 选项 | 作用 | Jev | pplx |
|---|---|---|---|
thetaMax | 用 max 所需的最低概率(0–1);发消息时、中途重判和派出 agent 的 effort 都用这个门槛 | 0.5 起点 | 0.47 按已存的评测回答校准 |
rejudgeEvery | 一轮进行中每隔几步重判一次(0–50);0 表示不按步数重判,派出 agent、启动 Workflow、加载 skill 时仍会重判 | 3 起点 | 3 起点 |
rejudgeSteps | 重判时决策模型读到的最近步数(1–16)。Jev 取上限,由 contextTokens 决定实际发多少 | 16 按 Jev 的上限 | 16 同 Jev |
rejudgeWaitMs | 重判的回答还没到时,下一步最多再等多久(0–8000 毫秒),然后沿用原来的 effort | 300 | 6000 |
thetaUp | 中途升档所需的最低置信度(0–1) | 0.3 按 AA 基准 | 0 按已存的评测回答校准 |
thetaDown | 中途降档所需的最低置信度(0–1,低于 thetaUp 时按 thetaUp 算),一次只降一档 | 0.55 按已存的评测回答扫出 | 0.55 同 Jev |
holdSteps | 升档之后多少步之内不降档(0–50) | 5 按 AA 基准 | 5 按 AA 基准 |
| 选项 | 作用 | Jev | pplx |
|---|---|---|---|
escalateAfter | 计入的工具调用失败满几次就问决策模型并升档(1–20) | 2 起点 | 2 起点 |
escalateMode | one-level 升一档,最高到 xhigh(决策模型自己有把握给更高时可以更高);max 直接升到 max | one-level 起点 | one-level 起点 |
escalateLimit | 一轮(或一个派出 agent)最多升几次(0–10);每次升档后失败计数清零 | 2 起点 | 2 起点 |
thetaExpected | 决策模型认为这些失败是预期内失败的概率达到多少就不升档(0–1) | 0.25 起点 | 0.25 沿用 Jev 的起点 |
escalateHaikuTo | 失败的 haiku agent 接着用哪个模型做:sonnet、opus、fable 或完整的模型 id;留空表示不换。你为这个 agent 点名的模型和排除的模型优先 | sonnet | sonnet |
| 选项 | 作用 | Jev | pplx |
|---|---|---|---|
agentFable | 打开后,决策模型可以为派出 agent(包括 Workflow 里的)选 fable,fable 比 opus 贵;你自己点名 fable 时不受这个开关限制 | false | false |
agentOverride | 主 agent 为派出的 agent 指定了模型时,决策模型的选择要达到这个置信度(0–1)才推翻它;Workflow 脚本里写了 model 时同样适用 | 0.6 起点 | 0.6 沿用 Jev 的起点 |
workflowMode | rewrite:把决定写进脚本再运行;return:第一次提交被拒绝并附上逐个 agent 的推荐,让主 agent 自己写进去,同一个 Workflow 第二次提交直接放行 | rewrite | rewrite |
| 选项 | 作用 | Jev | pplx |
|---|---|---|---|
skillsMax | 一条消息最多推荐几个 skill(0–10) | 3 | 3 |
skillsMinRelevance | 推荐一个 skill 所需的最低相关度(0–1):决策模型对「这个 skill 是否正好做这条消息要做的那种工作」回答「是」的概率 | 0.75 | 0.75 沿用 Jev |
skillsShortlist | 第二段补读 SKILL.md 开头、逐个判断的 skill 最多几个(1–10) | 4 起点 | 4 起点 |
skillsProfileModel | 写 skill 画像的模型,写别名(haiku)或完整的模型 id。通过你的 Claude Code 登录调用,算在你的用量里;换了模型,所有画像重写 | haiku | haiku |
skillsProfilesPerSession | 每次会话开始最多写几份还没有的画像(0–500);0 表示不写 | 30 | 30 |
skillsAlwaysListed | 一直留在主 agent 的 skill 列表里的 skill,写列表里的名字(同步来的 skill 带前缀,例如 anthropic-skills:pdf)。列表项,不在 /config 里 | 空 | 空 |
skillsNeverSuggested | 从不推荐给主 agent、也不提示你的 skill,写法同上。它们照常安装,Skill 工具仍能按名字加载,find_skill 也不返回它们。列表项,不在 /config 里 | 空 | 空 |
findSkillMax | find_skill 一次最多返回几个 skill(1–10) | 5 | 5 |
findSkillMinRelevance | find_skill 返回一个 skill 所需的最低相关度(0–1),比推荐的门槛低:这是主 agent 主动问的,它会自己看描述再决定 | 0.5 起点 | 0.5 沿用 Jev 的起点 |
| 选项 | 作用 | Jev | pplx |
|---|---|---|---|
unresolvedMaxAfter | 未解决次数达到它时,发消息时和中途重判时 effort 题的说明里多一条强提示(这项工作属于多次尝试都没解决的故障);0 表示不给,只带摘要和次数,最大 10。强提示不写档位,不改 pickEffort 和 thetaMax | 3 暂定 | 3 暂定 |
summaryModel | 在你本人的消息开始的那一轮结束后,在后台续写问题摘要的模型,写别名(haiku)或完整的模型 id。通过你的 Claude Code 登录调用,算在你的用量里;摘要是对话的转述,所以 contextMessages 为 0 或 unresolved 开关关着(默认)时不写 | haiku | haiku |
/dp/dp 打开依据面板;已经打开时关掉它
/dp log 打开依据面板(老习惯的别名,不会关掉它)
/dp status 是否开启、effort 的锁定状态、各项功能的开关
/dp on | off 总开关。关闭后不发任何决策请求,每一步都按引擎原样发出(锁定也不生效),脚部写 ○ dp 已关
/dp <功能> on | off 单项功能的开关,例如 /dp main-effort off(功能名见 /dp status 的列表)
/dp lock <档位> 把主 agent 的 effort 锁在 low、medium、high、xhigh 或 max,这一轮的每一步和之后的每一轮都用它,优先于决策
/dp unlock 解除锁定(也可以写 /dp lock off)
/dp log N 在对话里列出最近 N 次决策和理由,最新的在最后(最多 300;日志保留最近 20 轮、最多 300 条)
main-effort、unresolved、midturn-effort、dispatched-agents、workflow-agents、workflow-labels、escalation、skills、skill-profiles、find-skill 和 signals;另有 hook-block-failures,打开后你自己的 settings hook 拦下的调用也算失败,参与强制升档。unresolved 和 hook-block-failures 默认关闭,其余默认开着。/dp unlock 或会话结束时解除。锁定期间决策照常进行并记录,不想为此等决策模型的话,用 /dp main-effort off。/dp log N 显示每项功能记录的决策:做了什么决定、针对哪条消息、理由(决策模型给出的各档概率和置信度)。同样的内容也写进 debug log,不进入对话。/dp signals off 可以停止记录。/dp off;彻底停用就 claude plugin disable dispatch-pilot@alex-mods。局限
find_skill。 一条要写 PR 描述、却没有推荐 pr 的消息,有提示、没有提示、提示写得更主动,主 agent 都直接写了正文;被明确要求查时,它能找到 find_skill 并用上。所以漏掉的推荐,目前只靠发消息时的推荐。contextTokens 6000),发消息时带 skill 推荐的请求最多约 2.9 万 token;其余种类的 state 最多 24000(约 2.67 万 token),中途重判、派出 agent 和关着 skill 推荐的发消息请求,满了的话比以前多约 2.4 万 token,按每 1k 约 13 毫秒外推,最多多约 0.3 秒(同样是外推,没有量过)。已有的实测:Jev 处理 2.19 万 token(skill 第一段,带全部画像)时 p50 约 560 毫秒、p90 约 615 毫秒;另一次整体偏慢的运行里,这一段 218 条中有 24 条(约 11%)超过 1500 毫秒,这些消息的决策整个超时,effort 也没有经过路由。按每多 1k token 约多 13 毫秒外推,现在的 state 再多约 80 毫秒,慢的时段超过 1500 毫秒的消息会比 11% 更多;这是外推,没有在新默认值上量过。如果看板上「决策模型超时」变多,可以把 contextTokens 调小(每少 1k 约快 13 毫秒),或者把 timeoutMs 调大;skill 那一题本身有 2.2 万 token,想去掉这部分延迟只能 /dp skills off。起点):评测数据只够定下少数几项,其余的见下面的「还没有数据的事」。问题用什么语言写。 选 pplx 时所有问题都用英文写(eval v2:pplx 英文问法更准)。选 Jev 时,发消息时判断主 agent effort 的那一个问题用中文写;其余问题(一轮中途重判、派出 agent、Workflow 里的 agent、卡住时的强制升档、skill 推荐和 find_skill)都用英文。依据是 effort-submit 在现在的问法上的对比(Jev,各 1 次):用中文问,中文题 85.0%、英文题 89.0%;用英文问,79.0%、78.0%。同样的请求之前跑过 3 次,中文问法也都高约 8 个百分点。其余问题在现在的问法上没有中文问法的数据,所以都没有改。派出 agent 和 skill 的评测已经有中文问法的变体(models-hint-zh、profiles-zh),还没有运行。发消息时的请求里,中文的 effort 问题和英文的 skill 问题放在一起,这种混合的请求没有单独评测过。
中文和英文的差距。 门槛是中文准确率比英文低不超过 4 个百分点,正好低 4 个百分点也算通过。现在的配置(Jev,effort 用中文问)在 effort-submit 上那一次正好差 −4.0,同样的请求之前 3 次是 0、0、−1;而中文题的准确率比用英文问时高 6 个百分点。其余三套评测都在门槛之内(见「评测」),但都是改措辞之前测的。
派出 agent 的模型文字按 AA 的基准重写,并用真实的 Jev 调过。 重写后的第一次评测(models-hint,100 题中英各一遍)整体从 70%/68% 掉到 60%/58%(模型部分 82/81,effort 部分 64/62)。之后调了三轮文字和 sonnet 的 effort 下限(sonnet 从 high 降到 medium),最后一版是整体 69%/68%、模型部分 89%/91%、effort 部分 71%/71%,同一版重复一次是 69%/67%、89%/90%、71%/71%。和重写之前(70%/68%,85%/85%,74%/72%)比,模型部分更好,effort 部分低 3 到 1 个百分点:往上取一档和模型下限让偏高多了、gold 命中从 50%/48% 变成 44%/44%,因为数据集的 gold 是按「够用的最便宜档」标的,没有按 AA 重标。注意这三轮文字是对着这 100 题的错题调的,所以最后的模型部分是在调过的题上量的,对没见过的请求会低一些。逐轮的表、发消息和中途重判的真实对照在 [docs/research/aa-benchmarks-2026-10.md](../docs/research/aa-benchmarks-2026
hooks/dispatch-pilot.ts 54 lines1// Dispatch Pilot's entry: it only assembles. Each feature lives in its own
2// file under features/ and registers its own hooks; the core goes last.
3//
4// Order is nesting: what registers first is outermost and sees an event
5// first. Features go above the core so that on every event they share, each
6// feature has run before the core finishes the event (sends the decision
7// request, writes the step). To add a feature: one file, one line here,
8// above `registerCore` (DEVELOPMENT.md, 开发).
9
10import type { Register } from 'claude-code'
11import { registerScreens } from './board/screens.tsx'
12import { registerCore } from './core/core.ts'
13import { registerReport } from './core/report.ts'
14import { setup } from './core/setup.ts'
15import { registerControl } from './features/control.ts'
16import { registerDispatchedAgents } from './features/dispatched-agents.ts'
17import { registerEscalation } from './features/escalation.ts'
18import { registerFindSkill } from './features/find-skill.ts'
19import { registerMainEffort } from './features/main-effort.ts'
20import { registerMidturnEffort } from './features/midturn-effort.ts'
21import { registerUnresolved } from './features/unresolved.ts'
22import { registerSkillProfiles } from './features/skill-profiles.ts'
23import { registerSkills } from './features/skills.ts'
24import { registerWorkflowAgents } from './features/workflow-agents.ts'
25import { registerWorkflowLabels } from './features/workflow-labels.ts'
26
27export const register: Register = (on, options) => {
28 const ctx = setup(options)
29
30 // The decision report keeps the board's turn count: outermost, so a turn is counted before anything beneath reads the board.
31 registerReport(on)
32 // The screens draw the report's data (the band, the footer); no feature touches them.
33 registerScreens(on)
34
35 // Features, outermost first.
36 registerControl(on, ctx)
37 registerMainEffort(on, ctx)
38 registerUnresolved(on, ctx)
39 registerEscalation(on, ctx) // above the mid-turn feature: its raise is in the plan before a mid-turn answer is applied
40 registerMidturnEffort(on, ctx)
41 registerDispatchedAgents(on, ctx)
42 registerWorkflowAgents(on, ctx)
43 // Inside workflow-agents: a Workflow call is this feature's only once that feature has sent the script on, so the
44 // agents of other runs never wait while it is still deciding.
45 registerWorkflowLabels(on, ctx)
46 // Outside skills: at session start it writes the profiles of the skills that feature has just read.
47 registerSkillProfiles(on, ctx)
48 registerSkills(on, ctx)
49 registerFindSkill(on, ctx)
50
51 // The core: always last.
52 registerCore(on, ctx)
53}
54hooks/board/screens.tsx 178 lines1// The screens' hooks (ADR 0004: drawn from the decision report's data, never
2// written to by a feature): the band above the prompt (`AbovePrompt`), the
3// summary at the right end of the footer (`SessionMode`) and the rationale pane
4// (`Pane`, the pane `/dp` opens). One hook tree for every surface, branching on
5// `e.surface`: the terminal draws the full design (band.tsx, footer.tsx,
6// pane.tsx) with its Raster; every other surface gets the same trees in plain
7// text, without a Raster (the tag in the footer in at most 24 characters, the
8// log in a window short of Desktop's 2000 nodes), until the Svg versions.
9//
10// Each hook reads the board, the log and the agent the person picked from
11// $.state while it draws, so the host draws it again when they change; while
12// an agent runs it asks for one more frame every `TICK_MS` (the spinner, the
13// running clocks and ribbons). The band and the footer keep what other mods
14// draw in the same place: the engine's drawing beneath goes in first (spec
15// story 38); the pane is Dispatch Pilot's own.
16
17import type { EngineInterface, On, Timer } from 'claude-code'
18import { errorText } from '../decision/backend.ts'
19import { report, type NoticeIo } from '../core/report.ts'
20import { isOn, isShown, listSwitches, masterOn } from '../core/switches.ts'
21import { bandTree } from './band.tsx'
22import { FOOTER_COLUMNS, footerTree, TAG_CHARS } from './footer.tsx'
23import { paneTree } from './pane.tsx'
24import { nodeCount, showsText } from './kit.tsx'
25import { NO_PANE_STATE, PANE_COLUMNS, PANE_ID, PANE_TITLE, withFold } from './rationale.ts'
26import { screenView, TICK_MS, type AgentRow, type ScreenView } from './view.ts'
27
28const BOARD = { plugin: 'dispatch-pilot', key: 'board' } as const
29const DECISIONS = { plugin: 'dispatch-pilot', key: 'decisionLog' } as const
30const SELECTED = { plugin: 'dispatch-pilot', key: 'selected' } as const
31const PROFILES = { plugin: 'dispatch-pilot', key: 'skillProfiles' } as const
32const LOCK = { plugin: 'dispatch-pilot', key: 'lock' } as const
33const PANE_VIEW = { plugin: 'dispatch-pilot', key: 'paneView' } as const
34const COUNT = { plugin: 'dispatch-pilot', key: 'unresolved' } as const
35
36/** The log's windows for a surface that refuses a large tree, widest first: what the pane keeps of the log and of the re-decisions. */
37const WINDOWS = [
38 { entries: 40, mids: 20 },
39 { entries: 16, mids: 8 },
40 { entries: 4, mids: 3 },
41 { entries: 0, mids: 0 },
42] as const
43/** A tree this big is drawn again in the next window: Desktop refuses 2000. */
44const MOST_NODES = 1800
45
46/** The next frame asked for while an agent runs (module state: a hot reload cancels it with the old module, and the next draw asks again). */
47let ticking: Timer | null = null
48
49/** The screens' view, read from $.state now (a read while drawing subscribes the drawing to the value). */
50async function viewOf($: EngineInterface, isWorking: boolean): Promise<ScreenView> {
51 const [board, log, picked, now] = await Promise.all([$.state.get(BOARD), $.state.get(DECISIONS), $.state.get(SELECTED), $.clock.now()])
52 return screenView({
53 board: board.value ?? { turn: 0, nodes: [] },
54 log: log.value ?? [],
55 now,
56 selected: picked.value ?? null,
57 master: masterOn(),
58 isOn,
59 isShown,
60 isWorking,
61 })
62}
63
64/** One more frame in `TICK_MS` while something runs; none once nothing does. */
65function tick($: EngineInterface, view: ScreenView): void {
66 const moving = view.live && (view.counts.running > 0 || view.rows.some((row) => row.node.state === 'running'))
67 if (!moving || ticking !== null) return
68 try {
69 ticking = $.clock.after(TICK_MS, () => {
70 ticking = null
71 $.ui.invalidate('ui.render')
72 })
73 } catch {
74 // no timer: the screens move on at the board's next change
75 ticking = null
76 }
77}
78
79/** The agent picked, its card shown in the rationale pane (written from a press, never while drawing). */
80async function pick($: EngineInterface, row: AgentRow): Promise<void> {
81 await $.state.set(SELECTED, { turn: row.node.turn, id: row.node.id })
82}
83
84export function registerScreens(on: On): void {
85 // The band above the prompt.
86 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
87 if (e.props.hasSurvey) return next(e)
88 const own = await next(e)
89 try {
90 const t = $.ui.resolve(e)
91 // The time ribbons are the terminal's Raster; elsewhere the rows are the same without them.
92 const ribbon = e.surface === 'terminal' ? $.ui.resolve(e).Raster : undefined
93 const view = await viewOf($, e.props.isWorking)
94 tick($, view)
95 // A digit picks an agent and brings up its card: a press is the person's asking, so the pane is placed at
96 // any width; it opens without taking the keys, so the next digit still picks. One not placed is closed again,
97 // and the decision report tells the person why in a toast.
98 const select = async (row: AgentRow) => {
99 await pick($, row)
100 const opened = await $.ui.open({ id: PANE_ID, title: PANE_TITLE, closeOnEscape: true, columns: PANE_COLUMNS })
101 if (!opened.isPlaced) {
102 await $.ui.close({ id: PANE_ID })
103 const io: NoticeIo = { debug: (line) => $.ui.log(line, { to: 'debug' }), now: () => $.clock.now(), toast: (text) => $.ui.toast(text) }
104 await report(io, { unplaced: { reason: opened.reason } })
105 }
106 }
107 const tree = bandTree(t, view, { cols: e.props.bodyColumns, rows: e.props.maxRows }, { select: (row) => void select(row).catch(() => undefined) }, ribbon)
108 if (tree === null) return own
109 const { Box } = t
110 return (
111 <Box flexDirection="column">
112 {own}
113 {tree}
114 </Box>
115 )
116 } catch (error) {
117 // A band that cannot be drawn leaves the place to the others.
118 $.ui.log(`band not drawn: ${errorText(error)}`, { to: 'debug' })
119 return own
120 }
121 })
122
123 // The right end of the footer, beside the engine's modes.
124 on('ui.render', { component: 'SessionMode' }, async ($, e, next) => {
125 const own = await next(e)
126 try {
127 const t = $.ui.resolve(e)
128 const view = await viewOf($, false)
129 tick($, view)
130 // A column between the engine's modes (or another mod's drawing) and the tag, when there are any: it counts in what the tag adds.
131 const gap = showsText(own, e.props.modes.length > 0) ? 1 : 0
132 const tree = footerTree(t, view, (e.surface === 'terminal' ? FOOTER_COLUMNS : TAG_CHARS) - gap)
133 if (tree === null) return own
134 const { Box } = t
135 return (
136 <Box gap={gap}>
137 {own}
138 {tree}
139 </Box>
140 )
141 } catch (error) {
142 $.ui.log(`footer not drawn: ${errorText(error)}`, { to: 'debug' })
143 return own
144 }
145 })
146
147 // The rationale pane (`/dp`): the card of the agent picked, the decision log.
148 on('ui.render', { component: 'Pane', requestId: PANE_ID }, async ($, e, next) => {
149 try {
150 const t = $.ui.resolve(e)
151 const [view, log, profiles, lock, kept, count] = await Promise.all([viewOf($, false), $.state.get(DECISIONS), $.state.get(PROFILES), $.state.get(LOCK), $.state.get(PANE_VIEW), $.state.get(COUNT)])
152 tick($, view)
153 const entries = log.value ?? []
154 const state = kept.value ?? NO_PANE_STATE
155 const shown = isOn('skills') && isOn('skill-profiles') ? (profiles.value ?? null) : null
156 // The summary the session holds, while the feature that writes it is on.
157 const summary = isOn('unresolved') ? (count.value?.summary ?? null) : null
158 const input = { view, log: entries, profiles: shown, switches: listSwitches().map((s) => ({ name: s.name, on: s.on })), master: masterOn(), lock: lock.value ?? null, summary, state }
159 const act = {
160 pick: (row: AgentRow) => void pick($, row).catch(() => undefined),
161 fold: (turn: number, open: boolean) => void $.state.set(PANE_VIEW, withFold(state, turn, open, entries)).catch(() => undefined),
162 failures: (open: boolean) => void $.state.set(PANE_VIEW, { ...state, failures: open }).catch(() => undefined),
163 }
164 if (e.surface === 'terminal') return paneTree(t, input, e.props.bodyColumns, act, $.ui.resolve(e).Raster)
165 // Elsewhere: the same pane in text, in the widest window of the log that stays short of Desktop's 2000 nodes.
166 let tree = paneTree(t, { ...input, window: WINDOWS[0] }, e.props.bodyColumns, act)
167 for (const window of WINDOWS.slice(1)) {
168 if (nodeCount(tree) < MOST_NODES) break
169 tree = paneTree(t, { ...input, window }, e.props.bodyColumns, act)
170 }
171 return tree
172 } catch (error) {
173 $.ui.log(`rationale pane not drawn: ${errorText(error)}`, { to: 'debug' })
174 return next(e)
175 }
176 })
177}
178hooks/core/core.ts 182 lines1// The core: the innermost hooks of every event it shares with the features.
2// The entry registers it last, so on each event every feature's hook has run
3// (on the way down) before the core finishes the event:
4//
5// prompt.submit sends the message's ballot as up to two decision requests
6// (the effort question alone, the rest together; ADR 0005),
7// in parallel, hands each part its answers, then lets the
8// prompt in
9// turn.start opens the main turn's record, with the effort its prompt
10// was decided at
11// turn.step the one writer of effort (and of a non-main loop's model):
12// sends every step as the plan table says
13// classic.PreToolUse notes the calls a settings hook refused (core/outcomes.ts)
14// command.run notes the command the person runs: the prompt that follows
15// may be its command turn (core/commands.ts)
16//
17// The core owns the unmatched registration of these events; a feature always
18// registers them with a matcher (DEVELOPMENT.md, 开发).
19//
20// Switched off (`/dp off`, core/switches.ts) the core stands down: prompt.submit
21// asks nothing and turn.step sends every step as the engine made it.
22
23import type { HttpInit, On } from 'claude-code'
24import { messageText } from '../decision/context.ts'
25import { EFFORT_PART } from '../decision/effort.ts'
26import { SKILLS_PART } from '../decision/skills.ts'
27import { answersFor, type State } from '../decision/system-one.ts'
28import { messageRequest } from '../decision/turn-start.ts'
29import { type Asked, type BackendIo, describeAsked, errorText } from '../decision/backend.ts'
30import { collect, type Contribution, type PartOutcome } from './ballot.ts'
31import { forgetCommand, noteCommand, typedCommand } from './commands.ts'
32import { noteBlocked } from './outcomes.ts'
33import { newTurn, planStep, replace, takePending, turnKey, type Cell, type PendingDecision } from './plans.ts'
34import { messageLimits, type Ctx } from './setup.ts'
35import { reportStep, type StepIo } from './report.ts'
36import { masterOn } from './switches.ts'
37
38const PENDING = { plugin: 'dispatch-pilot', key: 'pending' } as const
39const TURNS = { plugin: 'dispatch-pilot', key: 'turns' } as const
40const AGENTS = { plugin: 'dispatch-pilot', key: 'agents' } as const
41const LOCK = { plugin: 'dispatch-pilot', key: 'lock' } as const
42const BOARD = { plugin: 'dispatch-pilot', key: 'board' } as const
43const DECISIONS = { plugin: 'dispatch-pilot', key: 'decisionLog' } as const
44const WORKFLOW_RUNS = { plugin: 'dispatch-pilot', key: 'workflowRuns' } as const
45const LABEL_RUNS = { plugin: 'dispatch-pilot', key: 'labelRuns' } as const
46const PPLX_RATE = { plugin: 'dispatch-pilot', key: 'pplxRate' } as const
47
48/** Prompts whose prompt.submit is letting them in right now: a turn that starts meanwhile is theirs. */
49const entering: string[] = []
50
51export function registerCore(on: On, ctx: Ctx): void {
52 on('prompt.submit', async ($, e, next) => {
53 // Every feature above has seen whether this is a command turn.
54 forgetCommand(e.text)
55 const ballot = collect(e.text)
56 // Switched off (/dp off): whatever was put in the ballot is not asked.
57 if (ballot.length === 0 || !masterOn()) return next(e)
58 const messages = ctx.config.context.messages > 0 ? await $.session.messages().catch(() => []) : []
59 const io: BackendIo = {
60 fetch: (url: string, init: HttpInit) => $.http.fetch(url, init),
61 sleep: (ms: number, signal: AbortSignal) => $.clock.sleep(ms, { signal }),
62 pace: { now: () => $.clock.now(), sentAt: { get: () => $.state.get(PPLX_RATE), set: (value, options) => $.state.set(PPLX_RATE, value, options) } },
63 }
64 // The main agent's effort goes in a request of its own (ADR 0005), the rest of the ballot in one. Both go out at once,
65 // each is answered or fails on its own, and each part reads the outcome of the request it was in.
66 const blocks = (
67 await Promise.all(
68 requestGroups(ballot).map(async (group) => {
69 const startedAt = await $.clock.now()
70 const ids = group.flatMap((part) => Object.keys(part.questions).map((id) => `${part.part}.${id}`)).join(', ')
71 let asked: Asked
72 let state: State = {}
73 try {
74 // The state's budget follows the request's longest question: the skills' question leaves less than any other.
75 const limits = messageLimits(ctx.config, group.some((part) => part.part === SKILLS_PART))
76 const request = messageRequest({ prompt: e.text, messages, limits, parts: group })
77 state = request.state
78 asked = await ctx.backend.ask(io, request, ctx.config.timeoutMs)
79 } catch (error) {
80 // A part's malformed questions: nothing was sent.
81 asked = { ok: false, failure: { kind: 'request', detail: errorText(error) } }
82 }
83 const ms = (await $.clock.now()) - startedAt
84 $.ui.log(`request [${ids}] to ${ctx.backend.name}: ${describeAsked(asked, ms)}`, { to: 'debug' })
85 return settleAll(group, asked, state)
86 }),
87 )
88 ).flat()
89 entering.push(e.text)
90 try {
91 return await next(blocks.length > 0 ? { ...e, context: [...(e.context ?? []), ...blocks] } : e)
92 } finally {
93 const at = entering.indexOf(e.text)
94 if (at >= 0) entering.splice(at, 1)
95 }
96 })
97
98 on('command.run', async ($, e, next) => {
99 noteCommand(e)
100 return next(e)
101 })
102
103 on('turn.start', async ($, e, next) => {
104 const pending: Cell<PendingDecision[]> = { get: () => $.state.get(PENDING), set: (value, options) => $.state.set(PENDING, value, options) }
105 // A command turn starts with the engine's command message: the prompt was the command as typed.
106 const started = typedCommand(e.text) ?? e.text
107 // Its own prompt's decision: the prompt now entering (whatever an inner
108 // hook made of its text), else a queued prompt with the turn's text.
109 const texts = entering.length === 1 ? [entering[0] as string, started] : [started]
110 const { before } = await replace(pending, (list) => takePending(list ?? [], texts).rest)
111 // What the landed write took out of the list it found.
112 const own = takePending(before ?? [], texts).taken
113 // The turn's message as the decision model read it (later decisions about the turn reuse it). A pending
114 // entry, decided or not, says the person's own message started the turn (only such a turn is re-decided);
115 // one marked `report` says a report did (decided at its start, not re-decided).
116 const prompt = messageText(started, ctx.config.contextByKind.rejudge)
117 await $.state.set({ ...TURNS, id: turnKey(e.turnId, undefined) }, newTurn(prompt, own?.effort ?? null, own !== null && own.report !== true))
118 return next(e)
119 })
120
121 // Beneath every feature's tool.call: which calls a PreToolUse settings hook refused (the call itself then only
122 // sees an error), for the features that count failures and read a loop's steps (core/outcomes.ts).
123 on('classic.PreToolUse', async ($, e, next) => {
124 const decided = await next(e)
125 if (decided.deny !== undefined && e.tool_use_id) noteBlocked(e.tool_use_id)
126 return decided
127 })
128
129 on('turn.step', async function* ($, e, next) {
130 // Switched off (/dp off): every step goes out as the engine made it, lock or no lock.
131 if (!masterOn()) return yield* next(e)
132 const agentId = e.agentId
133 const { value: turn } = await $.state.get({ ...TURNS, id: turnKey(e.turnId, agentId) })
134 const { value: agent } = agentId === undefined ? { value: undefined } : await $.state.get({ ...AGENTS, id: agentId })
135 const { value: lock = null } = agentId === undefined ? await $.state.get(LOCK) : { value: null }
136 const { step, source } = planStep(e, { lock, turn, agent })
137 // What the step goes out with, for the board: every loop's, the main agent's included.
138 const io: StepIo = {
139 board: { get: () => $.state.get(BOARD), set: (value, options) => $.state.set(BOARD, value, options) },
140 decisions: { get: () => $.state.get(DECISIONS), set: (value, options) => $.state.set(DECISIONS, value, options) },
141 debug: (line) => $.ui.log(line, { to: 'debug' }),
142 now: () => $.clock.now(),
143 toast: (text) => $.ui.toast(text),
144 agents: () => $.agent.list(),
145 runDirs: async () => {
146 const [noted, labelled] = await Promise.all([$.state.get(WORKFLOW_RUNS), $.state.get(LABEL_RUNS)])
147 return [...new Set([...(noted.value ?? []), ...(labelled.value ?? [])].map((run) => run.dir))]
148 },
149 journal: async (dir) => {
150 const path = `${dir}/journal.jsonl`
151 return (await $.fs.exists(path)) ? await $.fs.read(path).catch(() => null) : null
152 },
153 }
154 await reportStep(io, { ...(agentId === undefined ? {} : { agentId }), model: step.model, effort: step.effort, source })
155 return yield* next(step)
156 })
157}
158
159/**
160 * The requests a ballot goes out in: the main agent's effort question alone, every other part's together (the skills'
161 * question is the longest and sets a smaller state budget, so it cannot share a request with the effort question without
162 * cutting the conversation the effort question reads). A group no part is in is not made: no empty request.
163 */
164function requestGroups(ballot: readonly Contribution[]): Contribution[][] {
165 return [ballot.filter((part) => part.part === EFFORT_PART), ballot.filter((part) => part.part !== EFFORT_PART)].filter((group) => group.length > 0)
166}
167
168/** Hands each part its outcome (in parallel) and gathers the context blocks they return, in ballot order. */
169async function settleAll(ballot: readonly Contribution[], asked: Asked, state: State): Promise<string[]> {
170 const settled = await Promise.all(
171 ballot.map(async (contribution) => {
172 const outcome: PartOutcome = asked.ok ? { ok: true, answers: answersFor(contribution, asked.answers), state } : { ok: false, failure: asked.failure }
173 try {
174 return (await contribution.settle(outcome)) ?? []
175 } catch {
176 return []
177 }
178 }),
179 )
180 return settled.flat()
181}
182hooks/core/report.ts 1362 lines1// The decision report (「决定汇报」, ADR 0004): the one writer of what Dispatch
2// Pilot shows people. The board, the decision log, the debug log and the toasts
3// are all written here, from the same structured data in $.state
4// (types/index.d.ts: `board`, `decisionLog`, `skillProfiles`); the screens
5// (hooks/board/) only draw that data.
6//
7// Two entries, and nothing else, are for others to call:
8//
9// report(io, what) 记一条决定: a feature (or a screen) hands over one
10// thing to show, tagged by its kind (`Reported`):
11// `decision`, a decision it made or why it could not
12// make one (which feature, about which agent, the
13// outcome, why, the rules' working when there is one:
14// the decision log, the agent's node, the debug log; a
15// route that failed raises a toast); `decisions`,
16// several of one event (the agent() calls of a
17// Workflow script: one write each, one toast at most);
18// `tally`, what a feature counts as a loop goes (the
19// mid-turn re-decisions, a loop's failed calls and
20// forced raises: on the agent's node); `switched`, the
21// person flipped the mod or a feature (`/dp`: the
22// screens draw again); `profiles`, how writing the
23// session's skill profiles goes (#33); `unplaced`, a
24// pane the person asked for that the surface did not
25// place (a toast). `io` is what that kind needs of the
26// host (`IoOf`): a `ReportIo` for most.
27// reportStep(io, step) 记一步读数: the reading of one model request, the
28// model and effort an agent's step went out with,
29// whoever's loop it is. Only observes, never decides
30// anything (ADR 0003). The core calls it from its
31// `turn.step`, which knows what the step goes out with.
32//
33// The rest of a loop's life (spawned, ended) and the board's turn count the
34// module hears with hooks of its own (`registerReport`, which the entry file
35// registers first: `turn.start`, `agent.spawn`, `turn.complete`, changing nothing
36// in them); nobody calls those. What else is exported (`decisionLine`,
37// `appendEntry`, `startTurn`, `profileWhy`, the types) only shapes or reads data.
38//
39// Pure but for those hooks: `$` stays in the hook owner's file (it may not cross
40// an import), so the caller builds its io of closures, a `ReportIo` as:
41//
42// const io: ReportIo = {
43// board: { get: () => $.state.get(BOARD), set: (value, options) => $.state.set(BOARD, value, options) },
44// decisions: { get: () => $.state.get(DECISIONS), set: (value, options) => $.state.set(DECISIONS, value, options) },
45// debug: (line) => $.ui.log(line, { to: 'debug' }),
46// now: () => $.clock.now(),
47// toast: (text) => $.ui.toast(text),
48// }
49//
50// where `BOARD = { plugin: 'dispatch-pilot', key: 'board' } as const` and
51// `DECISIONS = { plugin: 'dispatch-pilot', key: 'decisionLog' } as const` are
52// the file's own literal refs (DEVELOPMENT.md, 开发). Neither entry ever
53// throws: a report that cannot be kept must not stop the decision or the step.
54
55import type { EngineInterface, On } from 'claude-code'
56import { errorText, failureLine, failureWords, type Failure } from '../decision/backend.ts'
57import { modelFamily, type AgentModel } from '../decision/dispatched-agent.ts'
58import { isEffort, type Effort } from '../decision/effort.ts'
59import type { BackendName } from './setup.ts'
60import type { UnresolvedChange, UnresolvedOption, UnresolvedThresholds } from '../decision/unresolved.ts'
61import { startedIn } from '../decision/workflow-labels.ts'
62import { type Cell, type EffortSource, update } from './plans.ts'
63
64const BOARD = { plugin: 'dispatch-pilot', key: 'board' } as const
65const DECISIONS = { plugin: 'dispatch-pilot', key: 'decisionLog' } as const
66const WORKFLOW_RUNS = { plugin: 'dispatch-pilot', key: 'workflowRuns' } as const
67const LABEL_RUNS = { plugin: 'dispatch-pilot', key: 'labelRuns' } as const
68
69// ---- the data (mirrors types/index.d.ts) --------------------------------------
70
71/** A model family, as a person says it (`/model`). */
72export type Model = AgentModel
73/** How a log entry reads: a decision made, a thing to watch, a failure, a plain record. */
74export type Tone = 'ok' | 'warn' | 'fail' | 'info'
75export type AgentState = 'queued' | 'running' | 'done' | 'failed'
76
77/** One step of the rules' working: its rule, whether it took effect, and the rule's own fields. */
78export type RuleStep = { rule: string; applied: boolean; [field: string]: string | number | boolean | null }
79
80/** A failed decision request, with the decision model that failed. */
81export type NodeFailure = Failure & { backend: string }
82
83/** One agent of one turn on the board. */
84export type BoardNode = {
85 turn: number
86 /**
87 * `main`, the agentId, or for a Workflow's agent not started yet the id of its call (`callNodeId`), for a
88 * Workflow not routed as a whole its tool call's id.
89 */
90 id: string
91 kind: 'main' | 'agent' | 'wf'
92 name: string
93 type: string
94 model?: Model
95 effort?: Effort | number
96 state: AgentState
97 t0: number
98 dur?: number
99 routed: boolean
100 locked?: true
101 why?: string
102 failure?: NodeFailure
103 decision?: number
104 /** The Workflow it belongs to (a Workflow agent, or the call of the script that stands for it before it starts). */
105 workflow?: { id: string; name: string }
106 /** The main agent's mid-turn re-decisions of the turn: steps made, decisions, level changes; `late`: an answer is not back, `failure`: the latest request failed. midturn-effort's. */
107 midturn?: { steps: number; judged: number; changed: number; late?: true; failure?: NodeFailure }
108 /** The failed tool calls, hook blocks and forced raises of the agent's loop (the escalation feature counts them); absent while none. */
109 counts?: { failed: number; blocked: number; raised: number; late?: true }
110}
111
112/** What an agent's step read: its model by family and its effort as it went out (absent: a model without effort). */
113export type Reading = { model?: Model; effort?: Effort | number }
114
115/**
116 * The id of the node a Workflow call of a script stands on until its agent starts (an agent of a Workflow has no id
117 * before): `<tool_use_id>#<index>`, the Workflow tool call's own id and the call's place in the script.
118 */
119export function callNodeId(workflow: string, index: number): string {
120 return `${workflow}#${index}`
121}
122
123/** The Workflow tool call a call's node (`callNodeId`) belongs to; null for any other node. */
124export function callWorkflowOf(id: string): string | null {
125 const at = id.indexOf('#')
126 return at < 0 ? null : id.slice(0, at)
127}
128
129/** One reading that differed from the agent's previous one in the same turn (types/index.d.ts, `board.changes`). */
130export type ReadingChange = {
131 turn: number
132 id: string
133 /** Seconds from the start of `turn`. */
134 at: number
135 /** The `n` of the decision log's last entry when it was read. */
136 after: number
137 from: Reading
138 to: Reading
139}
140
141/**
142 * What a feature met beside an agent's route that is neither a decision nor the route itself, for the band's
143 * event stream (types/index.d.ts, `board.notes`): a request that failed (`failed`, `why` its words), one that
144 * gave nothing to use (`skipped`, `why` the Skip), an answer not back in time (`late`, `why` '').
145 */
146export type BoardNote = {
147 turn: number
148 /** `main`, or the agentId it was about. */
149 id: string
150 feature: string
151 /** Seconds from the start of `turn`. */
152 at: number
153 /** The `n` of the decision log's last entry when it was written. */
154 after: number
155 kind: 'failed' | 'skipped' | 'late'
156 why: string
157}
158
159export type Board = {
160 turn: number
161 /** When the latest turns started (`$.clock.now()`), by turn. */
162 starts?: { turn: number; at: number }[]
163 changes?: ReadingChange[]
164 notes?: BoardNote[]
165 nodes: BoardNode[]
166}
167
168/** The board keeps at most this many notes (those of the nodes' turns). */
169const NOTES_KEPT = 50
170
171/** One decision as the log keeps it. */
172export type LogEntry = {
173 n: number
174 turn: number
175 /** Seconds from the start of `turn` to when it was reported (0 for a decision made before its turn started); absent in an entry an earlier version wrote. */
176 at?: number
177 feature: string
178 agent?: string
179 tone: Tone
180 outcome: string
181 subject: string
182 reason: string
183 probs?: Record<Effort, number>
184 conf?: number
185 trace?: RuleStep[]
186 floor?: { from: Effort; to: Effort; model: Model }
187 /** The model (by family) and effort an agent was decided to run with (the main agent: its effort only); the agent's node says what its steps went out with. */
188 model?: Model
189 effort?: Effort
190 mid?: { current: Effort; picked: Effort; result: Effort; threshold?: number; held?: string; remaining?: number }
191 forced?: ForcedRaise
192 skills?: SkillsPicked
193 counts?: { failed: number; blocked: number; raised: number }
194 /** A Workflow call's decision that was sent back to the main agent to write in (return mode). */
195 sentBack?: true
196 /** An unresolved-count judgement (`unresolved`). */
197 unresolved?: UnresolvedRecord
198 /** The strong hint was given with this decision's request (#41). */
199 hint?: HintRecord
200}
201
202/** The strong hint given with a request: the count it was given at, the setting it had reached, and `mid` when it was a mid-turn re-decision's. */
203export type HintRecord = { count: number; maxAfter: number; where?: 'mid' }
204
205/**
206 * What one message's answer to the unresolved question did to the count: the count before and after, the change, the
207 * option the answer leaned to most and every option's probability, the backend's confidence (on record only), and
208 * the two bars it was held to. The card draws it as it is; nothing is worked out again (ADR 0004).
209 */
210export type UnresolvedRecord = {
211 before: number
212 count: number
213 change: UnresolvedChange
214 top: UnresolvedOption
215 probs: Record<UnresolvedOption, number>
216 conf?: number
217 thresholds: UnresolvedThresholds
218}
219
220/** A raise the failures forced: the level (or, for a haiku agent, the model) it went from and to, and the level it holds the agent at least at. */
221export type ForcedRaise = { kind: 'effort' | 'model'; from: string; to: string; floor?: Effort }
222
223/** The skills a message (or a find_skill call) got: those suggested to the main agent, and those only the person can start (「可试 /x」), each with its relevance. */
224export type SkillsPicked = { suggest: { name: string; relevance: number }[]; try: { name: string; relevance: number }[] }
225
226/** The log keeps the entries of this many latest turns, and at most `LOG_ENTRIES` of them. */
227export const LOG_TURNS = 20
228export const LOG_ENTRIES = 300
229
230// ---- what a feature hands over ------------------------------------------------
231
232/** What every report says of who it is about. */
233type About = {
234 /** The feature's switch name; a report's decision adds ` (agent report)`. */
235 feature: string
236 /** `main`, or the agentId of the dispatched or Workflow agent. */
237 agent: string
238 /**
239 * `next`: made before the turn it is for has started (at `prompt.submit`, for a message that will start
240 * the turn), so it belongs to the turn to come; `current` (the default): the turn running now.
241 */
242 forTurn?: 'current' | 'next'
243 /** What it was about, for the log: the start of the message, an agent's label. */
244 subject?: string
245 /**
246 * How to start the agent's node when the board has none yet (the main agent's is known). `state`: `running` by
247 * default, `queued` for a decision made for a turn to come or for a Workflow call whose agent has not started.
248 * `workflow`: the Workflow the agent (or its call) belongs to.
249 */
250 node?: { kind: 'agent' | 'wf'; name: string; type: string; state?: AgentState; workflow?: { id: string; name: string } }
251 /** Whether it left the agent routed (a decision) or not (a failure): the node's `routed`. Left as it is when not given. */
252 routed?: boolean
253 /**
254 * The id of the node that stood for this agent before it started (a Workflow call, queued): the agent's node
255 * takes its place, and the decision made for the call (its `decision` link), when it has one.
256 */
257 replaces?: string
258 /**
259 * A decision beside the agent's route (a mid-turn re-decision, a forced raise, a skill suggestion): in the log,
260 * and not on the agent's node (its `decision` link, `routed` and `why` stay the route's).
261 */
262 aside?: true
263}
264
265/** A decision made. */
266export type Decided = About & {
267 /** What was decided, in a few words: `effort high`. */
268 outcome: string
269 /** Why: what the decision model said, the rule that applied. */
270 reason: string
271 tone?: Tone
272 probs?: Record<Effort, number>
273 conf?: number
274 trace?: RuleStep[]
275 floor?: LogEntry['floor']
276 mid?: LogEntry['mid']
277 /** The model (by family) and the effort it decided for an agent: in the log entry, since the node's own are what its steps read. */
278 model?: Model
279 effort?: Effort
280 /** A Workflow call's decision that was sent back to the main agent to write in (return mode). */
281 sentBack?: true
282 forced?: ForcedRaise
283 skills?: SkillsPicked
284 counts?: LogEntry['counts']
285 unresolved?: UnresolvedRecord
286 hint?: HintRecord
287}
288
289/** A decision the feature could not make: the decision request failed. Not an entry of the log; it says why on the agent's node. */
290export type NotDecided = About & { failure: NodeFailure }
291
292/**
293 * No decision was asked for, or none could be made without a failed request: why the agent (or the Workflow) runs
294 * as it was written, in a few words, on its node. Not a log entry. `asWritten`: by design, not a miss (the Workflow
295 * was already sent back once). `offBoard`: nothing to show at all, so nothing is written (a script with no agent()
296 * call).
297 */
298export type Left = About & { why: string; asWritten?: true; offBoard?: true }
299
300/**
301 * An agent whose decision was made earlier, for its call, has started: the agent is on the board as the one the
302 * decision routed. Not a log entry (the decision is, since it was made).
303 */
304export type Started = About & { started: true }
305
306/**
307 * Why a feature that was asked for something said nothing, when that is neither a decision nor a failed request:
308 * the answer said nothing about the skills (`unanswered`), the session's skills could not be read (`unread`),
309 * there was no skill to rate (`none`), an error of the mod's own (`error`; the debug log has it). Not in the log and
310 * on no node: a note on the board (`board.notes`, kind `skipped`), which the band's event stream draws.
311 */
312export type Skip = 'unanswered' | 'unread' | 'none' | 'error'
313export type Skipped = About & { skipped: Skip; aside: true }
314
315export type ReportedDecision = Decided | NotDecided | Left | Started | Skipped
316
317/** What a feature counts as its loop goes (`report`'s `tally`), not a decision of its own. */
318export type Tallied =
319 /** The main agent's turn as the mid-turn re-decision sees it. `quiet`: not re-decided yet this turn (nothing to show). `late`: the answer for this step is not back; `failure`: its request failed. */
320 | { feature: 'midturn-effort'; agent: 'main'; steps: number; judged: number; changed: number; quiet: boolean; late?: true; failure?: NodeFailure }
321 /** A loop's failed tool calls, hook blocks and forced raises. `turnStart`: the main agent's counts start over with a new turn. */
322 | { feature: 'escalation'; agent: string; failed: number; blocked: number; raised: number; late?: true; turnStart?: true }
323
324/** What one model request went out with, as the core sends it. */
325export type StepReading = {
326 /** Absent for the main agent. */
327 agentId?: string
328 /** The model id as sent. */
329 model: string
330 /** The effort as sent; absent for a model without effort. */
331 effort?: string | number
332 /** Where the effort came from: the person's lock, a plan, or the engine (not routed). */
333 source: EffortSource
334}
335
336/** An agent of the roster `$.agent.list()` answers, as far as a node needs it. */
337export type RosterAgent = { id: string; name?: string; description: string; type: string }
338
339/**
340 * What the readings need beyond `ReportIo`: the roster of dispatched agents, and the directories of the session's
341 * Workflow runs (to tell a Workflow's agent, and its label, from the journal). `stepIo($)` builds it.
342 */
343export type StepIo = ReportIo & {
344 /** `() => $.agent.list()`: dispatched agents only, a Workflow's are not in it. */
345 agents: () => Promise<readonly RosterAgent[]>
346 /** The run directories the mod noted (`workflowRuns`, `labelRuns`), oldest first. */
347 runDirs: () => Promise<readonly string[]>
348 /** A run's journal text; null when it is not there. */
349 journal: (dir: string) => Promise<string | null>
350}
351
352/** What the module needs of the host, as closures (see the top of the file). */
353export type ReportIo = {
354 board: Cell<Board>
355 decisions: Cell<LogEntry[]>
356 /** `(line) => $.ui.log(line, { to: 'debug' })` */
357 debug: (line: string) => void
358 /** `() => $.clock.now()`: when a decision, a reading or a note is reported. */
359 now: () => Promise<number>
360 /** The host's toast, as `toast: (text) => $.ui.toast(text)`: a route that failed. */
361 toast: (text: string) => void
362}
363
364// ---- entry one (记一条决定): what a feature hands over ----------------------------
365
366/** What `report` is handed, one kind at a time (see the top of the file). */
367export type Reported =
368 /** A decision made, or why none was (`ReportedDecision`). */
369 | { decision: ReportedDecision }
370 /** The decisions of one event, in the order made: one write to the log and one to the board, at most one toast. */
371 | { decisions: readonly ReportedDecision[] }
372 /** What a feature counts as a loop goes (`Tallied`). */
373 | { tally: Tallied }
374 /** The person flipped the mod or a feature (`/dp`). */
375 | { switched: Switched }
376 /** How writing the session's skill profiles goes (`ProfileEvent`). */
377 | { profiles: ProfileEvent }
378 /** The decision model in use is not the one asked for: no Perplexity key, so Jev (`DecisionModelEvent`). */
379 | { fellBack: DecisionModelEvent }
380 /** A pane the person asked for (a digit on the band) that the surface did not place, with the surface's reason; it was closed again. */
381 | { unplaced: { reason: string } }
382
383/** What a toast of its own needs of the host: the debug log, the clock and the toast. */
384export type NoticeIo = Pick<ReportIo, 'debug' | 'now' | 'toast'>
385
386/** What the decision model's fallback needs of the host: the board (for the turn), the decision log and the debug log. */
387export type DecisionModelIo = Pick<ReportIo, 'board' | 'decisions' | 'debug'>
388
389/** The host closures each kind of report needs: a switch only the redraw, the skill profiles their own state, a pane not placed the toast, the rest a `ReportIo`. */
390export type IoOf<R extends Reported> = R extends { switched: Switched }
391 ? SwitchIo
392 : R extends { profiles: ProfileEvent }
393 ? ProfilesIo
394 : R extends { fellBack: DecisionModelEvent }
395 ? DecisionModelIo
396 : R extends { unplaced: unknown }
397 ? NoticeIo
398 : ReportIo
399
400/**
401 * Entry one (记一条决定): reports what a feature hands over, by its kind; `io` is what that kind needs of the host
402 * (`IoOf`). Never throws.
403 */
404export async function report<R extends Reported>(io: IoOf<R>, what: R): Promise<void> {
405 const item: Reported = what
406 if ('decision' in item) return reportDecisions(io as ReportIo, [item.decision])
407 if ('decisions' in item) return reportDecisions(io as ReportIo, item.decisions)
408 if ('tally' in item) return reportTally(io as ReportIo, item.tally)
409 if ('switched' in item) return reportSwitch(io as SwitchIo, item.switched)
410 if ('profiles' in item) return reportProfiles(io as ProfilesIo, item.profiles)
411 if ('fellBack' in item) return reportFellBack(io as DecisionModelIo, item.fellBack)
412 return reportUnplaced(io as NoticeIo, item.unplaced)
413}
414
415/** The rationale pane was not placed, in the person's words: why (the surface's own reason), and where to look instead. */
416export function unplacedText(reason: string): string {
417 return `依据面板没有放出来(${reason}),已经关上;/dp log 10 在对话里列出最近 10 条决定`
418}
419
420/**
421 * A pane the person asked for that the surface did not place: the debug line, and a toast so they know (the `/dp`
422 * command says it in its answer instead), unless a toast went up within `TOAST_GAP_MS`. Never throws.
423 */
424async function reportUnplaced(io: NoticeIo, unplaced: { reason: string }): Promise<void> {
425 try {
426 io.debug(`rationale pane not placed (${unplaced.reason}), closed again`)
427 toastOnce(io, await io.now(), unplacedText(unplaced.reason))
428 } catch {
429 // a notice that cannot be given leaves the pane closed all the same
430 }
431}
432
433/**
434 * The decisions of one event (a single one, or the several agents of a Workflow): the debug log line and the
435 * decision log entry of each decision made, the agents' nodes on the board, a note for the band when a request
436 * beside a route came to nothing, and a toast when a route failed. One write to the log, one to the board, at
437 * most one toast; each is what it would make of it alone, in the order given. Never throws.
438 */
439async function reportDecisions(io: ReportIo, decisions: readonly ReportedDecision[]): Promise<void> {
440 if (decisions.length === 0) return
441 try {
442 // Which turn each is for. A board that cannot be read does not stop a decision from being logged (as turn 1).
443 const board = await read(io.board).catch((error: unknown) => {
444 io.debug(`board not read: ${errorText(error)}`)
445 return EMPTY
446 })
447 const now = await io.now()
448 const turnOf = (decision: ReportedDecision) => board.turn + (decision.forTurn === 'next' ? 1 : 0)
449 const made = decisions.filter(isDecided)
450 for (const decision of made) io.debug(decisionLine(decision))
451 // The number of each decision in the log, in the order made.
452 const numbers: number[] = []
453 if (made.length > 0) {
454 try {
455 await update(io.decisions, (list) => {
456 numbers.length = 0
457 let all = list ?? []
458 for (const decision of made) {
459 all = appendEntry(all, entryOf(decision, turnOf(decision), elapsed(board, turnOf(decision), now)))
460 numbers.push(all.at(-1)?.n ?? 0)
461 }
462 return all
463 })
464 } catch (error) {
465 numbers.length = 0
466 io.debug(`decision not kept for /dp log: ${errorText(error)}`)
467 }
468 }
469 // A request beside the route that came to nothing: a note for the band's event stream, after the log's last entry.
470 const noted = decisions.filter((decision) => 'skipped' in decision || ('failure' in decision && decision.aside === true))
471 const logged = noted.length === 0 ? 0 : (numbers.at(-1) ?? (await lastLogged(io)))
472 const notes = noted.map((decision): BoardNote => {
473 const of = { turn: turnOf(decision), id: decision.agent, feature: decision.feature, at: elapsed(board, turnOf(decision), now), after: logged }
474 return 'skipped' in decision ? { ...of, kind: 'skipped', why: decision.skipped } : { ...of, kind: 'failed', why: 'failure' in decision ? failureLine(decision.failure.backend, decision.failure) : '' }
475 })
476 let after = board
477 // A decision beside the agent's route is not on its node; a Workflow with nothing to show is not on the board.
478 const onBoard = (decision: ReportedDecision) => decision.aside !== true && !('skipped' in decision) && !('offBoard' in decision)
479 if (decisions.some(onBoard) || notes.length > 0) {
480 try {
481 after = await update(io.board, (current) => {
482 let next = current ?? EMPTY
483 let at = 0
484 for (const decision of decisions) {
485 const n = isDecided(decision) ? numbers[at++] : undefined
486 if (onBoard(decision)) next = withNode(next, turnOf(decision), decision as Decided | NotDecided | Left | Started, n)
487 }
488 return notes.length === 0 ? next : withNotes(next, notes)
489 })
490 } catch (error) {
491 io.debug(`decision not kept on the board: ${errorText(error)}`)
492 }
493 }
494 // The routes that failed: a toast, so the person notices; the board says the rest.
495 const failed = decisions.filter((decision): decision is NotDecided => 'failure' in decision && decision.aside !== true)
496 if (failed.length > 0) toastOnce(io, now, failedText(failed, after, turnOf))
497 } catch (error) {
498 io.debug(`decision not reported: ${errorText(error)}`)
499 }
500}
501
502/** A decision that was made (the others say why there is none, or that an agent started). */
503function isDecided(decision: ReportedDecision): decision is Decided {
504 return 'outcome' in decision
505}
506
507/** The `n` of the decision log's last entry (0: none, or the log cannot be read). */
508async function lastLogged(io: ReportIo): Promise<number> {
509 try {
510 return ((await io.decisions.get()).value ?? []).at(-1)?.n ?? 0
511 } catch {
512 return 0
513 }
514}
515
516/** The engine drops a plugin's second toast within this many milliseconds of its last one (measured on 2.1.289). */
517const TOAST_GAP_MS = 2000
518
519/** When the last toast was raised (module state: a hot reload forgets it, and the next toast may then be one the engine drops). */
520let toastedAt = Number.NEGATIVE_INFINITY
521
522/**
523 * Raises the toast unless one was raised within `TOAST_GAP_MS`, which the engine would drop: the failures that
524 * follow one closely are on the board all the same.
525 */
526function toastOnce(io: Pick<ReportIo, 'toast'>, now: number, text: string): void {
527 if (now - toastedAt < TOAST_GAP_MS) return
528 toastedAt = now
529 try {
530 io.toast(text)
531 } catch {
532 // a toast that cannot be shown is skipped: the board has it
533 }
534}
535
536/** The toast for the routes of one event that failed: who is not routed, and why (the first failure's words, then its details). */
537function failedText(failed: readonly NotDecided[], board: Board, turnOf: (decision: ReportedDecision) => number): string {
538 const first = failed[0] as NotDecided
539 const why = `${failureWords(first.failure)}(${failureLine(first.failure.backend, first.failure)})`
540 if (failed.length > 1) return `${failed.length} 个 agent 未路由:${why}`
541 if (first.agent === 'main') return `主 agent 未路由:${why}`
542 const name = board.nodes.find((node) => node.turn === turnOf(first) && node.id === first.agent)?.name ?? first.node?.name ?? first.agent
543 return `「${name.length > 24 ? `${name.slice(0, 23)}…` : name}」未路由:${why}`
544}
545
546/** The board with the notes added: those of turns older than the nodes' dropped, and the oldest past `NOTES_KEPT`. */
547function withNotes(board: Board, notes: readonly BoardNote[]): Board {
548 return { ...board, notes: [...(board.notes ?? []), ...notes].filter((note) => note.turn >= board.turn - 1).slice(-NOTES_KEPT) }
549}
550
551// ---- entry two (记一步读数): a step's reading ----------------------------------------
552
553/** How many writes may be lost to other writers in a row: a Workflow's agents step within milliseconds of each other. */
554const READING_ATTEMPTS = 32
555
556/**
557 * Entry two (记一步读数): reports the reading of one step, the model (by family) and the effort it
558 * went out with, whether the effort was routed, and that the agent is running.
559 * Reading never changes what is sent. The agent's node is made at its first
560 * step when the board has none (named from the roster of dispatched agents, or
561 * from the label its Workflow journal gives it; a loop neither knows is not an
562 * agent of the board), carries its start, and a reading that differs from the
563 * node's previous one adds a change event. Writes the board only when
564 * something is new. Never throws.
565 */
566export async function reportStep(io: StepIo, step: StepReading): Promise<void> {
567 const id = step.agentId ?? 'main'
568 const main = step.agentId === undefined
569 const family = modelFamily(step.model)
570 const effort = typeof step.effort === 'number' || isEffort(step.effort) ? step.effort : undefined
571 const reading: Reading = { ...(family === null ? {} : { model: family }), ...(effort === undefined ? {} : { effort }) }
572 // `source`: the engine's own effort goes out unless a plan sets it. An agent that was decided to run as it does has no plan
573 // to say so (haiku takes no effort, a Workflow script was written into): its reading is checked against its decision below.
574 let routed = step.source !== 'engine'
575 const locked = step.source === 'locked'
576 try {
577 const now = await io.now()
578 const seen = main ? now : seenAt(id, now)
579 const peek = await read(io.board)
580 const before = main ? mainNode(peek) : continuing(peek, id)
581 // A node that has not begun is named now (the roster may know it only from here on); one that has keeps its name.
582 const identity = main || (before !== undefined && begun(before)) ? null : await lookUp(io, id)
583 if (!main && before === undefined && identity === null) return
584 const prior = main ? undefined : (before ?? callNodeOf(peek, identity))
585 if (!routed && prior?.routed === true && prior.decision !== undefined) routed = await wentOutAsDecided(io, prior.decision, reading)
586 const logged = before !== undefined && differs(before, reading) ? (((await io.decisions.get()).value ?? []).at(-1)?.n ?? 0) : 0
587 await modify(
588 io.board,
589 (current) => {
590 const board = current ?? EMPTY
591 const old = main ? mainNode(board) : continuing(board, id)
592 if (old === undefined && identity === null && !main) return undefined
593 const turn = old?.turn ?? board.turn
594 const started = old !== undefined && begun(old)
595 // A Workflow's agent that starts takes the place of the node its call stood for, and what was decided for it.
596 const stood = old === undefined && !main ? callNodeOf(board, identity) : undefined
597 const base = old ?? { ...newNode(turn, id, identity ?? undefined, 'running'), ...carriedFrom(stood, routed) }
598 const named = identity !== null && !started ? { kind: identity.kind, name: identity.name, type: identity.type } : {}
599 const next: BoardNode = {
600 ...without(base, 'model', 'effort', 'locked', 'dur'),
601 ...named,
602 state: 'running',
603 t0: main ? 0 : started ? base.t0 : elapsed(board, turn, seen),
604 routed,
605 ...reading,
606 ...(locked ? { locked: true as const } : {}),
607 }
608 const changes = old !== undefined && differs(old, reading) ? [{ turn, id, at: elapsed(board, turn, now), after: logged, from: readingOf(old), to: reading }] : []
609 if (old !== undefined && changes.length === 0 && JSON.stringify(old) === JSON.stringify(next)) return undefined
610 return {
611 ...board,
612 ...(changes.length === 0 ? {} : { changes: [...(board.changes ?? []), ...changes].filter((change) => change.turn >= board.turn - 1) }),
613 nodes: [...board.nodes.filter((node) => node !== old && node !== stood), next],
614 }
615 },
616 READING_ATTEMPTS,
617 )
618 if (!main) firstSeen.delete(id)
619 } catch (error) {
620 io.debug(`reading not kept on the board: ${errorText(error)}`)
621 }
622}
623
624/** Why an agent's turn did not end in an answer, in the board's words (a node's `why` when it failed and nothing else said). */
625export const ENDED_WORDS = { error: '出错', refusal: '拒绝回答', aborted: '被中断' } as const
626
627/** How a loop's turn ended (`turn.complete`). */
628export type LoopEnd = { agentId?: string; reason: 'answer' | 'aborted' | 'refusal' | 'error'; durationMs: number }
629
630/**
631 * Reports the end of a loop's turn: done after an answer, else failed (and
632 * why, when nothing else said), for as long as it ran. The main agent's turn
633 * ends its own node; an agent's, its own, wherever its turn's node is. Never throws.
634 */
635async function reportEnd(io: StepIo, end: LoopEnd): Promise<void> {
636 const id = end.agentId ?? 'main'
637 const main = end.agentId === undefined
638 try {
639 const now = await io.now()
640 const peek = await read(io.board)
641 const known = main ? mainNode(peek) : continuing(peek, id)
642 // The loop is over: whatever its steps missed, it is looked up once more in full.
643 missed.delete(id)
644 const identity = main || known !== undefined ? null : (await identify(io, id, { roster: true, dirs: await runDirsOf(io) })).identity
645 if (!main && known === undefined && identity === null) return
646 const seen = main || known !== undefined ? now : (firstSeen.get(id) ?? now - end.durationMs)
647 await modify(
648 io.board,
649 (current) => {
650 const board = current ?? EMPTY
651 const old = main ? mainNode(board) : continuing(board, id)
652 if (old === undefined && identity === null && !main) return undefined
653 const turn = old?.turn ?? board.turn
654 const stood = old === undefined && !main ? callNodeOf(board, identity) : undefined
655 const base = old ?? { ...newNode(turn, id, identity ?? undefined, 'running'), ...carriedFrom(stood, true) }
656 const failed = end.reason !== 'answer'
657 const next: BoardNode = {
658 ...base,
659 ...(old === undefined && !main ? { t0: elapsed(board, turn, seen) } : {}),
660 state: failed ? 'failed' : 'done',
661 dur: Math.round(end.durationMs) / 1000,
662 ...(failed && base.why === undefined ? { why: end.reason === 'aborted' ? ENDED_WORDS.aborted : end.reason === 'refusal' ? ENDED_WORDS.refusal : ENDED_WORDS.error } : {}),
663 }
664 return { ...board, nodes: [...board.nodes.filter((node) => node !== old && node !== stood), next] }
665 },
666 READING_ATTEMPTS,
667 )
668 if (!main) firstSeen.delete(id)
669 } catch (error) {
670 io.debug(`end of loop not kept on the board: ${errorText(error)}`)
671 }
672}
673
674/**
675 * Reports a dispatched agent started: queued until its first step. Its node is
676 * made (or, when a decision about it made one, marked queued from now); the
677 * agent's own first step starts it. Never throws.
678 */
679async function reportSpawn(io: StepIo, spawn: { agentId: string; name?: string; description: string; type: string }): Promise<void> {
680 try {
681 const now = await io.now()
682 await modify(
683 io.board,
684 (current) => {
685 const board = current ?? EMPTY
686 const old = continuing(board, spawn.agentId)
687 if (old !== undefined && begun(old)) return undefined
688 const turn = old?.turn ?? board.turn
689 const name = spawn.name !== undefined && spawn.name !== '' ? spawn.name : spawn.description !== '' ? spawn.description : spawn.agentId
690 const base = old ?? newNode(turn, spawn.agentId, { kind: 'agent', name, type: spawn.type }, 'queued')
691 const next: BoardNode = { ...base, state: 'queued', t0: elapsed(board, turn, now) }
692 return { ...board, nodes: [...board.nodes.filter((node) => node !== old), next] }
693 },
694 READING_ATTEMPTS,
695 )
696 } catch (error) {
697 io.debug(`agent not kept on the board: ${errorText(error)}`)
698 }
699}
700
701// ---- what `report` does with a tally --------------------------------------------
702
703/**
704 * Reports what a feature counts as its loop goes, which is no decision: the
705 * mid-turn re-decision's steps, decisions and changes of the main agent's
706 * turn, and each loop's failed tool calls and forced raises. On the agent's
707 * node of the current turn (`midturn`, `counts`); an answer that is late, or a
708 * re-decision's request that failed, is also a note for the band's event
709 * stream. Writes the board only when the node changed. Never throws.
710 */
711async function reportTally(io: ReportIo, tally: Tallied): Promise<void> {
712 try {
713 const noteworthy = tally.late === true || ('failure' in tally && tally.failure !== undefined)
714 const now = await io.now()
715 const logged = noteworthy ? await lastLogged(io) : 0
716 await modify(io.board, (current) => withTally(current ?? EMPTY, tally, { now, logged }))
717 } catch (error) {
718 io.debug(`tally not kept on the board: ${errorText(error)}`)
719 }
720}
721
722// ---- what `report` does with a switch -------------------------------------------
723
724/** What the person flipped with `/dp`: the whole mod, or one feature. */
725export type Switched = { master: boolean } | { feature: string; on: boolean }
726
727/** What a switch needs of the host: `() => $.ui.invalidate('ui.render')`, the screens drawn again. */
728export type SwitchIo = { redraw: () => void }
729
730/**
731 * Reports a switch the person flipped (the control feature's `/dp`). The board
732 * and the decision log keep what they hold: the screens read the switches as
733 * they draw, and leave out what a feature that is off owns (its decisions, the
734 * parts of the nodes it writes, `defineSwitch`'s `parts`), so they are drawn
735 * again now. Never throws.
736 */
737function reportSwitch(io: SwitchIo, _change: Switched): void {
738 try {
739 io.redraw()
740 } catch {
741 // a redraw that cannot be asked for waits for the next change of the board
742 }
743}
744
745/**
746 * The board with the tally on its agent's node (the main agent's of the turn running, an agent's where its steps
747 * go on: `continuing`), and a note when an answer just went late or a request just failed; undefined when the
748 * node already says it, or there is nothing to say. A loop with no node is no agent of the board (the readings
749 * make nodes), the main agent's is made.
750 */
751function withTally(board: Board, tally: Tallied, when: { now: number; logged: number }): Board | undefined {
752 const main = tally.agent === 'main'
753 const old = main ? mainNode(board) : continuing(board, tally.agent)
754 if (old === undefined && !main) return undefined
755 let next: BoardNode
756 const notes: BoardNote[] = []
757 const note = (turn: number, kind: BoardNote['kind'], why: string) =>
758 notes.push({ turn, id: tally.agent, feature: tally.feature, at: elapsed(board, turn, when.now), after: when.logged, kind, why })
759 if (tally.feature === 'midturn-effort') {
760 if (tally.quiet) return undefined
761 const midturn = { steps: tally.steps, judged: tally.judged, changed: tally.changed, ...(tally.late === true ? { late: true as const } : {}), ...(tally.failure === undefined ? {} : { failure: tally.failure }) }
762 next = { ...(old ?? newNode(board.turn, tally.agent, undefined, 'running')), midturn }
763 if (tally.late === true && old?.midturn?.late !== true) note(next.turn, 'late', '')
764 if (tally.failure !== undefined && JSON.stringify(old?.midturn?.failure) !== JSON.stringify(tally.failure)) note(next.turn, 'failed', failureLine(tally.failure.backend, tally.failure))
765 } else {
766 const any = tally.failed + tally.blocked + tally.raised > 0
767 if (old === undefined && !any) return undefined
768 const base = old ?? newNode(board.turn, tally.agent, undefined, 'running')
769 next = any
770 ? { ...base, counts: { failed: tally.failed, blocked: tally.blocked, raised: tally.raised, ...(tally.late === true ? { late: true as const } : {}) } }
771 : without(base, 'counts')
772 if (tally.late === true && old?.counts?.late !== true) note(next.turn, 'late', '')
773 }
774 if (old !== undefined && JSON.stringify(old) === JSON.stringify(next)) return undefined
775 const changed = { ...board, nodes: [...board.nodes.filter((node) => node !== old), next] }
776 return notes.length === 0 ? changed : withNotes(changed, notes)
777}
778
779/** Who a loop with no node is, from what the engine and the Workflow journals say; null: not an agent the board shows. */
780type Identity = { kind: 'agent' | 'wf'; name: string; type: string }
781
782/**
783 * Who a loop is, by the roster (`roster`) and the journals of the run directories `dirs` (null: none read). `read`
784 * is false when the directories or a journal could not be read, so a miss may not be one.
785 */
786async function identify(io: StepIo, id: string, look: { roster: boolean; dirs: readonly string[] | null }): Promise<{ identity: Identity | null; read: boolean }> {
787 if (look.roster) {
788 try {
789 const info = (await io.agents()).find((agent) => agent.id === id)
790 if (info !== undefined) return { identity: { kind: 'agent', name: info.name !== undefined && info.name !== '' ? info.name : info.description !== '' ? info.description : id, type: info.type }, read: true }
791 } catch (error) {
792 io.debug(`agent roster not read: ${errorText(error)}`)
793 }
794 }
795 if (look.dirs === null) return { identity: null, read: false }
796 try {
797 // A Workflow's agents are in no roster: the journal of the run it started in has the label it started under.
798 for (const dir of [...look.dirs].reverse()) {
799 const journal = await io.journal(dir)
800 const start = journal === null ? null : startedIn(journal, id)
801 if (start !== null) return { identity: { kind: 'wf', name: start.label, type: 'workflow' }, read: true }
802 }
803 } catch (error) {
804 io.debug(`workflow journals not read: ${errorText(error)}`)
805 return { identity: null, read: false }
806 }
807 return { identity: null, read: true }
808}
809
810/** The session's Workflow run directories, oldest first; null when they cannot be read. */
811async function runDirsOf(io: StepIo): Promise<readonly string[] | null> {
812 try {
813 return await io.runDirs()
814 } catch (error) {
815 io.debug(`workflow journals not read: ${errorText(error)}`)
816 return null
817 }
818}
819
820/**
821 * The loops a step found no one to be, by id: how many of its steps have looked, and the run directories whose
822 * journals it was looked for in (null: they could not be read). Module state: a hot reload forgets it, and costs each
823 * such loop one full look-up more.
824 */
825const missed = new Map<string, { sightings: number; dirs: string | null }>()
826const MISSED_KEPT = 64
827
828/**
829 * `identify` for a step, at most as often as the answer can change: the first step of a loop no one knows (an engine
830 * fork) reads the roster and every run's journal, and its later steps would read them all again. After a miss, the
831 * journals are read again only once the session has another run directory (a Workflow's agent can step before its run
832 * is recorded), or after they could not be read; the roster at the loop's 2nd, 4th, 8th... step (an agent it lists late).
833 */
834async function lookUp(io: StepIo, id: string): Promise<Identity | null> {
835 const miss = missed.get(id)
836 const sightings = (miss?.sightings ?? 0) + 1
837 const roster = miss === undefined || (sightings & (sightings - 1)) === 0
838 const dirs = await runDirsOf(io)
839 const seen = dirs === null ? null : dirs.join('\n')
840 const journals = miss === undefined || seen === null || miss.dirs === null ? roster : seen !== miss.dirs
841 const found = roster || journals ? await identify(io, id, { roster, dirs: journals ? dirs : null }) : { identity: null, read: true }
842 // The newest miss goes last, so the oldest go first past `MISSED_KEPT`.
843 missed.delete(id)
844 if (found.identity !== null) return found.identity
845 missed.set(id, { sightings, dirs: !journals ? (miss?.dirs ?? null) : found.read ? seen : null })
846 for (const key of missed.keys()) if (missed.size > MISSED_KEPT) missed.delete(key)
847 return null
848}
849
850/**
851 * When a loop no node was made for was first seen (module state: a hot reload forgets it, and an agent only
852 * found later then starts at that step). A Workflow's first agent can step before its run is recorded, so its
853 * name comes with a later step, or with the end of its loop, and starts when it first stepped.
854 */
855const firstSeen = new Map<string, number>()
856const FIRST_SEEN_KEPT = 64
857
858function seenAt(id: string, now: number): number {
859 const seen = firstSeen.get(id)
860 if (seen !== undefined) return seen
861 firstSeen.set(id, now)
862 for (const key of firstSeen.keys()) if (firstSeen.size > FIRST_SEEN_KEPT) firstSeen.delete(key)
863 return now
864}
865
866/**
867 * The node a Workflow call of a script stands for before its agent starts (a decision was made for the call, or it
868 * was left as written): queued, one of the Workflow's calls, called what the agent's label is. Null when there is
869 * none, or the agent is not a Workflow's.
870 */
871function callNodeOf(board: Board, identity: Identity | null | undefined): BoardNode | undefined {
872 if (identity === null || identity === undefined || identity.kind !== 'wf') return undefined
873 return board.nodes.find((node) => node.kind === 'wf' && node.state === 'queued' && node.workflow !== undefined && callWorkflowOf(node.id) === node.workflow.id && node.name === identity.name)
874}
875
876/** What a Workflow's agent takes over from the node its call stood for: the Workflow, the decision made for it, and why it was left as written (when its steps go out unrouted). */
877function carriedFrom(stood: BoardNode | undefined, routed: boolean): Partial<BoardNode> {
878 if (stood === undefined) return {}
879 return {
880 ...(stood.workflow === undefined ? {} : { workflow: stood.workflow }),
881 ...(stood.decision === undefined ? {} : { decision: stood.decision }),
882 ...(routed || stood.why === undefined ? {} : { why: stood.why }),
883 ...(routed || stood.failure === undefined ? {} : { failure: stood.failure }),
884 }
885}
886
887/** Whether a reading is the model and effort the decision numbered `n` in the log decided for the agent (a decision that names none: no). */
888async function wentOutAsDecided(io: ReportIo, n: number, reading: Reading): Promise<boolean> {
889 try {
890 const entry = ((await io.decisions.get()).value ?? []).find((kept) => kept.n === n)
891 return entry?.model !== undefined && entry.model === reading.model && entry.effort === reading.effort
892 } catch {
893 return false
894 }
895}
896
897/** The main agent's node of the turn running. */
898function mainNode(board: Board): BoardNode | undefined {
899 return board.nodes.find((node) => node.turn === board.turn && node.id === 'main')
900}
901
902/**
903 * An agent's node that its steps go on in: its latest one, if it is of the turn
904 * running or has not ended (an agent still running when the next turn starts
905 * stays in the turn it began in); else it is a new row in this turn.
906 */
907function continuing(board: Board, id: string): BoardNode | undefined {
908 const latest = board.nodes.filter((node) => node.id === id).reduce<BoardNode | undefined>((best, node) => (best === undefined || node.turn >= best.turn ? node : best), undefined)
909 return latest !== undefined && (latest.turn === board.turn || latest.state === 'queued' || latest.state === 'running') ? latest : undefined
910}
911
912/** Whether the node has had a reading: its start is the first step's, and its name is settled. */
913function begun(node: BoardNode): boolean {
914 return node.state !== 'queued' && (node.id === 'main' || node.t0 > 0 || node.model !== undefined || node.effort !== undefined)
915}
916
917function readingOf(node: BoardNode): Reading {
918 return { ...(node.model === undefined ? {} : { model: node.model }), ...(node.effort === undefined ? {} : { effort: node.effort }) }
919}
920
921/** Whether the node has a previous reading, and it is not this one. */
922function differs(node: BoardNode, reading: Reading): boolean {
923 if (node.model === undefined && node.effort === undefined) return false
924 return node.model !== reading.model || node.effort !== reading.effort
925}
926
927/** Seconds from the start of `turn` to `ms` (0 when the turn's start is not known). */
928function elapsed(board: Board, turn: number, ms: number): number {
929 const start = board.starts?.find((kept) => kept.turn === turn)?.at
930 return start === undefined ? 0 : Math.max(0, Math.round(ms - start)) / 1000
931}
932
933// ---- the turn count and the loops' lifecycle ------------------------------------
934
935/** The board when a main turn starts at `now`: the count up by one, the nodes of older turns gone (the turn before it stays: it is folded into one line). */
936export function startTurn(board: Board | undefined, now: number): Board {
937 const turn = (board?.turn ?? 0) + 1
938 const kept = (of: number) => of >= turn - 1
939 return {
940 turn,
941 starts: [...(board?.starts ?? []).filter((start) => kept(start.turn)), { turn, at: now }],
942 ...(board?.changes === undefined ? {} : { changes: board.changes.filter((change) => kept(change.turn)) }),
943 ...(board?.notes === undefined ? {} : { notes: board.notes.filter((note) => kept(note.turn)) }),
944 nodes: (board?.nodes ?? []).filter((node) => kept(node.turn)),
945 }
946}
947
948/**
949 * What the module's own hooks need of the host, as closures over `$`. `$` is followed only into a function of
950 * the same file, so the core builds its `StepIo` itself (core.ts, `turn.step`): keep the two alike.
951 */
952function stepIo($: EngineInterface): StepIo {
953 return {
954 board: { get: () => $.state.get(BOARD), set: (value, options) => $.state.set(BOARD, value, options) },
955 decisions: { get: () => $.state.get(DECISIONS), set: (value, options) => $.state.set(DECISIONS, value, options) },
956 debug: (line) => $.ui.log(line, { to: 'debug' }),
957 now: () => $.clock.now(),
958 toast: (text) => $.ui.toast(text),
959 agents: () => $.agent.list(),
960 runDirs: async () => {
961 const [noted, labelled] = await Promise.all([$.state.get(WORKFLOW_RUNS), $.state.get(LABEL_RUNS)])
962 return [...new Set([...(noted.value ?? []), ...(labelled.value ?? [])].map((run) => run.dir))]
963 },
964 journal: async (dir) => {
965 const path = `${dir}/journal.jsonl`
966 return (await $.fs.exists(path)) ? await $.fs.read(path).catch(() => null) : null
967 },
968 }
969}
970
971/**
972 * The module's own hooks, which only watch (every event goes on as it is):
973 * a main turn starting (the board's turn count and its start time), a
974 * dispatched agent spawned (queued until its first step), a loop's turn ending
975 * (its node done or failed). A board that cannot be kept never stops anything.
976 */
977export function registerReport(on: On): void {
978 on('turn.start', { turnId: /(?:)/ }, async ($, e, next) => {
979 try {
980 const now = await $.clock.now()
981 await update<Board>({ get: () => $.state.get(BOARD), set: (value, options) => $.state.set(BOARD, value, options) }, (board) => startTurn(board, now))
982 } catch {
983 // the count is the board's own: a turn starts without it
984 }
985 return next(e)
986 })
987
988 on('agent.spawn', { tool_use_id: /(?:)/ }, async ($, e, next) => {
989 const result = await next(e)
990 if (result.deny !== undefined || result.agentId === undefined) return result
991 await reportSpawn(stepIo($), { agentId: result.agentId, ...(e.name === undefined ? {} : { name: e.name }), description: e.description, type: e.subagentType })
992 return result
993 })
994
995 on('turn.complete', { turnId: /(?:)/ }, async ($, e, next) => {
996 const result = await next(e)
997 await reportEnd(stepIo($), { ...(e.agentId === undefined ? {} : { agentId: e.agentId }), reason: e.reason, durationMs: e.durationMs })
998 return result
999 })
1000}
1001
1002// ---- the log ------------------------------------------------------------------
1003
1004/** A decision as one line: `effort high · "the message":why`. The debug log's line, and `/dp log N`'s after `#n feature:`. */
1005export function decisionLine(decision: { outcome: string; subject?: string; reason: string }): string {
1006 return `${decision.outcome}${decision.subject ? ` · ${decision.subject}` : ''}:${decision.reason}`
1007}
1008
1009/** The list with `entry` added last, numbered, the entries of turns older than the latest `LOG_TURNS` dropped, and then the oldest past `LOG_ENTRIES`. */
1010export function appendEntry(list: readonly LogEntry[], entry: Omit<LogEntry, 'n'>): LogEntry[] {
1011 const all = [...list, { n: (list.at(-1)?.n ?? 0) + 1, ...entry }]
1012 const latest = Math.max(...all.map((kept) => kept.turn))
1013 return all.filter((kept) => kept.turn > latest - LOG_TURNS).slice(-LOG_ENTRIES)
1014}
1015
1016function entryOf(decision: Decided, turn: number, at: number): Omit<LogEntry, 'n'> {
1017 return {
1018 turn,
1019 at,
1020 feature: decision.feature,
1021 agent: decision.agent,
1022 tone: decision.tone ?? 'ok',
1023 outcome: decision.outcome,
1024 subject: decision.subject ?? '',
1025 reason: decision.reason,
1026 ...(decision.probs === undefined ? {} : { probs: decision.probs }),
1027 ...(decision.conf === undefined ? {} : { conf: decision.conf }),
1028 ...(decision.trace === undefined ? {} : { trace: decision.trace }),
1029 ...(decision.floor === undefined ? {} : { floor: decision.floor }),
1030 ...(decision.model === undefined ? {} : { model: decision.model }),
1031 ...(decision.effort === undefined ? {} : { effort: decision.effort }),
1032 ...(decision.mid === undefined ? {} : { mid: decision.mid }),
1033 ...(decision.forced === undefined ? {} : { forced: decision.forced }),
1034 ...(decision.skills === undefined ? {} : { skills: decision.skills }),
1035 ...(decision.counts === undefined ? {} : { counts: decision.counts }),
1036 ...(decision.unresolved === undefined ? {} : { unresolved: decision.unresolved }),
1037 ...(decision.hint === undefined ? {} : { hint: decision.hint }),
1038 ...(decision.sentBack === true ? { sentBack: true as const } : {}),
1039 }
1040}
1041
1042// ---- the board ----------------------------------------------------------------
1043
1044const EMPTY: Board = { turn: 0, nodes: [] }
1045
1046function newNode(turn: number, id: string, node: About['node'], state: AgentState): BoardNode {
1047 if (id === 'main') return { turn, id, kind: 'main', name: '主 agent', type: 'main', state, t0: 0, routed: false }
1048 return { turn, id, kind: node?.kind ?? 'agent', name: node?.name ?? id, type: node?.type ?? 'agent', state: node?.state ?? state, t0: 0, routed: false, ...(node?.workflow === undefined ? {} : { workflow: node.workflow }) }
1049}
1050
1051/** The board with the decision on its agent's node of `turn` (the node made when there is none), `n` the decision's number in the log when it has one. */
1052function withNode(board: Board, turn: number, decision: Decided | NotDecided | Left | Started, n: number | undefined): Board {
1053 const old = board.nodes.find((node) => node.turn === turn && node.id === decision.agent)
1054 // The node that stood for the agent before it started gives it what was decided for it, and goes.
1055 const stood = old === undefined && decision.replaces !== undefined ? board.nodes.find((node) => node.id === decision.replaces) : undefined
1056 const inherited = stood?.decision === undefined ? {} : { decision: stood.decision }
1057 const base = old ?? { ...newNode(turn, decision.agent, decision.node, decision.forTurn === 'next' ? 'queued' : 'running'), ...inherited }
1058 const routed = decision.routed === undefined ? {} : { routed: decision.routed }
1059 let next: BoardNode
1060 if ('failure' in decision) {
1061 next = { ...base, ...routed, why: failureLine(decision.failure.backend, decision.failure), failure: decision.failure }
1062 } else if ('why' in decision) {
1063 next = { ...without(base, 'failure'), routed: false, why: decision.why }
1064 } else if ('started' in decision) {
1065 next = { ...without(base, 'why', 'failure'), routed: true, state: 'running' }
1066 } else {
1067 next = {
1068 ...(decision.routed === true ? without(base, 'why', 'failure') : base),
1069 ...routed,
1070 ...(n === undefined ? {} : { decision: n }),
1071 }
1072 }
1073 return { ...board, nodes: [...board.nodes.filter((node) => node !== old && node !== stood), next] }
1074}
1075
1076/** The value of a cell, or the empty board. */
1077async function read(cell: Cell<Board>): Promise<Board> {
1078 return (await cell.get()).value ?? EMPTY
1079}
1080
1081/** `update` that writes only when `change` has something to write (it returns undefined when the value stands). */
1082async function modify<T>(cell: Cell<T>, change: (current: T | undefined) => T | undefined, attempts = 8): Promise<void> {
1083 for (let attempt = 0; attempt < attempts; attempt++) {
1084 const { value, version } = await cell.get()
1085 const next = change(value)
1086 if (next === undefined) return
1087 if ((await cell.set(next, { ifVersion: version })).isSet) return
1088 }
1089 throw new Error(`$.state write lost to other writers ${attempts} times in a row`)
1090}
1091
1092/** The object without these keys (a $.state value holds no `undefined`). */
1093function without<T extends object, K extends keyof T>(value: T, ...keys: K[]): Omit<T, K> {
1094 const rest = { ...value }
1095 for (const key of keys) delete rest[key]
1096 return rest
1097}
1098
1099// ---- what `report` does with the decision model's fallback ------------------------------
1100
1101/** The decision model that decides is not the one asked for (ADR 0006): e.g. `asked` is pplx, `using` is Jev, since there is no Perplexity key and there is a TypeSafe one. */
1102export type DecisionModelEvent = { asked: BackendName; using: BackendName }
1103
1104/** What a person reads of each decision model: its name, and the settings that hold its key (the options' names and the environment variable they type). The one place both are written. */
1105const DECISION_MODELS: Readonly<Record<BackendName, { name: string; keys: string }>> = {
1106 pplx: { name: 'pplx', keys: 'perplexityApiKey 或环境变量 PERPLEXITY_API_KEY' },
1107 jev: { name: 'Jev', keys: 'typesafeApiKey' },
1108}
1109
1110/** The feature name of the entry the fallback leaves in the decision log. */
1111export const DECISION_MODEL_FEATURE = 'decision-model'
1112
1113/**
1114 * The session's one entry in the decision log saying why Jev decides although pplx is the default: at the turn the session
1115 * start belongs to, with the debug line. A session start again at the same turn (a hot reload) takes the place of the entry
1116 * before. Not on the board's nodes, in no band or footer, and no toast. Never throws.
1117 */
1118async function reportFellBack(io: DecisionModelIo, fell: DecisionModelEvent): Promise<void> {
1119 try {
1120 const asked = DECISION_MODELS[fell.asked]
1121 const using = DECISION_MODELS[fell.using]
1122 // A board that cannot be read puts the session start at turn 1, where the first message will be.
1123 const turn = Math.max(1, (await read(io.board).catch(() => EMPTY)).turn)
1124 const entry: Omit<LogEntry, 'n'> = {
1125 turn,
1126 feature: DECISION_MODEL_FEATURE,
1127 tone: 'info',
1128 outcome: `改用 ${using.name}`,
1129 subject: '',
1130 reason: `没有 ${asked.name} 的密钥(${asked.keys}),改用 ${using.name};补上密钥就用 ${asked.name},decisionModel 设成 ${using.name} 则不再提示`,
1131 }
1132 io.debug(`decision model: ${decisionLine(entry)}`)
1133 await update(io.decisions, (list) => {
1134 const kept = list ?? []
1135 const at = kept.findIndex((old) => old.feature === DECISION_MODEL_FEATURE && old.turn === turn)
1136 return at < 0 ? appendEntry(kept, entry) : kept.map((old, i) => (i === at ? { n: old.n, ...entry } : old))
1137 })
1138 } catch (error) {
1139 try {
1140 io.debug(`decision model fallback not reported: ${errorText(error)}`)
1141 } catch {
1142 // nowhere left to say it
1143 }
1144 }
1145}
1146
1147// ---- what `report` does with the skill profiles' events ----------------------------
1148
1149/** Why writing the profiles stopped for the session (types/index.d.ts, `skillProfiles.stop`). */
1150export type ProfilesStopReason = 'store-read' | 'model-refused' | 'api-error' | 'store-write' | 'off' | 'error'
1151
1152/** How writing the session's skill profiles went (#33; types/index.d.ts, `skillProfiles`). */
1153export type ProfilesState = {
1154 phase: 'writing' | 'done' | 'stopped'
1155 /** The turn the session start belongs to: the turn of its decision log entry. */
1156 turn: number
1157 model: string
1158 kept: number
1159 planned: number
1160 written: number
1161 failed: number
1162 deferred: number
1163 stop?: { reason: ProfilesStopReason; detail: string }
1164 failures: { name: string; reason: string }[]
1165}
1166
1167/** The failed skills the state names; `failed` still counts every one. */
1168export const PROFILES_FAILURES_KEPT = 50
1169
1170/** What the skill profiles' report needs of the host: the board (for the turn), the decision log, the debug log, and the state it keeps. */
1171export type ProfilesIo = Pick<ReportIo, 'board' | 'decisions' | 'debug'> & { profiles: Cell<ProfilesState> }
1172
1173/** Why the writing stopped, with what the debug line and the state say of it. */
1174export type ProfilesStop =
1175 | { reason: 'model-refused'; model: string; detail: string }
1176 | { reason: 'api-error'; skill: string; why: string; ms: number }
1177 | { reason: 'store-write'; skill: string; detail: string }
1178 | { reason: 'off' }
1179 | { reason: 'error'; detail: string }
1180
1181/**
1182 * What happens as the skill profiles are written (features/skill-profiles.ts), in the order it happens. `left`:
1183 * the skills lacking a profile that this session did not get to.
1184 */
1185export type ProfileEvent =
1186 /** The session start has read the catalog: `offered` skills can have one, `due` of them lack one; at most `perSession` are written. */
1187 | { event: 'start'; model: string; perSession: number; offered: number; due: number }
1188 /** The store cannot be read: nothing is kept or written. */
1189 | { event: 'unreadable'; model: string; offered: number }
1190 /** Another session wrote the profile meanwhile: it is kept. */
1191 | { event: 'found' }
1192 | { event: 'written'; skill: string; model: string; ms: number; input: number; output: number }
1193 /** The model gave no profile for the skill (`why`: the reply's reason); the writing goes on. */
1194 | { event: 'failed'; skill: string; why: string; ms: number }
1195 /** The reply for the skill is no profile; the writing goes on. */
1196 | { event: 'unfit'; skill: string; ms: number; text: string }
1197 /** The session's most is written; `left` is what no longer fits (the debug line counts what is not written, failed included). */
1198 | { event: 'quota'; perSession: number; left: number }
1199 /** The loop is over, as it should be. */
1200 | { event: 'finish'; left: number }hooks/core/setup.ts 559 lines1// The options the person set (userConfig), read once per load into what every
2// feature receives: `ctx`. Pure: no `$`.
3//
4// Every option is read here, and only here: its bounds and its fallback are
5// written once, and every feature (and the eval, which builds its requests
6// from the same Config) reads the value from `ctx.config`. The fallbacks are
7// the manifest's defaults (the engine fills those in; a test or the eval may
8// leave an option out), except for the options whose default depends on the
9// decision model: those have no default in the manifest, so the engine
10// passes nothing for them until the person sets one, and their defaults are
11// in BACKEND_DEFAULTS below, the one table the mod, the eval and
12// scripts/decide*.ts read them from.
13
14import type { PluginOptions } from 'claude-code'
15import type { Backend } from '../decision/backend.ts'
16import type { ContextLimits } from '../decision/context.ts'
17import { DEFAULT_AGENT_MODELS, type AgentModel, type DispatchAsk, type DispatchSettings } from '../decision/dispatched-agent.ts'
18import type { EffortAsk, EffortRules } from '../decision/effort.ts'
19import { RAISE_MODES, type RaiseMode } from '../decision/escalation.ts'
20import { jevBackend } from '../decision/jev.ts'
21import type { MidturnLimits, MidturnRules } from '../decision/midturn.ts'
22import { resolveModel, type ResolvedModel } from '../decision/model-ids.ts'
23import { pplxBackend } from '../decision/pplx.ts'
24import { rateLimited } from '../decision/pplx-rate.ts'
25import type { SkillPolicy } from '../decision/skills.ts'
26
27/** The decision models the person can choose (`decisionModel`): Perplexity's pplx-decider and TypeSafe's Jev. */
28export type BackendName = 'pplx' | 'jev'
29
30/**
31 * The decision model the person asked for: Jev when `decisionModel` names it, else pplx, the default (ADR 0006). Anything
32 * that names neither (unset, a typo, the `clef` of a 0.3.x configuration) reads as unset. Which one decides in the end also
33 * depends on the keys (`chooseBackend`).
34 */
35export function backendNameOf(options: PluginOptions): BackendName {
36 return options.decisionModel === 'jev' ? 'jev' : 'pplx'
37}
38
39/** Which decision model decides, and whether it is not the one asked for. */
40export type Choice = { backend: BackendName; fellBack: boolean }
41
42/**
43 * The decision model that decides (ADR 0006). Jev when the person asked for Jev. Otherwise pplx when there is a Perplexity key;
44 * with none, Jev when there is a TypeSafe key (`fellBack`: a 0.3.1 configuration keeps its routing); with neither, pplx, which
45 * then fails every request with the Perplexity key named as the one missing.
46 */
47export function chooseBackend(asked: BackendName, keys: { perplexity: string; typesafe: string }): Choice {
48 if (asked === 'jev') return { backend: 'jev', fellBack: false }
49 if (keys.perplexity !== '') return { backend: 'pplx', fellBack: false }
50 return keys.typesafe !== '' ? { backend: 'jev', fellBack: true } : { backend: 'pplx', fellBack: false }
51}
52
53/**
54 * A `decisionModel` that names a decision model since removed (Clef, 0.4.0), left in the person's settings from before.
55 * The engine reads a value outside the option's list as the option's default (and warns), so the mod is handed the
56 * default, never the old value: this finds it in the settings file itself (`pluginConfigs[dispatch-pilot@...].options`,
57 * as `$.settings.read({ source: 'user' })` returns them). Null when the settings name none.
58 */
59export function removedDecisionModel(settings: unknown): 'clef' | null {
60 const configs = (settings as { pluginConfigs?: unknown } | null | undefined)?.pluginConfigs
61 if (typeof configs !== 'object' || configs === null) return null
62 for (const [plugin, config] of Object.entries(configs)) {
63 if (!plugin.startsWith('dispatch-pilot@')) continue
64 const options = (config as { options?: unknown } | null)?.options
65 if (typeof options === 'object' && options !== null && (options as { decisionModel?: unknown }).decisionModel === 'clef') return 'clef'
66 }
67 return null
68}
69
70/** The options whose default depends on the decision model: the manifest gives them none. */
71export const PER_BACKEND_OPTIONS = [
72 'timeoutMs',
73 'contextMessages',
74 'contextTokens',
75 'rejudgeSteps',
76 'rejudgeWaitMs',
77 'thetaUp',
78 'thetaDown',
79 'thetaMax',
80 'thetaExpected',
81 'agentOverride',
82 'skillsMinRelevance',
83 'findSkillMinRelevance',
84] as const
85export type PerBackendOption = (typeof PER_BACKEND_OPTIONS)[number]
86
87/**
88 * The kinds of request whose state has a budget of its own (`contextTokens`
89 * reads by kind): a message that carries the skills' question (and find_skill's
90 * first stage) is `Config.context`; the rest are these. `rejudge` is a
91 * mid-turn re-decision and a stuck one.
92 */
93export const CONTEXT_KINDS = ['messagePlain', 'rejudge', 'agent', 'workflow'] as const
94export type ContextKind = (typeof CONTEXT_KINDS)[number]
95
96/**
97 * What one decision model brings: each option's default where the person left
98 * it unset, the most `contextTokens` and `contextMessages` read as, how every
99 * question is asked and the threshold for taking the level above.
100 *
101 * Adding a decision model means touching, beyond its entry here: `BackendName`; the backend in decision/ and its
102 * construction in `setup()` (`configs`, `backends`, and whether it passes through the rate limit); `backendNameOf` and
103 * `chooseBackend` (which key decides, and what it falls back to); `DECISION_MODELS` in core/report.ts (its name and key
104 * settings, for the fallback entry); `decisionModel`'s options in plugin.json, its key option and the README's
105 * configuration tables (a column each); and, for the eval, `backendFor` and `PRICES` in eval/node.ts, `--backend` in
106 * eval/run.ts and eval/eval-v2-flow.ts, and the BACKENDS list in eval/lib/docs.ts. A `Record<BackendName, ...>` the compiler
107 * checks (this table, `setup()`, PRICES, DECISION_MODELS); the rest it does not, and tests/docs-sync.test.ts only the README tables.
108 */
109export type BackendDefaults = Readonly<Record<PerBackendOption, number>> & {
110 /** `contextTokens` above this reads as this. */
111 contextTokensMax: number
112 /** `contextMessages` above this reads as this. */
113 contextMessagesMax: number
114 /**
115 * The least probability the level above the most probable one needs to be taken instead (`EffortRules.roundUp`, decision/effort.ts).
116 * Not an option: set from the stored answers of the decision model (eval/resummarize.ts), like the other values of the table.
117 */
118 roundUp: number
119 /**
120 * How the decision model is asked, language and kind of question (the eval's variables, `effort-submit` for the effort question
121 * beside a message): `turnStart` for that question, `other` for every other one (the mid-turn re-decision, a dispatched agent
122 * and a Workflow's, a stuck loop, the skills, the problem summary's text). Not options.
123 */
124 ask: { turnStart: EffortAsk; other: EffortAsk }
125 /**
126 * What the state of each other kind of request may take, in estimated tokens (`contextTokens` is the message that carries the
127 * skills' question). Not options: each is the most that kind's longest question leaves, worked out in DEVELOPMENT.md (配置,
128 * 「Jev 的上下文默认值怎么算」) and checked by tests/backend-defaults.test.ts. What the person sets as `contextTokens` holds for
129 * every kind, each taking the smaller of it and its own.
130 */
131 contextByKind: Readonly<Record<ContextKind, number>>
132 /**
133 * How long find_skill's two requests may take in all, in ms; null: as long
134 * as a message waits (`timeoutMs`). Not an option: set by the latency
135 * measured, not calibrated. A hook has 10 s of its own, and the timer's wait
136 * counts toward it (the mods reference, Limits): this leaves room for the rest.
137 */
138 findSkillWaitMs: number | null
139 /** Whether find_skill's first request offers a skill by its profile (else by its description); the second re-reads it by both. */
140 findSkillProfiles: boolean
141}
142
143/**
144 * Jev's: the values the eval of #4, #14, #15 and #16 set or kept (README, 配置 has the table; DEVELOPMENT.md, 配置 what each rests on),
145 * except the three that say how much Jev reads of the conversation: by the person's decision ("the more it is given, the
146 * better it judges"), those are set to what Jev accepts, not to what the eval ran (4 messages, 2000 tokens, 4 steps; the
147 * eval sets are too short to tell, and no comparison was run). Jev takes 64k tokens a request and 32k for "the state and
148 * the longest question". The longest question is the first stage of the skills (111 skills with profiles: about 16k tokens
149 * as the mod counts, 21.9k as Jev does), which shares its request with the state, so `contextTokens` is what that leaves,
150 * 6000 (DEVELOPMENT.md, 配置, 「Jev 的上下文默认值怎么算」; tests/backend-defaults.test.ts checks the sum). The two counts
151 * that fill it are at their most: `contextMessages` 32 and `rejudgeSteps` 16, so the token budget, not the count, ends what is sent.
152 */
153const JEV_DEFAULTS: BackendDefaults = {
154 timeoutMs: 1500,
155 contextMessages: 32,
156 contextMessagesMax: 32,
157 contextTokens: 6000,
158 contextTokensMax: 16000,
159 // Every other kind of request has questions of at most 700 tokens as Jev counts them (the effort question 421, a re-decision's
160 // 295 with its trouble, an agent's 457, nine of them 1,836): the state plus the longest of them within 28,800 (90% of 32k)
161 // allows about 25,400 estimated tokens, and 24000 is the round number below it. A Workflow batch (at most 8 agents' questions,
162 // 20,100 in all) and the whole request stay within 57,600 (90% of 64k) too.
163 contextByKind: { messagePlain: 24000, rejudge: 24000, agent: 24000, workflow: 24000 },
164 rejudgeSteps: 16,
165 // A step waits this long for a re-decision not yet back before it goes on at the level it had: Jev's mid-turn requests took about 280 ms
166 // at p50 and 350 ms at p90 (DEVELOPMENT.md, 配置), and they are sent when the step's tools start, so they are mostly back by then.
167 rejudgeWaitMs: 300,
168 // Raising is easy, lowering is hard (AA: a Sonnet 5.5 at medium scores 41 on the index and at high 47, at low 36; Terminal-Bench
169 // 20.7% at low against 43.9% at high): 0.4 to 0.3 for a raise, 0.6 to 0.75 for a lowering, 3 to 5 steps held after a raise.
170 // The lowering gate came back down to 0.55 in 0.2.3: with 0.75 the level sent was too high too often (too high 11.5/11.0% to
171 // 19.5/17.5%); a scan of the stored effort-midturn answers (eval/rescore.ts --theta-down) found none of 0.55 to 0.75 to bring it back, and 0.55 is the one with the least too high and too low together.
172 thetaUp: 0.3,
173 thetaDown: 0.55,
174 thetaMax: 0.5,
175 thetaExpected: 0.25,
176 agentOverride: 0.6,
177 // #16, both runs on the old wording: from 0.7 to 0.75 the Chinese-English gap of the skill suggestions went from -3.7 to -2.3 points.
178 skillsMinRelevance: 0.75,
179 findSkillMinRelevance: 0.5,
180 findSkillWaitMs: null,
181 findSkillProfiles: true,
182 // Raising is easy (DEVELOPMENT.md, 「按 AA 基准校正」): the level above the most probable one is taken from 0.3.
183 roundUp: 0.3,
184 // effort-submit on the current wording, one run each (2026-10-05): asked in Chinese, Chinese items 85% and English
185 // items 89%; asked in English, 79% and 78%. The questions asked later have no data in Chinese and stay in English.
186 ask: { turnStart: { language: 'zh', primitive: 'score' }, other: { language: 'en', primitive: 'score' } },
187}
188
189/**
190 * pplx-decider-v1.1-27b's: B′ of eval v2 (#45, ADR 0006: the state within 48000 tokens, every question in English, no problem
191 * summary), with the thresholds calibrated on its own stored answers (DEVELOPMENT.md, eval v2 的结果). It takes a 262k-token
192 * window, so what it reads is not cut to Jev's 32k: 2000 messages (the budget, not the count, ends what is sent) and 48000
193 * tokens for every request but the skills'; those two stages keep Jev's 6000, since the profiles of the skills are the same
194 * length for either model. A request takes seconds, not Jev's fraction of one, so a message waits 8000 ms (a hook has 10 s) and a
195 * step 6000 ms for a re-decision. What the eval did not calibrate for pplx (the failure bar, the agents' override, the skills'
196 * relevance) is Jev's.
197 */
198const PPLX_DEFAULTS: BackendDefaults = {
199 timeoutMs: 8000,
200 contextMessages: 2000,
201 contextMessagesMax: 2000,
202 contextTokens: 6000,
203 contextTokensMax: 48000,
204 contextByKind: { messagePlain: 48000, rejudge: 48000, agent: 48000, workflow: 48000 },
205 rejudgeSteps: 16,
206 // A re-decision takes seconds, not 300 ms: a step waits for it up to this long (the most the manifest allows is 8000).
207 rejudgeWaitMs: 6000,
208 // Calibrated offline on the stored pplx answers (eval v2, DEVELOPMENT.md): thetaMax 0.47 (0.48 sits on submit-034's p(max) of
209 // 0.480), thetaDown 0.55 as Jev's (from 0.6 up, the English score question sends the level too high more often than Jev).
210 // thetaUp 0 raises a level mid-turn on any answer above it (0.3 is the cautious alternative); it does not move the level a
211 // message is sent at, which thetaMax and roundUp decide.
212 thetaUp: 0,
213 thetaDown: 0.55,
214 thetaMax: 0.47,
215 thetaExpected: 0.25,
216 agentOverride: 0.6,
217 skillsMinRelevance: 0.75,
218 findSkillMinRelevance: 0.5,
219 // Find_skill's two requests wait as long as the profiles' question takes (a message's wait would be 8000 ms of a 10 s hook).
220 findSkillWaitMs: 6000,
221 findSkillProfiles: true,
222 // The level above the most probable one is taken from 0.45 (ADR 0006): the offline scan of the stored effort-submit answers
223 // (English questions) took the level sent too high from 18.5% to 13.0% with the recall of max unchanged (11 of 16); eval v2's
224 // level too low went from 12.5% to 15.0%.
225 roundUp: 0.45,
226 ask: { turnStart: { language: 'en', primitive: 'score' }, other: { language: 'en', primitive: 'score' } },
227}
228
229/** The defaults by decision model: the only place they are written. */
230export const BACKEND_DEFAULTS: Readonly<Record<BackendName, BackendDefaults>> = {
231 pplx: PPLX_DEFAULTS,
232 jev: JEV_DEFAULTS,
233}
234
235/**
236 * What the state of a message's decision request may take: the message that
237 * carries the skills' question (`withSkills`) has `context`, whose tokens
238 * leave room for that question; the effort question's request, and any other
239 * plain message, has `contextByKind.messagePlain` (ADR 0005). The core and the
240 * eval build the request's state from the same limits.
241 */
242export function messageLimits(config: Pick<Config, 'context' | 'contextByKind'>, withSkills: boolean): ContextLimits {
243 return withSkills ? config.context : { ...config.context, tokens: config.contextByKind.messagePlain }
244}
245
246/** The settings are also the effort rules' parameters (`EffortRules`: `thetaMax` and `roundUp`): `traceEffort(reading, config)`. */
247export type Config = EffortRules & {
248 /** The decision model the options were read for: its defaults (BACKEND_DEFAULTS) stand for what the person left unset. */
249 backend: BackendName
250 /**
251 * What came from that decision model's defaults (for the debug log, describeDefaults): the options the person left
252 * unset, with the value each took, and an option set above that model's most, read as the most.
253 */
254 defaults: { used: readonly (readonly [PerBackendOption, number])[]; capped: readonly { option: PerBackendOption; set: number; read: number }[] }
255 /** The TypeSafe key in the options (`typesafeApiKey`, trimmed; '' if unset). Not enumerable (see readConfig): it is not in a written-out or copied config. */
256 readonly typesafeApiKey: string
257 /** The Perplexity key in the options (`perplexityApiKey`, trimmed; '' if unset; not enumerable either): the environment's (`Secrets`) stands where this is empty. */
258 readonly perplexityApiKey: string
259 /** How many pplx requests the mod sends in a second at most (`pplxQps`; the account's limit): the rest wait their turn (decision/pplx-rate.ts). Jev has no such limit. */
260 pplxQps: number
261 /** How long a decision request may take before the prompt goes on without it. */
262 timeoutMs: number
263 /** How the decision model is asked (BACKEND_DEFAULTS ask): `turnStart` the effort question beside each message, `other` every other question (`ctx.ask`). */
264 ask: { turnStart: EffortAsk; other: EffortAsk }
265 /**
266 * What the decision model reads of the conversation: how many recent messages, how many tokens in all. The tokens are the
267 * budget of a message that carries the skills' question (and of find_skill's first stage); the other kinds of request have
268 * theirs in `contextByKind` (a mid-turn re-decision's in `midturn.limits`, too).
269 */
270 context: ContextLimits
271 /** The tokens the state of each other kind of request may take: the decision model's default for the kind, or what the person set if smaller. */
272 contextByKind: Readonly<Record<ContextKind, number>>
273 /** The main agent's effort decided again while a turn runs (#5); a stuck loop's re-decision (#7) reads its loop the same way. */
274 midturn: {
275 /** Re-decide at every step whose index is a multiple of this; 0 for never. */
276 every: number
277 /** How long a step waits for a re-decision (or a stuck one) not back yet. */
278 waitMs: number
279 /** What a re-decision reads: the latest steps, within the context budget. */
280 limits: MidturnLimits
281 /** How an answer moves the level: thetaUp, thetaDown, thetaMax, roundUp, holdSteps. */
282 rules: MidturnRules
283 }
284 /** Forced escalation (#7). */
285 escalation: {
286 /** Counted failures that make a loop stuck. */
287 after: number
288 mode: RaiseMode
289 /** Forced raises a loop gets at most. */
290 limit: number
291 /** The failures count as expected when the answer reaches this probability. */
292 thetaExpected: number
293 /** The model a failing haiku agent is switched to, resolved to a full id; null: never switch (left empty, or a name no model family has). */
294 haikuTo: ResolvedModel | null
295 /** `escalateHaikuTo` as the person wrote it, for the log when it names no model. */
296 haikuToWritten: string
297 }
298 /** Dispatched agents and a Workflow's agents (#6, #8, #9). */
299 agents: {
300 /** The models the decision model may choose, cheapest first (fable with `agentFable`). */
301 models: readonly AgentModel[]
302 /** The decision model's pick replaces the main agent's (or the script's) only at this confidence or above. */
303 thetaOverride: number
304 /** `rewrite` writes a Workflow's decisions into its script; `return` sends the Workflow back with them, once. */
305 workflowMode: 'rewrite' | 'return'
306 }
307 /** Skills (#10, #11, #12). */
308 skills: {
309 /** What a message is suggested: at most `max`, from `minRelevance`. */
310 suggest: SkillPolicy
311 /** What find_skill returns: at most `max`, from `minRelevance`. */
312 find: SkillPolicy
313 /** How long find_skill's two requests may take in all (the decision model's: BACKEND_DEFAULTS findSkillWaitMs). */
314 findWaitMs: number
315 /** Whether find_skill's first request offers skills by their profiles (the decision model's: findSkillProfiles). */
316 findByProfile: boolean
317 /** How many of stage one's best stage two re-reads. */
318 shortlist: number
319 /** Skills the main agent keeps in its listing (names as the listing spells them). */
320 alwaysListed: readonly string[]
321 /** Skills never offered, to the main agent or to the person. */
322 neverSuggested: readonly string[]
323 /** The model that writes skill profiles; part of every profile's key. */
324 profileModel: string
325 /** How many missing profiles a session start writes at most. */
326 profilesPerSession: number
327 }
328 /** The unresolved count and the problem summary (#39, #40). */
329 unresolved: {
330 /** The model that writes the problem summary after each of the person's turns (`summaryModel`). */
331 summaryModel: string
332 /** The count from which the effort question of a message and of a mid-turn re-decision carries the strong hint (`unresolvedMaxAfter`; 0: never). */
333 maxAfter: number
334 }
335}
336
337/**
338 * What the mod reads from the environment: `$.env.get` is asynchronous and needs `$`, which neither `setup` nor an import may
339 * have, so the session start (features/control.ts) reads it into this object, which `ctx.backend` reads when it asks. A
340 * value set in the options always comes first. Nothing here is ever written down: not to the debug log, the board or `$.state`.
341 */
342export type Secrets = {
343 /** `PERPLEXITY_API_KEY`, trimmed; '' while unread or unset. */
344 perplexityEnvKey: string
345}
346
347/** The Perplexity key in use: the options' (`perplexityApiKey`) first, else the environment's. '' for none. */
348export function perplexityKey(config: Pick<Config, 'perplexityApiKey'>, secrets: Secrets): string {
349 return config.perplexityApiKey !== '' ? config.perplexityApiKey : secrets.perplexityEnvKey
350}
351
352/**
353 * What the entry hands every feature's register: data and pure functions, never `$`.
354 *
355 * Which decision model decides is not known at register: the Perplexity key may be in the environment, which the session start
356 * reads (`secrets`). `config`, `backend`, `ask` and `fellBack` are therefore read afresh at each use and follow the choice: a
357 * feature that keeps a value from `ctx.config` when it registers keeps the wrong model's (the defaults differ), so it reads
358 * `ctx.config` where it acts.
359 */
360export type Ctx = {
361 /** The settings of the decision model that decides (`chooseBackend`). */
362 readonly config: Config
363 /** The decision model that decides. */
364 readonly backend: Backend
365 /** What the session start read from the environment (a key not in the options). */
366 readonly secrets: Secrets
367 /** How every question but the effort question beside a message is asked: the decision model's (`config.ask.other`). */
368 readonly ask: EffortAsk
369 /** Whether Jev decides only because there is no Perplexity key (a TypeSafe key is set): the session start records it. */
370 readonly fellBack: boolean
371}
372
373export function setup(options: PluginOptions, table: Readonly<Record<BackendName, BackendDefaults>> = BACKEND_DEFAULTS): Ctx {
374 const asked = backendNameOf(options)
375 // Both are worked out here, whichever decides: pure and cheap, and `chooseBackend` picks between them once the environment is read.
376 const configs: Record<BackendName, Config> = { pplx: readConfig(options, table, 'pplx'), jev: readConfig(options, table, 'jev') }
377 const secrets: Secrets = { perplexityEnvKey: '' }
378 const backends: Record<BackendName, Backend> = {
379 // Every pplx request goes through the rate limit (#51); Jev does not.
380 pplx: rateLimited(pplxBackend(() => perplexityKey(configs.pplx, secrets)), configs.pplx.pplxQps),
381 jev: jevBackend(configs.jev.typesafeApiKey),
382 }
383 const choice = () => chooseBackend(asked, { perplexity: perplexityKey(configs.pplx, secrets), typesafe: configs.pplx.typesafeApiKey })
384 return {
385 get config() {
386 return configs[choice().backend]
387 },
388 get backend() {
389 return backends[choice().backend]
390 },
391 secrets,
392 get ask() {
393 return configs[choice().backend].ask.other
394 },
395 get fellBack() {
396 return choice().fellBack
397 },
398 }
399}
400
401/**
402 * The person's options, each within its bounds; where one is missing or of
403 * the wrong type, the manifest's default, or for an option whose default
404 * depends on the decision model (PER_BACKEND_OPTIONS), that model's
405 * (`table`, BACKEND_DEFAULTS unless a test reads another). `contextTokens`
406 * and `contextMessages` read at most that model's most. `backend` is the
407 * decision model it reads for: what `decisionModel` asks for unless the
408 * keys decide otherwise (`setup` reads both).
409 */
410export function readConfig(options: PluginOptions, table: Readonly<Record<BackendName, BackendDefaults>> = BACKEND_DEFAULTS, backend: BackendName = backendNameOf(options)): Config {
411 const defaults = table[backend]
412 const used: (readonly [PerBackendOption, number])[] = []
413 const capped: { option: PerBackendOption; set: number; read: number }[] = []
414 /** A per-backend option: the person's value within [min, max], else the decision model's default (noted for the log). */
415 const own = (option: PerBackendOption, min: number, max: number, round = false) => {
416 const set = options[option]
417 if (typeof set !== 'number' || !Number.isFinite(set)) {
418 used.push([option, defaults[option]])
419 return defaults[option]
420 }
421 const read = numberIn(set, min, max, defaults[option])
422 const value = round ? Math.round(read) : read
423 if (set > max) capped.push({ option, set, read: value })
424 return value
425 }
426 const whole = (value: unknown, min: number, max: number, fallback: number) => Math.round(numberIn(value, min, max, fallback))
427 const thetaMax = own('thetaMax', 0, 1)
428 // `contextTokens` by kind of request: unset, each kind's own (BACKEND_DEFAULTS); set, the smaller of it and each kind's.
429 const contextTokens = own('contextTokens', 100, defaults.contextTokensMax, true)
430 const contextSet = typeof options.contextTokens === 'number' && Number.isFinite(options.contextTokens)
431 const byKind = (cap: number) => (contextSet ? Math.min(contextTokens, cap) : cap)
432 const context = { messages: own('contextMessages', 0, defaults.contextMessagesMax, true), tokens: byKind(defaults.contextTokens) }
433 const contextByKind = Object.fromEntries(CONTEXT_KINDS.map((kind) => [kind, byKind(defaults.contextByKind[kind])])) as Record<ContextKind, number>
434 const haikuToWritten = stringOf(options.escalateHaikuTo, 'sonnet').trim()
435 // A hook's own budget is 10 s and the timer's wait counts toward it.
436 const timeoutMs = own('timeoutMs', 200, 8000)
437 const config: Omit<Config, 'typesafeApiKey' | 'perplexityApiKey'> = {
438 backend,
439 defaults: { used, capped },
440 pplxQps: whole(options.pplxQps, 1, 50, 1),
441 timeoutMs,
442 ask: defaults.ask,
443 thetaMax,
444 roundUp: defaults.roundUp,
445 context,
446 contextByKind,
447 midturn: {
448 every: whole(options.rejudgeEvery, 0, 50, 3),
449 // A hook's own budget is 10 s and the timer's wait counts toward it, as for `timeoutMs`.
450 waitMs: own('rejudgeWaitMs', 0, 8000, true),
451 limits: { steps: own('rejudgeSteps', 1, 16, true), tokens: contextByKind.rejudge },
452 rules: {
453 thetaUp: own('thetaUp', 0, 1),
454 thetaDown: own('thetaDown', 0, 1),
455 thetaMax,
456 roundUp: defaults.roundUp,
457 holdSteps: whole(options.holdSteps, 0, 50, 5),
458 },
459 },
460 escalation: {
461 after: whole(options.escalateAfter, 1, 20, 2),
462 mode: RAISE_MODES.find((mode) => mode === options.escalateMode) ?? 'one-level',
463 limit: whole(options.escalateLimit, 0, 10, 2),
464 thetaExpected: own('thetaExpected', 0, 1),
465 haikuTo: haikuToWritten === '' ? null : resolveModel(haikuToWritten),
466 haikuToWritten,
467 },
468 agents: {
469 models: options.agentFable === true ? [...DEFAULT_AGENT_MODELS, 'fable'] : DEFAULT_AGENT_MODELS,
470 thetaOverride: own('agentOverride', 0, 1),
471 workflowMode: stringOf(options.workflowMode, 'rewrite') === 'return' ? 'return' : 'rewrite',
472 },
473 skills: {
474 suggest: { max: whole(options.skillsMax, 0, 10, 3), minRelevance: own('skillsMinRelevance', 0, 1) },
475 find: { max: whole(options.findSkillMax, 1, 10, 5), minRelevance: own('findSkillMinRelevance', 0, 1) },
476 findWaitMs: defaults.findSkillWaitMs ?? timeoutMs,
477 findByProfile: defaults.findSkillProfiles,
478 shortlist: whole(options.skillsShortlist, 1, 10, 4),
479 alwaysListed: namesOf(options.skillsAlwaysListed),
480 neverSuggested: namesOf(options.skillsNeverSuggested),
481 profileModel: stringOf(options.skillsProfileModel, DEFAULT_PROFILE_MODEL).trim() || DEFAULT_PROFILE_MODEL,
482 profilesPerSession: whole(options.skillsProfilesPerSession, 0, 500, 30),
483 },
484 unresolved: {
485 summaryModel: stringOf(options.summaryModel, DEFAULT_SUMMARY_MODEL).trim() || DEFAULT_SUMMARY_MODEL,
486 maxAfter: whole(options.unresolvedMaxAfter, 0, 10, 3),
487 },
488 }
489 // In the table's order, whatever order they were read in.
490 used.sort((a, b) => PER_BACKEND_OPTIONS.indexOf(a[0]) - PER_BACKEND_OPTIONS.indexOf(b[0]))
491 // The keys are not enumerable: `JSON.stringify(config)`, `{ ...config }` and `Object.keys` do not meet them, so a
492 // config written to a log, a result file or the board (the eval saves its settings) cannot carry a key by accident.
493 return Object.defineProperties(config, {
494 typesafeApiKey: { value: stringOf(options.typesafeApiKey, '').trim(), enumerable: false },
495 perplexityApiKey: { value: stringOf(options.perplexityApiKey, '').trim(), enumerable: false },
496 }) as Config
497}
498
499/**
500 * One debug-log line on what the decision model's defaults decided, e.g.
501 * `settings for jev: left unset, so jev's defaults: timeoutMs 1500, ...;
502 * skill suggestions on until /dp skills off; contextTokens 20000 reads as
503 * 16000, the most with jev`.
504 */
505export function describeDefaults(config: Pick<Config, 'backend' | 'defaults'>): string {
506 const { backend, defaults } = config
507 const used = defaults.used.length === 0 ? 'every option set' : `left unset, so ${backend}'s defaults: ${defaults.used.map(([option, value]) => `${option} ${value}`).join(', ')}`
508 const capped = defaults.capped.map((cap) => `; ${cap.option} ${cap.set} reads as ${cap.read}, the most with ${backend}`).join('')
509 return `settings for ${backend}: ${used}; skill suggestions on until /dp skills off${capped}`
510}
511
512/**
513 * How an agent's model and effort are decided (decision/dispatched-agent.ts):
514 * the same for a dispatched agent (#6) and a Workflow's agents (#8, #9), and
515 * for the eval; `ask` is how its questions are written (the eval's variants).
516 */
517export function dispatchSettings(ctx: { config: Pick<Config, 'agents' | 'thetaMax' | 'roundUp'>; ask: EffortAsk }, ask: Partial<DispatchAsk> = {}): DispatchSettings {
518 // Each read goes to `ctx`: a feature builds this when it registers, before the session start has settled the decision model.
519 return {
520 get models() {
521 return ctx.config.agents.models
522 },
523 get ask() {
524 return { ...ctx.ask, ...ask }
525 },
526 get thetaOverride() {
527 return ctx.config.agents.thetaOverride
528 },
529 get thetaMax() {
530 return ctx.config.thetaMax
531 },
532 get roundUp() {
533 return ctx.config.roundUp
534 },
535 }
536}
537
538/** The cheap model that writes skill profiles, unless the person names another (`skillsProfileModel`). */
539export const DEFAULT_PROFILE_MODEL = 'haiku'
540
541/** The cheap model that writes the problem summary, unless the person names another (`summaryModel`). */
542export const DEFAULT_SUMMARY_MODEL = 'haiku'
543
544/** A numeric option clamped to [min, max]; `fallback` when it is not a number. */
545export function numberIn(value: unknown, min: number, max: number, fallback: number): number {
546 return typeof value === 'number' && Number.isFinite(value) ? Math.min(max, Math.max(min, value)) : fallback
547}
548
549/** A string option; `fallback` when it is not a string. A sensitive one never set arrives as ''. */
550export function stringOf(value: unknown, fallback: string): string {
551 return typeof value === 'string' ? value : fallback
552}
553
554/** A list-of-names option (`multiple` in the manifest; a comma-separated string also reads): trimmed, empty ones dropped. */
555export function namesOf(value: unknown): string[] {
556 const items = Array.isArray(value) ? value : typeof value === 'string' ? value.split(',') : []
557 return items.flatMap((item) => (typeof item === 'string' && item.trim() !== '' ? [item.trim()] : []))
558}
559hooks/features/control.ts 211 lines1// Feature: the person's control over Dispatch Pilot, the `/dp` command.
2//
3// /dp opens the rationale pane (依据面板), or closes it when open
4// /dp log opens the same pane (the old habit's name for it)
5// /dp status what is on, the lock
6// /dp on | off the whole mod
7// /dp <name> on | off one feature (each registers its own, core/switches.ts)
8// /dp lock <effort> hold the main agent at that effort on every step
9// /dp unlock release it
10// /dp log N the last N decisions and why, in the conversation
11//
12// A pane the surface does not place (one that places no panes) is closed again
13// at once, and the answer says so (spec #22: no pane left open and waiting).
14//
15// What the person flips is kept in $.store and loaded back at session start.
16// The engine puts the plugin's name before a command's answer, so the texts
17// here do not start with one.
18//
19// It also records what the engine measures of the session (context fill, limit
20// percentages, cost) in the debug log. Only recorded: nothing reads them to
21// decide anything (spec, 可见性与控制; a later quota-saving mode may).
22
23import type { On, SessionMeasureInput } from 'claude-code'
24import { EFFORTS, isEffort, type Effort } from '../decision/effort.ts'
25import { decisionLine, LOG_ENTRIES, report, unplacedText, type DecisionModelIo, type LogEntry, type SwitchIo } from '../core/report.ts'
26import { describeDefaults, removedDecisionModel, type Ctx } from '../core/setup.ts'
27import { defineSwitch, isOn, listSwitches, loadOverrides, masterOn, overrides, parseOverrides, setMaster, setSwitch } from '../core/switches.ts'
28import { errorText } from '../decision/backend.ts'
29import { PANE_COLUMNS, PANE_ID, PANE_TITLE } from '../board/rationale.ts'
30
31const LOCK = { plugin: 'dispatch-pilot', key: 'lock' } as const
32const DECISIONS = { plugin: 'dispatch-pilot', key: 'decisionLog' } as const
33const BOARD = { plugin: 'dispatch-pilot', key: 'board' } as const
34/** The switches the person flipped, in $.store. */
35const SWITCHES_KEY = 'switches'
36
37export function registerControl(on: On, ctx: Ctx): void {
38 defineSwitch({ name: 'signals', info: '把上下文、限额和花费的读数记进 debug log,不参与任何决策' })
39
40 // Under a match-all matcher: other features set themselves up at session start too.
41 on('session.start', { cwd: /(?:)/ }, async ($, e, next) => {
42 // The person's switches first: the features beneath read them while the session starts.
43 loadOverrides(await $.store.get(SWITCHES_KEY).catch(() => undefined))
44 // The environment's Perplexity key (the options' comes first), for the requests that follow; never written down, only where it came from.
45 // Read before anything is said of the decision model: whether it is pplx or Jev depends on it (core/setup.ts, `Ctx`).
46 ctx.secrets.perplexityEnvKey = ((await $.env.get('PERPLEXITY_API_KEY').catch(() => undefined)) ?? '').trim()
47 // Which options the decision model's defaults decided (core/setup.ts BACKEND_DEFAULTS).
48 $.ui.log(describeDefaults(ctx.config), { to: 'debug' })
49 // Jev decides only for want of a Perplexity key: the decision log says so (ADR 0006), where the person looks for why.
50 if (ctx.fellBack) {
51 const io: DecisionModelIo = {
52 board: { get: () => $.state.get(BOARD), set: (value, options) => $.state.set(BOARD, value, options) },
53 decisions: { get: () => $.state.get(DECISIONS), set: (value, options) => $.state.set(DECISIONS, value, options) },
54 debug: (line) => $.ui.log(line, { to: 'debug' }),
55 }
56 await report(io, { fellBack: { asked: 'pplx', using: 'jev' } })
57 }
58 if (ctx.config.backend === 'pplx') $.ui.log(`pplx key: ${ctx.config.perplexityApiKey !== '' ? 'from the options' : ctx.secrets.perplexityEnvKey !== '' ? 'from PERPLEXITY_API_KEY' : 'not set'}`, { to: 'debug' })
59 // A decisionModel left over in the person's settings that names a model since removed: the engine reads it as the default.
60 const removed = removedDecisionModel(await $.settings.read({ source: 'user' }).catch(() => undefined))
61 if (removed !== null) $.ui.log(`decisionModel ${removed} 已移除,按没设处理`, { to: 'debug' })
62 const result = await next(e)
63 await $.command
64 .register({
65 name: 'dp',
66 description: 'Dispatch Pilot:依据面板、功能开关、effort 锁定、最近的决定',
67 argumentHint: '[log | status | on|off | <name> on|off | lock <effort> | unlock | log N]',
68 immediate: true,
69 })
70 .catch((error) => $.ui.log(`/dp was not registered: ${errorText(error)}`, { to: 'debug' }))
71 return result
72 })
73
74 // `window` is in every measurement: a matcher on a field that is always there.
75 on('session.measure', { context: { window: /(?:)/ } }, async ($, e, next) => {
76 if (isOn('signals')) {
77 try {
78 $.ui.log(describeSignals(e), { to: 'debug' })
79 } catch {
80 // a reading that cannot be written down is skipped
81 }
82 }
83 return next(e)
84 })
85
86 on('command.run', { command: 'dp' }, async ($, e) => {
87 const command = parseControl(e.args)
88 // The screens draw again: they leave out what a feature that is off owns.
89 const reporting: SwitchIo = { redraw: () => $.ui.invalidate('ui.render') }
90 // Keeps one switch as the person just flipped it, beside whatever another session saved meanwhile;
91 // says so when it cannot.
92 const save = async (name: string) => {
93 try {
94 const saved = parseOverrides(await $.store.get(SWITCHES_KEY))
95 const now = overrides()[name]
96 if (now === undefined) delete saved[name]
97 else saved[name] = now
98 await $.store.set(SWITCHES_KEY, saved)
99 return ''
100 } catch {
101 return '(没有保存:本地存储用不了,只在这个会话里有效)'
102 }
103 }
104 try {
105 switch (command.kind) {
106 case 'pane': {
107 // `/dp` toggles; `/dp log` only opens (and asks for the keys again).
108 if (command.toggle && (await $.ui.panes()).some((pane) => pane.id === PANE_ID)) {
109 await $.ui.close({ id: PANE_ID })
110 return { text: '依据面板已关闭' }
111 }
112 const opened = await $.ui.open({ id: PANE_ID, title: PANE_TITLE, focus: true, closeOnEscape: true, columns: PANE_COLUMNS })
113 if (opened.isPlaced) return { text: '依据面板已打开:p / n 翻看 agent,Esc 关闭' }
114 await $.ui.close({ id: PANE_ID })
115 return { text: unplacedText(opened.reason) }
116 }
117 case 'status': {
118 const { value: lock = null } = await $.state.get(LOCK)
119 return { text: describeStatus(lock) }
120 }
121 case 'master':
122 setMaster(command.on)
123 await report(reporting, { switched: { master: command.on } })
124 return { text: `Dispatch Pilot 已${command.on ? '打开' : '关闭'}${await save('master')}` }
125 case 'switch': {
126 const spec = listSwitches().find((s) => s.name === command.name)
127 if (spec === undefined || !setSwitch(command.name, command.on)) {
128 const names = listSwitches().map((s) => s.name).join(', ')
129 return { text: `没有叫「${command.name}」的开关(现有:${names})` }
130 }
131 await report(reporting, { switched: { feature: spec.name, on: command.on } })
132 return { text: `${spec.name} 已${command.on ? '打开' : '关闭'}:${spec.info}${await save(spec.name)}` }
133 }
134 case 'lock': {
135 await $.state.set(LOCK, command.effort)
136 const inert = masterOn() ? '' : '(Dispatch Pilot 关着:/dp on 才生效)'
137 return { text: `主 agent 已锁在 ${command.effort}:每一步都用这一档${inert};/dp unlock 解除` }
138 }
139 case 'unlock':
140 await $.state.set(LOCK, null)
141 return { text: '已解除锁定,effort 回到按决定走' }
142 case 'log': {
143 const { value: kept = [] } = await $.state.get(DECISIONS)
144 return { text: describeDecisions(kept.slice(-command.count)) }
145 }
146 case 'unknown':
147 return { text: `没看懂这条命令。\n${USAGE}` }
148 }
149 } catch (error) {
150 // Always answer: a failing command must not leave the person without a word.
151 return { text: `出错了(${errorText(error)})` }
152 }
153 })
154}
155
156type Control =
157 | { kind: 'pane'; toggle: boolean }
158 | { kind: 'status' }
159 | { kind: 'master'; on: boolean }
160 | { kind: 'switch'; name: string; on: boolean }
161 | { kind: 'lock'; effort: Effort }
162 | { kind: 'unlock' }
163 | { kind: 'log'; count: number }
164 | { kind: 'unknown' }
165
166/** The words after `/dp`, case and spacing aside. */
167function parseControl(args: string): Control {
168 const words = args.trim().toLowerCase().split(/\s+/).filter(Boolean)
169 const [first, second] = words
170 if (words.length === 0) return { kind: 'pane', toggle: true }
171 if (words.length === 1 && first === 'status') return { kind: 'status' }
172 if (words.length === 1 && (first === 'on' || first === 'off')) return { kind: 'master', on: first === 'on' }
173 if (words.length === 1 && first === 'unlock') return { kind: 'unlock' }
174 if (words.length === 2 && first === 'lock' && second === 'off') return { kind: 'unlock' }
175 if (words.length === 2 && first === 'lock' && isEffort(second)) return { kind: 'lock', effort: second }
176 if (words.length === 1 && first === 'log') return { kind: 'pane', toggle: false }
177 if (words.length === 2 && first === 'log' && /^[1-9]\d*$/.test(second ?? '')) return { kind: 'log', count: Math.min(LOG_ENTRIES, Number(second)) }
178 if (words.length === 2 && first !== undefined && (second === 'on' || second === 'off')) return { kind: 'switch', name: first, on: second === 'on' }
179 return { kind: 'unknown' }
180}
181
182const USAGE = `/dp(依据面板) | /dp status | /dp on|off | /dp <功能名> on|off | /dp lock <${EFFORTS.join('|')}> | /dp unlock | /dp log N`
183
184/** `/dp status`'s answer: whether the mod is on, the lock, each feature's switch. */
185function describeStatus(lock: Effort | null): string {
186 const switches = listSwitches()
187 const width = Math.max(0, ...switches.map((s) => s.name.length))
188 return [
189 `Dispatch Pilot ${masterOn() ? '开着' : '关着'}。effort 锁定:${lock ?? '没有'}。`,
190 '功能开关(/dp <功能名> on|off):',
191 ...switches.map((s) => ` ${s.on ? '开' : '关'} ${s.name.padEnd(width)} ${s.info}`),
192 USAGE,
193 ].join('\n')
194}
195
196/** `/dp log N`'s answer: the decisions, oldest first, each with its reason. */
197function describeDecisions(entries: readonly LogEntry[]): string {
198 if (entries.length === 0) return '还没有记下任何决定'
199 const header = `最近 ${entries.length} 条决定,最新的在最后`
200 return [header, ...entries.map((entry) => `#${entry.n} ${entry.feature}:${decisionLine(entry)}`)].join('\n')
201}
202
203/** One measurement as a debug-log line: what the engine reported, `n/a` for what it has no figure for. */
204function describeSignals(measure: SessionMeasureInput): string {
205 const { context, rateLimits, cost } = measure
206 const fill = context.percent === undefined ? `n/a (window ${context.window})` : `${context.percent}% (${context.tokens ?? '?'}/${context.window} tokens)`
207 const limits = rateLimits.length === 0 ? 'n/a' : rateLimits.map((limit) => `${limit.kind} ${limit.percentUsed}%${limit.resetsAt === undefined ? '' : ` resets ${limit.resetsAt}`}`).join(', ')
208 const spent = cost === undefined ? 'n/a' : `$${cost.usd.toFixed(4)}`
209 return `signals: context ${fill}; limits ${limits}; cost ${spent}; changed ${measure.changed.join(', ')}`
210}
211hooks/features/dispatched-agents.ts 177 lines1// Feature: each agent the main agent dispatches gets its own decision on its
2// model and effort, once, when it is spawned (agent.spawn). The model is
3// rewritten on the spawn; the effort goes into the plan table under the
4// agent's id, and the core's turn.step writer sends it on every step of that
5// agent. The main agent's own model and cache are untouched (ADR 0001).
6//
7// The decision reads the person's own words this turn (`said`), so a model
8// they name for the work wins over the decision model, which wins over the
9// main agent's pick. Agents dispatched together (one message, several Agent
10// calls) each get their own request and start as soon as their own answer is
11// in. A failed or late answer lets the agent start as the main agent asked.
12//
13// The main agent reads, after the Agent tool's result, the model and effort its
14// agent started with and whose choice the model was (#34): agent.spawn runs
15// inside the Agent tool call, under its id, so the spawn leaves the note for the
16// call to append.
17//
18// Its switch is `dispatched-agents` (`/dp dispatched-agents off`).
19
20import type { HttpInit, On } from 'claude-code'
21import { describeAsked, type BackendIo, type Failure } from '../decision/backend.ts'
22import { messageText } from '../decision/context.ts'
23import { decideDispatch, dispatchEvidence, dispatchNote, dispatchPart, dispatchReason, dispatchState, modelFamily, termsOf, type Dispatch, type DispatchNote } from '../decision/dispatched-agent.ts'
24import { quoteStart } from '../decision/redact.ts'
25import { renderSummary } from '../decision/summary.ts'
26import { answersFor, mergeParts } from '../decision/system-one.ts'
27import { update, type Cell } from '../core/plans.ts'
28import { isPersonsMessage } from '../core/prompts.ts'
29import { report, type ReportIo } from '../core/report.ts'
30import { dispatchSettings, type Ctx } from '../core/setup.ts'
31import { defineSwitch, isOn } from '../core/switches.ts'
32import { UNRESOLVED_SWITCH } from './unresolved.ts'
33
34const AGENTS = { plugin: 'dispatch-pilot', key: 'agents' } as const
35const SAID = { plugin: 'dispatch-pilot', key: 'said' } as const
36const BOARD = { plugin: 'dispatch-pilot', key: 'board' } as const
37const DECISIONS = { plugin: 'dispatch-pilot', key: 'decisionLog' } as const
38const COUNT = { plugin: 'dispatch-pilot', key: 'unresolved' } as const
39const PPLX_RATE = { plugin: 'dispatch-pilot', key: 'pplxRate' } as const
40
41/** The switch's name, in `/dp` and in the decision log. */
42const SWITCH = 'dispatched-agents'
43/**
44 * The notes spawns left for their Agent tool calls, by the call's tool_use_id: written by the agent.spawn inside the
45 * call, taken as the call returns. A module variable (CLAUDE.md): the kit's `$.state` does not show the outer hook
46 * what the nested spawn wrote (DEVELOPMENT.md 已实测, not yet measured on the engine), and a hot reload while an
47 * agent runs loses only that agent's note. At most MAX_NOTES are kept: past that (calls whose result never came
48 * back, or more agents running at once) the oldest go.
49 */
50const pendingNotes = new Map<string, string>()
51const MAX_NOTES = 32
52/** At most this many of the person's messages are kept for one turn. */
53const MAX_SAID = 8
54
55export function registerDispatchedAgents(on: On, ctx: Ctx): void {
56 defineSwitch({ name: SWITCH, info: '主 agent 派出 agent 时决定它的模型和 effort' })
57 const settings = dispatchSettings(ctx)
58
59 // The person's words this turn: a message sent while idle starts them
60 // afresh, one typed during the turn joins them. Other prompts (an agent's
61 // hand-back, a task notification) leave them as they are. Kept whatever the
62 // switches say (nothing is asked or changed here), so a decision made right
63 // after the person switches the feature back on still reads them.
64 on('prompt.submit', { text: /(?:)/ }, async ($, e, next) => {
65 const result = await next(e)
66 if (!isPersonsMessage(e) || result.drop !== undefined) return result
67 const words = messageText(e.text, ctx.config.context.tokens)
68 const said: Cell<string[]> = { get: () => $.state.get(SAID), set: (value, options) => $.state.set(SAID, value, options) }
69 await update(said, (list) => (e.turnId === undefined ? [words] : [...(list ?? []), words].slice(-MAX_SAID)))
70 return result
71 })
72
73 // The note the spawn left for its Agent tool call goes beside the tool's result; a call whose spawn left none (a fork,
74 // a teammate, a refused spawn) is answered as it is.
75 on('tool.call', { tool: 'Agent' }, async ($, e, next) => {
76 const result = await next(e)
77 const note = pendingNotes.get(e.tool_use_id)
78 pendingNotes.delete(e.tool_use_id)
79 if (note === undefined || result.deny !== undefined) return result
80 return { ...result, context: [...(result.context ?? []), note] }
81 })
82
83 on('agent.spawn', { tool_use_id: /(?:)/ }, async ($, e, next) => {
84 // A fork always runs on its parent's model. A teammate lives across many
85 // tasks that one decision at its spawn cannot see.
86 if (e.fork || e.isTeammate) return next(e)
87 /** Leaves the note for the Agent tool call this spawn belongs to. */
88 const leaveNote = (note: DispatchNote) => {
89 pendingNotes.set(e.tool_use_id, dispatchNote(note))
90 while (pendingNotes.size > MAX_NOTES) pendingNotes.delete(pendingNotes.keys().next().value as string)
91 }
92 if (!isOn(SWITCH)) {
93 const started = await next(e)
94 if (started.deny === undefined) leaveNote({ routed: false, started: started.model, why: 'off' })
95 return started
96 }
97 const { value: said = [] } = await $.state.get(SAID)
98 // The problem the main agent is on goes along as background, with no hint (#41): the summary and the count, while the
99 // unresolved switch is on and there are any.
100 const { value: held } = isOn(UNRESOLVED_SWITCH) ? await $.state.get(COUNT) : { value: undefined }
101 const dispatch: Dispatch = {
102 user_message: said.join('\n'),
103 agent_type: e.subagentType,
104 description: e.description,
105 prompt: e.prompt,
106 requested_model: e.model ?? null,
107 ...(held?.summary === undefined ? {} : { problem_summary: renderSummary(held.summary, ctx.ask.language) }),
108 ...(held === undefined || held.count === 0 ? {} : { unresolved_count: held.count }),
109 }
110 const part = dispatchPart(dispatch, settings)
111 const request = mergeParts(dispatchState(dispatch, ctx.config.contextByKind.agent), [part])
112 const io: BackendIo = {
113 fetch: (url: string, init: HttpInit) => $.http.fetch(url, init),
114 sleep: (ms: number, signal: AbortSignal) => $.clock.sleep(ms, { signal }),
115 pace: { now: () => $.clock.now(), sentAt: { get: () => $.state.get(PPLX_RATE), set: (value, options) => $.state.set(PPLX_RATE, value, options) } },
116 }
117 const reporting: ReportIo = {
118 board: { get: () => $.state.get(BOARD), set: (value, options) => $.state.set(BOARD, value, options) },
119 decisions: { get: () => $.state.get(DECISIONS), set: (value, options) => $.state.set(DECISIONS, value, options) },
120 debug: (line) => $.ui.log(line, { to: 'debug' }),
121 now: () => $.clock.now(),
122 toast: (text) => $.ui.toast(text),
123 }
124 const about = `${quoteStart(e.description)}(${e.subagentType})`
125 /** What every report of this agent says of it; its id is known once it has started (a spawn is not yet an agent). */
126 const reportOf = (agent: string) => ({ feature: SWITCH, agent, subject: about, node: { kind: 'agent' as const, name: e.name ?? (e.description === '' ? e.subagentType : e.description), type: e.subagentType } })
127 const startedAt = await $.clock.now()
128 const asked = await ctx.backend.ask(io, request, ctx.config.timeoutMs)
129 const ms = (await $.clock.now()) - startedAt
130 $.ui.log(`request [${Object.keys(request.questions).join(', ')}] to ${ctx.backend.name} for agent ${about}: ${describeAsked(asked, ms)}`, { to: 'debug' })
131 // The agent starts as the main agent asked; the board says why it was not routed, once it has an id to say it of.
132 // A spawn refused beneath started no agent: there is nothing to say of it.
133 const notRouted = async (failure: Failure) => {
134 const started = await next(e)
135 if (started.deny !== undefined) return started
136 await report(reporting, { decision: { ...reportOf(started.agentId ?? e.tool_use_id), routed: false, failure: { backend: ctx.backend.name, ...failure } } })
137 leaveNote({ routed: false, started: started.model, why: { failure, backend: ctx.backend.name } })
138 return started
139 }
140 if (!asked.ok) return notRouted(asked.failure)
141 const decision = decideDispatch(answersFor(part, asked.answers), dispatch, settings)
142 if (!decision.answered) return notRouted({ kind: 'parse', detail: 'no answer about the agent' })
143 // The main agent's pick, when it stands, goes on as the main agent wrote it.
144 const spawn = decision.model !== null && decision.model !== modelFamily(e.model) ? { ...e, model: decision.model } : e
145 const result = await next(spawn)
146 if (result.deny !== undefined) return result
147 leaveNote({ routed: true, decision, started: result.model, requested: modelFamily(e.model), thetaOverride: settings.thetaOverride })
148 // The agent has started: nothing after this may fail its spawn.
149 try {
150 // The effort and the person's terms go into the plan (what changes the agent later keeps to the
151 // terms); the model is set on the spawn, and a planned model would also pin every step against the
152 // engine's overload fallback.
153 const terms = termsOf(decision)
154 if (result.agentId !== undefined && (decision.effort !== null || terms !== null)) {
155 await $.state.set({ ...AGENTS, id: result.agentId }, { effort: decision.effort, floor: null, model: null, terms })
156 }
157 const family = decision.model ?? modelFamily(result.model)
158 const model = family ?? result.model
159 const outcome = decision.effort === null ? model : `${model} ${decision.effort}`
160 await report(reporting, {
161 decision: {
162 ...reportOf(result.agentId ?? e.tool_use_id),
163 routed: true,
164 outcome,
165 reason: dispatchReason(decision, modelFamily(e.model), settings.thetaOverride, '主 agent 指定的'),
166 ...dispatchEvidence(decision),
167 ...(family === null ? {} : { model: family }),
168 ...(decision.effort === null ? {} : { effort: decision.effort }),
169 },
170 })
171 } catch (error) {
172 $.ui.log(`agent ${about} started (${result.agentId ?? 'no id'}), but its plan was not recorded: ${String(error)}`, { to: 'debug' })
173 }
174 return result
175 })
176}
177hooks/features/escalation.ts 663 lines1// Feature: a loop whose tool calls keep failing goes up a level (#7).
2//
3// Every tool call of the main agent and of the agents it dispatches is told
4// apart as it ends (core/outcomes.ts), and each loop's failures are counted
5// here, in one record (`escalation`) that the mid-turn re-decision (#5) sends
6// in its counts too: a call the person refused is no failure, one a hook
7// blocked counts toward escalating only with `hook-block-failures` on, and
8// the counts start over whenever this feature deals with them (a forced
9// raise, failures found expected, or nothing left to raise).
10//
11// When a loop's counted failures reach `escalateAfter`, the decision model is
12// asked at once, as the call that reaches them ends (story 18): the mid-turn
13// effort question with the trouble, and whether the failures are what the
14// work expects (a test written to fail first, a search that finds nothing).
15// The loop's next step takes the answer, waiting `rejudgeWaitMs` at most, as
16// a mid-turn re-decision does; an answer later than that is taken at a later
17// step. Expected: the counts start over and the answer is an ordinary
18// re-decision. Otherwise (also when no answer comes) the loop goes up: one
19// level (at most to xhigh) or straight to max, at most `escalateLimit` times
20// per turn (per run, for an agent); a haiku agent, which takes no effort,
21// goes on as another model. The person's terms for an agent's work hold
22// (core/plans.ts `AgentPlan`).
23//
24// It sits above the mid-turn feature (the entry registers it first), so a
25// step's own mid-turn answer finds the raise already in the turn's plan and
26// cannot undercut it.
27
28import type { EngineInterface, HttpInit, On, TurnStepInput } from 'claude-code'
29import { describeAsked, errorText, failureLine, within, type BackendIo, type Failure } from '../decision/backend.ts'
30import { AGENT_MODELS, effortFloor, modelFamily, type AgentModel, type Terms } from '../decision/dispatched-agent.ts'
31import { briefOf, forcedTarget, readExpected, rowsFromTranscript, stepsFromRows, stuckRequest, traceRaise, troubleText, type RaiseMode, type TranscriptRow } from '../decision/escalation.ts'
32import { higherEffort, isEffort, probsOf, readEffort, readingText, type Effort, type EffortReading } from '../decision/effort.ts'
33import { contentLanguage, judgeMidturn, MIDTURN_LEVEL, midturnRecord, outcomeOf, verdictReason, type MidturnInput, type MidturnLimits, type MidturnPosition, type MidturnRules, type MidturnVerdict } from '../decision/midturn.ts'
34import { modelId, type ResolvedModel } from '../decision/model-ids.ts'
35import { quoteStart } from '../decision/redact.ts'
36import { answersFor } from '../decision/system-one.ts'
37import { startedIn } from '../decision/workflow-labels.ts'
38import { endedAs, noteEnded, wasBlocked } from '../core/outcomes.ts'
39import { floorHeld, forced, MAIN, newTurn, redecided, replace, turnKey, update, type AgentPlan, type Cell, type TurnRecord } from '../core/plans.ts'
40import { report, type Decided, type ReportIo } from '../core/report.ts'
41import type { Ctx } from '../core/setup.ts'
42import { defineSwitch, isOn, masterOn } from '../core/switches.ts'
43
44const ESCALATION = { plugin: 'dispatch-pilot', key: 'escalation' } as const
45const TURNS = { plugin: 'dispatch-pilot', key: 'turns' } as const
46const AGENTS = { plugin: 'dispatch-pilot', key: 'agents' } as const
47const LOCK = { plugin: 'dispatch-pilot', key: 'lock' } as const
48const RUNS = { plugin: 'dispatch-pilot', key: 'workflowRuns' } as const
49const BOARD = { plugin: 'dispatch-pilot', key: 'board' } as const
50const DECISIONS = { plugin: 'dispatch-pilot', key: 'decisionLog' } as const
51const PPLX_RATE = { plugin: 'dispatch-pilot', key: 'pplxRate' } as const
52
53/** At most this many Workflow runs' directories are kept. */
54const MAX_RUNS = 8
55
56/** The feature's switch (`/dp escalation on|off`). */
57const SWITCH = 'escalation'
58/** Whether a call one of the person's hooks blocked counts as a failure (`/dp hook-block-failures on|off`); off by default. */
59const BLOCKS_SWITCH = 'hook-block-failures'
60
61/** One loop's failed calls and forced raises (`escalation` in types/index.d.ts): the main agent's for one turn, an agent's for its run. */
62type LoopRecord = {
63 /** The main turn the counts belong to; '' for an agent. */
64 turnId: string
65 /** Calls that failed, and calls a hook refused, since the loop began. */
66 failures: number
67 hookBlocks: number
68 /** The counts when they last started over: what is counted is the rest. */
69 base: Counted
70 /** Forced raises so far (failures found expected are not one). */
71 raises: number
72 /** The loop's latest step, the engine's effort and model on it: what a re-decision asked as a call ends reads. */
73 step: number | null
74 engine: Effort | null
75 model: string | null
76 /** The step a stuck re-decision was last asked for; null before the first. */
77 askedFor: number | null
78 /** An agent's step at its latest raise (the main agent's is its turn's `raisedAt`); null before one. */
79 raisedAt: number | null
80 /** The escalation switch was off when the loop was last seen: what was counted meanwhile is written off once it is back on. */
81 paused: boolean
82}
83
84type Counted = { failures: number; hookBlocks: number }
85
86/** What the feature runs with: the shared ctx and the options it reads. */
87type Settings = {
88 ctx: Ctx
89 /** Counted failures that make a loop stuck. */
90 after: number
91 mode: RaiseMode
92 /** Forced raises a loop gets at most. */
93 limit: number
94 /** The failures count as expected when the answer reaches this probability. */
95 thetaExpected: number
96 /** The model a failing haiku agent is switched to, as a step names it (the engine takes no alias for a step); null for none. */
97 haikuTo: ResolvedModel | null
98 /** `escalateHaikuTo` as written, for the log when it names no model. */
99 haikuToWritten: string
100 /** The models agents may run on: a ruled-out escalateHaikuTo gives way to the next of them up. */
101 models: readonly AgentModel[]
102 /** The mid-turn rules: what a raise the decision model itself suggests needs, and how an ordinary re-decision moves the level. */
103 rules: MidturnRules
104 limits: MidturnLimits
105 /** How long a step waits for a stuck re-decision not back yet. */
106 waitMs: number
107}
108
109/** What the decision model made of a stuck loop. */
110type Stuck = {
111 /** The effort answer; null when it was not asked or not answered. */
112 reading: EffortReading | null
113 /** The probability that the failures were expected; null without an answer. */
114 expected: number | null
115 /** Why the question went unanswered, when it did. */
116 failure: Failure | null
117 /** No transcript of the loop could be read: nothing was asked. */
118 unread: boolean
119}
120
121/**
122 * A stuck re-decision on its way, or answered and not yet taken, by loop id:
123 * asked for step `forStep` (of turn `turnId`, for the main agent), about the
124 * counts in `covered` (what starts over once it is dealt with). `answer`
125 * never rejects; once it resolves, `settled` holds it and `ms` how long it took.
126 * Promises cannot live in $.state: a reload drops them, and the loop's next
127 * step asks again.
128 */
129type Asking = { forStep: number; turnId: string; covered: Counted; about: string; answer: Promise<Stuck>; settled: Stuck | null; ms: number }
130const asking = new Map<string, Asking>()
131
132export function registerEscalation(on: On, ctx: Ctx): void {
133 defineSwitch({ name: SWITCH, info: '工具调用接连失败的 agent,强制升高它的 effort', parts: ['counts'] })
134 defineSwitch({ name: BLOCKS_SWITCH, info: '判断要不要强制升档时,把被你的 hook 拦下的调用也算失败', default: false })
135 // Read at each use: which decision model's settings these are is settled at the session start (core/setup.ts, `Ctx`).
136 const settings: Settings = {
137 ctx,
138 get after() {
139 return ctx.config.escalation.after
140 },
141 get mode() {
142 return ctx.config.escalation.mode
143 },
144 get limit() {
145 return ctx.config.escalation.limit
146 },
147 get thetaExpected() {
148 return ctx.config.escalation.thetaExpected
149 },
150 get haikuTo() {
151 return ctx.config.escalation.haikuTo
152 },
153 get haikuToWritten() {
154 return ctx.config.escalation.haikuToWritten
155 },
156 get models() {
157 return ctx.config.agents.models
158 },
159 get rules() {
160 return ctx.config.midturn.rules
161 },
162 get limits() {
163 return ctx.config.midturn.limits
164 },
165 get waitMs() {
166 return ctx.config.midturn.waitMs
167 },
168 }
169
170 on('tool.call', { tool: /(?:)/ }, async ($, e, next) => {
171 const result = await next(e)
172 // The calls the model made, not another plugin's `$.tool.call`; counted while Dispatch Pilot is on, whatever this
173 // feature's switch says (the mid-turn re-decision sends the counts too).
174 if (next.origin.plugin !== 'engine' || !masterOn()) return result
175 try {
176 // A Workflow's run: its directory holds each of its agents' transcripts, which the engine keeps from the mod.
177 const run = e.tool === 'Workflow' ? launchedRun(result.result) : null
178 if (run !== null) {
179 const runs: Cell<{ runId: string; dir: string }[]> = { get: () => $.state.get(RUNS), set: (value, options) => $.state.set(RUNS, value, options) }
180 await update(runs, (list) => [...(list ?? []).filter((kept) => kept.runId !== run.runId), run].slice(-MAX_RUNS))
181 }
182 const ended = outcomeOf(e.tool, result, wasBlocked(e.tool_use_id))
183 noteEnded(e.tool_use_id, ended)
184 // A refusal by a plugin's tool.call hook (`{ deny }`) is nobody's failure and no settings hook's block: the
185 // Workflow feature's hand-back of a script is one. Only what the tool reported, and what the settings hooks refused, count.
186 if (typeof result.deny === 'string' || (ended !== 'failed' && ended !== 'blocked')) return result
187 const id = e.agentId ?? MAIN
188 const ref = { ...ESCALATION, id }
189 const cell: Cell<LoopRecord> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
190 const record = await update(cell, (r) => {
191 const loop = r ?? newLoop('')
192 return ended === 'failed' ? { ...loop, failures: loop.failures + 1 } : { ...loop, hookBlocks: loop.hookBlocks + 1 }
193 })
194 if (!isOn(SWITCH)) return result
195 await showCounts($, id, record, null)
196 // The call that reaches the threshold asks at once: the answer is for the loop's next step (once per step).
197 if (record.step !== null && record.askedFor !== record.step + 1) await launch($, settings, id, e.agentId, record, record.step + 1)
198 } catch (error) {
199 $.ui.log(`escalation: ${errorText(error)}`, { to: 'debug' })
200 }
201 return result
202 })
203
204 on('turn.step', { turnId: /(?:)/ }, async function* ($, e, next) {
205 if (masterOn()) {
206 try {
207 await atStep($, settings, e)
208 } catch (error) {
209 $.ui.log(`escalation: ${errorText(error)}`, { to: 'debug' })
210 }
211 }
212 return yield* next(e)
213 })
214}
215
216/**
217 * At a loop's step: keeps where the loop is (a new main turn starts its counts
218 * afresh), then takes the stuck re-decision meant for this step and applies
219 * it. With none on its way while the counted failures call for one (nothing
220 * to raise when the call ended, or a reload dropped the request), deals with
221 * them here: written off when nothing can be raised, else asked now.
222 */
223async function atStep($: EngineInterface, s: Settings, e: TurnStepInput): Promise<void> {
224 const main = e.agentId === undefined
225 const id = e.agentId ?? MAIN
226 const ref = { ...ESCALATION, id }
227 const cell: Cell<LoopRecord> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
228 const engine = isEffort(e.effort) ? e.effort : null
229 const on = isOn(SWITCH)
230 /** Whether a record belongs to an earlier main turn (or there is none): this step starts the turn's counts afresh. */
231 const startsTurn = (r: LoopRecord | undefined) => main && (r === undefined || r.turnId !== e.turnId)
232 const { before, after } = await replace(cell, (r) => {
233 const loop = startsTurn(r) || r === undefined ? newLoop(main ? e.turnId : '') : r
234 // Back on after being off: what was counted meanwhile is written off, so it is counted from here on.
235 const written = on && loop.paused ? { ...loop, base: { failures: loop.failures, hookBlocks: loop.hookBlocks }, paused: false } : loop
236 return { ...written, step: e.index, engine, model: e.model, paused: !on }
237 })
238 let record = after
239 if (startsTurn(before)) {
240 // A new main turn: what the one before counted, and asked, is over.
241 asking.delete(MAIN)
242 await showCounts($, MAIN, null, null)
243 }
244 if (!on) return
245
246 let pending = asking.get(id)
247 if (pending === undefined || (main && pending.turnId !== e.turnId) || pending.forStep > e.index) {
248 if (pending !== undefined) return
249 // No re-decision on its way, though the counted failures may call for one: nothing could be raised when
250 // the call ended, or a reload dropped the request. Dealt with here: written off, or asked now.
251 const counted = countedOf(record)
252 if (counted.failures + counted.hookBlocks < s.after || record.raises >= s.limit || record.askedFor === e.index || s.ctx.backend.configured === false) return
253 record = await update(cell, (r) => ({ ...(r ?? record), askedFor: e.index }))
254 const raise = await raiseOf($, s, id, e.agentId, record, e.index)
255 if (raise.kind === 'none') return
256 if (raise.kind === 'keep') {
257 const kept = await startOver($, cell, id, countedNow(record), false)
258 const brief = e.agentId === undefined ? '' : briefOf((await agentRows($, e.agentId)) ?? [])
259 await decide($, id, kept, { outcome: raise.outcome, about: aboutOf(e.agentId, brief, e.index, counted), reason: raise.reason, tone: 'info' })
260 return
261 }
262 await launch($, s, id, e.agentId, record, e.index)
263 pending = asking.get(id)
264 if (pending === undefined) return
265 }
266 const answer = pending.settled ?? (await within((ms, signal) => $.clock.sleep(ms, { signal }), pending.answer, s.waitMs, null))
267 if (answer === null) {
268 // Not back yet: the step goes as it is, and the answer is taken at a later step (as a mid-turn re-decision's).
269 await showCounts($, id, record, 'late')
270 return
271 }
272 asking.delete(id)
273 await apply($, s, e, cell, pending.covered, pending.about, answer)
274}
275
276/** What a stuck loop's re-decision would do, from where the loop stands. */
277type Raise =
278 /** Leave the loop be: the person's lock holds the main agent's effort, or the step takes no effort level. */
279 | { kind: 'none' }
280 /** Nothing to raise: the failures are written off, and the decision recorded. */
281 | { kind: 'keep'; outcome: string; reason: string }
282 /** A higher effort, from `current`. */
283 | { kind: 'level'; current: Effort; target: Effort }
284 /** Another model (a haiku agent: it takes no effort). */
285 | { kind: 'model'; from: string; to: ResolvedModel; note: string }
286
287/**
288 * What raising the loop would be at step `at`, by its plan: the main agent's
289 * turn (unless the person locked its effort), or an agent's effective model
290 * (the plan's, else the engine's) and the person's terms for its work.
291 */
292async function raiseOf($: EngineInterface, s: Settings, id: string, agentId: string | undefined, record: LoopRecord, at: number): Promise<Raise> {
293 const top = (current: Effort) => ({ kind: 'keep' as const, outcome: `effort ${current}(保持)`, reason: s.mode === 'max' ? '已经是 max' : '升一档最高到 xhigh' })
294 if (agentId === undefined) {
295 const { value: lock = null } = await $.state.get(LOCK)
296 if (lock !== null || record.engine === null) return { kind: 'none' }
297 const { value: turn } = await $.state.get({ ...TURNS, id: turnKey(record.turnId, undefined) })
298 const current = higherEffort(turn?.effort ?? record.engine, floorHeld(turn, at)) as Effort
299 const target = forcedTarget(current, s.mode)
300 return target === null ? top(current) : { kind: 'level', current, target }
301 }
302 const { value: planned } = await $.state.get({ ...AGENTS, id })
303 const model = planned?.model ?? record.model ?? ''
304 const family = modelFamily(model)
305 // A model the mod does not know, and that takes no effort level: nothing to raise.
306 if (family === null && record.engine === null) return { kind: 'none' }
307 if (family === 'haiku') {
308 const to = haikuSwitch(s, planned?.terms ?? null)
309 return 'why' in to ? { kind: 'keep', outcome: `model ${model}(保持)`, reason: to.why } : { kind: 'model', from: model, to: to.to, note: to.note }
310 }
311 const named = planned?.terms?.effort ?? null
312 if (named !== null) return { kind: 'keep', outcome: `effort ${named}(保持)`, reason: `${named} 是你给它点名的 effort` }
313 // Without a level of its own (moved off a model that takes none), its steps go at the engine's own for an agent: medium (measured on 2.1.289).
314 const current = higherEffort(planned?.effort ?? record.engine ?? 'medium', planned?.floor ?? null) as Effort
315 const target = forcedTarget(current, s.mode)
316 return target === null ? top(current) : { kind: 'level', current, target }
317}
318
319/**
320 * Sends the stuck re-decision for the loop's step `forStep`, when its counted
321 * failures call for one and there is something to raise; the answer waits in
322 * `asking` for that step. Once per step, one at a time per loop.
323 */
324async function launch($: EngineInterface, s: Settings, id: string, agentId: string | undefined, record: LoopRecord, forStep: number): Promise<void> {
325 const counted = countedOf(record)
326 if (counted.failures + counted.hookBlocks < s.after || record.raises >= s.limit || s.ctx.backend.configured === false || asking.has(id)) return
327 const raise = await raiseOf($, s, id, agentId, record, forStep)
328 if (raise.kind !== 'level' && raise.kind !== 'model') return
329 const ref = { ...ESCALATION, id }
330 const cell: Cell<LoopRecord> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
331 await update(cell, (r) => ({ ...(r ?? record), askedFor: forStep }))
332 // What the counts say since they last started over, and what they start over from once this is dealt with.
333 const since = { failures: record.failures - record.base.failures, hookBlocks: record.hookBlocks - record.base.hookBlocks }
334 const entry: Asking = { forStep, turnId: record.turnId, covered: countedNow(record), about: '', answer: Promise.resolve(UNREAD), settled: null, ms: 0 }
335
336 let input: MidturnInput
337 if (agentId === undefined) {
338 const { value: turn } = await $.state.get({ ...TURNS, id: turnKey(record.turnId, undefined) })
339 const plan = turn ?? newTurn('', null, false)
340 const rows = (await $.session.messages().catch(() => [])) as TranscriptRow[]
341 entry.about = aboutOf(undefined, '', forStep, counted)
342 input = {
343 message: plan.prompt,
344 step: forStep,
345 current_effort: raise.kind === 'level' ? raise.current : 'medium',
346 counts: { judgments: plan.decisions, changes: plan.changes, failures: since.failures, hook_blocks: since.hookBlocks },
347 recent_steps: stepsFromRows(rows, { language: contentLanguage(plan.prompt), ended: endedAs }),
348 trouble: troubleText(counted),
349 }
350 } else {
351 const rows = await agentRows($, agentId)
352 const brief = rows === null ? '' : briefOf(rows)
353 entry.about = aboutOf(agentId, brief, forStep, counted)
354 if (rows === null) {
355 // No transcript to read (a workflow agent whose run the mod did not record): raised without asking.
356 entry.settled = UNREAD
357 asking.set(id, entry)
358 return
359 }
360 input = {
361 message: brief,
362 step: forStep,
363 current_effort: raise.kind === 'level' ? raise.current : 'medium',
364 counts: { judgments: 1 + record.raises, changes: record.raises, failures: since.failures, hook_blocks: since.hookBlocks },
365 recent_steps: stepsFromRows(rows, { language: contentLanguage(brief), ended: endedAs }),
366 trouble: troubleText(counted),
367 }
368 }
369 // A haiku agent takes no effort: only whether its failures were expected is asked.
370 const { request, effortPart, expectedPart } = stuckRequest(input, { limits: s.limits, ask: s.ctx.ask, effort: raise.kind === 'level' })
371 const io: BackendIo = {
372 fetch: (url: string, init: HttpInit) => $.http.fetch(url, init),
373 sleep: (ms: number, signal: AbortSignal) => $.clock.sleep(ms, { signal }),
374 pace: { now: () => $.clock.now(), sentAt: { get: () => $.state.get(PPLX_RATE), set: (value, options) => $.state.set(PPLX_RATE, value, options) } },
375 }
376 const startedAt = await $.clock.now()
377 entry.answer = s.ctx.backend.ask(io, request, s.ctx.config.timeoutMs).then(async (asked): Promise<Stuck> => {
378 entry.ms = (await $.clock.now()) - startedAt
379 $.ui.log(`request [${Object.keys(request.questions).join(', ')}] for ${entry.about} to ${s.ctx.backend.name}: ${describeAsked(asked, entry.ms)}`, { to: 'debug' })
380 let stuck: Stuck = { reading: null, expected: null, failure: null, unread: false }
381 if (!asked.ok) stuck = { ...stuck, failure: asked.failure }
382 else {
383 const expected = readExpected(answersFor(expectedPart, asked.answers))
384 stuck = {
385 reading: effortPart === null ? null : readEffort(answersFor(effortPart, asked.answers)[MIDTURN_LEVEL]),
386 expected,
387 failure: expected === null ? { kind: 'parse', detail: 'no answer to the question' } : null,
388 unread: false,
389 }
390 }
391 entry.settled = stuck
392 return stuck
393 })
394 asking.set(id, entry)
395}
396
397/** The answer when nothing could be asked: no transcript of the agent to read. */
398const UNREAD: Stuck = { reading: null, expected: null, failure: null, unread: true }
399
400/**
401 * Applies a stuck re-decision's answer at step `e`, to the loop as it stands
402 * now (a re-decision may have moved it since the question went out): expected
403 * failures leave an ordinary re-decision; otherwise the loop goes up (or, when
404 * it can no longer, its failures are written off). Either way the counts the
405 * question covered start over.
406 */
407async function apply($: EngineInterface, s: Settings, e: TurnStepInput, cell: Cell<LoopRecord>, covered: Counted, about: string, answer: Stuck): Promise<void> {
408 const id = e.agentId ?? MAIN
409 const { value: record } = await cell.get()
410 if (record === undefined) return
411 const raise = await raiseOf($, s, id, e.agentId, record, e.index)
412 if (raise.kind === 'none') return
413 if (raise.kind === 'keep') {
414 const kept = await startOver($, cell, id, covered, false)
415 await decide($, id, kept, { outcome: raise.outcome, about, reason: raise.reason, tone: 'info' })
416 return
417 }
418 const p = answer.expected
419 if (p !== null && p >= s.thetaExpected) {
420 const why = `这些失败是预期内的(概率 ${p.toFixed(2)},预期内失败门槛 ${s.thetaExpected.toFixed(2)}),不强制升档`
421 if (raise.kind === 'model') {
422 const kept = await startOver($, cell, id, covered, false)
423 await decide($, id, kept, { outcome: `model ${raise.from}(保持)`, about, reason: why, tone: 'info' })
424 return
425 }
426 // Expected: nothing is forced; the answer's effort is an ordinary re-decision, as mid-turn.
427 const level = await redecide($, s, e, record, raise.current, answer.reading)
428 const kept = await startOver($, cell, id, covered, false)
429 await decide($, id, kept, {
430 outcome: `effort ${level.effort}${level.effort === raise.current ? '(保持)' : `(原 ${raise.current})`}`,
431 about,
432 reason: level.why === null ? why : `${why};${level.why}`,
433 tone: level.effort === raise.current ? 'info' : 'ok',
434 ...(level.working === undefined
435 ? {}
436 : {
437 probs: probsOf(level.working.reading),
438 conf: level.working.verdict.confidence,
439 trace: [...level.working.verdict.trace.pick.steps, ...level.working.verdict.trace.steps],
440 mid: midturnRecord(level.working.verdict, level.working.position, s.rules),
441 }),
442 })
443 return
444 }
445
446 if (raise.kind === 'model') {
447 const ref = { ...AGENTS, id }
448 const planCell: Cell<AgentPlan> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
449 await update(planCell, (r) => ({ ...(r ?? { effort: null, floor: null, model: null, terms: null }), model: raise.to.id }))
450 const raised = await startOver($, cell, id, covered, true, e.index)
451 await decide($, id, raised, {
452 outcome: `model ${raise.to.id}(原 ${raise.from})`,
453 about,
454 reason: `haiku 没有 effort 可升,改用 ${raise.to.id}${raise.note};${knownReason(s, answer)}`,
455 tone: 'warn',
456 forced: { kind: 'model', from: raise.from, to: raise.to.id },
457 })
458 return
459 }
460 const { level, steps } = traceRaise(answer.reading, { from: raise.current, target: raise.target, mode: s.mode }, s.rules)
461 if (e.agentId === undefined) {
462 // Raised from this step on; for holdSteps steps nothing lowers it below the forced level, then ordinary re-decisions take over.
463 const ref = { ...TURNS, id: turnKey(e.turnId, undefined) }
464 const turnCell: Cell<TurnRecord> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
465 await update(turnCell, (r) => forced(r ?? newTurn('', null, false), level, raise.target, e.index, s.rules.holdSteps))
466 } else {
467 // The raise holds for the rest of the agent's run: nothing re-decides an agent mid-run, so it goes into its effort.
468 const ref = { ...AGENTS, id }
469 const planCell: Cell<AgentPlan> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
470 await update(planCell, (r) => ({ ...(r ?? { effort: null, floor: null, model: null, terms: null }), effort: level }))
471 }
472 const raised = await startOver($, cell, id, covered, true, e.index)
473 await decide($, id, raised, {
474 outcome: `effort ${level}(原 ${raise.current})`,
475 about,
476 reason: raiseReason(s, answer, level, raise.target),
477 tone: 'warn',
478 forced: { kind: 'effort', from: raise.current, to: level, floor: raise.target },
479 trace: steps,
480 ...(answer.reading === null ? {} : { probs: probsOf(answer.reading), ...(answer.reading.confidence === null ? {} : { conf: answer.reading.confidence }) }),
481 })
482}
483
484/**
485 * An ordinary re-decision from a stuck loop's effort answer, its failures
486 * found expected: the main agent's turn by the mid-turn rules (a turn the
487 * person started, as mid-turn), an agent's plan the same way. The level it
488 * goes on at, and why (null when there was nothing to decide from).
489 */
490async function redecide(
491 $: EngineInterface,
492 s: Settings,
493 e: TurnStepInput,
494 record: LoopRecord,
495 current: Effort,
496 reading: EffortReading | null,
497): Promise<{ effort: Effort; why: string | null; working?: { reading: EffortReading; verdict: MidturnVerdict; position: MidturnPosition } }> {
498 if (reading === null) return { effort: current, why: null }
499 if (e.agentId === undefined) {
500 const ref = { ...TURNS, id: turnKey(e.turnId, undefined) }
501 const turnCell: Cell<TurnRecord> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
502 const { value: turn } = await turnCell.get()
503 if (turn?.person !== true) return { effort: current, why: null }
504 const position = { current, sinceRaise: turn.raisedAt == null ? null : e.index - turn.raisedAt, atLeast: floorHeld(turn, e.index) }
505 const verdict = judgeMidturn(reading, position, s.rules)
506 await update(turnCell, (r) => redecided(r ?? turn, current, verdict.effort, e.index))
507 return { effort: verdict.effort, why: `${readingText(reading)};${verdictReason(verdict, position, s.rules)}`, working: { reading, verdict, position } }
508 }
509 const ref = { ...AGENTS, id: e.agentId }
510 const planCell: Cell<AgentPlan> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
511 const { value: plan } = await planCell.get()
512 // A routed agent's re-decision goes no lower than its model's floor either (effortFloor): its effort was decided with it.
513 const floor = higherEffort(plan?.floor ?? null, plan?.effort == null ? null : effortFloor(modelFamily(plan.model ?? record.model ?? '')))
514 const position = { current, sinceRaise: record.raisedAt === null ? null : e.index - record.raisedAt, atLeast: floor }
515 const verdict = judgeMidturn(reading, position, s.rules)
516 if (verdict.effort !== current) await update(planCell, (r) => ({ ...(r ?? { effort: null, floor: null, model: null, terms: null }), effort: verdict.effort }))
517 return { effort: verdict.effort, why: `${readingText(reading)};${verdictReason(verdict, position, s.rules)}`, working: { reading, verdict, position } }
518}
519
520/**
521 * The model a failing haiku agent goes on as: `escalateHaikuTo`, unless the
522 * person ruled it out for the agent's work, then the next model up that agents
523 * may run on and the person did not rule out. None (with why) when the person
524 * named haiku, when escalateHaikuTo names no model, or when every model up is
525 * ruled out.
526 */
527function haikuSwitch(s: Settings, terms: Terms | null): { to: ResolvedModel; note: string } | { why: string } {
528 if (terms?.model === 'haiku') return { why: 'haiku 是你给它点名的模型' }
529 const to = s.haikuTo
530 if (to === null) return { why: s.haikuToWritten === '' ? '没有设置失败的 haiku 改用哪个模型' : `设置的改用模型 ${JSON.stringify(s.haikuToWritten)} 这个 mod 不认识` }
531 const banned = terms?.banned ?? []
532 if (!banned.includes(to.family)) return { to, note: '' }
533 const up = AGENT_MODELS.slice(AGENT_MODELS.indexOf(to.family) + 1).find((family) => s.models.includes(family) && !banned.includes(family))
534 if (up !== undefined) return { to: { family: up, id: modelId(up) }, note: `(${to.family} 被你排除了)` }
535 const above = AGENT_MODELS.filter((family) => family !== 'haiku' && banned.includes(family))
536 return { why: `haiku 以上的模型都被你排除了(${above.join('、')})` }
537}
538
539/** The run a Workflow tool's result says it launched: its id and its directory; null for none (refused, failed). */
540function launchedRun(result: unknown): { runId: string; dir: string } | null {
541 if (typeof result !== 'object' || result === null) return null
542 const { runId, transcriptDir } = result as { runId?: unknown; transcriptDir?: unknown }
543 return typeof runId === 'string' && typeof transcriptDir === 'string' ? { runId, dir: transcriptDir } : null
544}
545
546/**
547 * An agent's transcript: as the session gives it, or for a workflow's agent
548 * (the engine keeps it from the mod) from the run's directory on disk, found
549 * among the runs this feature noted; null when neither can be read.
550 */
551async function agentRows($: EngineInterface, agentId: string): Promise<TranscriptRow[] | null> {
552 const found: unknown = await $.session.messages({ agentId }).catch(() => null)
553 if (Array.isArray(found)) return found as TranscriptRow[]
554 const { value: runs = [] } = await $.state.get(RUNS)
555 for (const run of [...runs].reverse()) {
556 const journal = await $.fs.read(`${run.dir}/journal.jsonl`).catch(() => null)
557 if (journal === null || startedIn(journal, agentId) === null) continue
558 const text = await $.fs.read(`${run.dir}/agent-${agentId}.jsonl`).catch(() => null)
559 return text === null ? null : rowsFromTranscript(text)
560 }
561 return null
562}
563
564/** What a decision about the loop is about: its step and the failures, and for an agent its task (`brief`, '' when unknown). */
565function aboutOf(agentId: string | undefined, brief: string, at: number, counted: Counted): string {
566 const failed = `工具调用失败 ${counted.failures + counted.hookBlocks} 次`
567 if (agentId === undefined) return `第 ${at} 步(${failed})`
568 return `${brief === '' ? `agent ${agentId}` : `agent ${quoteStart(brief)}`},第 ${at} 步(${failed})`
569}
570
571/** What the decision report needs of the host: its two cells, the debug log, the clock and the toast. */
572function ioOf($: EngineInterface): ReportIo {
573 return {
574 board: { get: () => $.state.get(BOARD), set: (value, options) => $.state.set(BOARD, value, options) },
575 decisions: { get: () => $.state.get(DECISIONS), set: (value, options) => $.state.set(DECISIONS, value, options) },
576 debug: (line) => $.ui.log(line, { to: 'debug' }),
577 now: () => $.clock.now(),
578 toast: (text) => $.ui.toast(text),
579 }
580}
581
582/** What a decision of this feature says, besides who it is about and the loop's counts (`decide` adds them). */
583type Said = Pick<Decided, 'outcome' | 'reason' | 'tone' | 'probs' | 'conf' | 'trace' | 'floor' | 'mid' | 'forced'> & { about: string }
584
585/**
586 * Reports a decision of this feature (debug log, `/dp log`, the board data), about loop `id`
587 * (`main`, or an agent's id) and with its counts as they stand. Beside the loop's own route: not on its node.
588 */
589async function decide($: EngineInterface, id: string, record: LoopRecord, said: Said): Promise<void> {
590 const { about, ...rest } = said
591 await report(ioOf($), {
592 decision: {
593 ...rest,
594 feature: SWITCH,
595 agent: id,
596 aside: true,
597 subject: about,
598 counts: { failed: record.failures, blocked: record.hookBlocks, raised: record.raises },
599 },
600 })
601}
602
603/** Why a forced raise went where it did, for the decision log. */
604function raiseReason(s: Settings, answer: Stuck, level: Effort, target: Effort): string {
605 const how = s.mode === 'max' ? '强制升到 max' : '强制升一档'
606 const higher = level === target ? '' : ',回答自己的判断更高'
607 return `${how}${higher};${knownReason(s, answer)}${answer.reading === null ? '' : `;${readingText(answer.reading)}`}`
608}
609
610/** Whether the failures were known to be expected, for the decision log: why they were not. */
611function knownReason(s: Settings, answer: Stuck): string {
612 if (answer.unread) return '读不到这个 agent 的记录,没有问这些失败是不是预期内的'
613 if (answer.expected !== null) return `不是预期内的失败(概率 ${answer.expected.toFixed(2)},预期内失败门槛 ${s.thetaExpected.toFixed(2)})`
614 return `决策模型没有回答(${failureLine(s.ctx.backend.name, answer.failure ?? { kind: 'parse', detail: 'no answer to the question' })})`
615}
616
617function newLoop(turnId: string): LoopRecord {
618 return { turnId, failures: 0, hookBlocks: 0, base: { failures: 0, hookBlocks: 0 }, raises: 0, step: null, engine: null, model: null, askedFor: null, raisedAt: null, paused: false }
619}
620
621/** The failures counted toward escalating: since the counts last started over, hook blocks only when they count. */
622function countedOf(record: LoopRecord): Counted {
623 return {
624 failures: record.failures - record.base.failures,
625 hookBlocks: isOn(BLOCKS_SWITCH) ? record.hookBlocks - record.base.hookBlocks : 0,
626 }
627}
628
629/** The counts as they stand: where they start over from once the failures are dealt with. */
630function countedNow(record: LoopRecord): Counted {
631 return { failures: record.failures, hookBlocks: record.hookBlocks }
632}
633
634/** The loop's counts start over from `covered` (a raise counts as one), and are shown. */
635async function startOver($: EngineInterface, cell: Cell<LoopRecord>, id: string, covered: Counted, raised: boolean, at?: number): Promise<LoopRecord> {
636 const record = await update(cell, (r) => ({
637 ...(r ?? newLoop('')),
638 base: { failures: Math.max(covered.failures, r?.base.failures ?? 0), hookBlocks: Math.max(covered.hookBlocks, r?.base.hookBlocks ?? 0) },
639 raises: (r?.raises ?? 0) + (raised ? 1 : 0),
640 ...(raised && at !== undefined ? { raisedAt: at } : {}),
641 }))
642 await showCounts($, id, record, null)
643 return record
644}
645
646/**
647 * Reports a loop's counts (on the board's node of it): the main agent's, or an agent's (null record: the main
648 * agent's start over with a new turn). `note`: `late`, a note for the band's event stream too.
649 */
650async function showCounts($: EngineInterface, id: string, record: LoopRecord | null, note: 'late' | null): Promise<void> {
651 await report(ioOf($), {
652 tally: {
653 feature: 'escalation',
654 agent: id,
655 failed: record?.failures ?? 0,
656 blocked: record?.hookBlocks ?? 0,
657 raised: record?.raises ?? 0,
658 ...(note === null ? {} : { late: true as const }),
659 ...(record === null ? { turnStart: true as const } : {}),
660 },
661 })
662}
663hooks/features/find-skill.ts 234 lines1// Feature: find_skill (#12). Most skills are left out of the main agent's
2// listing (features/skills.ts, ADR 0002), so when, partway through a turn,
3// the work turns out to need a skill nobody suggested, the main agent can ask:
4// it calls the tool with a few words on the work, and the decision model rates
5// the session's skills against them, the same way it rates them beside a
6// message (the same ranker, `modRanker`, in the same two stages (#11), over the
7// same skills and their profiles, with the recent conversation read by the
8// same rules). Nothing is pushed: the skills come back only as the tool's
9// answer.
10//
11// Its switch is `find-skill` (`/dp find-skill off`), apart from `skills`
12// (spec #56). The tool stays registered whatever the switches say, with a
13// description that never changes (the tool list is part of the prompt cache);
14// switched off, it says so when called.
15
16import type { EngineInterface, HttpInit, On } from 'claude-code'
17import { type Asked, type BackendIo, describeAsked, errorText, failureText } from '../decision/backend.ts'
18import { turnStartState } from '../decision/context.ts'
19import { quoteStart } from '../decision/redact.ts'
20import { modRanker, pickSkills, skillOpening, type SkillPick, type SkillPolicy, type SkillRanking } from '../decision/skills.ts'
21import { answersFor, mergeParts, type DecisionRequest } from '../decision/system-one.ts'
22import { readSessionSkills } from '../core/profiles.ts'
23import { report, type ReportIo } from '../core/report.ts'
24import type { Ctx } from '../core/setup.ts'
25import { describeStages, rankingSettings, type CatalogSkill } from '../core/skills.ts'
26import { defineSwitch, isOn, masterOn } from '../core/switches.ts'
27
28const CATALOG = { plugin: 'dispatch-pilot', key: 'skillCatalog' } as const
29const BOARD = { plugin: 'dispatch-pilot', key: 'board' } as const
30const DECISIONS = { plugin: 'dispatch-pilot', key: 'decisionLog' } as const
31const PPLX_RATE = { plugin: 'dispatch-pilot', key: 'pplxRate' } as const
32
33/** The switch's name, in `/dp` and in the decision log. */
34const SWITCH = 'find-skill'
35/** The skills feature's switch for profiles (#11): off, skills are rated by their descriptions here too. */
36const PROFILES = 'skill-profiles'
37/** The tool's name; the model calls it as `mcp__dispatch-pilot__find_skill`. */
38const TOOL = 'find_skill'
39const TOOL_CALLED = 'mcp__dispatch-pilot__find_skill'
40
41/** What the model reads about the tool. Fixed: nothing of the session goes into it. */
42const DESCRIPTION =
43 "Searches this session's skills for the ones that fit a piece of work, and returns each one's exact name, its relevance (0 to 1) and what it is for. Use it when the work at hand might have a skill you have not been shown: a file format, a service or its tooling, or a way of working such as reviewing, planning, testing or releasing. Load a skill it returns with the Skill tool by that exact name. Skills only the user can start are never returned."
44const INPUT = {
45 type: 'object',
46 properties: {
47 query: {
48 type: 'string',
49 description: 'The kind of work you need a skill for, in a few words, such as "fill in a form in a PDF file" or "review a branch before merging".',
50 },
51 },
52 required: ['query'],
53 additionalProperties: false,
54}
55/** How an answer that brings no skill ends: the main agent's way on. */
56const CARRY_ON = 'Carry on without it, or load a skill you know with the Skill tool by its exact name.'
57
58/**
59 * The session's skills: as read earlier this session (the skills feature
60 * reads them at its start), else read now, with the profiles the store holds
61 * (#11), and kept for the session; null when the session cannot be read.
62 */
63async function sessionCatalog($: EngineInterface, model: string): Promise<CatalogSkill[] | null> {
64 const { value } = await $.state.get(CATALOG)
65 if (value) return value.skills
66 const found = await readSessionSkills(
67 {
68 commands: () => $.command.list(),
69 listed: async () => (await $.session.usage({ breakdown: 'summary' })).context.breakdown?.skills?.skillFrontmatter ?? [],
70 overrides: async (source) => (await $.settings.read({ source })).skillOverrides,
71 home: () => $.env.get('HOME'),
72 cwd: () => $.session.cwd(),
73 exists: (path) => $.fs.exists(path),
74 read: (path) => $.fs.read(path),
75 list: (path) => $.fs.list(path),
76 get: (key) => $.store.get(key),
77 },
78 model,
79 )
80 if (found !== null) await $.state.set(CATALOG, { skills: found.skills })
81 return found?.skills ?? null
82}
83
84/** One decision request for find_skill through the person's decision model, its outcome in the debug log. */
85async function askLogged($: EngineInterface, ctx: Ctx, what: string, about: string, request: DecisionRequest, timeoutMs: number): Promise<Asked> {
86 const io: BackendIo = {
87 fetch: (url: string, init: HttpInit) => $.http.fetch(url, init),
88 sleep: (ms: number, signal: AbortSignal) => $.clock.sleep(ms, { signal }),
89 pace: { now: () => $.clock.now(), sentAt: { get: () => $.state.get(PPLX_RATE), set: (value, options) => $.state.set(PPLX_RATE, value, options) } },
90 }
91 const startedAt = await $.clock.now()
92 const asked = await ctx.backend.ask(io, request, timeoutMs)
93 const ms = (await $.clock.now()) - startedAt
94 $.ui.log(`${what} [${Object.keys(request.questions).join(', ')}] to ${ctx.backend.name} ${about}: ${describeAsked(asked, ms)}`, { to: 'debug' })
95 return asked
96}
97
98/** The opening of a catalog skill's SKILL.md for the ranking's second stage; null when it has no file. */
99async function openingOf($: EngineInterface, catalog: readonly CatalogSkill[], name: string): Promise<string | null> {
100 const file = catalog.find((skill) => skill.name === name)?.file ?? null
101 return file === null ? null : skillOpening(await $.fs.read(file))
102}
103
104export function registerFindSkill(on: On, ctx: Ctx): void {
105 defineSwitch({ name: SWITCH, info: '回答主 agent 的 find_skill:挑出适合它说的那项工作的 skill' })
106
107 /** Skills never offered (the option the skills feature reads too). */
108 const neverSuggested = new Set(ctx.config.skills.neverSuggested)
109 const policy = (): SkillPolicy => ctx.config.skills.find
110 /** How the mod's ranker ranks: the settings it rates the skills beside each message with. */
111 const rankBy = rankingSettings(ctx)
112 /** The model whose profiles the skills are offered by (#11). */
113 const model = ctx.config.skills.profileModel
114
115 // Registered once every plugin is loaded, under a match-all matcher (other
116 // features set themselves up at session start too). Without a decision
117 // model nothing could rate the skills, and the main agent keeps its listing.
118 on('session.start', { cwd: /(?:)/ }, async ($, e, next) => {
119 const result = await next(e)
120 if (ctx.backend.configured === false) return result
121 try {
122 const { tool } = await $.tool.register({ name: TOOL, description: DESCRIPTION, inputSchema: INPUT })
123 if (tool !== TOOL_CALLED) $.ui.log(`find_skill is registered as ${tool}, but the hook answers ${TOOL_CALLED}: its calls will fail`, { to: 'debug' })
124 } catch (error) {
125 $.ui.log(`find_skill was not registered: ${errorText(error)}`, { to: 'debug' })
126 }
127 return result
128 })
129
130 // The model's call: answered here, never passed on. Whatever goes wrong,
131 // the answer says so at once (fail open: the turn goes on without a skill).
132 on('tool.call', { tool: TOOL_CALLED }, async ($, e) => {
133 if (!masterOn()) return { result: `Dispatch Pilot is switched off (/dp on turns it back on), so find_skill rated no skills. ${CARRY_ON}` }
134 if (!isOn(SWITCH)) return { result: `find_skill is switched off (/dp ${SWITCH} on turns it back on), so it rated no skills. ${CARRY_ON}` }
135 // A dispatched agent's listing is left whole (ADR 0002): it has every skill in view.
136 if (e.agentId !== undefined) {
137 return { result: "find_skill rates skills for the main agent, whose skill listing is cut short. Yours lists this session's skills: pick from it and load one with the Skill tool by its exact name." }
138 }
139 const query = typeof e.query === 'string' ? e.query.replace(/\s+/g, ' ').trim() : ''
140 if (query === '') return { result: 'find_skill needs a query: a few words on the kind of work you need a skill for.' }
141 const io: ReportIo = {
142 board: { get: () => $.state.get(BOARD), set: (value, options) => $.state.set(BOARD, value, options) },
143 decisions: { get: () => $.state.get(DECISIONS), set: (value, options) => $.state.set(DECISIONS, value, options) },
144 debug: (line) => $.ui.log(line, { to: 'debug' }),
145 now: () => $.clock.now(),
146 toast: (text) => $.ui.toast(text),
147 }
148 // The main agent's own call, beside its decision for the turn: in the log, not on its node.
149 const call = { feature: SWITCH, agent: 'main', aside: true as const }
150 try {
151 // The skills the main agent can load, as a message asks about them: their question of stage
152 // one is the very one beside a message (those only the person can start have one of their
153 // own, which is not asked: they never come back); each by its profile, as beside a message,
154 // unless profiles are switched off.
155 const known = await sessionCatalog($, model)
156 const candidates = known?.filter((skill) => skill.by === 'model' && !neverSuggested.has(skill.name)).map((skill) => (isOn(PROFILES) ? skill : { ...skill, profile: null })) ?? null
157 if (candidates === null) {
158 await report(io, { decision: { ...call, skipped: 'unread' as const } })
159 return { result: `find_skill could not read this session's skills. ${CARRY_ON}` }
160 }
161 const about = `for find_skill ${quoteStart(query)}`
162 const ranker = modRanker(
163 {
164 ask: (request, timeoutMs) => askLogged($, ctx, 'second request', about, request, timeoutMs),
165 opening: (option) => openingOf($, candidates, option.name),
166 },
167 rankBy,
168 )
169 // The first request offers each skill by its profile where the decision model can read them all in time
170 // (BACKEND_DEFAULTS findSkillProfiles), else by its description; the second re-reads by both.
171 const part = ranker.part(ctx.config.skills.findByProfile ? candidates : candidates.map((skill) => ({ ...skill, profile: null })))
172 if (part === null) {
173 await report(io, { decision: { ...call, skipped: 'none' as const } })
174 return { result: `This session has no skill that find_skill could return. ${CARRY_ON}` }
175 }
176
177 // The same state as beside a message, the work named in place of the message; the ranker's second
178 // request (#11) asks about the same. Both requests share one wait: the second gets what the first left
179 // of it. The wait is the decision model's (findWaitMs: a message's timeoutMs with Jev),
180 // within the hook's own 10 s.
181 const waitMs = ctx.config.skills.findWaitMs
182 const startedAt = await $.clock.now()
183 const messages = ctx.config.context.messages > 0 ? await $.session.messages().catch(() => []) : []
184 const request = mergeParts(turnStartState({ prompt: query, messages, limits: ctx.config.context }), [part])
185 const asked = await askLogged($, ctx, 'request', about, request, waitMs)
186 const left = waitMs - ((await $.clock.now()) - startedAt)
187 const ranked = asked.ok ? await ranker.rank(answersFor(part, asked.answers), candidates, { state: request.state, timeoutMs: left }) : null
188 const failure = !asked.ok ? asked.failure : ranked === null ? { kind: 'parse' as const, detail: 'no answer about the skills' } : ranked.failed
189 if (ranked === null || failure !== undefined) {
190 const lost = failure ?? { kind: 'parse' as const, detail: 'no answer about the skills' }
191 const why = failureText(ctx.backend.name, lost)
192 await report(io, { decision: { ...call, failure: { backend: ctx.backend.name, ...lost } } })
193 return { result: `find_skill could not rate the skills (${why}). ${CARRY_ON}` }
194 }
195 const ranking = ranked
196
197 // Only skills the main agent can load were asked about, and they alone can come back.
198 const { suggest } = pickSkills(ranking, candidates, policy())
199 const names = suggest.map((skill) => skill.name).join('、')
200 await report(io, {
201 decision: {
202 ...call,
203 subject: quoteStart(query),
204 outcome: suggest.length > 0 ? `查到 ${names}` : '没查到 skill',
205 reason: describeRanking(ranking, policy()),
206 tone: suggest.length > 0 ? 'ok' : 'info',
207 skills: { suggest: suggest.map(({ name, relevance }) => ({ name, relevance })), try: [] },
208 },
209 })
210 return { result: found(query, suggest, policy()) }
211 } catch (error) {
212 $.ui.log(`find_skill failed: ${errorText(error)}`, { to: 'debug' })
213 await report(io, { decision: { ...call, skipped: 'error' as const } })
214 return { result: `find_skill could not rate the skills (an error in Dispatch Pilot, written to the debug log). ${CARRY_ON}` }
215 }
216 })
217}
218
219/** Why: what each stage of the ranking said (`describeStages`), and the bar a skill had to reach. */
220function describeRanking(ranking: SkillRanking, policy: SkillPolicy): string {
221 return `${describeStages(ranking)};相关度 ${policy.minRelevance.toFixed(2)} 起返回,最多 ${policy.max} 个`
222}
223
224/** The answer: the skills that fit, each by name, relevance and description, most relevant first; or that none does. */
225function found(query: string, suggest: readonly SkillPick[], policy: SkillPolicy): string {
226 if (suggest.length === 0) {
227 return `No skill fits "${query}": none reached relevance ${policy.minRelevance.toFixed(2)}. Carry on without one, try other words for the work, or load a skill you know with the Skill tool by its exact name.`
228 }
229 return [
230 `Skills that fit "${query}", rated by Dispatch Pilot’s decision model (relevance 0 to 1), most relevant first. Load one with the Skill tool by its exact name if it fits the work:`,
231 ...suggest.map((skill) => `- ${skill.name} (relevance ${skill.relevance.toFixed(2)})${skill.description ? `: ${skill.description}` : ''}`),
232 ].join('\n')
233}
234hooks/features/main-effort.ts 197 lines1// Feature: the main agent's effort, decided each time the person sends a
2// message, and when a report starts a turn of its own: a dispatched agent's
3// hand-back or a background task's notice reaching the idle session (those
4// turns are usually a look at a result and the next dispatch; the session's own
5// effort would be a waste). It adds the effort question to the message's ballot
6// (the core sends it); the answer becomes the effort of the turn the message
7// starts, on every step of that turn (the core's turn.step writer). A report's
8// text is read as the message; its turn is not re-decided mid-turn.
9//
10// A command turn (#19: `/implement #19`, a skill or markdown command the
11// person typed) is decided like their message: the decision model reads the
12// command as typed and what the command is for (core/commands.ts), never the
13// prompt it expands to.
14//
15// A message typed while a turn runs (`e.turnId`) is delivered into that turn
16// at its next step (a `queued_command` attachment; measured on 2.1.289), so
17// its decision takes the running turn from then on. It also waits as pending,
18// in case the turn ends first and the message starts a turn of its own.
19
20import type { EngineInterface, On } from 'claude-code'
21import { EFFORTS, LEVEL, probsOf, readEffort, readingText, traceEffort, type Effort, type EffortReading } from '../decision/effort.ts'
22import { quoteStart } from '../decision/redact.ts'
23import { turnStartPart } from '../decision/turn-start.ts'
24import { givesHint, judgeUnresolved, readUnresolved, UNRESOLVED } from '../decision/unresolved.ts'
25import { contribute, type PartOutcome } from '../core/ballot.ts'
26import { commandOf, commandState } from '../core/commands.ts'
27import { hintDecision, keptCount, keptSummary, moveCount, unresolvedDecision, type CountCell } from '../core/unresolved.ts'
28import { addPending, revise, turnKey, update, type Cell, type PendingDecision, type TurnRecord } from '../core/plans.ts'
29import { isPersonsMessage, startsReportTurn } from '../core/prompts.ts'
30import { report, type HintRecord, type ReportIo, type UnresolvedRecord } from '../core/report.ts'
31import type { Ctx } from '../core/setup.ts'
32import { defineSwitch, isOn } from '../core/switches.ts'
33import { UNRESOLVED_SWITCH } from './unresolved.ts'
34
35const PENDING = { plugin: 'dispatch-pilot', key: 'pending' } as const
36const TURNS = { plugin: 'dispatch-pilot', key: 'turns' } as const
37const BOARD = { plugin: 'dispatch-pilot', key: 'board' } as const
38const DECISIONS = { plugin: 'dispatch-pilot', key: 'decisionLog' } as const
39const CATALOG = { plugin: 'dispatch-pilot', key: 'skillCatalog' } as const
40const COUNT = { plugin: 'dispatch-pilot', key: 'unresolved' } as const
41
42/**
43 * What a command turn's command is for: its skill in the session's catalog as
44 * the skills feature read it (by its profile while profiles are on), else the
45 * command as `$.command.list()` describes it.
46 */
47async function describeCommand($: EngineInterface, name: string): Promise<Readonly<Record<string, string>> | null> {
48 const { value: catalog } = await $.state.get(CATALOG)
49 const found = catalog?.skills.find((skill) => skill.name === name)
50 const skill = found && !isOn('skill-profiles') ? { ...found, profile: null } : found
51 const listed = skill?.description ? undefined : (await $.command.list().catch(() => [])).find((command) => command.name === name)?.description
52 return commandState(name, skill, listed)
53}
54
55/** The unresolved count and summary in `$.state`, as a cell. */
56function countCell($: EngineInterface): CountCell {
57 return { get: () => $.state.get(COUNT), set: (value, options) => $.state.set(COUNT, value, options) }
58}
59
60export function registerMainEffort(on: On, ctx: Ctx): void {
61 defineSwitch({ name: 'main-effort', info: '发消息时决定主 agent 的 effort' })
62
63 on('prompt.submit', { text: /(?:)/ }, async ($, e, next) => {
64 // The person's own message, or a report that starts a turn of its own (a dispatched agent's hand-back, a task notice).
65 const handBack = !isPersonsMessage(e) && startsReportTurn(e)
66 if ((!isPersonsMessage(e) && !handBack) || !isOn('main-effort')) return next(e)
67 const pending: Cell<PendingDecision[]> = { get: () => $.state.get(PENDING), set: (value, options) => $.state.set(PENDING, value, options) }
68 let added: PendingDecision | null = null
69
70 /** The message waits for its turn, decided or not: the turn it starts is the person's own (mid-turn re-decisions are for such turns). */
71 const wait = async (effort: Effort | null) => {
72 const entry: PendingDecision = { text: e.text, effort, at: await $.clock.now(), ...(handBack ? { report: true as const } : {}) }
73 await update(pending, (list) => addPending(list ?? [], entry))
74 added = entry
75 }
76
77 const ran = handBack ? null : commandOf(e.text)
78 const command = ran === null ? null : await describeCommand($, ran.command)
79
80 // The person's own message also asks whether it says the problem they are on is still not solved (the unresolved
81 // count); a report that starts a turn is no word of theirs, and the switch can leave the question out. It travels
82 // in this part, so it is in the effort request with the 24000-token state it reads (ADR 0005).
83 const counting = !handBack && isOn(UNRESOLVED_SWITCH)
84 // The summary there is: a write that is not done yet is no reason to wait, the decision uses the one before it.
85 const summary = counting ? await keptSummary(countCell($)).catch(() => null) : null
86 // The count before this message (its own answer comes in the same request), and with it the strong hint once it has
87 // reached the setting: a sentence in the effort question about the kind of work, the level still the decision model's (ADR 0005).
88 const count = counting ? await keptCount(countCell($)).catch(() => 0) : 0
89 const maxAfter = ctx.config.unresolved.maxAfter
90 const hint: HintRecord | undefined = givesHint(count, maxAfter) ? { count, maxAfter } : undefined
91 // About the main agent of the turn this message starts, or of the one running when it was typed into it.
92 const forTurn = e.turnId === undefined ? ('next' as const) : ('current' as const)
93
94 /** The effort this message's turn goes out at, and the board's word for it. */
95 const decideEffort = async (outcome: PartOutcome, io: ReportIo, unresolved?: UnresolvedRecord): Promise<Effort | null> => {
96 const about = { feature: handBack ? 'main-effort (agent report)' : 'main-effort', agent: 'main', forTurn, subject: quoteStart(e.text) }
97 if (!outcome.ok) {
98 await report(io, { decision: { ...about, routed: false, failure: { backend: ctx.backend.name, ...outcome.failure } } })
99 await wait(null)
100 return null
101 }
102 const reading = readEffort(outcome.answers[LEVEL])
103 if (reading === null) {
104 await report(io, { decision: { ...about, routed: false, failure: { backend: ctx.backend.name, kind: 'parse', detail: 'no effort answer' } } })
105 await wait(null)
106 return null
107 }
108 // pickEffort's rules with their working: the board shows the steps, never recomputes them (#23).
109 const { effort, steps } = traceEffort(reading, ctx.config)
110 await report(io, {
111 decision: {
112 ...about,
113 routed: true,
114 outcome: `effort ${effort}`,
115 // The level as data: what reads the log (the band, the pane) never reads it out of the words.
116 effort,
117 reason: describeReading(reading, effort, ctx.config.thetaMax),
118 probs: probsOf(reading),
119 ...(reading.confidence === null ? {} : { conf: reading.confidence }),
120 trace: steps,
121 // What this message did to the unresolved count: the card draws it with the effort's.
122 ...(unresolved === undefined ? {} : { unresolved }),
123 // The card says the hint was given.
124 ...(hint === undefined ? {} : { hint }),
125 },
126 })
127 await wait(effort)
128 const running = e.turnId
129 if (running !== undefined) {
130 const ref = { ...TURNS, id: turnKey(running, undefined) }
131 const turn: Cell<TurnRecord> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
132 await update(turn, (record) => revise(record, effort))
133 }
134 return effort
135 }
136
137 /**
138 * What the answer to the unresolved question does to the count, which is moved here: what the log and the card say
139 * of it; null when there is no answer or the count cannot be kept (the debug log says so, the count stays).
140 */
141 const countUnresolved = async (answered: Extract<PartOutcome, { ok: true }>) => {
142 const reading = readUnresolved(answered.answers[UNRESOLVED])
143 if (reading === null) {
144 $.ui.log(`unresolved for ${quoteStart(e.text)}: no answer, the count stays`, { to: 'debug' })
145 return null
146 }
147 const judged = judgeUnresolved(reading)
148 try {
149 return unresolvedDecision(judged, await moveCount(countCell($), judged.change))
150 } catch (error) {
151 $.ui.log(`unresolved count not kept: ${error instanceof Error ? error.message : String(error)}`, { to: 'debug' })
152 return null
153 }
154 }
155
156 contribute(e.text, {
157 // Asked as the decision model's table says for these questions (Chinese with Jev); the other questions keep ctx.ask's.
158 ...turnStartPart({ ask: ctx.config.ask.turnStart, unresolved: counting, command, summary, count, maxAfter }),
159 settle: async (outcome) => {
160 const io: ReportIo = {
161 board: { get: () => $.state.get(BOARD), set: (value, options) => $.state.set(BOARD, value, options) },
162 decisions: { get: () => $.state.get(DECISIONS), set: (value, options) => $.state.set(DECISIONS, value, options) },
163 debug: (line) => $.ui.log(line, { to: 'debug' }),
164 now: () => $.clock.now(),
165 toast: (text) => $.ui.toast(text),
166 }
167 // The count first, so the effort's decision can say what this message did to it; neither depends on the other
168 // (either answer can be missing, and the effort is decided and kept whatever becomes of the count).
169 const counted = counting && outcome.ok ? await countUnresolved(outcome) : null
170 const effort = await decideEffort(outcome, io, counted?.unresolved)
171 // A count that moved is a decision of its own in the log (an answer that left it as it was is on the effort's card only).
172 if (counted?.unresolved !== undefined && counted.unresolved.before !== counted.unresolved.count) {
173 await report(io, { decision: { feature: UNRESOLVED_SWITCH, agent: 'main', aside: true, forTurn, subject: quoteStart(e.text), ...counted } })
174 }
175 // Each hint given is a decision of its own in the log: the count, that it was given, the level that came of it.
176 if (hint !== undefined) await report(io, { decision: { feature: UNRESOLVED_SWITCH, agent: 'main', aside: true, forTurn, subject: quoteStart(e.text), ...hintDecision(hint, effort) } })
177 },
178 })
179
180 const result = await next(e)
181 // Refused beneath: no turn will take this decision.
182 const withdrawn = added as PendingDecision | null
183 if (result.drop !== undefined && withdrawn !== null) {
184 await update(pending, (list) => (list ?? []).filter((entry) => !(entry.text === withdrawn.text && entry.at === withdrawn.at)))
185 }
186 return result
187 })
188}
189
190/** Why a level was picked: every level's probability, `max` held back below thetaMax when it was the most likely, and the backend's confidence. */
191function describeReading(reading: EffortReading, picked: Effort, thetaMax: number): string {
192 const p = reading.probabilities
193 const max = p[EFFORTS.length - 1] ?? 0
194 const held = picked !== 'max' && p.every((other) => other <= max)
195 return readingText(reading, held ? `max 的概率没到 max 门槛 ${thetaMax.toFixed(2)}` : undefined)
196}
197hooks/features/midturn-effort.ts 391 lines1// Feature: the main agent's effort, decided again while a turn runs (#5).
2//
3// Every N steps of a turn, and when the main agent dispatches an agent,
4// starts a Workflow or loads a skill, the decision model is asked again how
5// much reasoning the rest of the work needs. The question goes out the moment
6// a tool call of the main agent starts (`tool.call`), so it is answered while
7// the tool runs; the next step (`turn.step`, a layer above the core) takes
8// the answer, waits briefly when it is not back yet, and writes the turn's
9// plan, which the core then sends. Raising needs thetaUp; lowering needs
10// thetaDown, goes one level at a time and waits holdSteps after a raise (a
11// forced one too: the escalation feature, #7, marks it in the turn's plan).
12//
13// The counts the question carries (failed calls, hook blocks) are the
14// escalation feature's: one counter for the turn, since the counts last
15// started over.
16//
17// Measured on 2.1.289: the engine runs a step's tool calls while the response
18// still streams, so `tool.call` fires inside the step, before the step's
19// stream ends. This layer keeps the text of the step as it streams.
20
21import type { EngineInterface, HttpInit, On } from 'claude-code'
22import { describeAsked, errorText, within, type Asked, type BackendIo } from '../decision/backend.ts'
23import { messageText } from '../decision/context.ts'
24import { higherEffort, isEffort, probsOf, readEffort, readingText, type Effort } from '../decision/effort.ts'
25import {
26 contentLanguage,
27 judgeMidturn,
28 midturnEffortPart,
29 midturnRecord,
30 midturnState,
31 MIDTURN_LEVEL,
32 outcomeOf,
33 resultLine,
34 toolDetail,
35 verdictReason,
36 type MidturnCounts,
37 type MidturnInput,
38 type MidturnLimits,
39 type MidturnRules,
40 type Outcome,
41} from '../decision/midturn.ts'
42import { renderSummary } from '../decision/summary.ts'
43import { answersFor, mergeParts } from '../decision/system-one.ts'
44import { givesHint } from '../decision/unresolved.ts'
45import { wasBlocked } from '../core/outcomes.ts'
46import { floorHeld, MAIN, redecided, turnKey, update, type Cell, type TurnRecord } from '../core/plans.ts'
47import { hintDecision } from '../core/unresolved.ts'
48import { report, type HintRecord, type NodeFailure, type ReportIo } from '../core/report.ts'
49import type { Ctx } from '../core/setup.ts'
50import { defineSwitch, isOn } from '../core/switches.ts'
51import { UNRESOLVED_SWITCH } from './unresolved.ts'
52
53const TURNS = { plugin: 'dispatch-pilot', key: 'turns' } as const
54const MIDTURN = { plugin: 'dispatch-pilot', key: 'midturn' } as const
55const MAIN_STEP = { plugin: 'dispatch-pilot', key: 'mainStep' } as const
56const FAILURES = { plugin: 'dispatch-pilot', key: 'escalation', id: MAIN } as const
57const LOCK = { plugin: 'dispatch-pilot', key: 'lock' } as const
58const BOARD = { plugin: 'dispatch-pilot', key: 'board' } as const
59const DECISIONS = { plugin: 'dispatch-pilot', key: 'decisionLog' } as const
60const COUNT = { plugin: 'dispatch-pilot', key: 'unresolved' } as const
61const PPLX_RATE = { plugin: 'dispatch-pilot', key: 'pplxRate' } as const
62
63/** The feature's switch (`/dp midturn-effort on|off`). */
64const SWITCH = 'midturn-effort'
65
66type ToolEnd = { name: string; detail: string; outcome: Exclude<Outcome, 'running'> }
67type StepRecord = { index: number; text: string; tools: ToolEnd[] }
68/** The feature's own record of a main turn (`midturn` in types/index.d.ts). */
69type MidturnRecord = {
70 steps: number
71 engine: Effort | null
72 askedFor: number | null
73 recent: StepRecord[]
74}
75
76/** Steps kept in a turn's record. */
77const MAX_RECENT = 16
78
79/** Calls that start a new phase of the work (dispatch an agent, start a Workflow, load a skill): each re-decides. */
80const PHASE_TOOLS: ReadonlySet<string> = new Set(['Agent', 'Task', 'Workflow', 'Skill'])
81
82/** What the feature runs with: the shared ctx and the options it reads. */
83type Settings = {
84 ctx: Ctx
85 /** Re-decide at every step whose index is a multiple of this; 0 for never. */
86 every: number
87 rules: MidturnRules
88 /** How long a step waits for a re-decision not back yet. */
89 waitMs: number
90 limits: MidturnLimits
91}
92
93/** A step of the main agent: its turn and its index. */
94type MainStep = { turnId: string; index: number }
95/** A tool call as it starts: its name and what it works on. */
96type Starting = { name: string; detail: string }
97
98/**
99 * A re-decision on its way: asked for step `forStep` (`reason` says why);
100 * `answer` never rejects, and once it resolves `settled` holds it and `ms`
101 * how long it took.
102 */
103type InFlight = { forStep: number; reason: string; answer: Promise<Asked>; settled: Asked | null; ms: number; hint?: HintRecord }
104/** Re-decisions on their way, by turn key. Promises cannot live in $.state: a reload drops them (the step then keeps its effort). */
105const inFlight = new Map<string, InFlight>()
106/** The text of the main step streaming now, by turn key. */
107const streamed = new Map<string, { index: number; block: number; text: string }>()
108
109export function registerMidturnEffort(on: On, ctx: Ctx): void {
110 defineSwitch({ name: SWITCH, info: '一轮进行中重新判断主 agent 的 effort', parts: ['midturn'] })
111 // Read at each use: which decision model's settings these are is settled at the session start (core/setup.ts, `Ctx`).
112 const settings: Settings = {
113 ctx,
114 get every() {
115 return ctx.config.midturn.every
116 },
117 get rules() {
118 return ctx.config.midturn.rules
119 },
120 get waitMs() {
121 return ctx.config.midturn.waitMs
122 },
123 get limits() {
124 return ctx.config.midturn.limits
125 },
126 }
127
128 on('tool.call', { tool: /(?:)/ }, async ($, e, next) => {
129 // The main agent's own calls only: not a dispatched agent's, nor another plugin's $.tool.call.
130 if (e.agentId !== undefined || next.origin.plugin !== 'engine' || !isOn(SWITCH)) return next(e)
131 const detail = toolDetail(e)
132 let step: MainStep | null = null
133 try {
134 step = (await $.state.get(MAIN_STEP)).value ?? null
135 if (step !== null) await launch($, settings, step, { name: e.tool, detail })
136 } catch (error) {
137 $.ui.log(`midturn: ${errorText(error)}`, { to: 'debug' })
138 }
139 const result = await next(e)
140 if (step !== null) {
141 try {
142 const ended: ToolEnd = { name: e.tool, detail, outcome: outcomeOf(e.tool, result, wasBlocked(e.tool_use_id)) }
143 const at = step.index
144 const ref = { ...MIDTURN, id: turnKey(step.turnId, undefined) }
145 const cell: Cell<MidturnRecord> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
146 await update(cell, (r) => withTool(r ?? newRecord(), at, ended))
147 } catch (error) {
148 $.ui.log(`midturn: ${errorText(error)}`, { to: 'debug' })
149 }
150 }
151 return result
152 })
153
154 on('turn.step', { turnId: /(?:)/ }, async function* ($, e, next) {
155 if (e.agentId !== undefined || !isOn(SWITCH)) return yield* next(e)
156 const key = turnKey(e.turnId, undefined)
157 if (e.index === 0) {
158 // A new turn: what is left of earlier turns (a text, an answer never taken) is dropped.
159 for (const old of [...streamed.keys()]) if (old !== key) streamed.delete(old)
160 for (const old of [...inFlight.keys()]) if (old !== key) inFlight.delete(old)
161 }
162 try {
163 const note = await takeAnswer($, settings, e, key)
164 const engine = isEffort(e.effort) ? e.effort : null
165 await $.state.set(MAIN_STEP, { turnId: e.turnId, index: e.index })
166 const ref = { ...MIDTURN, id: key }
167 const cell: Cell<MidturnRecord> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
168 const record = await update(cell, (r) => ({ ...(r === undefined || e.index === 0 ? newRecord() : r), steps: e.index + 1, engine }))
169 const { value: turn } = await $.state.get({ ...TURNS, id: key })
170 // Quiet until the turn has been re-decided once (so a short turn shows nothing).
171 await report(ioOf($), {
172 tally: {
173 feature: 'midturn-effort',
174 agent: 'main',
175 steps: record.steps,
176 judged: turn?.decisions ?? 0,
177 changed: turn?.changes ?? 0,
178 quiet: record.askedFor === null || turn === undefined,
179 ...(note === null ? {} : 'late' in note ? { late: true as const } : { failure: note.failure }),
180 },
181 })
182 } catch (error) {
183 $.ui.log(`midturn: ${errorText(error)}`, { to: 'debug' })
184 }
185 streamed.set(key, { index: e.index, block: -1, text: '' })
186 const stream = next(e)
187 for await (const chunk of stream) {
188 if (chunk.kind === 'text') {
189 const live = streamed.get(key)
190 if (live !== undefined && live.index === e.index) {
191 live.text += live.block !== -1 && live.block !== chunk.index ? `\n${chunk.text}` : chunk.text
192 live.block = chunk.index
193 }
194 }
195 yield chunk
196 }
197 const result = await stream.result
198 try {
199 const text = messageText(result.answer, settings.limits.tokens)
200 const ref = { ...MIDTURN, id: key }
201 const cell: Cell<MidturnRecord> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
202 await update(cell, (r) => withText(r ?? newRecord(), e.index, text))
203 } catch (error) {
204 $.ui.log(`midturn: ${errorText(error)}`, { to: 'debug' })
205 }
206 return result
207 })
208}
209
210/**
211 * Asks again about the turn's effort, for the step after `step`, when a call
212 * starting is a reason to: a phase tool, or a step index that is a multiple of
213 * `every`. Once per step; the answer is left for that step to take.
214 */
215async function launch($: EngineInterface, s: Settings, step: MainStep, starting: Starting): Promise<void> {
216 const upcoming = step.index + 1
217 const key = turnKey(step.turnId, undefined)
218 const [{ value: turn }, { value: record }, { value: lock = null }, { value: failures }, sentAt] = await Promise.all([
219 $.state.get({ ...TURNS, id: key }),
220 $.state.get({ ...MIDTURN, id: key }),
221 $.state.get(LOCK),
222 $.state.get(FAILURES),
223 $.clock.now(),
224 ])
225 // Only a turn the person's own message started is re-decided, whether its start was decided or not (that request
226 // may have failed: then from the session's own effort); never without a decision model set up (every request would
227 // fail at once).
228 if (turn === undefined || turn.person !== true || lock !== null || record === undefined || record.engine === null || s.ctx.backend.configured === false) return
229 const reason = PHASE_TOOLS.has(starting.name) ? starting.name : s.every > 0 && upcoming % s.every === 0 ? `每 ${s.every} 步` : null
230 if (reason === null || record.askedFor === upcoming || inFlight.get(key)?.forStep === upcoming) return
231 // The problem the person is on, as of now (this turn's own message has moved the count already): the summary and the
232 // count go along, and the strong hint once the count has reached the setting (#41). Nothing of it with the switch off.
233 const { value: held } = isOn(UNRESOLVED_SWITCH) ? await $.state.get(COUNT) : { value: undefined }
234 const count = held?.count ?? 0
235 const summary = held?.summary
236 const maxAfter = s.ctx.config.unresolved.maxAfter
237 const hint: HintRecord | undefined = givesHint(count, maxAfter) ? { count, maxAfter, where: 'mid' } : undefined
238 const input: MidturnInput = {
239 message: turn.prompt,
240 step: upcoming,
241 current_effort: higherEffort(turn.effort ?? record.engine, floorHeld(turn, upcoming)) as Effort,
242 counts: countsOf(turn, step.turnId, failures),
243 recent_steps: summaries(record, step.index, key, starting, contentLanguage(turn.prompt)),
244 ...(summary === undefined ? {} : { problem_summary: renderSummary(summary, s.ctx.ask.language) }),
245 ...(count > 0 ? { unresolved_count: count } : {}),
246 }
247 const request = mergeParts(midturnState(input, s.limits), [midturnEffortPart(s.ctx.ask, { hint: hint !== undefined })])
248 const io: BackendIo = {
249 fetch: (url: string, init: HttpInit) => $.http.fetch(url, init),
250 sleep: (ms: number, signal: AbortSignal) => $.clock.sleep(ms, { signal }),
251 pace: { now: () => $.clock.now(), sentAt: { get: () => $.state.get(PPLX_RATE), set: (value, options) => $.state.set(PPLX_RATE, value, options) } },
252 }
253 const asking = s.ctx.backend.ask(io, request, s.ctx.config.timeoutMs)
254 const entry: InFlight = { forStep: upcoming, reason, answer: asking, settled: null, ms: 0, ...(hint === undefined ? {} : { hint }) }
255 entry.answer = asking.then(async (answered) => {
256 entry.ms = (await $.clock.now()) - sentAt
257 entry.settled = answered
258 return answered
259 })
260 inFlight.set(key, entry)
261 const ref = { ...MIDTURN, id: key }
262 const cell: Cell<MidturnRecord> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
263 await update(cell, (r) => ({ ...(r ?? newRecord()), askedFor: upcoming }))
264}
265
266/**
267 * The turn's counts as the decision model reads them: its decisions and level
268 * changes, and the failed calls and hook blocks the escalation feature counted
269 * for it, since its counts last started over (a forced raise, or failures
270 * found expected).
271 */
272function countsOf(turn: TurnRecord, turnId: string, failures: { turnId: string; failures: number; hookBlocks: number; base: { failures: number; hookBlocks: number } } | undefined): MidturnCounts {
273 const counted = failures !== undefined && failures.turnId === turnId
274 return {
275 judgments: turn.decisions,
276 changes: turn.changes,
277 failures: counted ? failures.failures - failures.base.failures : 0,
278 hook_blocks: counted ? failures.hookBlocks - failures.base.hookBlocks : 0,
279 }
280}
281
282/**
283 * Takes the re-decision meant for this step, if any, and writes what it
284 * decides into the turn's plan. An answer not back yet gets `waitMs` more;
285 * still none, the step keeps the effort it had and the answer stays for a
286 * later step. Resolves to what the tally should add: `late`, or the failed
287 * request that is why there is no answer; null otherwise.
288 */
289async function takeAnswer($: EngineInterface, s: Settings, e: { index: number; effort?: unknown }, key: string): Promise<{ late: true } | { failure: NodeFailure } | null> {
290 const pending = inFlight.get(key)
291 if (pending === undefined || pending.forStep > e.index) return null
292 let asked = pending.settled
293 if (asked === null) {
294 asked = await within((ms, signal) => $.clock.sleep(ms, { signal }), pending.answer, s.waitMs, null)
295 }
296 if (asked === null) return { late: true }
297 inFlight.delete(key)
298 $.ui.log(`request [midturn.${MIDTURN_LEVEL}] for step ${pending.forStep} (${pending.reason}) to ${s.ctx.backend.name}: ${describeAsked(asked, pending.ms)}`, { to: 'debug' })
299 const engine = isEffort(e.effort) ? e.effort : null
300 const [{ value: turn }, { value: lock = null }] = await Promise.all([$.state.get({ ...TURNS, id: key }), $.state.get(LOCK)])
301 if (turn === undefined || lock !== null || engine === null) return null
302 const floor = floorHeld(turn, e.index)
303 const current = higherEffort(turn.effort ?? engine, floor) as Effort
304 const reading = asked.ok ? readEffort(answersFor(midturnEffortPart(s.ctx.ask), asked.answers)[MIDTURN_LEVEL]) : null
305 /** The hint given with this request is a decision of its own in the log: the level the decision model came to, when it did. */
306 const logHint = async (picked: Effort | null) => {
307 if (pending.hint !== undefined) await report(ioOf($), { decision: { feature: UNRESOLVED_SWITCH, agent: 'main', aside: true, subject: `第 ${e.index} 步(${pending.reason})`, ...hintDecision(pending.hint, picked) } })
308 }
309 if (reading === null) {
310 await logHint(null)
311 return { failure: { backend: s.ctx.backend.name, ...(asked.ok ? { kind: 'parse' as const, detail: 'no effort answer' } : asked.failure) } }
312 }
313
314 const ref = { ...TURNS, id: key }
315 const turnCell: Cell<TurnRecord> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
316 const sinceRaise = turn.raisedAt == null ? null : e.index - turn.raisedAt
317 const position = { current, sinceRaise, atLeast: floor }
318 const verdict = judgeMidturn(reading, position, s.rules)
319 await update(turnCell, (r) => redecided(r ?? turn, current, verdict.effort, e.index))
320 // Beside the turn's own decision (the node's link to it stays), with the rules' working: the answer's pick, then the move.
321 await report(ioOf($), {
322 decision: {
323 feature: SWITCH,
324 agent: 'main',
325 aside: true,
326 subject: `第 ${e.index} 步(${pending.reason})`,
327 outcome: `effort ${verdict.effort}${verdict.effort === current ? '(保持)' : `(原 ${current})`}`,
328 reason: `${readingText(reading)};${verdictReason(verdict, position, s.rules)}`,
329 tone: verdict.effort === current ? 'info' : 'ok',
330 probs: probsOf(reading),
331 conf: verdict.confidence,
332 trace: [...verdict.trace.pick.steps, ...verdict.trace.steps],
333 mid: midturnRecord(verdict, position, s.rules),
334 },
335 })
336 await logHint(verdict.picked)
337 return null
338}
339
340function newRecord(): MidturnRecord {
341 return { steps: 0, engine: null, askedFor: null, recent: [] }
342}
343
344/**
345 * The latest steps as the decision model reads them, oldest first: each
346 * step's text and its tool calls with how they ended. Step `index` comes
347 * last, its text as streamed so far, the call `starting` in it marked running.
348 */
349function summaries(record: MidturnRecord, index: number, key: string, starting: Starting, language: 'en' | 'zh'): MidturnInput['recent_steps'] {
350 const live = streamed.get(key)
351 const now = record.recent.find((step) => step.index === index)
352 const line = (t: { name: string; detail: string; outcome: Outcome }) => ({ name: t.name, result: resultLine(t.outcome, t.detail, language) })
353 return [
354 ...record.recent.filter((step) => step.index < index).map((step) => ({ assistant_text: step.text, tools: step.tools.map(line) })),
355 {
356 assistant_text: live?.index === index ? live.text : (now?.text ?? ''),
357 tools: [...(now?.tools ?? []).map(line), line({ ...starting, outcome: 'running' })],
358 },
359 ]
360}
361
362/** The record with a tool call's ending added to its step. */
363function withTool(record: MidturnRecord, index: number, ended: ToolEnd): MidturnRecord {
364 const at = record.recent.findIndex((step) => step.index === index)
365 const step = at >= 0 ? (record.recent[at] as StepRecord) : { index, text: '', tools: [] }
366 const recent = [...record.recent]
367 if (at >= 0) recent[at] = { ...step, tools: [...step.tools, ended] }
368 else recent.push({ ...step, tools: [ended] })
369 return { ...record, recent: recent.sort((a, b) => a.index - b.index).slice(-MAX_RECENT) }
370}
371
372/** The record with a step's text, once the step has streamed whole. */
373function withText(record: MidturnRecord, index: number, text: string): MidturnRecord {
374 const at = record.recent.findIndex((step) => step.index === index)
375 const recent = [...record.recent]
376 if (at >= 0) recent[at] = { ...(recent[at] as StepRecord), text }
377 else recent.push({ index, text, tools: [] })
378 return { ...record, recent: recent.sort((a, b) => a.index - b.index).slice(-MAX_RECENT) }
379}
380
381/** What the decision report needs of the host: its two cells, the debug log, the clock and the toast. */
382function ioOf($: EngineInterface): ReportIo {
383 return {
384 board: { get: () => $.state.get(BOARD), set: (value, options) => $.state.set(BOARD, value, options) },
385 decisions: { get: () => $.state.get(DECISIONS), set: (value, options) => $.state.set(DECISIONS, value, options) },
386 debug: (line) => $.ui.log(line, { to: 'debug' }),
387 now: () => $.clock.now(),
388 toast: (text) => $.ui.toast(text),
389 }
390}
391hooks/features/unresolved.ts 161 lines1// Feature: the unresolved count and the problem summary (#39, #40; GLOSSARY 未解决次数,
2// 问题摘要). Each of the person's messages is asked, in the effort request, whether it
3// says the problem they and the main agent are on is still not solved, is solved, or is
4// another one; the answer moves the count (core/unresolved.ts). The question and the
5// reading of its answer are in features/main-effort.ts, which owns the effort part of
6// the ballot they travel in (one contribution per part); this file owns the switch,
7// the count's start-over and the summary.
8//
9// The summary is a record of the problem for the decision model to read beside the
10// conversation. When a turn that the person's own message started ends, a cheap model
11// continues it, in the background: the hook does not wait, and a message that comes
12// before it is written is decided with the summary as it was (main-effort.ts reads it).
13// Writes go one after the other, each continuing what the one before wrote. A write
14// that fails (the model errs, times out, answers something that is no summary) or that
15// a reload of the mod lost leaves the summary as it was and is told to the decision
16// report; one the count's start-over overtook is dropped when it lands.
17//
18// The count is per session: `/clear` and a new session start it over (`session.end`
19// fires for both; `session.start` also fires on a hot reload, which must keep it),
20// a compaction keeps it. Switched off, the question is not asked, the count stays as it
21// is and no summary is written.
22
23import type { EngineInterface, ModelCompleteResult, On } from 'claude-code'
24import { errorText } from '../decision/backend.ts'
25import { clipToTokens } from '../decision/context.ts'
26import { quoteStart, redactSecrets } from '../decision/redact.ts'
27import { readSummary, SUMMARY_MAX_REPLY, SUMMARY_SYSTEM, SUMMARY_TIMEOUT_MS, summaryPrompt, turnTools, type TurnInput } from '../decision/summary.ts'
28import { turnKey } from '../core/plans.ts'
29import { report, type ReportIo } from '../core/report.ts'
30import type { Ctx } from '../core/setup.ts'
31import { defineSwitch, isOn } from '../core/switches.ts'
32import { clearCount, dropSummary, landSummary, lostSummaries, queueSummary, summaryFailure, summaryFor, type CountCell, type SummaryFailure } from '../core/unresolved.ts'
33
34const COUNT = { plugin: 'dispatch-pilot', key: 'unresolved' } as const
35const TURNS = { plugin: 'dispatch-pilot', key: 'turns' } as const
36const BOARD = { plugin: 'dispatch-pilot', key: 'board' } as const
37const DECISIONS = { plugin: 'dispatch-pilot', key: 'decisionLog' } as const
38
39/** The switch's name: `/dp unresolved on|off`. */
40export const UNRESOLVED_SWITCH = 'unresolved'
41
42/** The writes of this load of the mod: after the one before it (they are chained), and the turns they are for. A hot reload empties both. */
43let chain: Promise<void> = Promise.resolve()
44const running = new Set<string>()
45/** The engine refused the summary's model: no more is asked of it until the conversation ends (a reload asks again). */
46let refused = false
47
48/** What a write is given when the turn ends: the turn, and what the cheap model is shown of it. */
49type Write = { turn: string; subject: string; shown: Omit<TurnInput, 'previous'> }
50
51export function registerUnresolved(on: On, ctx: Ctx): void {
52 // Off until the person turns it on (#48, ADR 0006); what they set by hand is in $.store and wins.
53 defineSwitch({ name: UNRESOLVED_SWITCH, info: '每条消息判断同一个问题是否仍未解决,数未解决的次数;每轮结束后在后台写问题摘要(默认关)', default: false })
54
55 on('session.end', { reason: /(?:)/ }, async ($, e, next) => {
56 refused = false
57 try {
58 await clearCount(countCell($))
59 } catch (error) {
60 $.ui.log(`unresolved count not cleared: ${error instanceof Error ? error.message : String(error)}`, { to: 'debug' })
61 }
62 return next(e)
63 })
64
65 // A reload of the mod loses the writes in flight (the engine drops their calls with the old module): the state still lists them.
66 on('session.start', { cwd: /(?:)/ }, async ($, e, next) => {
67 const result = await next(e)
68 try {
69 for (const turn of await lostSummaries(countCell($), running)) await tell($, turn, 'lost', ctx.config.unresolved.summaryModel)
70 } catch (error) {
71 $.ui.log(`summaries lost to a reload not told: ${errorText(error)}`, { to: 'debug' })
72 }
73 return result
74 })
75
76 // A turn the person's own message started ends: its summary is written, in the background.
77 on('turn.complete', { turnId: /(?:)/ }, async ($, e, next) => {
78 const result = await next(e)
79 // The summary is a paraphrase of the conversation: none is written when the person sends the decision model none (`contextMessages` 0).
80 if (e.agentId !== undefined || refused || !isOn(UNRESOLVED_SWITCH) || ctx.backend.configured === false || ctx.config.context.messages <= 0) return result
81 try {
82 const { value: turn } = await $.state.get({ ...TURNS, id: turnKey(e.turnId, undefined) })
83 // Only the person's own message: a hand-back or a task notice that started the turn is no word of theirs.
84 if (turn === undefined || !turn.person) return result
85 const messages = await $.session.messages().catch(() => [])
86 const write: Write = { turn: e.turnId, subject: quoteStart(turn.prompt), shown: { person: turn.prompt, reply: e.answer, tools: turnTools(messages) } }
87 await queueSummary(countCell($), e.turnId)
88 running.add(e.turnId)
89 chain = chain.then(() => writeSummary($, ctx, write)).catch(() => undefined)
90 } catch (error) {
91 $.ui.log(`summary not queued: ${errorText(error)}`, { to: 'debug' })
92 }
93 return result
94 })
95}
96
97function countCell($: EngineInterface): CountCell {
98 return { get: () => $.state.get(COUNT), set: (value, options) => $.state.set(COUNT, value, options) }
99}
100
101/** The decision report's closures over this hook's `$`. */
102function reportIo($: EngineInterface): ReportIo {
103 return {
104 board: { get: () => $.state.get(BOARD), set: (value, options) => $.state.set(BOARD, value, options) },
105 decisions: { get: () => $.state.get(DECISIONS), set: (value, options) => $.state.set(DECISIONS, value, options) },
106 debug: (line) => $.ui.log(line, { to: 'debug' }),
107 now: () => $.clock.now(),
108 toast: (text) => $.ui.toast(text),
109 }
110}
111
112/** A write that did not make a summary, told to the decision report (the one that was there stays). `turn` only names the write in the debug log. */
113async function tell($: EngineInterface, turn: string, failure: SummaryFailure, model: string, subject = ''): Promise<void> {
114 $.ui.log(`summary for turn ${turn}: ${failure === 'lost' ? 'lost to a reload' : failure}`, { to: 'debug' })
115 await report(reportIo($), { decision: { feature: UNRESOLVED_SWITCH, agent: 'main', aside: true, subject, ...summaryFailure(failure, model) } })
116}
117
118/**
119 * Continues the summary with one turn, with the cheap model, once the writes before it are done. The summary it
120 * continues is read when it starts, so it holds what those wrote. Never throws.
121 */
122async function writeSummary($: EngineInterface, ctx: Ctx, write: Write): Promise<void> {
123 const cell = countCell($)
124 const model = ctx.config.unresolved.summaryModel
125 try {
126 const { wanted, previous } = await summaryFor(cell, write.turn)
127 // Cleared since the turn ended (solved, another problem, /clear), or switched off meanwhile: nothing to write.
128 if (!wanted || !isOn(UNRESOLVED_SWITCH) || refused) {
129 await dropSummary(cell, write.turn)
130 return
131 }
132 let reply: ModelCompleteResult
133 try {
134 reply = await $.model.complete({ model, system: SUMMARY_SYSTEM, prompt: summaryPrompt({ ...write.shown, previous }), maxTokens: SUMMARY_MAX_REPLY, timeoutMs: SUMMARY_TIMEOUT_MS })
135 } catch (error) {
136 refused = true
137 await dropSummary(cell, write.turn)
138 await tell($, write.turn, { reason: 'refused', detail: errorText(error) }, model, write.subject)
139 return
140 }
141 if (!reply.isAnswered) {
142 await dropSummary(cell, write.turn)
143 await tell($, write.turn, reply.reason === 'api-error' ? { reason: 'api-error', status: reply.status, error: reply.error } : { reason: reply.reason }, model, write.subject)
144 return
145 }
146 const summary = readSummary(reply.text)
147 if (summary === null) {
148 await dropSummary(cell, write.turn)
149 await tell($, write.turn, { reason: 'unfit', text: clipToTokens(redactSecrets(reply.text).replace(/\s+/g, ' ').trim(), 60) }, model, write.subject)
150 return
151 }
152 const kept = await landSummary(cell, write.turn, summary)
153 $.ui.log(`summary for turn ${write.turn}: ${kept ? 'written' : 'dropped, it was cleared meanwhile'} (${reply.usage.output_tokens} tokens from ${model})`, { to: 'debug' })
154 } catch (error) {
155 // The session went away under it, or a bug: either way the session is not held up.
156 $.ui.log(`summary for turn ${write.turn} not written: ${errorText(error)}`, { to: 'debug' })
157 } finally {
158 running.delete(write.turn)
159 }
160}
161