SLOPSHOPPER

dispatch-pilot

Dispatch Pilot:一个在 Claude 之外的决策模型(默认 Perplexity 的 pplx,没有 Perplexity key 时退回 TypeSafe 的 Jev,也可以固定用 Jev)在你每次发消息时决定主 agent 这一轮的…

newpanebandspinnerguardcommand
★ 2v0.4.0MITupdated 2026-10-08alexcz-a11y/claude-mods/dispatch-pilot
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · dispatch-pilot
│ ┃ 依据 ✕ › fix the failing auth╭────────────────────────────────────────────╮ │ ┃ ◆ 依据 · 主 agent │ dispatch-pilot │ │ ┃ 功能(/dp <名字> on|off) signals 开 main- ● dispatch-pilot: sett│ 主 agent 未路由:决策模型没配好(pplx:没有填 │ │ ┃ [ ‹ 上一个 (p) ] [ 下一个 › (n) ] band 上按 ● dispatch-pilot: pplx│ perplexityApiKey 或 PERPLEXITY_API_KEY) │ │ ┃ ╭─────────────────────────────────────────── ⏺ Read(src/auth.ts) ╰────────────────────────────────────────────╯ │ ┃ │ ◆ 主 agent — ▱▱ ⎿ Read 6 lines │ ┃ │ 状态 未路由 · 决策模型没配好 ⏺ Update(src/auth.ts) │ ┃ │ 类型 没有配好决策模型的密钥或账号,或者 ⎿ Added 2 lines, removed 1 line │ ┃ │ 后端 pplx ⏺ Bash(bun test) │ ┃ │ 细节 pplx:没有填 perplexityApiKey 或 P ⎿ 3 pass, 1 fail │ ┃ │ 排查 debug log(claude --debug-file │ ┃ │ <路径>)里有这次请求的那一行,写着 ● Done. refresh now rejects expired claims and logs an audit event. │ ┃ │ 久、怎么失败的 │ ┃ ╰─────────────────────────────────────────── ✻ Worked for 42s · done 4:20 PM │ ┃ │ ┃ ◆ 决策日志 还没有决定 › /dp │ ⎿ dispatch-pilot: 依据面板已打开:p / n 翻看 agent,Esc 关闭 │ ● dispatch-pilot: skills: no decision model is set up, so the main ag │ ● dispatch-pilot: request [effort.level] to pplx: config: no Perplexi │ │ ⟨Claude Code's own drawing⟩ ◆ 第 1 轮 0:42 · 主 agent · 未路由 · 决策模型没配好 · /dp log 看依据 ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts ⟨Claude Code's own drawing⟩ ◆

Draws

Band
⟨Claude Code's own drawing⟩ ◆ 第 1 轮 0:42 · 主 agent · 未路由 · 决策模型没配好 · /dp log 看依据
Pane · 依据
◆ 依据 · 主 agent 1/1 功能(/dp <名字> on|off) signals 开 main-effort 开 unres [ ‹ 上一个 (p) ] [ 下一个 › (n) ] band 上按 0-9 直接选 ╭─────────────────────────────────────────────────────────── │ ◆ 主 agent — ▱▱▱▱▱ — │ 状态 未路由 · 决策模型没配好 │ 类型 没有配好决策模型的密钥或账号,或者密钥被拒绝 │ 后端 pplx │ 细节 pplx:没有填 perplexityApiKey 或 PERPLEXITY_API_KE │ 排查 debug log(claude --debug-file │ <路径>)里有这次请求的那一行,写着它发了什么、等了 │ 久、怎么失败的 ╰─────────────────────────────────────────────────────────── ◆ 决策日志 还没有决定
README

Dispatch Pilot

Dispatch Pilot 是 alex-mods marketplace 里的一个 Claude Code mod。它在 Claude 之外调用一个决策模型(默认是 Perplexity 的 pplx;没有 Perplexity 的 key、但有 TypeSafe 的 key 时退回 TypeSafe 的 Jev,也可以指定用 Jev),替你决定 Claude Code 怎么干活:主 agent 每一轮用哪档 effort,派出的 agent 用哪个模型和哪档 effort,主 agent 该看哪几个 skill。

主 agent 的模型从不改变,所以 prompt cache 不受影响。这只在 Claude Code 订阅下成立,见「要求」。

它做什么

决策模型只负责给出判断,不参与回答你。每个判断的结果写在看板上,依据记进决策日志,在依据面板里看(/dp),也可以用 /dp log N 在对话里列出。

  • 主 agent 的 effort。 你每发一条消息,Dispatch Pilot 把这条消息和最近几条消息发给决策模型,问它这项工作需要多少逐步推理,得到 low、medium、high、xhigh、max 五档各自的概率。取概率最高的一档,并列时取较高的一档;如果高一档的概率也有 0.3 以上,就再往上取一档(升高容易:低估一档要付出质量,高估一档只多花 token);max 只在它自己的概率达到 thetaMax 时才用。这一轮的每个模型请求都按这一档发出。一轮进行中,每隔几步(rejudgeEvery),以及主 agent 派出 agent、启动 Workflow 或加载 skill 时,会再判断一次(中途重判);升档和降档都有防抖:升档要置信度达到 thetaUp(0.3),降档要达到 thetaDown(0.75)而且一次只降一档,升档之后 holdSteps(5)步之内不降。派出 agent 交回结果、后台任务结束,会话空闲时它们会开始主 agent 的新一轮,这样的一轮也走同样的判断:决策模型读到的是那条报告的文字,不推荐 skill,也不在这一轮中途重判(会话自己的 xhigh 用在只是看一眼结果上是浪费);报告送进一轮正在进行的对话时不另外判断。其他会话的消息、插件自己发的消息和空消息都不判断。你输入的 skill 或 markdown 命令(例如 /implement #19)开始的一轮也照常判断:决策模型读你输入的命令和参数,以及这个命令是做什么的,不读命令展开后的正文;参数里的点名照样算;这样的一轮不推荐 skill。/dp、/clear 这类本地命令不开始一轮,也就不判断。
  • 派出 agent 的模型和 effort。 主 agent 用 Agent 工具派出一个 agent 时,决策模型为它选模型(默认在 haiku、sonnet、opus 中选,打开 agentFable 后加入 fable)和 effort;选了 haiku 就不设 effort,haiku 不支持。三个模型的分工按 Artificial Analysis 的基准定(见下面的「模型和 effort 的依据」):haiku 只做一两步就能跑完、结果只需要收集起来并按要求排版的只读查找,sonnet 承担大多数执行类工作(终端、需求明确的代码修改、自动化、调研汇报),opus 做需要审慎判断或细微错误代价高的工作(安全、并发、涉及钱、迁移、生产)、难推理、设计、原因未知的 bug、科学或算法类代码,以及结论取决于记忆中的事实、而且无法在仓库或文档里查证的调研;fable 在 AA 上没有领先 Opus 的地方,价格是 2.5 倍,所以仍默认关闭。决策出来的 effort 不低于所选模型的下限:sonnet 和 opus 至少 medium,haiku 不带 effort。你在消息里点名的模型或 effort(「用 opus」「effort 开 low」)一定照办,下限和往上取的一档都不会改你点名的值,你排除的模型(「别用 opus」)不会用(决策请求整体失败时除外,见「局限和待评测」)。主 agent 自己为这个 agent 指定了模型时,只有决策模型选了别的、而且置信度达到 agentOverride 才推翻它。
  • Workflow 里的 agent。 主 agent 提交 Workflow 脚本时,对脚本里每个 agent() 调用点做同样的判断,把模型和 effort 写进脚本再运行,并告诉主 agent 写了什么;workflowMode 选 return 时改为退回脚本,附上逐个调用的推荐,让主 agent 自己写进去。脚本写不进去的(用 scriptPath 或 name 提交、恢复的运行、读不了的脚本),在每个 agent 启动时按它的 label 设置,这叫兜底。
  • 卡住时强制升档。 主 agent 或派出 agent 的工具调用接连失败 escalateAfter 次时,再问决策模型一次,把它的 effort 升一档(escalateMode 选 max 则直接升到 max);haiku 没有 effort 可升,改用 sonnet 接着做(escalateHaikuTo)。这些失败本来就在意料之中的,例如先写下、要看它红的测试,或者没找到东西而以非零退出的搜索,不升档;是不是预期内失败由决策模型判断,不靠关键词。你自己拒绝的调用从不算失败。
  • 未解决次数。 你本人的每条消息(包括斜杠命令)发出时,决策模型还回答一道三选一的题:这条消息是在说你和主 agent 最近在处理的那个问题仍未解决、已经解决,还是换了新问题或无关;不靠关键词,「再看看」或贴一段同样的报错也认得出。「仍未解决」把握够(暂定 0.5)次数加一,「已经解决」或「新问题或无关」把握更高(暂定 0.7)才清零,都没到次数不变。次数存在会话里,/clear 和新会话清零,/compact 和热重载保留;agent 交回结果、后台任务通知开始的一轮不问。依据卡片显示次数和这一次的结论。次数和下面的摘要一起作为决策模型判断 effort 的依据,见「强提示」。这项功能(unresolved 开关,次数、摘要、强提示共用)默认关闭:评测里它对默认的决策模型没有提升,所以不问这道题、不数次数、不写摘要、不给强提示。/dp unresolved on 打开,对 pplx 和 Jev 都生效;/dp unresolved off 再关。开关存在 $.store,你手动设过的状态升级后保持原样;0.3.1 里它是默认开着的,开着时没有保存过任何设置,所以从 0.3.1 升级上来、一直开着的,现在变成了关闭,要用请自己打开。
  • 问题摘要。 你本人的消息开始的那一轮结束后,一个便宜的模型(summaryModel,默认 haiku,用你的 Claude 登录和用量)在后台续写一份简短的摘要:问题是什么、试过哪些做法、现在到了哪一步,不超过 500 token,只记试过什么,不判断成没成功;你的下一条消息说「仍未解决」时,最后一次尝试标上「未解决」。决策模型发消息时读它,借此看到超出上下文预算的更早几轮;它从不等摘要写完,来不及时用上一份。写失败(出错、超时)保留旧摘要并在决策日志里记一条。问题解决、换了问题、/clear 和新会话时和次数一起清空,/compact 和热重载保留。依据面板里能看摘要全文。交给写摘要模型的内容先脱敏;contextMessages 为 0 时不写(摘要是对话的转述)。默认关闭,随 unresolved 开关一起打开和关闭。
  • 强提示。 次数达到 unresolvedMaxAfter(默认 3,0 到 10,0 表示不给)时,发消息时和中途重判时 effort 题的说明里多一条:这项工作属于多次尝试都没解决的故障。它只描述情境:不写档位的名字或序号,pickEffort 的规则和 thetaMax 都不变,用不用最高一档仍由决策模型结合摘要和对话自己决定。次数和摘要也进这两种请求的 state(unresolved_count、problem_summary);派出 agent 的决定读摘要和次数当背景,但从不加强提示(「查一个文件」这类子任务不该因此判到最高一档)。发消息时带的是这条消息之前的次数(这条消息自己的结论要同一个请求的回答才知道),所以默认设置下第四次说「仍未解决」的那条消息的中途重判先读到强提示,下一条消息的请求再带上。每次给了强提示,决策日志记一条(次数、已给强提示、决策模型判出的档位),依据卡片显示「已给强提示」;看板不新增事件。unresolved 开关关着(默认)时这些都不带。
  • skill。 主 agent 不再读完整的 skill 列表(装的 skill 多时这一段很长),读到的是一句固定的提示。改由决策模型在你发消息时,从本会话的 skill 里挑出相关的几个,连同名字、描述和相关度附在消息后面交给主 agent;只能由你触发的 skill 不推荐给主 agent,只在看板上提示你(「可试 /x」)。一轮进行中,主 agent 还可以用 find_skill 工具按几个词查 skill。skill 本身和 Skill 工具都不变,主 agent 仍然可以按名字加载任何 skill。
  • 失败时放行。 决策模型超时(timeoutMs)、出错、回答无法解析,或者没有配密钥时,消息照常进入,不额外等待,这一轮用会话自己的 effort,看板上写明原因,并弹一个 toast。

本文提到的配置项(例如 escalateAfter、timeoutMs),默认值都在「配置」的表里。每项功能都可以用 /dp 单独关掉,也可以整个 mod 一起关(见「控制:/dp」)。完整的行为规则(各种优先级、各种失败情形、看板和日志的写法)见 DEVELOPMENT.md 的「它做什么」。

模型和 effort 的依据

模型的分工、往上取一档、中途的门槛和各模型的 effort 下限,都按 Artificial Analysis(AA)Intelligence Index v4.3.2 的十个分项校正过。几个要点:Sonnet 5.5 的分数随 effort 掉得很快(Terminal-Bench:low 20.7%、medium 29.8%、high 43.9%、max 63.6%;指数 36、41、47、56),所以低估 sonnet 的代价大;Opus 5.5 在 medium 就有 51;Sonnet 与 Opus 在终端、自动化、知识工作上持平或略高,拉开的是事实知识(Omniscience 32 对 46)、难推理(HLE 差 6.4)和科学代码(SciCode 差 5.9);Haiku 4.5 在 Terminal-Bench 上是 0%。数字、来源和已存评测回答按新规则离线重算的结果见 docs/research/aa-benchmarks-2026-10.md。

看板

看板是 Dispatch Pilot 常驻的界面,只显示读数和结论,依据放在依据面板里。给人看的文字都是中文,模型名和 effort 档名照 /model、/effort 的写法。

  • prompt 上方的 band。 一轮进行中,顶上一条状态条:第几轮、用时、agent 的运行 / 完成 / 失败 / 排队数、Workflow 的进度、中途重判的次数。下面每个 agent 一行:主 agent 在最前,其余按开始的先后,各带数字键(0 是主 agent,1 到 9 按先后)、名字、模型和 effort、状态,以及一条时间色带(同一根时间轴,看得出谁和谁并行)。状态格写运行或完成了多久、失败、排队,或者「未路由 · 原因」。再下面是这一轮的事件流:每个决定、中途重判和强制升档(一次改档只算一条,写从哪档到哪档)、skill 推荐和「可试 /x」、find_skill 的查询,以及请求失败、回答迟到。一轮结束后 band 折成一行:主 agent 的模型和 effort、这一档是怎么来的、「可试 /x」。引擎给 band 的行数不到 4 行时,退成一行摘要。
  • 脚部右端的摘要。 连同和左边隔开的一列,不超过 12 列:状态符号、主 agent 的模型和 effort、+N 个运行中的 agent;放不下时 effort 先写短(xhi),再省掉模型。
  • 未路由。 一轮或一个 agent 的模型和 effort 照引擎原样发出、没有经过 Dispatch Pilot 的决定,看板上总写原因:决策模型超时、连不上、繁忙、被限速、额度用完、密钥被拒绝、没配好、回答读不懂,或者功能已关。路由失败时还会弹一个 toast,写谁没路由和原因。
  • 选中一个 agent。 在空的 prompt 里按它的数字键,依据面板就打开在它的卡片上。
  • 别的端。 Desktop 等不支持时间色带和概率条的端画同样内容的纯文字版:band 的每行没有时间色带,脚部是一个不超过 24 个字符的标签(同样连隔开的那一列算在内)。没在 Desktop 上亲眼看过。
  • 关掉的功能(/dp <功能> off)的决定、事件和计数不出现在 band 和脚部;/dp off 时 band 上没有 Dispatch Pilot 的内容,脚部写 ○ dp 已关。claude -p 没有界面,内容写进 debug log。

依据面板

/dp 打开依据面板,再输一次或按 Esc 关掉;/dp log 也是打开它。全屏且终端够宽时它停靠在对话右边(约 75 列),否则在 prompt 上方。面板窄时文字换行缩进,不截断。

  • 顶部(灰字):每项功能的开关状态(名字加「开」或「关」,关着的用灰色,默认关的 hook-block-failures 也在内)、/dp lock 的锁定,以及本会话 skill 画像的情况(skill 画像:保留 52 · 新写 3 · 失败 1;写的过程中是「生成中 2/5」;停写时写原因;有失败时按 f 列出失败的 skill 和原因)。
  • 依据卡片。 「‹ 上一个 (p) / 下一个 › (n)」按 band 的顺序翻看 agent。卡片是选中 agent 的依据:状态;没路由时几个字的原因,再写失败的类型、决策模型、细节和去 debug log 哪里看;给它的模型(派出 agent 和 Workflow 里的 agent)和理由(主 agent 也有);各档 effort 的概率;规则推演,从决策模型给的概率走到最终的 effort,每一步写它做了什么(最可能的一档、max 门槛、上取一档、模型下限、下限或强制升档);结果。主 agent 的卡片还注明发消息时的置信度只记录、不参与选档,并列出这一轮的强制升档和每次中途重判(建议档、当前档、带门槛刻度的置信条,和结论:升档、降档、降档被拦、防抖中还差几步、保持)。卡片只画决定时存下的推演,从不重新计算。
  • 决策日志。 按轮分组,最新的一轮在前,每条带编号、结果(颜色和 ✔ · ⚠ ✘ 符号;每轮的标题按同样的符号计数)、模型和 effort、是哪项功能针对什么、各档概率和理由。每轮有一个字母键(从 a 起,跳过 p、n、f)折叠或展开,最新两轮默认展开。日志保留最近 20 轮、最多 300 条。
  • 面板的按键只在面板拿到键盘时有用(/dp 打开时会拿键盘;在 band 上按数字键打开的不抢键盘,可以接着按别的数字)。面板没被放出来时(比如接着的端不放面板)会立即关掉,并说明原因(/dp 在回复里说,按数字键打开的弹一个 toast),这时用 /dp log 10 在对话里看。

发给决策模型的内容

  • 你这条消息,加上之前最近的几条消息(最多 contextMessages 条,总长度不超过 contextTokens)。主 agent 的 effort 题单独发一个请求,能读到的最近对话比和 skill 推荐一起发时长得多(Jev 默认 24000 token,对 6000);skill 推荐的请求和它并行发出,不增加等待。每条只有文字和调用过的工具名,不包含文件内容和工具输出。一轮中途重判时发的是这一轮最近几步(rejudgeSteps)的摘要:主 agent 写的文字、调用的工具和一句话结果,同样不含文件内容、写入的内容和工具输出。
  • 问题摘要(写好之后,发消息时的 effort 请求里):不超过 500 token,是 summaryModel 对你和主 agent 这个问题的转述,输入给它的内容先脱敏;它占 state 的预算,最近的对话让给它。
  • 派出 agent 和 Workflow 里的 agent:主 agent 写给它的任务(prompt)、描述、agent 类型,加上你这一轮说的话。
  • 打开 skill 推荐时:本会话每个 skill 的名字和画像(还没有画像的用描述),以及排在前面的几个 skill 的描述、画像和 SKILL.md 开头约 700 个字符。skill 画像由你自己的 Claude 登录写(模型是 skillsProfileModel),用你的用量,不经过决策模型的提供方。
  • 发送前对常见的 secret 格式脱敏,替换成 [REDACTED]:各家的 API key 和 token、password=... 这类赋值、URL 里的密码、私钥和 JWT。
  • 请求发往 TypeSafe(api.typesafe.ai)。
  • 每次请求的结果和每个决定都写进 debug log(claude --debug-file <路径>),不进入会话。

开销

  • 决策请求。 评测里 800 个 effort 请求共 614,292 input token(平均约 770 个),Jev 约 0.026 美元。评测的上下文很短;Jev 的 state 现在最多约 6.7k token,一个这样的 effort 请求不到 0.0003 美元(Jev 只按输入计费,每百万 token 0.042 美元),带 skill 推荐的约 2.9 万 token,约 0.0012 美元。
  • skill 推荐。 打开后,每条消息的第一个请求还带着本会话每个 skill 的名字和画像:111 个 skill 都写好画像时约 2.19 万 input token(不带画像约 8.6k);第一段分到 0.1 以上的 skill 才会发第二个请求,约 1.3k。作为交换,隐藏 skill 列表每个会话省下约 6.6k input token(本机 66 个 skill 时实测),换成的提示只有 360 个字符。
  • skill 画像用你自己的 Claude 登录写(模型是 skillsProfileModel),算在你的用量里:每份约 2k 输入和 200 输出 token,每个 SKILL.md 版本只写一次,每次会话开始最多写 skillsProfilesPerSession 份。
  • 问题摘要同样用你自己的 Claude 登录写(模型是 summaryModel),算在你的用量里:你本人的消息开始的每一轮结束后一次,输入是上一份摘要加这一轮(你的话、主 agent 的最终回复、工具汇总,各自截断),输出不超过 500 token。
  • claude plugin details dispatch-pilot@alex-mods(2.1.289)显示 0 个组件、常驻开销约 0 token:它看不到 mod 在运行时附加和替换的内容。

要求

  • Claude Code 2.1.287 及以上,mod 在 Claude Code 里默认启用。0.4.0 测试用的是 Claude Code 2.1.294(默认的 pplx 配置跑通一轮真实流程,见 #53),此前版本用的是 2.1.291(订阅登录,看板和依据面板在终端里看过)和 jev-1.13.0;缓存和评测的实测是在 2.1.289 上做的;跑 eval/ 和 scripts/ 里的 Node 脚本用的是 Node 26.5,用 mod 本身不需要 Node。
  • 只支持 Claude Code 订阅(ADR 0001,docs/adr/0001-main-agent-effort-only-no-model-switch.md)。Dispatch Pilot 每轮、每一步都改主 agent 的 effort,但从不改它的模型:换模型必然让 prompt cache 失效,而在订阅下,同一个模型内切换 effort 保留缓存(2.1.289 上 Opus 5.5 加订阅实测)。Bedrock、Vertex 和各种网关上切换 effort 会让缓存失效,不在支持范围内。Claude Code 的文档只点名了 Opus 5.5、Sonnet 5.5 和 Fable 5.1 保留缓存,其他大多数模型上每档 effort 各有一份缓存,切换会重算整段请求(见 docs/research/decision-models-and-caching.md 的 3.4)。
  • 一个决策模型的账号。 默认的决策模型是 pplx,要 Perplexity 的 API key(Decisions API);Jev 要 TypeSafe 的 API key。没有 Perplexity 的 key、但有 TypeSafe 的 key 时,自动退回 Jev(0.3.1 的老用户升级后不会失去路由,依据面板的决策日志里会有一条说明原因);两种 key 都没有时 mod 不发任何请求,每一轮都按会话自己的 effort 走,看板上写明缺的是 Perplexity 的 key。

安装

claude plugin marketplace add alexcz-a11y/claude-mods
claude plugin install dispatch-pilot@alex-mods --scope user

想改 mod 或者试未发布的改动时,也可以从本地克隆添加:claude plugin marketplace add ./claude-mods(相对路径要以 ./ 或 ../ 开头,否则会被当成 GitHub 仓库)。从本地目录添加的 marketplace,Claude Code 直接从那个目录加载 mod,克隆里检出的是哪个版本,用的就是哪个版本。

在 shell 里装好的 mod,下次启动 Claude Code 时才加载:装好后重启 Claude Code,或者在已经开着的会话里运行 /reload-plugins。在会话里运行 /plugin,看到 1 mod active · dispatch-pilot 说明 mod 已经加载;运行 /dp 会打开依据面板,/dp status 列出各项功能的开关。接着给它一个决策模型的密钥。

pplx(默认)。 给它 Perplexity 的 API key,有两种给法,同时给了就用第 1 种:

  1. 在 Claude Code 的会话里运行 /plugin configure dispatch-pilot@alex-mods,在弹出的配置对话框里填 perplexityApiKey;输入会被遮住。在会话里用 /plugin 安装时也会弹出这个对话框;在 shell 里用 claude plugin install 安装不会弹,装好后用这条命令,或者用 stdin 传进去(需要 jq):
   read -rs PERPLEXITY_API_KEY && export PERPLEXITY_API_KEY   # 粘贴密钥后回车,不回显,值也不会留在 shell 历史里
   jq -n '{perplexityApiKey: env.PERPLEXITY_API_KEY}' | claude plugin configure dispatch-pilot@alex-mods --values-stdin
  1. 设环境变量 PERPLEXITY_API_KEY(启动 Claude Code 的环境里要有)。Claude Code Desktop 读不到敏感的 userConfig(#35),用环境变量就绕开了这个问题。

不要用 claude plugin install --config perplexityApiKey=... 传密钥,也不要用 jq --arg:它们的值会出现在进程参数里,同一台机器上的 ps 看得到。

--values-stdin 读一个 JSON 对象,值都是单行字符串,没写到的选项保持原值。保存之后要重启 Claude Code 才生效,命令也会提示 Configuration saved. Restart Claude Code to apply it.。不带参数运行 claude plugin configure dispatch-pilot@alex-mods 会列出所有选项,并标出哪些还没有设置。

key 只放在请求头里,不会出现在 debug log、看板、依据面板和 $.state 里;会话开始时 debug log 只写 key 从哪里来(选项还是环境变量),不写 key 本身。

Jev。 想用 Jev,就把 decisionModel 设成 jev(在 /config 里选,或写进 pluginConfigs),再填 TypeSafe 的 API key typesafeApiKey,填法同上(配置对话框,或 stdin 的 JSON 里写 typesafeApiKey;环境变量 TYPESAFE_API_KEY 只用来放进 stdin,mod 本身不读它)。设成 jev 就始终用 Jev,哪怕也有 Perplexity 的 key。

只有 TypeSafe 的 key 的老用户不用改任何配置:decisionModel 没设(或设的是 pplx、已移除的 clef、拼错的值)又没有 Perplexity 的 key 时,用 Jev,每个会话的决策日志(/dp 打开的依据面板)里有一条「改用 Jev」写明原因。以后想换成 pplx,只要填上 Perplexity 的 key。

decisionModel 在 /config 里是下拉选择,有 pplx(默认)和 jev 两项;填了别的值(包括已移除的 clef),按没设处理,也就是默认的 pplx(没有 Perplexity 的 key 时按上面的规则退回 Jev)。

更新

从 GitHub 添加 marketplace 的,在 shell 里运行:

claude plugin marketplace update alex-mods
claude plugin update dispatch-pilot@alex-mods

第一条刷新 marketplace 的列表,第二条更新 mod。更新之后重启 Claude Code,或者在开着的会话里运行 /reload-plugins。

claude plugin update 看的是版本号(.claude-plugin/plugin.json 的 version):版本号和你装的一样,它就回答已经是最新版本(is already at the latest version),不换掉本机的副本,哪怕分支上已经有新的提交。新的版本会升版本号。这个 marketplace 默认不自动更新;要打开,在会话里运行 /plugin,到 Marketplaces 里选 alex-mods,再选 Enable auto-update。自动更新同样只在版本号变了时才换。

从本地克隆安装的,不用 claude plugin update:在克隆里取新的提交(例如 git -C claude-mods pull),然后重启 Claude Code 或运行 /reload-plugins。Claude Code 直接从克隆加载 mod,不看版本号。

从 jev-pilot 切换

Dispatch Pilot 和 jev-pilot 不能共存:两者都在 turn.step 上改主 agent 的 effort,会互相覆盖。已经装了 jev-pilot 的话,按这个顺序切换:

  1. 先在 jev-pilot 还启用的会话里运行 /jev-pilot:setup restore。 它把 /jev-pilot:setup 改过的 skill 设置(settings 里的 skillOverrides)恢复成原样,并删掉备份。这条命令是 jev-pilot 自己的,只有它加载着才能用,所以要在停用它之前运行。不恢复的话,那些 skill 会一直对主 agent 隐藏,Dispatch Pilot 也推荐不了它们。没有运行过 /jev-pilot:setup 的话,这一步可以跳过。
  2. 停用 jev-pilot:claude plugin disable jev-pilot@jev-pilot。
  3. 重启 Claude Code:已经开着的会话还带着 jev-pilot,要重启。
  4. 之后只用 claude 启动,不要用 claude-jev。 jev-pilot 的 marketplace 安装停用之后,claude-jev 启动器会改用 --plugin-dir 加载它在 ~/.claude/plugins/marketplaces/jev-pilot 里的本地副本,jev-pilot 又被加载进来(还会起它自己的 router),两者就又撞在一起了。

配置

选项的值存在两个地方:

  • 敏感项(typesafeApiKey、perplexityApiKey)存进平台的安全凭据存储(Claude Code 文档的说法),输入时被遮住,不在 /config 里出现。用配置对话框(/plugin configure dispatch-pilot@alex-mods),或者 claude plugin configure ... --values-stdin 填,见「安装」。
  • 其他选项存在 user settings 的 pluginConfigs 下,在 /config 面板里一项一行,可以直接改(需要 Claude Code 2.1.269 及以上)。两个列表项 skillsAlwaysListed 和 skillsNeverSuggested 不在 /config 里出现,要在 ~/.claude/settings.json 里写成字符串数组。项目和本地 settings 里的 pluginConfigs 会被 Claude Code 忽略。
{
  "pluginConfigs": {
    "dispatch-pilot@alex-mods": {
      "options": { "timeoutMs": 2500, "skillsAlwaysListed": ["anthropic-skills:pdf"] }
    }
  }
}

选项留空就取决策模型的默认值:timeoutMs、contextMessages、contextTokens、thetaMax、thetaUp、thetaDown、rejudgeSteps、rejudgeWaitMs、thetaExpected、agentOverride、skillsMinRelevance、findSkillMinRelevance 这 12 项在 manifest 里没有默认值,所以 /config 里显示为空,你不设时 Dispatch Pilot 取表里你选的决策模型的那一列。你自己设了某一项,就用你设的值。会话开始时 debug log 会写一行哪些选项用了默认值。

另有几个值也随决策模型而定,但不是选项,/config 里改不了,也不在下面的表里:往上取一档的门槛(高一档的概率达到它,就往上取一档;Jev 是 0.3,pplx 是 0.45),问题怎么问(语言和题型:Jev 发消息时判 effort 的那一题用中文,其余所有问题用英文;pplx 所有问题都用英文;都是 Score),以及 contextMessages 的上限(Jev 是 32,pplx 是 2000)。它们和上面那些默认值一起写在 core/setup.ts 的 BACKEND_DEFAULTS,评测和 mod 读的是同一张表。依据卡片的规则推演里,「上取一档」一步写的就是这里的门槛。

pplx 的默认值(B′,ADR 0006,docs/adr/0006-dispatch-pilot-pplx-default-decision-model.md):pplx 的窗口大得多,不受 Jev 那 32k 的限制,所以发消息时的 effort 题、中途重判、派出 agent 和 Workflow 的 state 各取 48000 token、最近 2000 条消息(由 token 预算截断,不是条数),skill 的两段排序(它们带着每个 skill 的画像)仍是 6000,因为 B′ 没有评测过 skill 题。它回答要几秒(eval v2 里 48000 token 的请求约 5 秒),所以发消息等 8000 毫秒(hook 自己的上限是 10 秒),中途重判和 find_skill 各等 6000 毫秒。所有问题都用英文问。thetaMax 0.47、thetaUp 0、thetaDown 0.55 和往上取一档的 0.45 是按已存的 pplx 回答离线定的(DEVELOPMENT.md 的「eval v2 的结果」):0.45 让日常消息判高从 18.5% 降到 13.0%,max 的召回不变。thetaExpected、agentOverride 和两个相关度门槛没有为 pplx 校准过,沿用 Jev 的值。用户自己设的值总是优先。

Jev 的默认值是按 Jev 的上下文上限配置的。 Jev 一个请求最多收 64k token,其中 state 加上最长的那一道题不能超过 32k。最长的题是发消息时 skill 推荐的第一段(111 个 skill 带画像时 Jev 计约 2.19 万 token),所以带 skill 题的请求(发消息时打开了 skill 推荐,以及 find_skill 的第一段)的 contextTokens 取 6000:state 最多约 6.7k(Jev 的计数),加上那一题约 28.6k,在 32k 的 90% 以内,整个请求在 64k 之内,而且这个值正好让每个 skill 的画像不被裁短。其余的请求没有这么长的题(最长的不到 700 token),所以按种类各取一个更大的值:关着 skill 推荐的发消息、中途重判(包括卡住时的)、派出 agent、一批 Workflow 调用,state 最多 24000(Jev 的计数约 26.7k,加上最长的题仍在 32k 的 90% 以内)。你自己设了 contextTokens,所有种类都用它,但每个种类各取它和自己的上限里较小的一个(设 16000 的话,带 skill 题的请求仍是 6000,其余是 16000)。contextMessages 和 rejudgeSteps 取 manifest 允许的最大值(32 条、16 步),让 token 预算而不是条数决定发多少,旧的消息、步骤放不下就整条丢掉。仍然只发文字和工具名,不发工具的输入和输出。这样配置是为了让 Jev 看到尽量多的信息;它没有评测过:评测用的是 4 条、2000 个 token、4 步,更多上下文是否真的判得更准没有数据,延迟见「局限和待评测」。计算过程见 DEVELOPMENT.md 的「Jev 的上下文默认值怎么算」。

下表的 Jev 和 pplx 两列是各自的默认值。起点:这个默认值是暂定的起点,还没有评测数据支持。按 Jev 的上限:按上面的算法取 Jev 能接受的最大值,没有评测。按 AA 基准:按 Artificial Analysis 的基准校正过(见「模型和 effort 的依据」),没有在这套评测集上量过。没有标注的,是按测到的数据定的,或者本来就不需要校准。

决策模型和密钥

选项作用Jevpplx
decisionModel决策模型,在 /config 里是下拉选择:pplx(Perplexity 的 pplx-decider-v1.1-27b)或 jev(TypeSafe 的 Jev)。默认 pplx;填了别的值(包括已移除的 clef),按没设处理。设成 jev 就始终用 Jev;其他情况有 Perplexity 的 key 用 pplx,没有但有 TypeSafe 的 key 就退回 Jev(决策日志里记一条原因),两种都没有就不发请求pplxpplx
typesafeApiKeyTypeSafe 的 API key,选 jev 时用;没有 Perplexity 的 key 时默认的 pplx 退回 Jev,用的也是它。敏感项空空
perplexityApiKeyPerplexity 的 API key,默认的决策模型 pplx 用它。敏感项。为空时读环境变量 PERPLEXITY_API_KEY,两处都有就用这里的;Perplexity 和 TypeSafe 的 key 都没有时不发任何请求。key 只放在请求头里,不进日志、看板和 $.state空空
pplxQps每秒最多向 Perplexity 发几个请求(1–50),只对 pplx 有用,Jev 不受它限制。Tier 0 账户限 1 QPS,一条消息又发 2 到 3 个请求,所以超出的在 mod 里排队,发消息时的 effort 题最先发,排队的时间算进该请求自己的等待。遇到 429 时,剩余等待时间不少于 Retry-After 加约 5 秒就按 Retry-After 重试一次,不够就不经路由地放行,看板写「被限速」。升级了账户就调高1 仅 pplx 用1

等多久,读多少

选项作用Jevpplx
timeoutMs一条消息等决策模型的最长时间(200–8000 毫秒),超过就不经路由地放行15008000
contextMessages随你的消息一起发的最近消息条数,从 0 到决策模型能接受的条数(Jev 是 32,pplx 是 2000)。两个模型都默认取上限,由 contextTokens 决定实际发多少32 按 Jev 的上限2000
contextTokens发给决策模型的 state 的 token 预算(100–16000,pplx 到 48000):你的消息加上最近的消息,按发出去的样子数(连同字段名和转义),中英文按同一个尺度数,旧消息先丢。按请求的种类取:默认值带 skill 题的请求 6000,其余 Jev 24000、pplx 48000,包括发消息时单独发的 effort 请求(见上面);你设了就每个种类取它和自己上限里较小的6000 按 Jev 的上限,其余种类 240006000,其余种类 48000

一轮中的 effort

选项作用Jevpplx
thetaMax用 max 所需的最低概率(0–1);发消息时、中途重判和派出 agent 的 effort 都用这个门槛0.5 起点0.47 按已存的评测回答校准
rejudgeEvery一轮进行中每隔几步重判一次(0–50);0 表示不按步数重判,派出 agent、启动 Workflow、加载 skill 时仍会重判3 起点3 起点
rejudgeSteps重判时决策模型读到的最近步数(1–16)。Jev 取上限,由 contextTokens 决定实际发多少16 按 Jev 的上限16 同 Jev
rejudgeWaitMs重判的回答还没到时,下一步最多再等多久(0–8000 毫秒),然后沿用原来的 effort3006000
thetaUp中途升档所需的最低置信度(0–1)0.3 按 AA 基准0 按已存的评测回答校准
thetaDown中途降档所需的最低置信度(0–1,低于 thetaUp 时按 thetaUp 算),一次只降一档0.55 按已存的评测回答扫出0.55 同 Jev
holdSteps升档之后多少步之内不降档(0–50)5 按 AA 基准5 按 AA 基准

卡住时强制升档

选项作用Jevpplx
escalateAfter计入的工具调用失败满几次就问决策模型并升档(1–20)2 起点2 起点
escalateModeone-level 升一档,最高到 xhigh(决策模型自己有把握给更高时可以更高);max 直接升到 maxone-level 起点one-level 起点
escalateLimit一轮(或一个派出 agent)最多升几次(0–10);每次升档后失败计数清零2 起点2 起点
thetaExpected决策模型认为这些失败是预期内失败的概率达到多少就不升档(0–1)0.25 起点0.25 沿用 Jev 的起点
escalateHaikuTo失败的 haiku agent 接着用哪个模型做:sonnet、opus、fable 或完整的模型 id;留空表示不换。你为这个 agent 点名的模型和排除的模型优先sonnetsonnet

派出 agent 和 Workflow

选项作用Jevpplx
agentFable打开后,决策模型可以为派出 agent(包括 Workflow 里的)选 fable,fable 比 opus 贵;你自己点名 fable 时不受这个开关限制falsefalse
agentOverride主 agent 为派出的 agent 指定了模型时,决策模型的选择要达到这个置信度(0–1)才推翻它;Workflow 脚本里写了 model 时同样适用0.6 起点0.6 沿用 Jev 的起点
workflowModerewrite:把决定写进脚本再运行;return:第一次提交被拒绝并附上逐个 agent 的推荐,让主 agent 自己写进去,同一个 Workflow 第二次提交直接放行rewriterewrite

skill

选项作用Jevpplx
skillsMax一条消息最多推荐几个 skill(0–10)33
skillsMinRelevance推荐一个 skill 所需的最低相关度(0–1):决策模型对「这个 skill 是否正好做这条消息要做的那种工作」回答「是」的概率0.750.75 沿用 Jev
skillsShortlist第二段补读 SKILL.md 开头、逐个判断的 skill 最多几个(1–10)4 起点4 起点
skillsProfileModel写 skill 画像的模型,写别名(haiku)或完整的模型 id。通过你的 Claude Code 登录调用,算在你的用量里;换了模型,所有画像重写haikuhaiku
skillsProfilesPerSession每次会话开始最多写几份还没有的画像(0–500);0 表示不写3030
skillsAlwaysListed一直留在主 agent 的 skill 列表里的 skill,写列表里的名字(同步来的 skill 带前缀,例如 anthropic-skills:pdf)。列表项,不在 /config 里空空
skillsNeverSuggested从不推荐给主 agent、也不提示你的 skill,写法同上。它们照常安装,Skill 工具仍能按名字加载,find_skill 也不返回它们。列表项,不在 /config 里空空
findSkillMaxfind_skill 一次最多返回几个 skill(1–10)55
findSkillMinRelevancefind_skill 返回一个 skill 所需的最低相关度(0–1),比推荐的门槛低:这是主 agent 主动问的,它会自己看描述再决定0.5 起点0.5 沿用 Jev 的起点

问题摘要

选项作用Jevpplx
unresolvedMaxAfter未解决次数达到它时,发消息时和中途重判时 effort 题的说明里多一条强提示(这项工作属于多次尝试都没解决的故障);0 表示不给,只带摘要和次数,最大 10。强提示不写档位,不改 pickEffort 和 thetaMax3 暂定3 暂定
summaryModel在你本人的消息开始的那一轮结束后,在后台续写问题摘要的模型,写别名(haiku)或完整的模型 id。通过你的 Claude Code 登录调用,算在你的用量里;摘要是对话的转述,所以 contextMessages 为 0 或 unresolved 开关关着(默认)时不写haikuhaiku

控制:/dp

/dp                    打开依据面板;已经打开时关掉它
/dp log                打开依据面板(老习惯的别名,不会关掉它)
/dp status             是否开启、effort 的锁定状态、各项功能的开关
/dp on | off           总开关。关闭后不发任何决策请求,每一步都按引擎原样发出(锁定也不生效),脚部写 ○ dp 已关
/dp <功能> on | off    单项功能的开关,例如 /dp main-effort off(功能名见 /dp status 的列表)
/dp lock <档位>        把主 agent 的 effort 锁在 low、medium、high、xhigh 或 max,这一轮的每一步和之后的每一轮都用它,优先于决策
/dp unlock             解除锁定(也可以写 /dp lock off)
/dp log N              在对话里列出最近 N 次决策和理由,最新的在最后(最多 300;日志保留最近 20 轮、最多 300 条)
  • 命令在一轮进行中也立即执行:锁定或解锁从下一步起生效。
  • 功能的开关是 main-effort、unresolved、midturn-effort、dispatched-agents、workflow-agents、workflow-labels、escalation、skills、skill-profiles、find-skill 和 signals;另有 hook-block-failures,打开后你自己的 settings hook 拦下的调用也算失败,参与强制升档。unresolved 和 hook-block-failures 默认关闭,其余默认开着。
  • 开关保存下来,下次启动会话时还是你离开时的样子;只保存和默认值不同的开关,所以一个新增的功能默认开着。
  • 锁定只在当前会话里有效,/dp unlock 或会话结束时解除。锁定期间决策照常进行并记录,不想为此等决策模型的话,用 /dp main-effort off。
  • 依据面板的决策日志和 /dp log N 显示每项功能记录的决策:做了什么决定、针对哪条消息、理由(决策模型给出的各档概率和置信度)。同样的内容也写进 debug log,不进入对话。
  • 信号(上下文占用、5 小时和 7 天限额的百分比、会话花费)每次变化时写一行进 debug log,只是记录,不参与任何决策;/dp signals off 可以停止记录。
  • 暂时不想用就 /dp off;彻底停用就 claude plugin disable dispatch-pilot@alex-mods。

局限和待评测

局限

  • 只支持 Claude Code 订阅,见「要求」。
  • 决策请求整体失败时,你排除的模型仍可能按主 agent 的指定启动。 哪些模型被排除,是决策模型从你的话里读出来的,没有它的回答就无从知道;mod 不改用关键词去猜,免得把只是提到某个模型的话当成排除。这是有意的取舍。
  • fork 出来的 agent(它总是用父 agent 的模型)和 agent team 的 teammate(它会长期存在、处理很多任务,派出时的一次判断看不到这些任务)不处理。
  • 主 agent 在推荐漏掉时不会自己去用 find_skill。 一条要写 PR 描述、却没有推荐 pr 的消息,有提示、没有提示、提示写得更主动,主 agent 都直接写了正文;被明确要求查时,它能找到 find_skill 并用上。所以漏掉的推荐,目前只靠发消息时的推荐。
  • Jev 对同一个 key 的并发请求像是依次处理。Workflow 里 prompt 是数据的调用(fan-out)在 agent 启动时当场判断,几个 agent 几毫秒内一起启动,排在后面的可能超时,那些 agent 按引擎原样启动。
  • Jev 的请求比以前大,延迟会多一点。 默认值把 Jev 的 state 放到最多约 6.7k token(带 skill 题的请求,contextTokens 6000),发消息时带 skill 推荐的请求最多约 2.9 万 token;其余种类的 state 最多 24000(约 2.67 万 token),中途重判、派出 agent 和关着 skill 推荐的发消息请求,满了的话比以前多约 2.4 万 token,按每 1k 约 13 毫秒外推,最多多约 0.3 秒(同样是外推,没有量过)。已有的实测:Jev 处理 2.19 万 token(skill 第一段,带全部画像)时 p50 约 560 毫秒、p90 约 615 毫秒;另一次整体偏慢的运行里,这一段 218 条中有 24 条(约 11%)超过 1500 毫秒,这些消息的决策整个超时,effort 也没有经过路由。按每多 1k token 约多 13 毫秒外推,现在的 state 再多约 80 毫秒,慢的时段超过 1500 毫秒的消息会比 11% 更多;这是外推,没有在新默认值上量过。如果看板上「决策模型超时」变多,可以把 contextTokens 调小(每少 1k 约快 13 毫秒),或者把 timeoutMs 调大;skill 那一题本身有 2.2 万 token,想去掉这部分延迟只能 /dp skills off。
  • 大多数默认值是暂定的起点(表里标了 起点):评测数据只够定下少数几项,其余的见下面的「还没有数据的事」。

问题用什么语言写。 选 pplx 时所有问题都用英文写(eval v2:pplx 英文问法更准)。选 Jev 时,发消息时判断主 agent effort 的那一个问题用中文写;其余问题(一轮中途重判、派出 agent、Workflow 里的 agent、卡住时的强制升档、skill 推荐和 find_skill)都用英文。依据是 effort-submit 在现在的问法上的对比(Jev,各 1 次):用中文问,中文题 85.0%、英文题 89.0%;用英文问,79.0%、78.0%。同样的请求之前跑过 3 次,中文问法也都高约 8 个百分点。其余问题在现在的问法上没有中文问法的数据,所以都没有改。派出 agent 和 skill 的评测已经有中文问法的变体(models-hint-zh、profiles-zh),还没有运行。发消息时的请求里,中文的 effort 问题和英文的 skill 问题放在一起,这种混合的请求没有单独评测过。

中文和英文的差距。 门槛是中文准确率比英文低不超过 4 个百分点,正好低 4 个百分点也算通过。现在的配置(Jev,effort 用中文问)在 effort-submit 上那一次正好差 −4.0,同样的请求之前 3 次是 0、0、−1;而中文题的准确率比用英文问时高 6 个百分点。其余三套评测都在门槛之内(见「评测」),但都是改措辞之前测的。

派出 agent 的模型文字按 AA 的基准重写,并用真实的 Jev 调过。 重写后的第一次评测(models-hint,100 题中英各一遍)整体从 70%/68% 掉到 60%/58%(模型部分 82/81,effort 部分 64/62)。之后调了三轮文字和 sonnet 的 effort 下限(sonnet 从 high 降到 medium),最后一版是整体 69%/68%、模型部分 89%/91%、effort 部分 71%/71%,同一版重复一次是 69%/67%、89%/90%、71%/71%。和重写之前(70%/68%,85%/85%,74%/72%)比,模型部分更好,effort 部分低 3 到 1 个百分点:往上取一档和模型下限让偏高多了、gold 命中从 50%/48% 变成 44%/44%,因为数据集的 gold 是按「够用的最便宜档」标的,没有按 AA 重标。注意这三轮文字是对着这 100 题的错题调的,所以最后的模型部分是在调过的题上量的,对没见过的请求会低一些。逐轮的表、发消息和中途重判的真实对照在 [docs/research/aa-benchmarks-2026-10.md](../docs/research/aa-benchmarks-2026

Source 52 files
hooks/dispatch-pilot.ts 54 lines
1// Dispatch Pilot's entry: it only assembles. Each feature lives in its own
2// file under features/ and registers its own hooks; the core goes last.
3//
4// Order is nesting: what registers first is outermost and sees an event
5// first. Features go above the core so that on every event they share, each
6// feature has run before the core finishes the event (sends the decision
7// request, writes the step). To add a feature: one file, one line here,
8// above `registerCore` (DEVELOPMENT.md, 开发).
9
10import type { Register } from 'claude-code'
11import { registerScreens } from './board/screens.tsx'
12import { registerCore } from './core/core.ts'
13import { registerReport } from './core/report.ts'
14import { setup } from './core/setup.ts'
15import { registerControl } from './features/control.ts'
16import { registerDispatchedAgents } from './features/dispatched-agents.ts'
17import { registerEscalation } from './features/escalation.ts'
18import { registerFindSkill } from './features/find-skill.ts'
19import { registerMainEffort } from './features/main-effort.ts'
20import { registerMidturnEffort } from './features/midturn-effort.ts'
21import { registerUnresolved } from './features/unresolved.ts'
22import { registerSkillProfiles } from './features/skill-profiles.ts'
23import { registerSkills } from './features/skills.ts'
24import { registerWorkflowAgents } from './features/workflow-agents.ts'
25import { registerWorkflowLabels } from './features/workflow-labels.ts'
26
27export const register: Register = (on, options) => {
28  const ctx = setup(options)
29
30  // The decision report keeps the board's turn count: outermost, so a turn is counted before anything beneath reads the board.
31  registerReport(on)
32  // The screens draw the report's data (the band, the footer); no feature touches them.
33  registerScreens(on)
34
35  // Features, outermost first.
36  registerControl(on, ctx)
37  registerMainEffort(on, ctx)
38  registerUnresolved(on, ctx)
39  registerEscalation(on, ctx) // above the mid-turn feature: its raise is in the plan before a mid-turn answer is applied
40  registerMidturnEffort(on, ctx)
41  registerDispatchedAgents(on, ctx)
42  registerWorkflowAgents(on, ctx)
43  // Inside workflow-agents: a Workflow call is this feature's only once that feature has sent the script on, so the
44  // agents of other runs never wait while it is still deciding.
45  registerWorkflowLabels(on, ctx)
46  // Outside skills: at session start it writes the profiles of the skills that feature has just read.
47  registerSkillProfiles(on, ctx)
48  registerSkills(on, ctx)
49  registerFindSkill(on, ctx)
50
51  // The core: always last.
52  registerCore(on, ctx)
53}
54
hooks/board/screens.tsx 178 lines
1// The screens' hooks (ADR 0004: drawn from the decision report's data, never
2// written to by a feature): the band above the prompt (`AbovePrompt`), the
3// summary at the right end of the footer (`SessionMode`) and the rationale pane
4// (`Pane`, the pane `/dp` opens). One hook tree for every surface, branching on
5// `e.surface`: the terminal draws the full design (band.tsx, footer.tsx,
6// pane.tsx) with its Raster; every other surface gets the same trees in plain
7// text, without a Raster (the tag in the footer in at most 24 characters, the
8// log in a window short of Desktop's 2000 nodes), until the Svg versions.
9//
10// Each hook reads the board, the log and the agent the person picked from
11// $.state while it draws, so the host draws it again when they change; while
12// an agent runs it asks for one more frame every `TICK_MS` (the spinner, the
13// running clocks and ribbons). The band and the footer keep what other mods
14// draw in the same place: the engine's drawing beneath goes in first (spec
15// story 38); the pane is Dispatch Pilot's own.
16
17import type { EngineInterface, On, Timer } from 'claude-code'
18import { errorText } from '../decision/backend.ts'
19import { report, type NoticeIo } from '../core/report.ts'
20import { isOn, isShown, listSwitches, masterOn } from '../core/switches.ts'
21import { bandTree } from './band.tsx'
22import { FOOTER_COLUMNS, footerTree, TAG_CHARS } from './footer.tsx'
23import { paneTree } from './pane.tsx'
24import { nodeCount, showsText } from './kit.tsx'
25import { NO_PANE_STATE, PANE_COLUMNS, PANE_ID, PANE_TITLE, withFold } from './rationale.ts'
26import { screenView, TICK_MS, type AgentRow, type ScreenView } from './view.ts'
27
28const BOARD = { plugin: 'dispatch-pilot', key: 'board' } as const
29const DECISIONS = { plugin: 'dispatch-pilot', key: 'decisionLog' } as const
30const SELECTED = { plugin: 'dispatch-pilot', key: 'selected' } as const
31const PROFILES = { plugin: 'dispatch-pilot', key: 'skillProfiles' } as const
32const LOCK = { plugin: 'dispatch-pilot', key: 'lock' } as const
33const PANE_VIEW = { plugin: 'dispatch-pilot', key: 'paneView' } as const
34const COUNT = { plugin: 'dispatch-pilot', key: 'unresolved' } as const
35
36/** The log's windows for a surface that refuses a large tree, widest first: what the pane keeps of the log and of the re-decisions. */
37const WINDOWS = [
38  { entries: 40, mids: 20 },
39  { entries: 16, mids: 8 },
40  { entries: 4, mids: 3 },
41  { entries: 0, mids: 0 },
42] as const
43/** A tree this big is drawn again in the next window: Desktop refuses 2000. */
44const MOST_NODES = 1800
45
46/** The next frame asked for while an agent runs (module state: a hot reload cancels it with the old module, and the next draw asks again). */
47let ticking: Timer | null = null
48
49/** The screens' view, read from $.state now (a read while drawing subscribes the drawing to the value). */
50async function viewOf($: EngineInterface, isWorking: boolean): Promise<ScreenView> {
51  const [board, log, picked, now] = await Promise.all([$.state.get(BOARD), $.state.get(DECISIONS), $.state.get(SELECTED), $.clock.now()])
52  return screenView({
53    board: board.value ?? { turn: 0, nodes: [] },
54    log: log.value ?? [],
55    now,
56    selected: picked.value ?? null,
57    master: masterOn(),
58    isOn,
59    isShown,
60    isWorking,
61  })
62}
63
64/** One more frame in `TICK_MS` while something runs; none once nothing does. */
65function tick($: EngineInterface, view: ScreenView): void {
66  const moving = view.live && (view.counts.running > 0 || view.rows.some((row) => row.node.state === 'running'))
67  if (!moving || ticking !== null) return
68  try {
69    ticking = $.clock.after(TICK_MS, () => {
70      ticking = null
71      $.ui.invalidate('ui.render')
72    })
73  } catch {
74    // no timer: the screens move on at the board's next change
75    ticking = null
76  }
77}
78
79/** The agent picked, its card shown in the rationale pane (written from a press, never while drawing). */
80async function pick($: EngineInterface, row: AgentRow): Promise<void> {
81  await $.state.set(SELECTED, { turn: row.node.turn, id: row.node.id })
82}
83
84export function registerScreens(on: On): void {
85  // The band above the prompt.
86  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
87    if (e.props.hasSurvey) return next(e)
88    const own = await next(e)
89    try {
90      const t = $.ui.resolve(e)
91      // The time ribbons are the terminal's Raster; elsewhere the rows are the same without them.
92      const ribbon = e.surface === 'terminal' ? $.ui.resolve(e).Raster : undefined
93      const view = await viewOf($, e.props.isWorking)
94      tick($, view)
95      // A digit picks an agent and brings up its card: a press is the person's asking, so the pane is placed at
96      // any width; it opens without taking the keys, so the next digit still picks. One not placed is closed again,
97      // and the decision report tells the person why in a toast.
98      const select = async (row: AgentRow) => {
99        await pick($, row)
100        const opened = await $.ui.open({ id: PANE_ID, title: PANE_TITLE, closeOnEscape: true, columns: PANE_COLUMNS })
101        if (!opened.isPlaced) {
102          await $.ui.close({ id: PANE_ID })
103          const io: NoticeIo = { debug: (line) => $.ui.log(line, { to: 'debug' }), now: () => $.clock.now(), toast: (text) => $.ui.toast(text) }
104          await report(io, { unplaced: { reason: opened.reason } })
105        }
106      }
107      const tree = bandTree(t, view, { cols: e.props.bodyColumns, rows: e.props.maxRows }, { select: (row) => void select(row).catch(() => undefined) }, ribbon)
108      if (tree === null) return own
109      const { Box } = t
110      return (
111        <Box flexDirection="column">
112          {own}
113          {tree}
114        </Box>
115      )
116    } catch (error) {
117      // A band that cannot be drawn leaves the place to the others.
118      $.ui.log(`band not drawn: ${errorText(error)}`, { to: 'debug' })
119      return own
120    }
121  })
122
123  // The right end of the footer, beside the engine's modes.
124  on('ui.render', { component: 'SessionMode' }, async ($, e, next) => {
125    const own = await next(e)
126    try {
127      const t = $.ui.resolve(e)
128      const view = await viewOf($, false)
129      tick($, view)
130      // A column between the engine's modes (or another mod's drawing) and the tag, when there are any: it counts in what the tag adds.
131      const gap = showsText(own, e.props.modes.length > 0) ? 1 : 0
132      const tree = footerTree(t, view, (e.surface === 'terminal' ? FOOTER_COLUMNS : TAG_CHARS) - gap)
133      if (tree === null) return own
134      const { Box } = t
135      return (
136        <Box gap={gap}>
137          {own}
138          {tree}
139        </Box>
140      )
141    } catch (error) {
142      $.ui.log(`footer not drawn: ${errorText(error)}`, { to: 'debug' })
143      return own
144    }
145  })
146
147  // The rationale pane (`/dp`): the card of the agent picked, the decision log.
148  on('ui.render', { component: 'Pane', requestId: PANE_ID }, async ($, e, next) => {
149    try {
150      const t = $.ui.resolve(e)
151      const [view, log, profiles, lock, kept, count] = await Promise.all([viewOf($, false), $.state.get(DECISIONS), $.state.get(PROFILES), $.state.get(LOCK), $.state.get(PANE_VIEW), $.state.get(COUNT)])
152      tick($, view)
153      const entries = log.value ?? []
154      const state = kept.value ?? NO_PANE_STATE
155      const shown = isOn('skills') && isOn('skill-profiles') ? (profiles.value ?? null) : null
156      // The summary the session holds, while the feature that writes it is on.
157      const summary = isOn('unresolved') ? (count.value?.summary ?? null) : null
158      const input = { view, log: entries, profiles: shown, switches: listSwitches().map((s) => ({ name: s.name, on: s.on })), master: masterOn(), lock: lock.value ?? null, summary, state }
159      const act = {
160        pick: (row: AgentRow) => void pick($, row).catch(() => undefined),
161        fold: (turn: number, open: boolean) => void $.state.set(PANE_VIEW, withFold(state, turn, open, entries)).catch(() => undefined),
162        failures: (open: boolean) => void $.state.set(PANE_VIEW, { ...state, failures: open }).catch(() => undefined),
163      }
164      if (e.surface === 'terminal') return paneTree(t, input, e.props.bodyColumns, act, $.ui.resolve(e).Raster)
165      // Elsewhere: the same pane in text, in the widest window of the log that stays short of Desktop's 2000 nodes.
166      let tree = paneTree(t, { ...input, window: WINDOWS[0] }, e.props.bodyColumns, act)
167      for (const window of WINDOWS.slice(1)) {
168        if (nodeCount(tree) < MOST_NODES) break
169        tree = paneTree(t, { ...input, window }, e.props.bodyColumns, act)
170      }
171      return tree
172    } catch (error) {
173      $.ui.log(`rationale pane not drawn: ${errorText(error)}`, { to: 'debug' })
174      return next(e)
175    }
176  })
177}
178
hooks/core/core.ts 182 lines
1// The core: the innermost hooks of every event it shares with the features.
2// The entry registers it last, so on each event every feature's hook has run
3// (on the way down) before the core finishes the event:
4//
5//   prompt.submit  sends the message's ballot as up to two decision requests
6//                  (the effort question alone, the rest together; ADR 0005),
7//                  in parallel, hands each part its answers, then lets the
8//                  prompt in
9//   turn.start     opens the main turn's record, with the effort its prompt
10//                  was decided at
11//   turn.step      the one writer of effort (and of a non-main loop's model):
12//                  sends every step as the plan table says
13//   classic.PreToolUse  notes the calls a settings hook refused (core/outcomes.ts)
14//   command.run    notes the command the person runs: the prompt that follows
15//                  may be its command turn (core/commands.ts)
16//
17// The core owns the unmatched registration of these events; a feature always
18// registers them with a matcher (DEVELOPMENT.md, 开发).
19//
20// Switched off (`/dp off`, core/switches.ts) the core stands down: prompt.submit
21// asks nothing and turn.step sends every step as the engine made it.
22
23import type { HttpInit, On } from 'claude-code'
24import { messageText } from '../decision/context.ts'
25import { EFFORT_PART } from '../decision/effort.ts'
26import { SKILLS_PART } from '../decision/skills.ts'
27import { answersFor, type State } from '../decision/system-one.ts'
28import { messageRequest } from '../decision/turn-start.ts'
29import { type Asked, type BackendIo, describeAsked, errorText } from '../decision/backend.ts'
30import { collect, type Contribution, type PartOutcome } from './ballot.ts'
31import { forgetCommand, noteCommand, typedCommand } from './commands.ts'
32import { noteBlocked } from './outcomes.ts'
33import { newTurn, planStep, replace, takePending, turnKey, type Cell, type PendingDecision } from './plans.ts'
34import { messageLimits, type Ctx } from './setup.ts'
35import { reportStep, type StepIo } from './report.ts'
36import { masterOn } from './switches.ts'
37
38const PENDING = { plugin: 'dispatch-pilot', key: 'pending' } as const
39const TURNS = { plugin: 'dispatch-pilot', key: 'turns' } as const
40const AGENTS = { plugin: 'dispatch-pilot', key: 'agents' } as const
41const LOCK = { plugin: 'dispatch-pilot', key: 'lock' } as const
42const BOARD = { plugin: 'dispatch-pilot', key: 'board' } as const
43const DECISIONS = { plugin: 'dispatch-pilot', key: 'decisionLog' } as const
44const WORKFLOW_RUNS = { plugin: 'dispatch-pilot', key: 'workflowRuns' } as const
45const LABEL_RUNS = { plugin: 'dispatch-pilot', key: 'labelRuns' } as const
46const PPLX_RATE = { plugin: 'dispatch-pilot', key: 'pplxRate' } as const
47
48/** Prompts whose prompt.submit is letting them in right now: a turn that starts meanwhile is theirs. */
49const entering: string[] = []
50
51export function registerCore(on: On, ctx: Ctx): void {
52  on('prompt.submit', async ($, e, next) => {
53    // Every feature above has seen whether this is a command turn.
54    forgetCommand(e.text)
55    const ballot = collect(e.text)
56    // Switched off (/dp off): whatever was put in the ballot is not asked.
57    if (ballot.length === 0 || !masterOn()) return next(e)
58    const messages = ctx.config.context.messages > 0 ? await $.session.messages().catch(() => []) : []
59    const io: BackendIo = {
60      fetch: (url: string, init: HttpInit) => $.http.fetch(url, init),
61      sleep: (ms: number, signal: AbortSignal) => $.clock.sleep(ms, { signal }),
62      pace: { now: () => $.clock.now(), sentAt: { get: () => $.state.get(PPLX_RATE), set: (value, options) => $.state.set(PPLX_RATE, value, options) } },
63    }
64    // The main agent's effort goes in a request of its own (ADR 0005), the rest of the ballot in one. Both go out at once,
65    // each is answered or fails on its own, and each part reads the outcome of the request it was in.
66    const blocks = (
67      await Promise.all(
68        requestGroups(ballot).map(async (group) => {
69          const startedAt = await $.clock.now()
70          const ids = group.flatMap((part) => Object.keys(part.questions).map((id) => `${part.part}.${id}`)).join(', ')
71          let asked: Asked
72          let state: State = {}
73          try {
74            // The state's budget follows the request's longest question: the skills' question leaves less than any other.
75            const limits = messageLimits(ctx.config, group.some((part) => part.part === SKILLS_PART))
76            const request = messageRequest({ prompt: e.text, messages, limits, parts: group })
77            state = request.state
78            asked = await ctx.backend.ask(io, request, ctx.config.timeoutMs)
79          } catch (error) {
80            // A part's malformed questions: nothing was sent.
81            asked = { ok: false, failure: { kind: 'request', detail: errorText(error) } }
82          }
83          const ms = (await $.clock.now()) - startedAt
84          $.ui.log(`request [${ids}] to ${ctx.backend.name}: ${describeAsked(asked, ms)}`, { to: 'debug' })
85          return settleAll(group, asked, state)
86        }),
87      )
88    ).flat()
89    entering.push(e.text)
90    try {
91      return await next(blocks.length > 0 ? { ...e, context: [...(e.context ?? []), ...blocks] } : e)
92    } finally {
93      const at = entering.indexOf(e.text)
94      if (at >= 0) entering.splice(at, 1)
95    }
96  })
97
98  on('command.run', async ($, e, next) => {
99    noteCommand(e)
100    return next(e)
101  })
102
103  on('turn.start', async ($, e, next) => {
104    const pending: Cell<PendingDecision[]> = { get: () => $.state.get(PENDING), set: (value, options) => $.state.set(PENDING, value, options) }
105    // A command turn starts with the engine's command message: the prompt was the command as typed.
106    const started = typedCommand(e.text) ?? e.text
107    // Its own prompt's decision: the prompt now entering (whatever an inner
108    // hook made of its text), else a queued prompt with the turn's text.
109    const texts = entering.length === 1 ? [entering[0] as string, started] : [started]
110    const { before } = await replace(pending, (list) => takePending(list ?? [], texts).rest)
111    // What the landed write took out of the list it found.
112    const own = takePending(before ?? [], texts).taken
113    // The turn's message as the decision model read it (later decisions about the turn reuse it). A pending
114    // entry, decided or not, says the person's own message started the turn (only such a turn is re-decided);
115    // one marked `report` says a report did (decided at its start, not re-decided).
116    const prompt = messageText(started, ctx.config.contextByKind.rejudge)
117    await $.state.set({ ...TURNS, id: turnKey(e.turnId, undefined) }, newTurn(prompt, own?.effort ?? null, own !== null && own.report !== true))
118    return next(e)
119  })
120
121  // Beneath every feature's tool.call: which calls a PreToolUse settings hook refused (the call itself then only
122  // sees an error), for the features that count failures and read a loop's steps (core/outcomes.ts).
123  on('classic.PreToolUse', async ($, e, next) => {
124    const decided = await next(e)
125    if (decided.deny !== undefined && e.tool_use_id) noteBlocked(e.tool_use_id)
126    return decided
127  })
128
129  on('turn.step', async function* ($, e, next) {
130    // Switched off (/dp off): every step goes out as the engine made it, lock or no lock.
131    if (!masterOn()) return yield* next(e)
132    const agentId = e.agentId
133    const { value: turn } = await $.state.get({ ...TURNS, id: turnKey(e.turnId, agentId) })
134    const { value: agent } = agentId === undefined ? { value: undefined } : await $.state.get({ ...AGENTS, id: agentId })
135    const { value: lock = null } = agentId === undefined ? await $.state.get(LOCK) : { value: null }
136    const { step, source } = planStep(e, { lock, turn, agent })
137    // What the step goes out with, for the board: every loop's, the main agent's included.
138    const io: StepIo = {
139      board: { get: () => $.state.get(BOARD), set: (value, options) => $.state.set(BOARD, value, options) },
140      decisions: { get: () => $.state.get(DECISIONS), set: (value, options) => $.state.set(DECISIONS, value, options) },
141      debug: (line) => $.ui.log(line, { to: 'debug' }),
142      now: () => $.clock.now(),
143      toast: (text) => $.ui.toast(text),
144      agents: () => $.agent.list(),
145      runDirs: async () => {
146        const [noted, labelled] = await Promise.all([$.state.get(WORKFLOW_RUNS), $.state.get(LABEL_RUNS)])
147        return [...new Set([...(noted.value ?? []), ...(labelled.value ?? [])].map((run) => run.dir))]
148      },
149      journal: async (dir) => {
150        const path = `${dir}/journal.jsonl`
151        return (await $.fs.exists(path)) ? await $.fs.read(path).catch(() => null) : null
152      },
153    }
154    await reportStep(io, { ...(agentId === undefined ? {} : { agentId }), model: step.model, effort: step.effort, source })
155    return yield* next(step)
156  })
157}
158
159/**
160 * The requests a ballot goes out in: the main agent's effort question alone, every other part's together (the skills'
161 * question is the longest and sets a smaller state budget, so it cannot share a request with the effort question without
162 * cutting the conversation the effort question reads). A group no part is in is not made: no empty request.
163 */
164function requestGroups(ballot: readonly Contribution[]): Contribution[][] {
165  return [ballot.filter((part) => part.part === EFFORT_PART), ballot.filter((part) => part.part !== EFFORT_PART)].filter((group) => group.length > 0)
166}
167
168/** Hands each part its outcome (in parallel) and gathers the context blocks they return, in ballot order. */
169async function settleAll(ballot: readonly Contribution[], asked: Asked, state: State): Promise<string[]> {
170  const settled = await Promise.all(
171    ballot.map(async (contribution) => {
172      const outcome: PartOutcome = asked.ok ? { ok: true, answers: answersFor(contribution, asked.answers), state } : { ok: false, failure: asked.failure }
173      try {
174        return (await contribution.settle(outcome)) ?? []
175      } catch {
176        return []
177      }
178    }),
179  )
180  return settled.flat()
181}
182
hooks/core/report.ts 1362 lines
1// The decision report (「决定汇报」, ADR 0004): the one writer of what Dispatch
2// Pilot shows people. The board, the decision log, the debug log and the toasts
3// are all written here, from the same structured data in $.state
4// (types/index.d.ts: `board`, `decisionLog`, `skillProfiles`); the screens
5// (hooks/board/) only draw that data.
6//
7// Two entries, and nothing else, are for others to call:
8//
9//   report(io, what)        记一条决定: a feature (or a screen) hands over one
10//                           thing to show, tagged by its kind (`Reported`):
11//                           `decision`, a decision it made or why it could not
12//                           make one (which feature, about which agent, the
13//                           outcome, why, the rules' working when there is one:
14//                           the decision log, the agent's node, the debug log; a
15//                           route that failed raises a toast); `decisions`,
16//                           several of one event (the agent() calls of a
17//                           Workflow script: one write each, one toast at most);
18//                           `tally`, what a feature counts as a loop goes (the
19//                           mid-turn re-decisions, a loop's failed calls and
20//                           forced raises: on the agent's node); `switched`, the
21//                           person flipped the mod or a feature (`/dp`: the
22//                           screens draw again); `profiles`, how writing the
23//                           session's skill profiles goes (#33); `unplaced`, a
24//                           pane the person asked for that the surface did not
25//                           place (a toast). `io` is what that kind needs of the
26//                           host (`IoOf`): a `ReportIo` for most.
27//   reportStep(io, step)    记一步读数: the reading of one model request, the
28//                           model and effort an agent's step went out with,
29//                           whoever's loop it is. Only observes, never decides
30//                           anything (ADR 0003). The core calls it from its
31//                           `turn.step`, which knows what the step goes out with.
32//
33// The rest of a loop's life (spawned, ended) and the board's turn count the
34// module hears with hooks of its own (`registerReport`, which the entry file
35// registers first: `turn.start`, `agent.spawn`, `turn.complete`, changing nothing
36// in them); nobody calls those. What else is exported (`decisionLine`,
37// `appendEntry`, `startTurn`, `profileWhy`, the types) only shapes or reads data.
38//
39// Pure but for those hooks: `$` stays in the hook owner's file (it may not cross
40// an import), so the caller builds its io of closures, a `ReportIo` as:
41//
42//   const io: ReportIo = {
43//     board: { get: () => $.state.get(BOARD), set: (value, options) => $.state.set(BOARD, value, options) },
44//     decisions: { get: () => $.state.get(DECISIONS), set: (value, options) => $.state.set(DECISIONS, value, options) },
45//     debug: (line) => $.ui.log(line, { to: 'debug' }),
46//     now: () => $.clock.now(),
47//     toast: (text) => $.ui.toast(text),
48//   }
49//
50// where `BOARD = { plugin: 'dispatch-pilot', key: 'board' } as const` and
51// `DECISIONS = { plugin: 'dispatch-pilot', key: 'decisionLog' } as const` are
52// the file's own literal refs (DEVELOPMENT.md, 开发). Neither entry ever
53// throws: a report that cannot be kept must not stop the decision or the step.
54
55import type { EngineInterface, On } from 'claude-code'
56import { errorText, failureLine, failureWords, type Failure } from '../decision/backend.ts'
57import { modelFamily, type AgentModel } from '../decision/dispatched-agent.ts'
58import { isEffort, type Effort } from '../decision/effort.ts'
59import type { BackendName } from './setup.ts'
60import type { UnresolvedChange, UnresolvedOption, UnresolvedThresholds } from '../decision/unresolved.ts'
61import { startedIn } from '../decision/workflow-labels.ts'
62import { type Cell, type EffortSource, update } from './plans.ts'
63
64const BOARD = { plugin: 'dispatch-pilot', key: 'board' } as const
65const DECISIONS = { plugin: 'dispatch-pilot', key: 'decisionLog' } as const
66const WORKFLOW_RUNS = { plugin: 'dispatch-pilot', key: 'workflowRuns' } as const
67const LABEL_RUNS = { plugin: 'dispatch-pilot', key: 'labelRuns' } as const
68
69// ---- the data (mirrors types/index.d.ts) --------------------------------------
70
71/** A model family, as a person says it (`/model`). */
72export type Model = AgentModel
73/** How a log entry reads: a decision made, a thing to watch, a failure, a plain record. */
74export type Tone = 'ok' | 'warn' | 'fail' | 'info'
75export type AgentState = 'queued' | 'running' | 'done' | 'failed'
76
77/** One step of the rules' working: its rule, whether it took effect, and the rule's own fields. */
78export type RuleStep = { rule: string; applied: boolean; [field: string]: string | number | boolean | null }
79
80/** A failed decision request, with the decision model that failed. */
81export type NodeFailure = Failure & { backend: string }
82
83/** One agent of one turn on the board. */
84export type BoardNode = {
85  turn: number
86  /**
87   * `main`, the agentId, or for a Workflow's agent not started yet the id of its call (`callNodeId`), for a
88   * Workflow not routed as a whole its tool call's id.
89   */
90  id: string
91  kind: 'main' | 'agent' | 'wf'
92  name: string
93  type: string
94  model?: Model
95  effort?: Effort | number
96  state: AgentState
97  t0: number
98  dur?: number
99  routed: boolean
100  locked?: true
101  why?: string
102  failure?: NodeFailure
103  decision?: number
104  /** The Workflow it belongs to (a Workflow agent, or the call of the script that stands for it before it starts). */
105  workflow?: { id: string; name: string }
106  /** The main agent's mid-turn re-decisions of the turn: steps made, decisions, level changes; `late`: an answer is not back, `failure`: the latest request failed. midturn-effort's. */
107  midturn?: { steps: number; judged: number; changed: number; late?: true; failure?: NodeFailure }
108  /** The failed tool calls, hook blocks and forced raises of the agent's loop (the escalation feature counts them); absent while none. */
109  counts?: { failed: number; blocked: number; raised: number; late?: true }
110}
111
112/** What an agent's step read: its model by family and its effort as it went out (absent: a model without effort). */
113export type Reading = { model?: Model; effort?: Effort | number }
114
115/**
116 * The id of the node a Workflow call of a script stands on until its agent starts (an agent of a Workflow has no id
117 * before): `<tool_use_id>#<index>`, the Workflow tool call's own id and the call's place in the script.
118 */
119export function callNodeId(workflow: string, index: number): string {
120  return `${workflow}#${index}`
121}
122
123/** The Workflow tool call a call's node (`callNodeId`) belongs to; null for any other node. */
124export function callWorkflowOf(id: string): string | null {
125  const at = id.indexOf('#')
126  return at < 0 ? null : id.slice(0, at)
127}
128
129/** One reading that differed from the agent's previous one in the same turn (types/index.d.ts, `board.changes`). */
130export type ReadingChange = {
131  turn: number
132  id: string
133  /** Seconds from the start of `turn`. */
134  at: number
135  /** The `n` of the decision log's last entry when it was read. */
136  after: number
137  from: Reading
138  to: Reading
139}
140
141/**
142 * What a feature met beside an agent's route that is neither a decision nor the route itself, for the band's
143 * event stream (types/index.d.ts, `board.notes`): a request that failed (`failed`, `why` its words), one that
144 * gave nothing to use (`skipped`, `why` the Skip), an answer not back in time (`late`, `why` '').
145 */
146export type BoardNote = {
147  turn: number
148  /** `main`, or the agentId it was about. */
149  id: string
150  feature: string
151  /** Seconds from the start of `turn`. */
152  at: number
153  /** The `n` of the decision log's last entry when it was written. */
154  after: number
155  kind: 'failed' | 'skipped' | 'late'
156  why: string
157}
158
159export type Board = {
160  turn: number
161  /** When the latest turns started (`$.clock.now()`), by turn. */
162  starts?: { turn: number; at: number }[]
163  changes?: ReadingChange[]
164  notes?: BoardNote[]
165  nodes: BoardNode[]
166}
167
168/** The board keeps at most this many notes (those of the nodes' turns). */
169const NOTES_KEPT = 50
170
171/** One decision as the log keeps it. */
172export type LogEntry = {
173  n: number
174  turn: number
175  /** Seconds from the start of `turn` to when it was reported (0 for a decision made before its turn started); absent in an entry an earlier version wrote. */
176  at?: number
177  feature: string
178  agent?: string
179  tone: Tone
180  outcome: string
181  subject: string
182  reason: string
183  probs?: Record<Effort, number>
184  conf?: number
185  trace?: RuleStep[]
186  floor?: { from: Effort; to: Effort; model: Model }
187  /** The model (by family) and effort an agent was decided to run with (the main agent: its effort only); the agent's node says what its steps went out with. */
188  model?: Model
189  effort?: Effort
190  mid?: { current: Effort; picked: Effort; result: Effort; threshold?: number; held?: string; remaining?: number }
191  forced?: ForcedRaise
192  skills?: SkillsPicked
193  counts?: { failed: number; blocked: number; raised: number }
194  /** A Workflow call's decision that was sent back to the main agent to write in (return mode). */
195  sentBack?: true
196  /** An unresolved-count judgement (`unresolved`). */
197  unresolved?: UnresolvedRecord
198  /** The strong hint was given with this decision's request (#41). */
199  hint?: HintRecord
200}
201
202/** The strong hint given with a request: the count it was given at, the setting it had reached, and `mid` when it was a mid-turn re-decision's. */
203export type HintRecord = { count: number; maxAfter: number; where?: 'mid' }
204
205/**
206 * What one message's answer to the unresolved question did to the count: the count before and after, the change, the
207 * option the answer leaned to most and every option's probability, the backend's confidence (on record only), and
208 * the two bars it was held to. The card draws it as it is; nothing is worked out again (ADR 0004).
209 */
210export type UnresolvedRecord = {
211  before: number
212  count: number
213  change: UnresolvedChange
214  top: UnresolvedOption
215  probs: Record<UnresolvedOption, number>
216  conf?: number
217  thresholds: UnresolvedThresholds
218}
219
220/** A raise the failures forced: the level (or, for a haiku agent, the model) it went from and to, and the level it holds the agent at least at. */
221export type ForcedRaise = { kind: 'effort' | 'model'; from: string; to: string; floor?: Effort }
222
223/** The skills a message (or a find_skill call) got: those suggested to the main agent, and those only the person can start (「可试 /x」), each with its relevance. */
224export type SkillsPicked = { suggest: { name: string; relevance: number }[]; try: { name: string; relevance: number }[] }
225
226/** The log keeps the entries of this many latest turns, and at most `LOG_ENTRIES` of them. */
227export const LOG_TURNS = 20
228export const LOG_ENTRIES = 300
229
230// ---- what a feature hands over ------------------------------------------------
231
232/** What every report says of who it is about. */
233type About = {
234  /** The feature's switch name; a report's decision adds ` (agent report)`. */
235  feature: string
236  /** `main`, or the agentId of the dispatched or Workflow agent. */
237  agent: string
238  /**
239   * `next`: made before the turn it is for has started (at `prompt.submit`, for a message that will start
240   * the turn), so it belongs to the turn to come; `current` (the default): the turn running now.
241   */
242  forTurn?: 'current' | 'next'
243  /** What it was about, for the log: the start of the message, an agent's label. */
244  subject?: string
245  /**
246   * How to start the agent's node when the board has none yet (the main agent's is known). `state`: `running` by
247   * default, `queued` for a decision made for a turn to come or for a Workflow call whose agent has not started.
248   * `workflow`: the Workflow the agent (or its call) belongs to.
249   */
250  node?: { kind: 'agent' | 'wf'; name: string; type: string; state?: AgentState; workflow?: { id: string; name: string } }
251  /** Whether it left the agent routed (a decision) or not (a failure): the node's `routed`. Left as it is when not given. */
252  routed?: boolean
253  /**
254   * The id of the node that stood for this agent before it started (a Workflow call, queued): the agent's node
255   * takes its place, and the decision made for the call (its `decision` link), when it has one.
256   */
257  replaces?: string
258  /**
259   * A decision beside the agent's route (a mid-turn re-decision, a forced raise, a skill suggestion): in the log,
260   * and not on the agent's node (its `decision` link, `routed` and `why` stay the route's).
261   */
262  aside?: true
263}
264
265/** A decision made. */
266export type Decided = About & {
267  /** What was decided, in a few words: `effort high`. */
268  outcome: string
269  /** Why: what the decision model said, the rule that applied. */
270  reason: string
271  tone?: Tone
272  probs?: Record<Effort, number>
273  conf?: number
274  trace?: RuleStep[]
275  floor?: LogEntry['floor']
276  mid?: LogEntry['mid']
277  /** The model (by family) and the effort it decided for an agent: in the log entry, since the node's own are what its steps read. */
278  model?: Model
279  effort?: Effort
280  /** A Workflow call's decision that was sent back to the main agent to write in (return mode). */
281  sentBack?: true
282  forced?: ForcedRaise
283  skills?: SkillsPicked
284  counts?: LogEntry['counts']
285  unresolved?: UnresolvedRecord
286  hint?: HintRecord
287}
288
289/** A decision the feature could not make: the decision request failed. Not an entry of the log; it says why on the agent's node. */
290export type NotDecided = About & { failure: NodeFailure }
291
292/**
293 * No decision was asked for, or none could be made without a failed request: why the agent (or the Workflow) runs
294 * as it was written, in a few words, on its node. Not a log entry. `asWritten`: by design, not a miss (the Workflow
295 * was already sent back once). `offBoard`: nothing to show at all, so nothing is written (a script with no agent()
296 * call).
297 */
298export type Left = About & { why: string; asWritten?: true; offBoard?: true }
299
300/**
301 * An agent whose decision was made earlier, for its call, has started: the agent is on the board as the one the
302 * decision routed. Not a log entry (the decision is, since it was made).
303 */
304export type Started = About & { started: true }
305
306/**
307 * Why a feature that was asked for something said nothing, when that is neither a decision nor a failed request:
308 * the answer said nothing about the skills (`unanswered`), the session's skills could not be read (`unread`),
309 * there was no skill to rate (`none`), an error of the mod's own (`error`; the debug log has it). Not in the log and
310 * on no node: a note on the board (`board.notes`, kind `skipped`), which the band's event stream draws.
311 */
312export type Skip = 'unanswered' | 'unread' | 'none' | 'error'
313export type Skipped = About & { skipped: Skip; aside: true }
314
315export type ReportedDecision = Decided | NotDecided | Left | Started | Skipped
316
317/** What a feature counts as its loop goes (`report`'s `tally`), not a decision of its own. */
318export type Tallied =
319  /** The main agent's turn as the mid-turn re-decision sees it. `quiet`: not re-decided yet this turn (nothing to show). `late`: the answer for this step is not back; `failure`: its request failed. */
320  | { feature: 'midturn-effort'; agent: 'main'; steps: number; judged: number; changed: number; quiet: boolean; late?: true; failure?: NodeFailure }
321  /** A loop's failed tool calls, hook blocks and forced raises. `turnStart`: the main agent's counts start over with a new turn. */
322  | { feature: 'escalation'; agent: string; failed: number; blocked: number; raised: number; late?: true; turnStart?: true }
323
324/** What one model request went out with, as the core sends it. */
325export type StepReading = {
326  /** Absent for the main agent. */
327  agentId?: string
328  /** The model id as sent. */
329  model: string
330  /** The effort as sent; absent for a model without effort. */
331  effort?: string | number
332  /** Where the effort came from: the person's lock, a plan, or the engine (not routed). */
333  source: EffortSource
334}
335
336/** An agent of the roster `$.agent.list()` answers, as far as a node needs it. */
337export type RosterAgent = { id: string; name?: string; description: string; type: string }
338
339/**
340 * What the readings need beyond `ReportIo`: the roster of dispatched agents, and the directories of the session's
341 * Workflow runs (to tell a Workflow's agent, and its label, from the journal). `stepIo($)` builds it.
342 */
343export type StepIo = ReportIo & {
344  /** `() => $.agent.list()`: dispatched agents only, a Workflow's are not in it. */
345  agents: () => Promise<readonly RosterAgent[]>
346  /** The run directories the mod noted (`workflowRuns`, `labelRuns`), oldest first. */
347  runDirs: () => Promise<readonly string[]>
348  /** A run's journal text; null when it is not there. */
349  journal: (dir: string) => Promise<string | null>
350}
351
352/** What the module needs of the host, as closures (see the top of the file). */
353export type ReportIo = {
354  board: Cell<Board>
355  decisions: Cell<LogEntry[]>
356  /** `(line) => $.ui.log(line, { to: 'debug' })` */
357  debug: (line: string) => void
358  /** `() => $.clock.now()`: when a decision, a reading or a note is reported. */
359  now: () => Promise<number>
360  /** The host's toast, as `toast: (text) => $.ui.toast(text)`: a route that failed. */
361  toast: (text: string) => void
362}
363
364// ---- entry one (记一条决定): what a feature hands over ----------------------------
365
366/** What `report` is handed, one kind at a time (see the top of the file). */
367export type Reported =
368  /** A decision made, or why none was (`ReportedDecision`). */
369  | { decision: ReportedDecision }
370  /** The decisions of one event, in the order made: one write to the log and one to the board, at most one toast. */
371  | { decisions: readonly ReportedDecision[] }
372  /** What a feature counts as a loop goes (`Tallied`). */
373  | { tally: Tallied }
374  /** The person flipped the mod or a feature (`/dp`). */
375  | { switched: Switched }
376  /** How writing the session's skill profiles goes (`ProfileEvent`). */
377  | { profiles: ProfileEvent }
378  /** The decision model in use is not the one asked for: no Perplexity key, so Jev (`DecisionModelEvent`). */
379  | { fellBack: DecisionModelEvent }
380  /** A pane the person asked for (a digit on the band) that the surface did not place, with the surface's reason; it was closed again. */
381  | { unplaced: { reason: string } }
382
383/** What a toast of its own needs of the host: the debug log, the clock and the toast. */
384export type NoticeIo = Pick<ReportIo, 'debug' | 'now' | 'toast'>
385
386/** What the decision model's fallback needs of the host: the board (for the turn), the decision log and the debug log. */
387export type DecisionModelIo = Pick<ReportIo, 'board' | 'decisions' | 'debug'>
388
389/** The host closures each kind of report needs: a switch only the redraw, the skill profiles their own state, a pane not placed the toast, the rest a `ReportIo`. */
390export type IoOf<R extends Reported> = R extends { switched: Switched }
391  ? SwitchIo
392  : R extends { profiles: ProfileEvent }
393    ? ProfilesIo
394    : R extends { fellBack: DecisionModelEvent }
395      ? DecisionModelIo
396      : R extends { unplaced: unknown }
397        ? NoticeIo
398        : ReportIo
399
400/**
401 * Entry one (记一条决定): reports what a feature hands over, by its kind; `io` is what that kind needs of the host
402 * (`IoOf`). Never throws.
403 */
404export async function report<R extends Reported>(io: IoOf<R>, what: R): Promise<void> {
405  const item: Reported = what
406  if ('decision' in item) return reportDecisions(io as ReportIo, [item.decision])
407  if ('decisions' in item) return reportDecisions(io as ReportIo, item.decisions)
408  if ('tally' in item) return reportTally(io as ReportIo, item.tally)
409  if ('switched' in item) return reportSwitch(io as SwitchIo, item.switched)
410  if ('profiles' in item) return reportProfiles(io as ProfilesIo, item.profiles)
411  if ('fellBack' in item) return reportFellBack(io as DecisionModelIo, item.fellBack)
412  return reportUnplaced(io as NoticeIo, item.unplaced)
413}
414
415/** The rationale pane was not placed, in the person's words: why (the surface's own reason), and where to look instead. */
416export function unplacedText(reason: string): string {
417  return `依据面板没有放出来(${reason}),已经关上;/dp log 10 在对话里列出最近 10 条决定`
418}
419
420/**
421 * A pane the person asked for that the surface did not place: the debug line, and a toast so they know (the `/dp`
422 * command says it in its answer instead), unless a toast went up within `TOAST_GAP_MS`. Never throws.
423 */
424async function reportUnplaced(io: NoticeIo, unplaced: { reason: string }): Promise<void> {
425  try {
426    io.debug(`rationale pane not placed (${unplaced.reason}), closed again`)
427    toastOnce(io, await io.now(), unplacedText(unplaced.reason))
428  } catch {
429    // a notice that cannot be given leaves the pane closed all the same
430  }
431}
432
433/**
434 * The decisions of one event (a single one, or the several agents of a Workflow): the debug log line and the
435 * decision log entry of each decision made, the agents' nodes on the board, a note for the band when a request
436 * beside a route came to nothing, and a toast when a route failed. One write to the log, one to the board, at
437 * most one toast; each is what it would make of it alone, in the order given. Never throws.
438 */
439async function reportDecisions(io: ReportIo, decisions: readonly ReportedDecision[]): Promise<void> {
440  if (decisions.length === 0) return
441  try {
442    // Which turn each is for. A board that cannot be read does not stop a decision from being logged (as turn 1).
443    const board = await read(io.board).catch((error: unknown) => {
444      io.debug(`board not read: ${errorText(error)}`)
445      return EMPTY
446    })
447    const now = await io.now()
448    const turnOf = (decision: ReportedDecision) => board.turn + (decision.forTurn === 'next' ? 1 : 0)
449    const made = decisions.filter(isDecided)
450    for (const decision of made) io.debug(decisionLine(decision))
451    // The number of each decision in the log, in the order made.
452    const numbers: number[] = []
453    if (made.length > 0) {
454      try {
455        await update(io.decisions, (list) => {
456          numbers.length = 0
457          let all = list ?? []
458          for (const decision of made) {
459            all = appendEntry(all, entryOf(decision, turnOf(decision), elapsed(board, turnOf(decision), now)))
460            numbers.push(all.at(-1)?.n ?? 0)
461          }
462          return all
463        })
464      } catch (error) {
465        numbers.length = 0
466        io.debug(`decision not kept for /dp log: ${errorText(error)}`)
467      }
468    }
469    // A request beside the route that came to nothing: a note for the band's event stream, after the log's last entry.
470    const noted = decisions.filter((decision) => 'skipped' in decision || ('failure' in decision && decision.aside === true))
471    const logged = noted.length === 0 ? 0 : (numbers.at(-1) ?? (await lastLogged(io)))
472    const notes = noted.map((decision): BoardNote => {
473      const of = { turn: turnOf(decision), id: decision.agent, feature: decision.feature, at: elapsed(board, turnOf(decision), now), after: logged }
474      return 'skipped' in decision ? { ...of, kind: 'skipped', why: decision.skipped } : { ...of, kind: 'failed', why: 'failure' in decision ? failureLine(decision.failure.backend, decision.failure) : '' }
475    })
476    let after = board
477    // A decision beside the agent's route is not on its node; a Workflow with nothing to show is not on the board.
478    const onBoard = (decision: ReportedDecision) => decision.aside !== true && !('skipped' in decision) && !('offBoard' in decision)
479    if (decisions.some(onBoard) || notes.length > 0) {
480      try {
481        after = await update(io.board, (current) => {
482          let next = current ?? EMPTY
483          let at = 0
484          for (const decision of decisions) {
485            const n = isDecided(decision) ? numbers[at++] : undefined
486            if (onBoard(decision)) next = withNode(next, turnOf(decision), decision as Decided | NotDecided | Left | Started, n)
487          }
488          return notes.length === 0 ? next : withNotes(next, notes)
489        })
490      } catch (error) {
491        io.debug(`decision not kept on the board: ${errorText(error)}`)
492      }
493    }
494    // The routes that failed: a toast, so the person notices; the board says the rest.
495    const failed = decisions.filter((decision): decision is NotDecided => 'failure' in decision && decision.aside !== true)
496    if (failed.length > 0) toastOnce(io, now, failedText(failed, after, turnOf))
497  } catch (error) {
498    io.debug(`decision not reported: ${errorText(error)}`)
499  }
500}
501
502/** A decision that was made (the others say why there is none, or that an agent started). */
503function isDecided(decision: ReportedDecision): decision is Decided {
504  return 'outcome' in decision
505}
506
507/** The `n` of the decision log's last entry (0: none, or the log cannot be read). */
508async function lastLogged(io: ReportIo): Promise<number> {
509  try {
510    return ((await io.decisions.get()).value ?? []).at(-1)?.n ?? 0
511  } catch {
512    return 0
513  }
514}
515
516/** The engine drops a plugin's second toast within this many milliseconds of its last one (measured on 2.1.289). */
517const TOAST_GAP_MS = 2000
518
519/** When the last toast was raised (module state: a hot reload forgets it, and the next toast may then be one the engine drops). */
520let toastedAt = Number.NEGATIVE_INFINITY
521
522/**
523 * Raises the toast unless one was raised within `TOAST_GAP_MS`, which the engine would drop: the failures that
524 * follow one closely are on the board all the same.
525 */
526function toastOnce(io: Pick<ReportIo, 'toast'>, now: number, text: string): void {
527  if (now - toastedAt < TOAST_GAP_MS) return
528  toastedAt = now
529  try {
530    io.toast(text)
531  } catch {
532    // a toast that cannot be shown is skipped: the board has it
533  }
534}
535
536/** The toast for the routes of one event that failed: who is not routed, and why (the first failure's words, then its details). */
537function failedText(failed: readonly NotDecided[], board: Board, turnOf: (decision: ReportedDecision) => number): string {
538  const first = failed[0] as NotDecided
539  const why = `${failureWords(first.failure)}(${failureLine(first.failure.backend, first.failure)})`
540  if (failed.length > 1) return `${failed.length} 个 agent 未路由:${why}`
541  if (first.agent === 'main') return `主 agent 未路由:${why}`
542  const name = board.nodes.find((node) => node.turn === turnOf(first) && node.id === first.agent)?.name ?? first.node?.name ?? first.agent
543  return `「${name.length > 24 ? `${name.slice(0, 23)}…` : name}」未路由:${why}`
544}
545
546/** The board with the notes added: those of turns older than the nodes' dropped, and the oldest past `NOTES_KEPT`. */
547function withNotes(board: Board, notes: readonly BoardNote[]): Board {
548  return { ...board, notes: [...(board.notes ?? []), ...notes].filter((note) => note.turn >= board.turn - 1).slice(-NOTES_KEPT) }
549}
550
551// ---- entry two (记一步读数): a step's reading ----------------------------------------
552
553/** How many writes may be lost to other writers in a row: a Workflow's agents step within milliseconds of each other. */
554const READING_ATTEMPTS = 32
555
556/**
557 * Entry two (记一步读数): reports the reading of one step, the model (by family) and the effort it
558 * went out with, whether the effort was routed, and that the agent is running.
559 * Reading never changes what is sent. The agent's node is made at its first
560 * step when the board has none (named from the roster of dispatched agents, or
561 * from the label its Workflow journal gives it; a loop neither knows is not an
562 * agent of the board), carries its start, and a reading that differs from the
563 * node's previous one adds a change event. Writes the board only when
564 * something is new. Never throws.
565 */
566export async function reportStep(io: StepIo, step: StepReading): Promise<void> {
567  const id = step.agentId ?? 'main'
568  const main = step.agentId === undefined
569  const family = modelFamily(step.model)
570  const effort = typeof step.effort === 'number' || isEffort(step.effort) ? step.effort : undefined
571  const reading: Reading = { ...(family === null ? {} : { model: family }), ...(effort === undefined ? {} : { effort }) }
572  // `source`: the engine's own effort goes out unless a plan sets it. An agent that was decided to run as it does has no plan
573  // to say so (haiku takes no effort, a Workflow script was written into): its reading is checked against its decision below.
574  let routed = step.source !== 'engine'
575  const locked = step.source === 'locked'
576  try {
577    const now = await io.now()
578    const seen = main ? now : seenAt(id, now)
579    const peek = await read(io.board)
580    const before = main ? mainNode(peek) : continuing(peek, id)
581    // A node that has not begun is named now (the roster may know it only from here on); one that has keeps its name.
582    const identity = main || (before !== undefined && begun(before)) ? null : await lookUp(io, id)
583    if (!main && before === undefined && identity === null) return
584    const prior = main ? undefined : (before ?? callNodeOf(peek, identity))
585    if (!routed && prior?.routed === true && prior.decision !== undefined) routed = await wentOutAsDecided(io, prior.decision, reading)
586    const logged = before !== undefined && differs(before, reading) ? (((await io.decisions.get()).value ?? []).at(-1)?.n ?? 0) : 0
587    await modify(
588      io.board,
589      (current) => {
590        const board = current ?? EMPTY
591        const old = main ? mainNode(board) : continuing(board, id)
592        if (old === undefined && identity === null && !main) return undefined
593        const turn = old?.turn ?? board.turn
594        const started = old !== undefined && begun(old)
595        // A Workflow's agent that starts takes the place of the node its call stood for, and what was decided for it.
596        const stood = old === undefined && !main ? callNodeOf(board, identity) : undefined
597        const base = old ?? { ...newNode(turn, id, identity ?? undefined, 'running'), ...carriedFrom(stood, routed) }
598        const named = identity !== null && !started ? { kind: identity.kind, name: identity.name, type: identity.type } : {}
599        const next: BoardNode = {
600          ...without(base, 'model', 'effort', 'locked', 'dur'),
601          ...named,
602          state: 'running',
603          t0: main ? 0 : started ? base.t0 : elapsed(board, turn, seen),
604          routed,
605          ...reading,
606          ...(locked ? { locked: true as const } : {}),
607        }
608        const changes = old !== undefined && differs(old, reading) ? [{ turn, id, at: elapsed(board, turn, now), after: logged, from: readingOf(old), to: reading }] : []
609        if (old !== undefined && changes.length === 0 && JSON.stringify(old) === JSON.stringify(next)) return undefined
610        return {
611          ...board,
612          ...(changes.length === 0 ? {} : { changes: [...(board.changes ?? []), ...changes].filter((change) => change.turn >= board.turn - 1) }),
613          nodes: [...board.nodes.filter((node) => node !== old && node !== stood), next],
614        }
615      },
616      READING_ATTEMPTS,
617    )
618    if (!main) firstSeen.delete(id)
619  } catch (error) {
620    io.debug(`reading not kept on the board: ${errorText(error)}`)
621  }
622}
623
624/** Why an agent's turn did not end in an answer, in the board's words (a node's `why` when it failed and nothing else said). */
625export const ENDED_WORDS = { error: '出错', refusal: '拒绝回答', aborted: '被中断' } as const
626
627/** How a loop's turn ended (`turn.complete`). */
628export type LoopEnd = { agentId?: string; reason: 'answer' | 'aborted' | 'refusal' | 'error'; durationMs: number }
629
630/**
631 * Reports the end of a loop's turn: done after an answer, else failed (and
632 * why, when nothing else said), for as long as it ran. The main agent's turn
633 * ends its own node; an agent's, its own, wherever its turn's node is. Never throws.
634 */
635async function reportEnd(io: StepIo, end: LoopEnd): Promise<void> {
636  const id = end.agentId ?? 'main'
637  const main = end.agentId === undefined
638  try {
639    const now = await io.now()
640    const peek = await read(io.board)
641    const known = main ? mainNode(peek) : continuing(peek, id)
642    // The loop is over: whatever its steps missed, it is looked up once more in full.
643    missed.delete(id)
644    const identity = main || known !== undefined ? null : (await identify(io, id, { roster: true, dirs: await runDirsOf(io) })).identity
645    if (!main && known === undefined && identity === null) return
646    const seen = main || known !== undefined ? now : (firstSeen.get(id) ?? now - end.durationMs)
647    await modify(
648      io.board,
649      (current) => {
650        const board = current ?? EMPTY
651        const old = main ? mainNode(board) : continuing(board, id)
652        if (old === undefined && identity === null && !main) return undefined
653        const turn = old?.turn ?? board.turn
654        const stood = old === undefined && !main ? callNodeOf(board, identity) : undefined
655        const base = old ?? { ...newNode(turn, id, identity ?? undefined, 'running'), ...carriedFrom(stood, true) }
656        const failed = end.reason !== 'answer'
657        const next: BoardNode = {
658          ...base,
659          ...(old === undefined && !main ? { t0: elapsed(board, turn, seen) } : {}),
660          state: failed ? 'failed' : 'done',
661          dur: Math.round(end.durationMs) / 1000,
662          ...(failed && base.why === undefined ? { why: end.reason === 'aborted' ? ENDED_WORDS.aborted : end.reason === 'refusal' ? ENDED_WORDS.refusal : ENDED_WORDS.error } : {}),
663        }
664        return { ...board, nodes: [...board.nodes.filter((node) => node !== old && node !== stood), next] }
665      },
666      READING_ATTEMPTS,
667    )
668    if (!main) firstSeen.delete(id)
669  } catch (error) {
670    io.debug(`end of loop not kept on the board: ${errorText(error)}`)
671  }
672}
673
674/**
675 * Reports a dispatched agent started: queued until its first step. Its node is
676 * made (or, when a decision about it made one, marked queued from now); the
677 * agent's own first step starts it. Never throws.
678 */
679async function reportSpawn(io: StepIo, spawn: { agentId: string; name?: string; description: string; type: string }): Promise<void> {
680  try {
681    const now = await io.now()
682    await modify(
683      io.board,
684      (current) => {
685        const board = current ?? EMPTY
686        const old = continuing(board, spawn.agentId)
687        if (old !== undefined && begun(old)) return undefined
688        const turn = old?.turn ?? board.turn
689        const name = spawn.name !== undefined && spawn.name !== '' ? spawn.name : spawn.description !== '' ? spawn.description : spawn.agentId
690        const base = old ?? newNode(turn, spawn.agentId, { kind: 'agent', name, type: spawn.type }, 'queued')
691        const next: BoardNode = { ...base, state: 'queued', t0: elapsed(board, turn, now) }
692        return { ...board, nodes: [...board.nodes.filter((node) => node !== old), next] }
693      },
694      READING_ATTEMPTS,
695    )
696  } catch (error) {
697    io.debug(`agent not kept on the board: ${errorText(error)}`)
698  }
699}
700
701// ---- what `report` does with a tally --------------------------------------------
702
703/**
704 * Reports what a feature counts as its loop goes, which is no decision: the
705 * mid-turn re-decision's steps, decisions and changes of the main agent's
706 * turn, and each loop's failed tool calls and forced raises. On the agent's
707 * node of the current turn (`midturn`, `counts`); an answer that is late, or a
708 * re-decision's request that failed, is also a note for the band's event
709 * stream. Writes the board only when the node changed. Never throws.
710 */
711async function reportTally(io: ReportIo, tally: Tallied): Promise<void> {
712  try {
713    const noteworthy = tally.late === true || ('failure' in tally && tally.failure !== undefined)
714    const now = await io.now()
715    const logged = noteworthy ? await lastLogged(io) : 0
716    await modify(io.board, (current) => withTally(current ?? EMPTY, tally, { now, logged }))
717  } catch (error) {
718    io.debug(`tally not kept on the board: ${errorText(error)}`)
719  }
720}
721
722// ---- what `report` does with a switch -------------------------------------------
723
724/** What the person flipped with `/dp`: the whole mod, or one feature. */
725export type Switched = { master: boolean } | { feature: string; on: boolean }
726
727/** What a switch needs of the host: `() => $.ui.invalidate('ui.render')`, the screens drawn again. */
728export type SwitchIo = { redraw: () => void }
729
730/**
731 * Reports a switch the person flipped (the control feature's `/dp`). The board
732 * and the decision log keep what they hold: the screens read the switches as
733 * they draw, and leave out what a feature that is off owns (its decisions, the
734 * parts of the nodes it writes, `defineSwitch`'s `parts`), so they are drawn
735 * again now. Never throws.
736 */
737function reportSwitch(io: SwitchIo, _change: Switched): void {
738  try {
739    io.redraw()
740  } catch {
741    // a redraw that cannot be asked for waits for the next change of the board
742  }
743}
744
745/**
746 * The board with the tally on its agent's node (the main agent's of the turn running, an agent's where its steps
747 * go on: `continuing`), and a note when an answer just went late or a request just failed; undefined when the
748 * node already says it, or there is nothing to say. A loop with no node is no agent of the board (the readings
749 * make nodes), the main agent's is made.
750 */
751function withTally(board: Board, tally: Tallied, when: { now: number; logged: number }): Board | undefined {
752  const main = tally.agent === 'main'
753  const old = main ? mainNode(board) : continuing(board, tally.agent)
754  if (old === undefined && !main) return undefined
755  let next: BoardNode
756  const notes: BoardNote[] = []
757  const note = (turn: number, kind: BoardNote['kind'], why: string) =>
758    notes.push({ turn, id: tally.agent, feature: tally.feature, at: elapsed(board, turn, when.now), after: when.logged, kind, why })
759  if (tally.feature === 'midturn-effort') {
760    if (tally.quiet) return undefined
761    const midturn = { steps: tally.steps, judged: tally.judged, changed: tally.changed, ...(tally.late === true ? { late: true as const } : {}), ...(tally.failure === undefined ? {} : { failure: tally.failure }) }
762    next = { ...(old ?? newNode(board.turn, tally.agent, undefined, 'running')), midturn }
763    if (tally.late === true && old?.midturn?.late !== true) note(next.turn, 'late', '')
764    if (tally.failure !== undefined && JSON.stringify(old?.midturn?.failure) !== JSON.stringify(tally.failure)) note(next.turn, 'failed', failureLine(tally.failure.backend, tally.failure))
765  } else {
766    const any = tally.failed + tally.blocked + tally.raised > 0
767    if (old === undefined && !any) return undefined
768    const base = old ?? newNode(board.turn, tally.agent, undefined, 'running')
769    next = any
770      ? { ...base, counts: { failed: tally.failed, blocked: tally.blocked, raised: tally.raised, ...(tally.late === true ? { late: true as const } : {}) } }
771      : without(base, 'counts')
772    if (tally.late === true && old?.counts?.late !== true) note(next.turn, 'late', '')
773  }
774  if (old !== undefined && JSON.stringify(old) === JSON.stringify(next)) return undefined
775  const changed = { ...board, nodes: [...board.nodes.filter((node) => node !== old), next] }
776  return notes.length === 0 ? changed : withNotes(changed, notes)
777}
778
779/** Who a loop with no node is, from what the engine and the Workflow journals say; null: not an agent the board shows. */
780type Identity = { kind: 'agent' | 'wf'; name: string; type: string }
781
782/**
783 * Who a loop is, by the roster (`roster`) and the journals of the run directories `dirs` (null: none read). `read`
784 * is false when the directories or a journal could not be read, so a miss may not be one.
785 */
786async function identify(io: StepIo, id: string, look: { roster: boolean; dirs: readonly string[] | null }): Promise<{ identity: Identity | null; read: boolean }> {
787  if (look.roster) {
788    try {
789      const info = (await io.agents()).find((agent) => agent.id === id)
790      if (info !== undefined) return { identity: { kind: 'agent', name: info.name !== undefined && info.name !== '' ? info.name : info.description !== '' ? info.description : id, type: info.type }, read: true }
791    } catch (error) {
792      io.debug(`agent roster not read: ${errorText(error)}`)
793    }
794  }
795  if (look.dirs === null) return { identity: null, read: false }
796  try {
797    // A Workflow's agents are in no roster: the journal of the run it started in has the label it started under.
798    for (const dir of [...look.dirs].reverse()) {
799      const journal = await io.journal(dir)
800      const start = journal === null ? null : startedIn(journal, id)
801      if (start !== null) return { identity: { kind: 'wf', name: start.label, type: 'workflow' }, read: true }
802    }
803  } catch (error) {
804    io.debug(`workflow journals not read: ${errorText(error)}`)
805    return { identity: null, read: false }
806  }
807  return { identity: null, read: true }
808}
809
810/** The session's Workflow run directories, oldest first; null when they cannot be read. */
811async function runDirsOf(io: StepIo): Promise<readonly string[] | null> {
812  try {
813    return await io.runDirs()
814  } catch (error) {
815    io.debug(`workflow journals not read: ${errorText(error)}`)
816    return null
817  }
818}
819
820/**
821 * The loops a step found no one to be, by id: how many of its steps have looked, and the run directories whose
822 * journals it was looked for in (null: they could not be read). Module state: a hot reload forgets it, and costs each
823 * such loop one full look-up more.
824 */
825const missed = new Map<string, { sightings: number; dirs: string | null }>()
826const MISSED_KEPT = 64
827
828/**
829 * `identify` for a step, at most as often as the answer can change: the first step of a loop no one knows (an engine
830 * fork) reads the roster and every run's journal, and its later steps would read them all again. After a miss, the
831 * journals are read again only once the session has another run directory (a Workflow's agent can step before its run
832 * is recorded), or after they could not be read; the roster at the loop's 2nd, 4th, 8th... step (an agent it lists late).
833 */
834async function lookUp(io: StepIo, id: string): Promise<Identity | null> {
835  const miss = missed.get(id)
836  const sightings = (miss?.sightings ?? 0) + 1
837  const roster = miss === undefined || (sightings & (sightings - 1)) === 0
838  const dirs = await runDirsOf(io)
839  const seen = dirs === null ? null : dirs.join('\n')
840  const journals = miss === undefined || seen === null || miss.dirs === null ? roster : seen !== miss.dirs
841  const found = roster || journals ? await identify(io, id, { roster, dirs: journals ? dirs : null }) : { identity: null, read: true }
842  // The newest miss goes last, so the oldest go first past `MISSED_KEPT`.
843  missed.delete(id)
844  if (found.identity !== null) return found.identity
845  missed.set(id, { sightings, dirs: !journals ? (miss?.dirs ?? null) : found.read ? seen : null })
846  for (const key of missed.keys()) if (missed.size > MISSED_KEPT) missed.delete(key)
847  return null
848}
849
850/**
851 * When a loop no node was made for was first seen (module state: a hot reload forgets it, and an agent only
852 * found later then starts at that step). A Workflow's first agent can step before its run is recorded, so its
853 * name comes with a later step, or with the end of its loop, and starts when it first stepped.
854 */
855const firstSeen = new Map<string, number>()
856const FIRST_SEEN_KEPT = 64
857
858function seenAt(id: string, now: number): number {
859  const seen = firstSeen.get(id)
860  if (seen !== undefined) return seen
861  firstSeen.set(id, now)
862  for (const key of firstSeen.keys()) if (firstSeen.size > FIRST_SEEN_KEPT) firstSeen.delete(key)
863  return now
864}
865
866/**
867 * The node a Workflow call of a script stands for before its agent starts (a decision was made for the call, or it
868 * was left as written): queued, one of the Workflow's calls, called what the agent's label is. Null when there is
869 * none, or the agent is not a Workflow's.
870 */
871function callNodeOf(board: Board, identity: Identity | null | undefined): BoardNode | undefined {
872  if (identity === null || identity === undefined || identity.kind !== 'wf') return undefined
873  return board.nodes.find((node) => node.kind === 'wf' && node.state === 'queued' && node.workflow !== undefined && callWorkflowOf(node.id) === node.workflow.id && node.name === identity.name)
874}
875
876/** What a Workflow's agent takes over from the node its call stood for: the Workflow, the decision made for it, and why it was left as written (when its steps go out unrouted). */
877function carriedFrom(stood: BoardNode | undefined, routed: boolean): Partial<BoardNode> {
878  if (stood === undefined) return {}
879  return {
880    ...(stood.workflow === undefined ? {} : { workflow: stood.workflow }),
881    ...(stood.decision === undefined ? {} : { decision: stood.decision }),
882    ...(routed || stood.why === undefined ? {} : { why: stood.why }),
883    ...(routed || stood.failure === undefined ? {} : { failure: stood.failure }),
884  }
885}
886
887/** Whether a reading is the model and effort the decision numbered `n` in the log decided for the agent (a decision that names none: no). */
888async function wentOutAsDecided(io: ReportIo, n: number, reading: Reading): Promise<boolean> {
889  try {
890    const entry = ((await io.decisions.get()).value ?? []).find((kept) => kept.n === n)
891    return entry?.model !== undefined && entry.model === reading.model && entry.effort === reading.effort
892  } catch {
893    return false
894  }
895}
896
897/** The main agent's node of the turn running. */
898function mainNode(board: Board): BoardNode | undefined {
899  return board.nodes.find((node) => node.turn === board.turn && node.id === 'main')
900}
901
902/**
903 * An agent's node that its steps go on in: its latest one, if it is of the turn
904 * running or has not ended (an agent still running when the next turn starts
905 * stays in the turn it began in); else it is a new row in this turn.
906 */
907function continuing(board: Board, id: string): BoardNode | undefined {
908  const latest = board.nodes.filter((node) => node.id === id).reduce<BoardNode | undefined>((best, node) => (best === undefined || node.turn >= best.turn ? node : best), undefined)
909  return latest !== undefined && (latest.turn === board.turn || latest.state === 'queued' || latest.state === 'running') ? latest : undefined
910}
911
912/** Whether the node has had a reading: its start is the first step's, and its name is settled. */
913function begun(node: BoardNode): boolean {
914  return node.state !== 'queued' && (node.id === 'main' || node.t0 > 0 || node.model !== undefined || node.effort !== undefined)
915}
916
917function readingOf(node: BoardNode): Reading {
918  return { ...(node.model === undefined ? {} : { model: node.model }), ...(node.effort === undefined ? {} : { effort: node.effort }) }
919}
920
921/** Whether the node has a previous reading, and it is not this one. */
922function differs(node: BoardNode, reading: Reading): boolean {
923  if (node.model === undefined && node.effort === undefined) return false
924  return node.model !== reading.model || node.effort !== reading.effort
925}
926
927/** Seconds from the start of `turn` to `ms` (0 when the turn's start is not known). */
928function elapsed(board: Board, turn: number, ms: number): number {
929  const start = board.starts?.find((kept) => kept.turn === turn)?.at
930  return start === undefined ? 0 : Math.max(0, Math.round(ms - start)) / 1000
931}
932
933// ---- the turn count and the loops' lifecycle ------------------------------------
934
935/** The board when a main turn starts at `now`: the count up by one, the nodes of older turns gone (the turn before it stays: it is folded into one line). */
936export function startTurn(board: Board | undefined, now: number): Board {
937  const turn = (board?.turn ?? 0) + 1
938  const kept = (of: number) => of >= turn - 1
939  return {
940    turn,
941    starts: [...(board?.starts ?? []).filter((start) => kept(start.turn)), { turn, at: now }],
942    ...(board?.changes === undefined ? {} : { changes: board.changes.filter((change) => kept(change.turn)) }),
943    ...(board?.notes === undefined ? {} : { notes: board.notes.filter((note) => kept(note.turn)) }),
944    nodes: (board?.nodes ?? []).filter((node) => kept(node.turn)),
945  }
946}
947
948/**
949 * What the module's own hooks need of the host, as closures over `$`. `$` is followed only into a function of
950 * the same file, so the core builds its `StepIo` itself (core.ts, `turn.step`): keep the two alike.
951 */
952function stepIo($: EngineInterface): StepIo {
953  return {
954    board: { get: () => $.state.get(BOARD), set: (value, options) => $.state.set(BOARD, value, options) },
955    decisions: { get: () => $.state.get(DECISIONS), set: (value, options) => $.state.set(DECISIONS, value, options) },
956    debug: (line) => $.ui.log(line, { to: 'debug' }),
957    now: () => $.clock.now(),
958    toast: (text) => $.ui.toast(text),
959    agents: () => $.agent.list(),
960    runDirs: async () => {
961      const [noted, labelled] = await Promise.all([$.state.get(WORKFLOW_RUNS), $.state.get(LABEL_RUNS)])
962      return [...new Set([...(noted.value ?? []), ...(labelled.value ?? [])].map((run) => run.dir))]
963    },
964    journal: async (dir) => {
965      const path = `${dir}/journal.jsonl`
966      return (await $.fs.exists(path)) ? await $.fs.read(path).catch(() => null) : null
967    },
968  }
969}
970
971/**
972 * The module's own hooks, which only watch (every event goes on as it is):
973 * a main turn starting (the board's turn count and its start time), a
974 * dispatched agent spawned (queued until its first step), a loop's turn ending
975 * (its node done or failed). A board that cannot be kept never stops anything.
976 */
977export function registerReport(on: On): void {
978  on('turn.start', { turnId: /(?:)/ }, async ($, e, next) => {
979    try {
980      const now = await $.clock.now()
981      await update<Board>({ get: () => $.state.get(BOARD), set: (value, options) => $.state.set(BOARD, value, options) }, (board) => startTurn(board, now))
982    } catch {
983      // the count is the board's own: a turn starts without it
984    }
985    return next(e)
986  })
987
988  on('agent.spawn', { tool_use_id: /(?:)/ }, async ($, e, next) => {
989    const result = await next(e)
990    if (result.deny !== undefined || result.agentId === undefined) return result
991    await reportSpawn(stepIo($), { agentId: result.agentId, ...(e.name === undefined ? {} : { name: e.name }), description: e.description, type: e.subagentType })
992    return result
993  })
994
995  on('turn.complete', { turnId: /(?:)/ }, async ($, e, next) => {
996    const result = await next(e)
997    await reportEnd(stepIo($), { ...(e.agentId === undefined ? {} : { agentId: e.agentId }), reason: e.reason, durationMs: e.durationMs })
998    return result
999  })
1000}
1001
1002// ---- the log ------------------------------------------------------------------
1003
1004/** A decision as one line: `effort high · "the message":why`. The debug log's line, and `/dp log N`'s after `#n feature:`. */
1005export function decisionLine(decision: { outcome: string; subject?: string; reason: string }): string {
1006  return `${decision.outcome}${decision.subject ? ` · ${decision.subject}` : ''}:${decision.reason}`
1007}
1008
1009/** The list with `entry` added last, numbered, the entries of turns older than the latest `LOG_TURNS` dropped, and then the oldest past `LOG_ENTRIES`. */
1010export function appendEntry(list: readonly LogEntry[], entry: Omit<LogEntry, 'n'>): LogEntry[] {
1011  const all = [...list, { n: (list.at(-1)?.n ?? 0) + 1, ...entry }]
1012  const latest = Math.max(...all.map((kept) => kept.turn))
1013  return all.filter((kept) => kept.turn > latest - LOG_TURNS).slice(-LOG_ENTRIES)
1014}
1015
1016function entryOf(decision: Decided, turn: number, at: number): Omit<LogEntry, 'n'> {
1017  return {
1018    turn,
1019    at,
1020    feature: decision.feature,
1021    agent: decision.agent,
1022    tone: decision.tone ?? 'ok',
1023    outcome: decision.outcome,
1024    subject: decision.subject ?? '',
1025    reason: decision.reason,
1026    ...(decision.probs === undefined ? {} : { probs: decision.probs }),
1027    ...(decision.conf === undefined ? {} : { conf: decision.conf }),
1028    ...(decision.trace === undefined ? {} : { trace: decision.trace }),
1029    ...(decision.floor === undefined ? {} : { floor: decision.floor }),
1030    ...(decision.model === undefined ? {} : { model: decision.model }),
1031    ...(decision.effort === undefined ? {} : { effort: decision.effort }),
1032    ...(decision.mid === undefined ? {} : { mid: decision.mid }),
1033    ...(decision.forced === undefined ? {} : { forced: decision.forced }),
1034    ...(decision.skills === undefined ? {} : { skills: decision.skills }),
1035    ...(decision.counts === undefined ? {} : { counts: decision.counts }),
1036    ...(decision.unresolved === undefined ? {} : { unresolved: decision.unresolved }),
1037    ...(decision.hint === undefined ? {} : { hint: decision.hint }),
1038    ...(decision.sentBack === true ? { sentBack: true as const } : {}),
1039  }
1040}
1041
1042// ---- the board ----------------------------------------------------------------
1043
1044const EMPTY: Board = { turn: 0, nodes: [] }
1045
1046function newNode(turn: number, id: string, node: About['node'], state: AgentState): BoardNode {
1047  if (id === 'main') return { turn, id, kind: 'main', name: '主 agent', type: 'main', state, t0: 0, routed: false }
1048  return { turn, id, kind: node?.kind ?? 'agent', name: node?.name ?? id, type: node?.type ?? 'agent', state: node?.state ?? state, t0: 0, routed: false, ...(node?.workflow === undefined ? {} : { workflow: node.workflow }) }
1049}
1050
1051/** The board with the decision on its agent's node of `turn` (the node made when there is none), `n` the decision's number in the log when it has one. */
1052function withNode(board: Board, turn: number, decision: Decided | NotDecided | Left | Started, n: number | undefined): Board {
1053  const old = board.nodes.find((node) => node.turn === turn && node.id === decision.agent)
1054  // The node that stood for the agent before it started gives it what was decided for it, and goes.
1055  const stood = old === undefined && decision.replaces !== undefined ? board.nodes.find((node) => node.id === decision.replaces) : undefined
1056  const inherited = stood?.decision === undefined ? {} : { decision: stood.decision }
1057  const base = old ?? { ...newNode(turn, decision.agent, decision.node, decision.forTurn === 'next' ? 'queued' : 'running'), ...inherited }
1058  const routed = decision.routed === undefined ? {} : { routed: decision.routed }
1059  let next: BoardNode
1060  if ('failure' in decision) {
1061    next = { ...base, ...routed, why: failureLine(decision.failure.backend, decision.failure), failure: decision.failure }
1062  } else if ('why' in decision) {
1063    next = { ...without(base, 'failure'), routed: false, why: decision.why }
1064  } else if ('started' in decision) {
1065    next = { ...without(base, 'why', 'failure'), routed: true, state: 'running' }
1066  } else {
1067    next = {
1068      ...(decision.routed === true ? without(base, 'why', 'failure') : base),
1069      ...routed,
1070      ...(n === undefined ? {} : { decision: n }),
1071    }
1072  }
1073  return { ...board, nodes: [...board.nodes.filter((node) => node !== old && node !== stood), next] }
1074}
1075
1076/** The value of a cell, or the empty board. */
1077async function read(cell: Cell<Board>): Promise<Board> {
1078  return (await cell.get()).value ?? EMPTY
1079}
1080
1081/** `update` that writes only when `change` has something to write (it returns undefined when the value stands). */
1082async function modify<T>(cell: Cell<T>, change: (current: T | undefined) => T | undefined, attempts = 8): Promise<void> {
1083  for (let attempt = 0; attempt < attempts; attempt++) {
1084    const { value, version } = await cell.get()
1085    const next = change(value)
1086    if (next === undefined) return
1087    if ((await cell.set(next, { ifVersion: version })).isSet) return
1088  }
1089  throw new Error(`$.state write lost to other writers ${attempts} times in a row`)
1090}
1091
1092/** The object without these keys (a $.state value holds no `undefined`). */
1093function without<T extends object, K extends keyof T>(value: T, ...keys: K[]): Omit<T, K> {
1094  const rest = { ...value }
1095  for (const key of keys) delete rest[key]
1096  return rest
1097}
1098
1099// ---- what `report` does with the decision model's fallback ------------------------------
1100
1101/** The decision model that decides is not the one asked for (ADR 0006): e.g. `asked` is pplx, `using` is Jev, since there is no Perplexity key and there is a TypeSafe one. */
1102export type DecisionModelEvent = { asked: BackendName; using: BackendName }
1103
1104/** What a person reads of each decision model: its name, and the settings that hold its key (the options' names and the environment variable they type). The one place both are written. */
1105const DECISION_MODELS: Readonly<Record<BackendName, { name: string; keys: string }>> = {
1106  pplx: { name: 'pplx', keys: 'perplexityApiKey 或环境变量 PERPLEXITY_API_KEY' },
1107  jev: { name: 'Jev', keys: 'typesafeApiKey' },
1108}
1109
1110/** The feature name of the entry the fallback leaves in the decision log. */
1111export const DECISION_MODEL_FEATURE = 'decision-model'
1112
1113/**
1114 * The session's one entry in the decision log saying why Jev decides although pplx is the default: at the turn the session
1115 * start belongs to, with the debug line. A session start again at the same turn (a hot reload) takes the place of the entry
1116 * before. Not on the board's nodes, in no band or footer, and no toast. Never throws.
1117 */
1118async function reportFellBack(io: DecisionModelIo, fell: DecisionModelEvent): Promise<void> {
1119  try {
1120    const asked = DECISION_MODELS[fell.asked]
1121    const using = DECISION_MODELS[fell.using]
1122    // A board that cannot be read puts the session start at turn 1, where the first message will be.
1123    const turn = Math.max(1, (await read(io.board).catch(() => EMPTY)).turn)
1124    const entry: Omit<LogEntry, 'n'> = {
1125      turn,
1126      feature: DECISION_MODEL_FEATURE,
1127      tone: 'info',
1128      outcome: `改用 ${using.name}`,
1129      subject: '',
1130      reason: `没有 ${asked.name} 的密钥(${asked.keys}),改用 ${using.name};补上密钥就用 ${asked.name},decisionModel 设成 ${using.name} 则不再提示`,
1131    }
1132    io.debug(`decision model: ${decisionLine(entry)}`)
1133    await update(io.decisions, (list) => {
1134      const kept = list ?? []
1135      const at = kept.findIndex((old) => old.feature === DECISION_MODEL_FEATURE && old.turn === turn)
1136      return at < 0 ? appendEntry(kept, entry) : kept.map((old, i) => (i === at ? { n: old.n, ...entry } : old))
1137    })
1138  } catch (error) {
1139    try {
1140      io.debug(`decision model fallback not reported: ${errorText(error)}`)
1141    } catch {
1142      // nowhere left to say it
1143    }
1144  }
1145}
1146
1147// ---- what `report` does with the skill profiles' events ----------------------------
1148
1149/** Why writing the profiles stopped for the session (types/index.d.ts, `skillProfiles.stop`). */
1150export type ProfilesStopReason = 'store-read' | 'model-refused' | 'api-error' | 'store-write' | 'off' | 'error'
1151
1152/** How writing the session's skill profiles went (#33; types/index.d.ts, `skillProfiles`). */
1153export type ProfilesState = {
1154  phase: 'writing' | 'done' | 'stopped'
1155  /** The turn the session start belongs to: the turn of its decision log entry. */
1156  turn: number
1157  model: string
1158  kept: number
1159  planned: number
1160  written: number
1161  failed: number
1162  deferred: number
1163  stop?: { reason: ProfilesStopReason; detail: string }
1164  failures: { name: string; reason: string }[]
1165}
1166
1167/** The failed skills the state names; `failed` still counts every one. */
1168export const PROFILES_FAILURES_KEPT = 50
1169
1170/** What the skill profiles' report needs of the host: the board (for the turn), the decision log, the debug log, and the state it keeps. */
1171export type ProfilesIo = Pick<ReportIo, 'board' | 'decisions' | 'debug'> & { profiles: Cell<ProfilesState> }
1172
1173/** Why the writing stopped, with what the debug line and the state say of it. */
1174export type ProfilesStop =
1175  | { reason: 'model-refused'; model: string; detail: string }
1176  | { reason: 'api-error'; skill: string; why: string; ms: number }
1177  | { reason: 'store-write'; skill: string; detail: string }
1178  | { reason: 'off' }
1179  | { reason: 'error'; detail: string }
1180
1181/**
1182 * What happens as the skill profiles are written (features/skill-profiles.ts), in the order it happens. `left`:
1183 * the skills lacking a profile that this session did not get to.
1184 */
1185export type ProfileEvent =
1186  /** The session start has read the catalog: `offered` skills can have one, `due` of them lack one; at most `perSession` are written. */
1187  | { event: 'start'; model: string; perSession: number; offered: number; due: number }
1188  /** The store cannot be read: nothing is kept or written. */
1189  | { event: 'unreadable'; model: string; offered: number }
1190  /** Another session wrote the profile meanwhile: it is kept. */
1191  | { event: 'found' }
1192  | { event: 'written'; skill: string; model: string; ms: number; input: number; output: number }
1193  /** The model gave no profile for the skill (`why`: the reply's reason); the writing goes on. */
1194  | { event: 'failed'; skill: string; why: string; ms: number }
1195  /** The reply for the skill is no profile; the writing goes on. */
1196  | { event: 'unfit'; skill: string; ms: number; text: string }
1197  /** The session's most is written; `left` is what no longer fits (the debug line counts what is not written, failed included). */
1198  | { event: 'quota'; perSession: number; left: number }
1199  /** The loop is over, as it should be. */
1200  | { event: 'finish'; left: number }
hooks/core/setup.ts 559 lines
1// The options the person set (userConfig), read once per load into what every
2// feature receives: `ctx`. Pure: no `$`.
3//
4// Every option is read here, and only here: its bounds and its fallback are
5// written once, and every feature (and the eval, which builds its requests
6// from the same Config) reads the value from `ctx.config`. The fallbacks are
7// the manifest's defaults (the engine fills those in; a test or the eval may
8// leave an option out), except for the options whose default depends on the
9// decision model: those have no default in the manifest, so the engine
10// passes nothing for them until the person sets one, and their defaults are
11// in BACKEND_DEFAULTS below, the one table the mod, the eval and
12// scripts/decide*.ts read them from.
13
14import type { PluginOptions } from 'claude-code'
15import type { Backend } from '../decision/backend.ts'
16import type { ContextLimits } from '../decision/context.ts'
17import { DEFAULT_AGENT_MODELS, type AgentModel, type DispatchAsk, type DispatchSettings } from '../decision/dispatched-agent.ts'
18import type { EffortAsk, EffortRules } from '../decision/effort.ts'
19import { RAISE_MODES, type RaiseMode } from '../decision/escalation.ts'
20import { jevBackend } from '../decision/jev.ts'
21import type { MidturnLimits, MidturnRules } from '../decision/midturn.ts'
22import { resolveModel, type ResolvedModel } from '../decision/model-ids.ts'
23import { pplxBackend } from '../decision/pplx.ts'
24import { rateLimited } from '../decision/pplx-rate.ts'
25import type { SkillPolicy } from '../decision/skills.ts'
26
27/** The decision models the person can choose (`decisionModel`): Perplexity's pplx-decider and TypeSafe's Jev. */
28export type BackendName = 'pplx' | 'jev'
29
30/**
31 * The decision model the person asked for: Jev when `decisionModel` names it, else pplx, the default (ADR 0006). Anything
32 * that names neither (unset, a typo, the `clef` of a 0.3.x configuration) reads as unset. Which one decides in the end also
33 * depends on the keys (`chooseBackend`).
34 */
35export function backendNameOf(options: PluginOptions): BackendName {
36  return options.decisionModel === 'jev' ? 'jev' : 'pplx'
37}
38
39/** Which decision model decides, and whether it is not the one asked for. */
40export type Choice = { backend: BackendName; fellBack: boolean }
41
42/**
43 * The decision model that decides (ADR 0006). Jev when the person asked for Jev. Otherwise pplx when there is a Perplexity key;
44 * with none, Jev when there is a TypeSafe key (`fellBack`: a 0.3.1 configuration keeps its routing); with neither, pplx, which
45 * then fails every request with the Perplexity key named as the one missing.
46 */
47export function chooseBackend(asked: BackendName, keys: { perplexity: string; typesafe: string }): Choice {
48  if (asked === 'jev') return { backend: 'jev', fellBack: false }
49  if (keys.perplexity !== '') return { backend: 'pplx', fellBack: false }
50  return keys.typesafe !== '' ? { backend: 'jev', fellBack: true } : { backend: 'pplx', fellBack: false }
51}
52
53/**
54 * A `decisionModel` that names a decision model since removed (Clef, 0.4.0), left in the person's settings from before.
55 * The engine reads a value outside the option's list as the option's default (and warns), so the mod is handed the
56 * default, never the old value: this finds it in the settings file itself (`pluginConfigs[dispatch-pilot@...].options`,
57 * as `$.settings.read({ source: 'user' })` returns them). Null when the settings name none.
58 */
59export function removedDecisionModel(settings: unknown): 'clef' | null {
60  const configs = (settings as { pluginConfigs?: unknown } | null | undefined)?.pluginConfigs
61  if (typeof configs !== 'object' || configs === null) return null
62  for (const [plugin, config] of Object.entries(configs)) {
63    if (!plugin.startsWith('dispatch-pilot@')) continue
64    const options = (config as { options?: unknown } | null)?.options
65    if (typeof options === 'object' && options !== null && (options as { decisionModel?: unknown }).decisionModel === 'clef') return 'clef'
66  }
67  return null
68}
69
70/** The options whose default depends on the decision model: the manifest gives them none. */
71export const PER_BACKEND_OPTIONS = [
72  'timeoutMs',
73  'contextMessages',
74  'contextTokens',
75  'rejudgeSteps',
76  'rejudgeWaitMs',
77  'thetaUp',
78  'thetaDown',
79  'thetaMax',
80  'thetaExpected',
81  'agentOverride',
82  'skillsMinRelevance',
83  'findSkillMinRelevance',
84] as const
85export type PerBackendOption = (typeof PER_BACKEND_OPTIONS)[number]
86
87/**
88 * The kinds of request whose state has a budget of its own (`contextTokens`
89 * reads by kind): a message that carries the skills' question (and find_skill's
90 * first stage) is `Config.context`; the rest are these. `rejudge` is a
91 * mid-turn re-decision and a stuck one.
92 */
93export const CONTEXT_KINDS = ['messagePlain', 'rejudge', 'agent', 'workflow'] as const
94export type ContextKind = (typeof CONTEXT_KINDS)[number]
95
96/**
97 * What one decision model brings: each option's default where the person left
98 * it unset, the most `contextTokens` and `contextMessages` read as, how every
99 * question is asked and the threshold for taking the level above.
100 *
101 * Adding a decision model means touching, beyond its entry here: `BackendName`; the backend in decision/ and its
102 * construction in `setup()` (`configs`, `backends`, and whether it passes through the rate limit); `backendNameOf` and
103 * `chooseBackend` (which key decides, and what it falls back to); `DECISION_MODELS` in core/report.ts (its name and key
104 * settings, for the fallback entry); `decisionModel`'s options in plugin.json, its key option and the README's
105 * configuration tables (a column each); and, for the eval, `backendFor` and `PRICES` in eval/node.ts, `--backend` in
106 * eval/run.ts and eval/eval-v2-flow.ts, and the BACKENDS list in eval/lib/docs.ts. A `Record<BackendName, ...>` the compiler
107 * checks (this table, `setup()`, PRICES, DECISION_MODELS); the rest it does not, and tests/docs-sync.test.ts only the README tables.
108 */
109export type BackendDefaults = Readonly<Record<PerBackendOption, number>> & {
110  /** `contextTokens` above this reads as this. */
111  contextTokensMax: number
112  /** `contextMessages` above this reads as this. */
113  contextMessagesMax: number
114  /**
115   * The least probability the level above the most probable one needs to be taken instead (`EffortRules.roundUp`, decision/effort.ts).
116   * Not an option: set from the stored answers of the decision model (eval/resummarize.ts), like the other values of the table.
117   */
118  roundUp: number
119  /**
120   * How the decision model is asked, language and kind of question (the eval's variables, `effort-submit` for the effort question
121   * beside a message): `turnStart` for that question, `other` for every other one (the mid-turn re-decision, a dispatched agent
122   * and a Workflow's, a stuck loop, the skills, the problem summary's text). Not options.
123   */
124  ask: { turnStart: EffortAsk; other: EffortAsk }
125  /**
126   * What the state of each other kind of request may take, in estimated tokens (`contextTokens` is the message that carries the
127   * skills' question). Not options: each is the most that kind's longest question leaves, worked out in DEVELOPMENT.md (配置,
128   * 「Jev 的上下文默认值怎么算」) and checked by tests/backend-defaults.test.ts. What the person sets as `contextTokens` holds for
129   * every kind, each taking the smaller of it and its own.
130   */
131  contextByKind: Readonly<Record<ContextKind, number>>
132  /**
133   * How long find_skill's two requests may take in all, in ms; null: as long
134   * as a message waits (`timeoutMs`). Not an option: set by the latency
135   * measured, not calibrated. A hook has 10 s of its own, and the timer's wait
136   * counts toward it (the mods reference, Limits): this leaves room for the rest.
137   */
138  findSkillWaitMs: number | null
139  /** Whether find_skill's first request offers a skill by its profile (else by its description); the second re-reads it by both. */
140  findSkillProfiles: boolean
141}
142
143/**
144 * Jev's: the values the eval of #4, #14, #15 and #16 set or kept (README, 配置 has the table; DEVELOPMENT.md, 配置 what each rests on),
145 * except the three that say how much Jev reads of the conversation: by the person's decision ("the more it is given, the
146 * better it judges"), those are set to what Jev accepts, not to what the eval ran (4 messages, 2000 tokens, 4 steps; the
147 * eval sets are too short to tell, and no comparison was run). Jev takes 64k tokens a request and 32k for "the state and
148 * the longest question". The longest question is the first stage of the skills (111 skills with profiles: about 16k tokens
149 * as the mod counts, 21.9k as Jev does), which shares its request with the state, so `contextTokens` is what that leaves,
150 * 6000 (DEVELOPMENT.md, 配置, 「Jev 的上下文默认值怎么算」; tests/backend-defaults.test.ts checks the sum). The two counts
151 * that fill it are at their most: `contextMessages` 32 and `rejudgeSteps` 16, so the token budget, not the count, ends what is sent.
152 */
153const JEV_DEFAULTS: BackendDefaults = {
154  timeoutMs: 1500,
155  contextMessages: 32,
156  contextMessagesMax: 32,
157  contextTokens: 6000,
158  contextTokensMax: 16000,
159  // Every other kind of request has questions of at most 700 tokens as Jev counts them (the effort question 421, a re-decision's
160  // 295 with its trouble, an agent's 457, nine of them 1,836): the state plus the longest of them within 28,800 (90% of 32k)
161  // allows about 25,400 estimated tokens, and 24000 is the round number below it. A Workflow batch (at most 8 agents' questions,
162  // 20,100 in all) and the whole request stay within 57,600 (90% of 64k) too.
163  contextByKind: { messagePlain: 24000, rejudge: 24000, agent: 24000, workflow: 24000 },
164  rejudgeSteps: 16,
165  // A step waits this long for a re-decision not yet back before it goes on at the level it had: Jev's mid-turn requests took about 280 ms
166  // at p50 and 350 ms at p90 (DEVELOPMENT.md, 配置), and they are sent when the step's tools start, so they are mostly back by then.
167  rejudgeWaitMs: 300,
168  // Raising is easy, lowering is hard (AA: a Sonnet 5.5 at medium scores 41 on the index and at high 47, at low 36; Terminal-Bench
169  // 20.7% at low against 43.9% at high): 0.4 to 0.3 for a raise, 0.6 to 0.75 for a lowering, 3 to 5 steps held after a raise.
170  // The lowering gate came back down to 0.55 in 0.2.3: with 0.75 the level sent was too high too often (too high 11.5/11.0% to
171  // 19.5/17.5%); a scan of the stored effort-midturn answers (eval/rescore.ts --theta-down) found none of 0.55 to 0.75 to bring it back, and 0.55 is the one with the least too high and too low together.
172  thetaUp: 0.3,
173  thetaDown: 0.55,
174  thetaMax: 0.5,
175  thetaExpected: 0.25,
176  agentOverride: 0.6,
177  // #16, both runs on the old wording: from 0.7 to 0.75 the Chinese-English gap of the skill suggestions went from -3.7 to -2.3 points.
178  skillsMinRelevance: 0.75,
179  findSkillMinRelevance: 0.5,
180  findSkillWaitMs: null,
181  findSkillProfiles: true,
182  // Raising is easy (DEVELOPMENT.md, 「按 AA 基准校正」): the level above the most probable one is taken from 0.3.
183  roundUp: 0.3,
184  // effort-submit on the current wording, one run each (2026-10-05): asked in Chinese, Chinese items 85% and English
185  // items 89%; asked in English, 79% and 78%. The questions asked later have no data in Chinese and stay in English.
186  ask: { turnStart: { language: 'zh', primitive: 'score' }, other: { language: 'en', primitive: 'score' } },
187}
188
189/**
190 * pplx-decider-v1.1-27b's: B′ of eval v2 (#45, ADR 0006: the state within 48000 tokens, every question in English, no problem
191 * summary), with the thresholds calibrated on its own stored answers (DEVELOPMENT.md, eval v2 的结果). It takes a 262k-token
192 * window, so what it reads is not cut to Jev's 32k: 2000 messages (the budget, not the count, ends what is sent) and 48000
193 * tokens for every request but the skills'; those two stages keep Jev's 6000, since the profiles of the skills are the same
194 * length for either model. A request takes seconds, not Jev's fraction of one, so a message waits 8000 ms (a hook has 10 s) and a
195 * step 6000 ms for a re-decision. What the eval did not calibrate for pplx (the failure bar, the agents' override, the skills'
196 * relevance) is Jev's.
197 */
198const PPLX_DEFAULTS: BackendDefaults = {
199  timeoutMs: 8000,
200  contextMessages: 2000,
201  contextMessagesMax: 2000,
202  contextTokens: 6000,
203  contextTokensMax: 48000,
204  contextByKind: { messagePlain: 48000, rejudge: 48000, agent: 48000, workflow: 48000 },
205  rejudgeSteps: 16,
206  // A re-decision takes seconds, not 300 ms: a step waits for it up to this long (the most the manifest allows is 8000).
207  rejudgeWaitMs: 6000,
208  // Calibrated offline on the stored pplx answers (eval v2, DEVELOPMENT.md): thetaMax 0.47 (0.48 sits on submit-034's p(max) of
209  // 0.480), thetaDown 0.55 as Jev's (from 0.6 up, the English score question sends the level too high more often than Jev).
210  // thetaUp 0 raises a level mid-turn on any answer above it (0.3 is the cautious alternative); it does not move the level a
211  // message is sent at, which thetaMax and roundUp decide.
212  thetaUp: 0,
213  thetaDown: 0.55,
214  thetaMax: 0.47,
215  thetaExpected: 0.25,
216  agentOverride: 0.6,
217  skillsMinRelevance: 0.75,
218  findSkillMinRelevance: 0.5,
219  // Find_skill's two requests wait as long as the profiles' question takes (a message's wait would be 8000 ms of a 10 s hook).
220  findSkillWaitMs: 6000,
221  findSkillProfiles: true,
222  // The level above the most probable one is taken from 0.45 (ADR 0006): the offline scan of the stored effort-submit answers
223  // (English questions) took the level sent too high from 18.5% to 13.0% with the recall of max unchanged (11 of 16); eval v2's
224  // level too low went from 12.5% to 15.0%.
225  roundUp: 0.45,
226  ask: { turnStart: { language: 'en', primitive: 'score' }, other: { language: 'en', primitive: 'score' } },
227}
228
229/** The defaults by decision model: the only place they are written. */
230export const BACKEND_DEFAULTS: Readonly<Record<BackendName, BackendDefaults>> = {
231  pplx: PPLX_DEFAULTS,
232  jev: JEV_DEFAULTS,
233}
234
235/**
236 * What the state of a message's decision request may take: the message that
237 * carries the skills' question (`withSkills`) has `context`, whose tokens
238 * leave room for that question; the effort question's request, and any other
239 * plain message, has `contextByKind.messagePlain` (ADR 0005). The core and the
240 * eval build the request's state from the same limits.
241 */
242export function messageLimits(config: Pick<Config, 'context' | 'contextByKind'>, withSkills: boolean): ContextLimits {
243  return withSkills ? config.context : { ...config.context, tokens: config.contextByKind.messagePlain }
244}
245
246/** The settings are also the effort rules' parameters (`EffortRules`: `thetaMax` and `roundUp`): `traceEffort(reading, config)`. */
247export type Config = EffortRules & {
248  /** The decision model the options were read for: its defaults (BACKEND_DEFAULTS) stand for what the person left unset. */
249  backend: BackendName
250  /**
251   * What came from that decision model's defaults (for the debug log, describeDefaults): the options the person left
252   * unset, with the value each took, and an option set above that model's most, read as the most.
253   */
254  defaults: { used: readonly (readonly [PerBackendOption, number])[]; capped: readonly { option: PerBackendOption; set: number; read: number }[] }
255  /** The TypeSafe key in the options (`typesafeApiKey`, trimmed; '' if unset). Not enumerable (see readConfig): it is not in a written-out or copied config. */
256  readonly typesafeApiKey: string
257  /** The Perplexity key in the options (`perplexityApiKey`, trimmed; '' if unset; not enumerable either): the environment's (`Secrets`) stands where this is empty. */
258  readonly perplexityApiKey: string
259  /** How many pplx requests the mod sends in a second at most (`pplxQps`; the account's limit): the rest wait their turn (decision/pplx-rate.ts). Jev has no such limit. */
260  pplxQps: number
261  /** How long a decision request may take before the prompt goes on without it. */
262  timeoutMs: number
263  /** How the decision model is asked (BACKEND_DEFAULTS ask): `turnStart` the effort question beside each message, `other` every other question (`ctx.ask`). */
264  ask: { turnStart: EffortAsk; other: EffortAsk }
265  /**
266   * What the decision model reads of the conversation: how many recent messages, how many tokens in all. The tokens are the
267   * budget of a message that carries the skills' question (and of find_skill's first stage); the other kinds of request have
268   * theirs in `contextByKind` (a mid-turn re-decision's in `midturn.limits`, too).
269   */
270  context: ContextLimits
271  /** The tokens the state of each other kind of request may take: the decision model's default for the kind, or what the person set if smaller. */
272  contextByKind: Readonly<Record<ContextKind, number>>
273  /** The main agent's effort decided again while a turn runs (#5); a stuck loop's re-decision (#7) reads its loop the same way. */
274  midturn: {
275    /** Re-decide at every step whose index is a multiple of this; 0 for never. */
276    every: number
277    /** How long a step waits for a re-decision (or a stuck one) not back yet. */
278    waitMs: number
279    /** What a re-decision reads: the latest steps, within the context budget. */
280    limits: MidturnLimits
281    /** How an answer moves the level: thetaUp, thetaDown, thetaMax, roundUp, holdSteps. */
282    rules: MidturnRules
283  }
284  /** Forced escalation (#7). */
285  escalation: {
286    /** Counted failures that make a loop stuck. */
287    after: number
288    mode: RaiseMode
289    /** Forced raises a loop gets at most. */
290    limit: number
291    /** The failures count as expected when the answer reaches this probability. */
292    thetaExpected: number
293    /** The model a failing haiku agent is switched to, resolved to a full id; null: never switch (left empty, or a name no model family has). */
294    haikuTo: ResolvedModel | null
295    /** `escalateHaikuTo` as the person wrote it, for the log when it names no model. */
296    haikuToWritten: string
297  }
298  /** Dispatched agents and a Workflow's agents (#6, #8, #9). */
299  agents: {
300    /** The models the decision model may choose, cheapest first (fable with `agentFable`). */
301    models: readonly AgentModel[]
302    /** The decision model's pick replaces the main agent's (or the script's) only at this confidence or above. */
303    thetaOverride: number
304    /** `rewrite` writes a Workflow's decisions into its script; `return` sends the Workflow back with them, once. */
305    workflowMode: 'rewrite' | 'return'
306  }
307  /** Skills (#10, #11, #12). */
308  skills: {
309    /** What a message is suggested: at most `max`, from `minRelevance`. */
310    suggest: SkillPolicy
311    /** What find_skill returns: at most `max`, from `minRelevance`. */
312    find: SkillPolicy
313    /** How long find_skill's two requests may take in all (the decision model's: BACKEND_DEFAULTS findSkillWaitMs). */
314    findWaitMs: number
315    /** Whether find_skill's first request offers skills by their profiles (the decision model's: findSkillProfiles). */
316    findByProfile: boolean
317    /** How many of stage one's best stage two re-reads. */
318    shortlist: number
319    /** Skills the main agent keeps in its listing (names as the listing spells them). */
320    alwaysListed: readonly string[]
321    /** Skills never offered, to the main agent or to the person. */
322    neverSuggested: readonly string[]
323    /** The model that writes skill profiles; part of every profile's key. */
324    profileModel: string
325    /** How many missing profiles a session start writes at most. */
326    profilesPerSession: number
327  }
328  /** The unresolved count and the problem summary (#39, #40). */
329  unresolved: {
330    /** The model that writes the problem summary after each of the person's turns (`summaryModel`). */
331    summaryModel: string
332    /** The count from which the effort question of a message and of a mid-turn re-decision carries the strong hint (`unresolvedMaxAfter`; 0: never). */
333    maxAfter: number
334  }
335}
336
337/**
338 * What the mod reads from the environment: `$.env.get` is asynchronous and needs `$`, which neither `setup` nor an import may
339 * have, so the session start (features/control.ts) reads it into this object, which `ctx.backend` reads when it asks. A
340 * value set in the options always comes first. Nothing here is ever written down: not to the debug log, the board or `$.state`.
341 */
342export type Secrets = {
343  /** `PERPLEXITY_API_KEY`, trimmed; '' while unread or unset. */
344  perplexityEnvKey: string
345}
346
347/** The Perplexity key in use: the options' (`perplexityApiKey`) first, else the environment's. '' for none. */
348export function perplexityKey(config: Pick<Config, 'perplexityApiKey'>, secrets: Secrets): string {
349  return config.perplexityApiKey !== '' ? config.perplexityApiKey : secrets.perplexityEnvKey
350}
351
352/**
353 * What the entry hands every feature's register: data and pure functions, never `$`.
354 *
355 * Which decision model decides is not known at register: the Perplexity key may be in the environment, which the session start
356 * reads (`secrets`). `config`, `backend`, `ask` and `fellBack` are therefore read afresh at each use and follow the choice: a
357 * feature that keeps a value from `ctx.config` when it registers keeps the wrong model's (the defaults differ), so it reads
358 * `ctx.config` where it acts.
359 */
360export type Ctx = {
361  /** The settings of the decision model that decides (`chooseBackend`). */
362  readonly config: Config
363  /** The decision model that decides. */
364  readonly backend: Backend
365  /** What the session start read from the environment (a key not in the options). */
366  readonly secrets: Secrets
367  /** How every question but the effort question beside a message is asked: the decision model's (`config.ask.other`). */
368  readonly ask: EffortAsk
369  /** Whether Jev decides only because there is no Perplexity key (a TypeSafe key is set): the session start records it. */
370  readonly fellBack: boolean
371}
372
373export function setup(options: PluginOptions, table: Readonly<Record<BackendName, BackendDefaults>> = BACKEND_DEFAULTS): Ctx {
374  const asked = backendNameOf(options)
375  // Both are worked out here, whichever decides: pure and cheap, and `chooseBackend` picks between them once the environment is read.
376  const configs: Record<BackendName, Config> = { pplx: readConfig(options, table, 'pplx'), jev: readConfig(options, table, 'jev') }
377  const secrets: Secrets = { perplexityEnvKey: '' }
378  const backends: Record<BackendName, Backend> = {
379    // Every pplx request goes through the rate limit (#51); Jev does not.
380    pplx: rateLimited(pplxBackend(() => perplexityKey(configs.pplx, secrets)), configs.pplx.pplxQps),
381    jev: jevBackend(configs.jev.typesafeApiKey),
382  }
383  const choice = () => chooseBackend(asked, { perplexity: perplexityKey(configs.pplx, secrets), typesafe: configs.pplx.typesafeApiKey })
384  return {
385    get config() {
386      return configs[choice().backend]
387    },
388    get backend() {
389      return backends[choice().backend]
390    },
391    secrets,
392    get ask() {
393      return configs[choice().backend].ask.other
394    },
395    get fellBack() {
396      return choice().fellBack
397    },
398  }
399}
400
401/**
402 * The person's options, each within its bounds; where one is missing or of
403 * the wrong type, the manifest's default, or for an option whose default
404 * depends on the decision model (PER_BACKEND_OPTIONS), that model's
405 * (`table`, BACKEND_DEFAULTS unless a test reads another). `contextTokens`
406 * and `contextMessages` read at most that model's most. `backend` is the
407 * decision model it reads for: what `decisionModel` asks for unless the
408 * keys decide otherwise (`setup` reads both).
409 */
410export function readConfig(options: PluginOptions, table: Readonly<Record<BackendName, BackendDefaults>> = BACKEND_DEFAULTS, backend: BackendName = backendNameOf(options)): Config {
411  const defaults = table[backend]
412  const used: (readonly [PerBackendOption, number])[] = []
413  const capped: { option: PerBackendOption; set: number; read: number }[] = []
414  /** A per-backend option: the person's value within [min, max], else the decision model's default (noted for the log). */
415  const own = (option: PerBackendOption, min: number, max: number, round = false) => {
416    const set = options[option]
417    if (typeof set !== 'number' || !Number.isFinite(set)) {
418      used.push([option, defaults[option]])
419      return defaults[option]
420    }
421    const read = numberIn(set, min, max, defaults[option])
422    const value = round ? Math.round(read) : read
423    if (set > max) capped.push({ option, set, read: value })
424    return value
425  }
426  const whole = (value: unknown, min: number, max: number, fallback: number) => Math.round(numberIn(value, min, max, fallback))
427  const thetaMax = own('thetaMax', 0, 1)
428  // `contextTokens` by kind of request: unset, each kind's own (BACKEND_DEFAULTS); set, the smaller of it and each kind's.
429  const contextTokens = own('contextTokens', 100, defaults.contextTokensMax, true)
430  const contextSet = typeof options.contextTokens === 'number' && Number.isFinite(options.contextTokens)
431  const byKind = (cap: number) => (contextSet ? Math.min(contextTokens, cap) : cap)
432  const context = { messages: own('contextMessages', 0, defaults.contextMessagesMax, true), tokens: byKind(defaults.contextTokens) }
433  const contextByKind = Object.fromEntries(CONTEXT_KINDS.map((kind) => [kind, byKind(defaults.contextByKind[kind])])) as Record<ContextKind, number>
434  const haikuToWritten = stringOf(options.escalateHaikuTo, 'sonnet').trim()
435  // A hook's own budget is 10 s and the timer's wait counts toward it.
436  const timeoutMs = own('timeoutMs', 200, 8000)
437  const config: Omit<Config, 'typesafeApiKey' | 'perplexityApiKey'> = {
438    backend,
439    defaults: { used, capped },
440    pplxQps: whole(options.pplxQps, 1, 50, 1),
441    timeoutMs,
442    ask: defaults.ask,
443    thetaMax,
444    roundUp: defaults.roundUp,
445    context,
446    contextByKind,
447    midturn: {
448      every: whole(options.rejudgeEvery, 0, 50, 3),
449      // A hook's own budget is 10 s and the timer's wait counts toward it, as for `timeoutMs`.
450      waitMs: own('rejudgeWaitMs', 0, 8000, true),
451      limits: { steps: own('rejudgeSteps', 1, 16, true), tokens: contextByKind.rejudge },
452      rules: {
453        thetaUp: own('thetaUp', 0, 1),
454        thetaDown: own('thetaDown', 0, 1),
455        thetaMax,
456        roundUp: defaults.roundUp,
457        holdSteps: whole(options.holdSteps, 0, 50, 5),
458      },
459    },
460    escalation: {
461      after: whole(options.escalateAfter, 1, 20, 2),
462      mode: RAISE_MODES.find((mode) => mode === options.escalateMode) ?? 'one-level',
463      limit: whole(options.escalateLimit, 0, 10, 2),
464      thetaExpected: own('thetaExpected', 0, 1),
465      haikuTo: haikuToWritten === '' ? null : resolveModel(haikuToWritten),
466      haikuToWritten,
467    },
468    agents: {
469      models: options.agentFable === true ? [...DEFAULT_AGENT_MODELS, 'fable'] : DEFAULT_AGENT_MODELS,
470      thetaOverride: own('agentOverride', 0, 1),
471      workflowMode: stringOf(options.workflowMode, 'rewrite') === 'return' ? 'return' : 'rewrite',
472    },
473    skills: {
474      suggest: { max: whole(options.skillsMax, 0, 10, 3), minRelevance: own('skillsMinRelevance', 0, 1) },
475      find: { max: whole(options.findSkillMax, 1, 10, 5), minRelevance: own('findSkillMinRelevance', 0, 1) },
476      findWaitMs: defaults.findSkillWaitMs ?? timeoutMs,
477      findByProfile: defaults.findSkillProfiles,
478      shortlist: whole(options.skillsShortlist, 1, 10, 4),
479      alwaysListed: namesOf(options.skillsAlwaysListed),
480      neverSuggested: namesOf(options.skillsNeverSuggested),
481      profileModel: stringOf(options.skillsProfileModel, DEFAULT_PROFILE_MODEL).trim() || DEFAULT_PROFILE_MODEL,
482      profilesPerSession: whole(options.skillsProfilesPerSession, 0, 500, 30),
483    },
484    unresolved: {
485      summaryModel: stringOf(options.summaryModel, DEFAULT_SUMMARY_MODEL).trim() || DEFAULT_SUMMARY_MODEL,
486      maxAfter: whole(options.unresolvedMaxAfter, 0, 10, 3),
487    },
488  }
489  // In the table's order, whatever order they were read in.
490  used.sort((a, b) => PER_BACKEND_OPTIONS.indexOf(a[0]) - PER_BACKEND_OPTIONS.indexOf(b[0]))
491  // The keys are not enumerable: `JSON.stringify(config)`, `{ ...config }` and `Object.keys` do not meet them, so a
492  // config written to a log, a result file or the board (the eval saves its settings) cannot carry a key by accident.
493  return Object.defineProperties(config, {
494    typesafeApiKey: { value: stringOf(options.typesafeApiKey, '').trim(), enumerable: false },
495    perplexityApiKey: { value: stringOf(options.perplexityApiKey, '').trim(), enumerable: false },
496  }) as Config
497}
498
499/**
500 * One debug-log line on what the decision model's defaults decided, e.g.
501 * `settings for jev: left unset, so jev's defaults: timeoutMs 1500, ...;
502 * skill suggestions on until /dp skills off; contextTokens 20000 reads as
503 * 16000, the most with jev`.
504 */
505export function describeDefaults(config: Pick<Config, 'backend' | 'defaults'>): string {
506  const { backend, defaults } = config
507  const used = defaults.used.length === 0 ? 'every option set' : `left unset, so ${backend}'s defaults: ${defaults.used.map(([option, value]) => `${option} ${value}`).join(', ')}`
508  const capped = defaults.capped.map((cap) => `; ${cap.option} ${cap.set} reads as ${cap.read}, the most with ${backend}`).join('')
509  return `settings for ${backend}: ${used}; skill suggestions on until /dp skills off${capped}`
510}
511
512/**
513 * How an agent's model and effort are decided (decision/dispatched-agent.ts):
514 * the same for a dispatched agent (#6) and a Workflow's agents (#8, #9), and
515 * for the eval; `ask` is how its questions are written (the eval's variants).
516 */
517export function dispatchSettings(ctx: { config: Pick<Config, 'agents' | 'thetaMax' | 'roundUp'>; ask: EffortAsk }, ask: Partial<DispatchAsk> = {}): DispatchSettings {
518  // Each read goes to `ctx`: a feature builds this when it registers, before the session start has settled the decision model.
519  return {
520    get models() {
521      return ctx.config.agents.models
522    },
523    get ask() {
524      return { ...ctx.ask, ...ask }
525    },
526    get thetaOverride() {
527      return ctx.config.agents.thetaOverride
528    },
529    get thetaMax() {
530      return ctx.config.thetaMax
531    },
532    get roundUp() {
533      return ctx.config.roundUp
534    },
535  }
536}
537
538/** The cheap model that writes skill profiles, unless the person names another (`skillsProfileModel`). */
539export const DEFAULT_PROFILE_MODEL = 'haiku'
540
541/** The cheap model that writes the problem summary, unless the person names another (`summaryModel`). */
542export const DEFAULT_SUMMARY_MODEL = 'haiku'
543
544/** A numeric option clamped to [min, max]; `fallback` when it is not a number. */
545export function numberIn(value: unknown, min: number, max: number, fallback: number): number {
546  return typeof value === 'number' && Number.isFinite(value) ? Math.min(max, Math.max(min, value)) : fallback
547}
548
549/** A string option; `fallback` when it is not a string. A sensitive one never set arrives as ''. */
550export function stringOf(value: unknown, fallback: string): string {
551  return typeof value === 'string' ? value : fallback
552}
553
554/** A list-of-names option (`multiple` in the manifest; a comma-separated string also reads): trimmed, empty ones dropped. */
555export function namesOf(value: unknown): string[] {
556  const items = Array.isArray(value) ? value : typeof value === 'string' ? value.split(',') : []
557  return items.flatMap((item) => (typeof item === 'string' && item.trim() !== '' ? [item.trim()] : []))
558}
559
hooks/features/control.ts 211 lines
1// Feature: the person's control over Dispatch Pilot, the `/dp` command.
2//
3//   /dp                  opens the rationale pane (依据面板), or closes it when open
4//   /dp log              opens the same pane (the old habit's name for it)
5//   /dp status           what is on, the lock
6//   /dp on | off         the whole mod
7//   /dp <name> on | off  one feature (each registers its own, core/switches.ts)
8//   /dp lock <effort>    hold the main agent at that effort on every step
9//   /dp unlock           release it
10//   /dp log N            the last N decisions and why, in the conversation
11//
12// A pane the surface does not place (one that places no panes) is closed again
13// at once, and the answer says so (spec #22: no pane left open and waiting).
14//
15// What the person flips is kept in $.store and loaded back at session start.
16// The engine puts the plugin's name before a command's answer, so the texts
17// here do not start with one.
18//
19// It also records what the engine measures of the session (context fill, limit
20// percentages, cost) in the debug log. Only recorded: nothing reads them to
21// decide anything (spec, 可见性与控制; a later quota-saving mode may).
22
23import type { On, SessionMeasureInput } from 'claude-code'
24import { EFFORTS, isEffort, type Effort } from '../decision/effort.ts'
25import { decisionLine, LOG_ENTRIES, report, unplacedText, type DecisionModelIo, type LogEntry, type SwitchIo } from '../core/report.ts'
26import { describeDefaults, removedDecisionModel, type Ctx } from '../core/setup.ts'
27import { defineSwitch, isOn, listSwitches, loadOverrides, masterOn, overrides, parseOverrides, setMaster, setSwitch } from '../core/switches.ts'
28import { errorText } from '../decision/backend.ts'
29import { PANE_COLUMNS, PANE_ID, PANE_TITLE } from '../board/rationale.ts'
30
31const LOCK = { plugin: 'dispatch-pilot', key: 'lock' } as const
32const DECISIONS = { plugin: 'dispatch-pilot', key: 'decisionLog' } as const
33const BOARD = { plugin: 'dispatch-pilot', key: 'board' } as const
34/** The switches the person flipped, in $.store. */
35const SWITCHES_KEY = 'switches'
36
37export function registerControl(on: On, ctx: Ctx): void {
38  defineSwitch({ name: 'signals', info: '把上下文、限额和花费的读数记进 debug log,不参与任何决策' })
39
40  // Under a match-all matcher: other features set themselves up at session start too.
41  on('session.start', { cwd: /(?:)/ }, async ($, e, next) => {
42    // The person's switches first: the features beneath read them while the session starts.
43    loadOverrides(await $.store.get(SWITCHES_KEY).catch(() => undefined))
44    // The environment's Perplexity key (the options' comes first), for the requests that follow; never written down, only where it came from.
45    // Read before anything is said of the decision model: whether it is pplx or Jev depends on it (core/setup.ts, `Ctx`).
46    ctx.secrets.perplexityEnvKey = ((await $.env.get('PERPLEXITY_API_KEY').catch(() => undefined)) ?? '').trim()
47    // Which options the decision model's defaults decided (core/setup.ts BACKEND_DEFAULTS).
48    $.ui.log(describeDefaults(ctx.config), { to: 'debug' })
49    // Jev decides only for want of a Perplexity key: the decision log says so (ADR 0006), where the person looks for why.
50    if (ctx.fellBack) {
51      const io: DecisionModelIo = {
52        board: { get: () => $.state.get(BOARD), set: (value, options) => $.state.set(BOARD, value, options) },
53        decisions: { get: () => $.state.get(DECISIONS), set: (value, options) => $.state.set(DECISIONS, value, options) },
54        debug: (line) => $.ui.log(line, { to: 'debug' }),
55      }
56      await report(io, { fellBack: { asked: 'pplx', using: 'jev' } })
57    }
58    if (ctx.config.backend === 'pplx') $.ui.log(`pplx key: ${ctx.config.perplexityApiKey !== '' ? 'from the options' : ctx.secrets.perplexityEnvKey !== '' ? 'from PERPLEXITY_API_KEY' : 'not set'}`, { to: 'debug' })
59    // A decisionModel left over in the person's settings that names a model since removed: the engine reads it as the default.
60    const removed = removedDecisionModel(await $.settings.read({ source: 'user' }).catch(() => undefined))
61    if (removed !== null) $.ui.log(`decisionModel ${removed} 已移除,按没设处理`, { to: 'debug' })
62    const result = await next(e)
63    await $.command
64      .register({
65        name: 'dp',
66        description: 'Dispatch Pilot:依据面板、功能开关、effort 锁定、最近的决定',
67        argumentHint: '[log | status | on|off | <name> on|off | lock <effort> | unlock | log N]',
68        immediate: true,
69      })
70      .catch((error) => $.ui.log(`/dp was not registered: ${errorText(error)}`, { to: 'debug' }))
71    return result
72  })
73
74  // `window` is in every measurement: a matcher on a field that is always there.
75  on('session.measure', { context: { window: /(?:)/ } }, async ($, e, next) => {
76    if (isOn('signals')) {
77      try {
78        $.ui.log(describeSignals(e), { to: 'debug' })
79      } catch {
80        // a reading that cannot be written down is skipped
81      }
82    }
83    return next(e)
84  })
85
86  on('command.run', { command: 'dp' }, async ($, e) => {
87    const command = parseControl(e.args)
88    // The screens draw again: they leave out what a feature that is off owns.
89    const reporting: SwitchIo = { redraw: () => $.ui.invalidate('ui.render') }
90    // Keeps one switch as the person just flipped it, beside whatever another session saved meanwhile;
91    // says so when it cannot.
92    const save = async (name: string) => {
93      try {
94        const saved = parseOverrides(await $.store.get(SWITCHES_KEY))
95        const now = overrides()[name]
96        if (now === undefined) delete saved[name]
97        else saved[name] = now
98        await $.store.set(SWITCHES_KEY, saved)
99        return ''
100      } catch {
101        return '(没有保存:本地存储用不了,只在这个会话里有效)'
102      }
103    }
104    try {
105      switch (command.kind) {
106        case 'pane': {
107          // `/dp` toggles; `/dp log` only opens (and asks for the keys again).
108          if (command.toggle && (await $.ui.panes()).some((pane) => pane.id === PANE_ID)) {
109            await $.ui.close({ id: PANE_ID })
110            return { text: '依据面板已关闭' }
111          }
112          const opened = await $.ui.open({ id: PANE_ID, title: PANE_TITLE, focus: true, closeOnEscape: true, columns: PANE_COLUMNS })
113          if (opened.isPlaced) return { text: '依据面板已打开:p / n 翻看 agent,Esc 关闭' }
114          await $.ui.close({ id: PANE_ID })
115          return { text: unplacedText(opened.reason) }
116        }
117        case 'status': {
118          const { value: lock = null } = await $.state.get(LOCK)
119          return { text: describeStatus(lock) }
120        }
121        case 'master':
122          setMaster(command.on)
123          await report(reporting, { switched: { master: command.on } })
124          return { text: `Dispatch Pilot 已${command.on ? '打开' : '关闭'}${await save('master')}` }
125        case 'switch': {
126          const spec = listSwitches().find((s) => s.name === command.name)
127          if (spec === undefined || !setSwitch(command.name, command.on)) {
128            const names = listSwitches().map((s) => s.name).join(', ')
129            return { text: `没有叫「${command.name}」的开关(现有:${names})` }
130          }
131          await report(reporting, { switched: { feature: spec.name, on: command.on } })
132          return { text: `${spec.name} 已${command.on ? '打开' : '关闭'}:${spec.info}${await save(spec.name)}` }
133        }
134        case 'lock': {
135          await $.state.set(LOCK, command.effort)
136          const inert = masterOn() ? '' : '(Dispatch Pilot 关着:/dp on 才生效)'
137          return { text: `主 agent 已锁在 ${command.effort}:每一步都用这一档${inert};/dp unlock 解除` }
138        }
139        case 'unlock':
140          await $.state.set(LOCK, null)
141          return { text: '已解除锁定,effort 回到按决定走' }
142        case 'log': {
143          const { value: kept = [] } = await $.state.get(DECISIONS)
144          return { text: describeDecisions(kept.slice(-command.count)) }
145        }
146        case 'unknown':
147          return { text: `没看懂这条命令。\n${USAGE}` }
148      }
149    } catch (error) {
150      // Always answer: a failing command must not leave the person without a word.
151      return { text: `出错了(${errorText(error)})` }
152    }
153  })
154}
155
156type Control =
157  | { kind: 'pane'; toggle: boolean }
158  | { kind: 'status' }
159  | { kind: 'master'; on: boolean }
160  | { kind: 'switch'; name: string; on: boolean }
161  | { kind: 'lock'; effort: Effort }
162  | { kind: 'unlock' }
163  | { kind: 'log'; count: number }
164  | { kind: 'unknown' }
165
166/** The words after `/dp`, case and spacing aside. */
167function parseControl(args: string): Control {
168  const words = args.trim().toLowerCase().split(/\s+/).filter(Boolean)
169  const [first, second] = words
170  if (words.length === 0) return { kind: 'pane', toggle: true }
171  if (words.length === 1 && first === 'status') return { kind: 'status' }
172  if (words.length === 1 && (first === 'on' || first === 'off')) return { kind: 'master', on: first === 'on' }
173  if (words.length === 1 && first === 'unlock') return { kind: 'unlock' }
174  if (words.length === 2 && first === 'lock' && second === 'off') return { kind: 'unlock' }
175  if (words.length === 2 && first === 'lock' && isEffort(second)) return { kind: 'lock', effort: second }
176  if (words.length === 1 && first === 'log') return { kind: 'pane', toggle: false }
177  if (words.length === 2 && first === 'log' && /^[1-9]\d*$/.test(second ?? '')) return { kind: 'log', count: Math.min(LOG_ENTRIES, Number(second)) }
178  if (words.length === 2 && first !== undefined && (second === 'on' || second === 'off')) return { kind: 'switch', name: first, on: second === 'on' }
179  return { kind: 'unknown' }
180}
181
182const USAGE = `/dp(依据面板) | /dp status | /dp on|off | /dp <功能名> on|off | /dp lock <${EFFORTS.join('|')}> | /dp unlock | /dp log N`
183
184/** `/dp status`'s answer: whether the mod is on, the lock, each feature's switch. */
185function describeStatus(lock: Effort | null): string {
186  const switches = listSwitches()
187  const width = Math.max(0, ...switches.map((s) => s.name.length))
188  return [
189    `Dispatch Pilot ${masterOn() ? '开着' : '关着'}。effort 锁定:${lock ?? '没有'}。`,
190    '功能开关(/dp <功能名> on|off):',
191    ...switches.map((s) => `  ${s.on ? '开' : '关'}  ${s.name.padEnd(width)}  ${s.info}`),
192    USAGE,
193  ].join('\n')
194}
195
196/** `/dp log N`'s answer: the decisions, oldest first, each with its reason. */
197function describeDecisions(entries: readonly LogEntry[]): string {
198  if (entries.length === 0) return '还没有记下任何决定'
199  const header = `最近 ${entries.length} 条决定,最新的在最后`
200  return [header, ...entries.map((entry) => `#${entry.n} ${entry.feature}:${decisionLine(entry)}`)].join('\n')
201}
202
203/** One measurement as a debug-log line: what the engine reported, `n/a` for what it has no figure for. */
204function describeSignals(measure: SessionMeasureInput): string {
205  const { context, rateLimits, cost } = measure
206  const fill = context.percent === undefined ? `n/a (window ${context.window})` : `${context.percent}% (${context.tokens ?? '?'}/${context.window} tokens)`
207  const limits = rateLimits.length === 0 ? 'n/a' : rateLimits.map((limit) => `${limit.kind} ${limit.percentUsed}%${limit.resetsAt === undefined ? '' : ` resets ${limit.resetsAt}`}`).join(', ')
208  const spent = cost === undefined ? 'n/a' : `$${cost.usd.toFixed(4)}`
209  return `signals: context ${fill}; limits ${limits}; cost ${spent}; changed ${measure.changed.join(', ')}`
210}
211
hooks/features/dispatched-agents.ts 177 lines
1// Feature: each agent the main agent dispatches gets its own decision on its
2// model and effort, once, when it is spawned (agent.spawn). The model is
3// rewritten on the spawn; the effort goes into the plan table under the
4// agent's id, and the core's turn.step writer sends it on every step of that
5// agent. The main agent's own model and cache are untouched (ADR 0001).
6//
7// The decision reads the person's own words this turn (`said`), so a model
8// they name for the work wins over the decision model, which wins over the
9// main agent's pick. Agents dispatched together (one message, several Agent
10// calls) each get their own request and start as soon as their own answer is
11// in. A failed or late answer lets the agent start as the main agent asked.
12//
13// The main agent reads, after the Agent tool's result, the model and effort its
14// agent started with and whose choice the model was (#34): agent.spawn runs
15// inside the Agent tool call, under its id, so the spawn leaves the note for the
16// call to append.
17//
18// Its switch is `dispatched-agents` (`/dp dispatched-agents off`).
19
20import type { HttpInit, On } from 'claude-code'
21import { describeAsked, type BackendIo, type Failure } from '../decision/backend.ts'
22import { messageText } from '../decision/context.ts'
23import { decideDispatch, dispatchEvidence, dispatchNote, dispatchPart, dispatchReason, dispatchState, modelFamily, termsOf, type Dispatch, type DispatchNote } from '../decision/dispatched-agent.ts'
24import { quoteStart } from '../decision/redact.ts'
25import { renderSummary } from '../decision/summary.ts'
26import { answersFor, mergeParts } from '../decision/system-one.ts'
27import { update, type Cell } from '../core/plans.ts'
28import { isPersonsMessage } from '../core/prompts.ts'
29import { report, type ReportIo } from '../core/report.ts'
30import { dispatchSettings, type Ctx } from '../core/setup.ts'
31import { defineSwitch, isOn } from '../core/switches.ts'
32import { UNRESOLVED_SWITCH } from './unresolved.ts'
33
34const AGENTS = { plugin: 'dispatch-pilot', key: 'agents' } as const
35const SAID = { plugin: 'dispatch-pilot', key: 'said' } as const
36const BOARD = { plugin: 'dispatch-pilot', key: 'board' } as const
37const DECISIONS = { plugin: 'dispatch-pilot', key: 'decisionLog' } as const
38const COUNT = { plugin: 'dispatch-pilot', key: 'unresolved' } as const
39const PPLX_RATE = { plugin: 'dispatch-pilot', key: 'pplxRate' } as const
40
41/** The switch's name, in `/dp` and in the decision log. */
42const SWITCH = 'dispatched-agents'
43/**
44 * The notes spawns left for their Agent tool calls, by the call's tool_use_id: written by the agent.spawn inside the
45 * call, taken as the call returns. A module variable (CLAUDE.md): the kit's `$.state` does not show the outer hook
46 * what the nested spawn wrote (DEVELOPMENT.md 已实测, not yet measured on the engine), and a hot reload while an
47 * agent runs loses only that agent's note. At most MAX_NOTES are kept: past that (calls whose result never came
48 * back, or more agents running at once) the oldest go.
49 */
50const pendingNotes = new Map<string, string>()
51const MAX_NOTES = 32
52/** At most this many of the person's messages are kept for one turn. */
53const MAX_SAID = 8
54
55export function registerDispatchedAgents(on: On, ctx: Ctx): void {
56  defineSwitch({ name: SWITCH, info: '主 agent 派出 agent 时决定它的模型和 effort' })
57  const settings = dispatchSettings(ctx)
58
59  // The person's words this turn: a message sent while idle starts them
60  // afresh, one typed during the turn joins them. Other prompts (an agent's
61  // hand-back, a task notification) leave them as they are. Kept whatever the
62  // switches say (nothing is asked or changed here), so a decision made right
63  // after the person switches the feature back on still reads them.
64  on('prompt.submit', { text: /(?:)/ }, async ($, e, next) => {
65    const result = await next(e)
66    if (!isPersonsMessage(e) || result.drop !== undefined) return result
67    const words = messageText(e.text, ctx.config.context.tokens)
68    const said: Cell<string[]> = { get: () => $.state.get(SAID), set: (value, options) => $.state.set(SAID, value, options) }
69    await update(said, (list) => (e.turnId === undefined ? [words] : [...(list ?? []), words].slice(-MAX_SAID)))
70    return result
71  })
72
73  // The note the spawn left for its Agent tool call goes beside the tool's result; a call whose spawn left none (a fork,
74  // a teammate, a refused spawn) is answered as it is.
75  on('tool.call', { tool: 'Agent' }, async ($, e, next) => {
76    const result = await next(e)
77    const note = pendingNotes.get(e.tool_use_id)
78    pendingNotes.delete(e.tool_use_id)
79    if (note === undefined || result.deny !== undefined) return result
80    return { ...result, context: [...(result.context ?? []), note] }
81  })
82
83  on('agent.spawn', { tool_use_id: /(?:)/ }, async ($, e, next) => {
84    // A fork always runs on its parent's model. A teammate lives across many
85    // tasks that one decision at its spawn cannot see.
86    if (e.fork || e.isTeammate) return next(e)
87    /** Leaves the note for the Agent tool call this spawn belongs to. */
88    const leaveNote = (note: DispatchNote) => {
89      pendingNotes.set(e.tool_use_id, dispatchNote(note))
90      while (pendingNotes.size > MAX_NOTES) pendingNotes.delete(pendingNotes.keys().next().value as string)
91    }
92    if (!isOn(SWITCH)) {
93      const started = await next(e)
94      if (started.deny === undefined) leaveNote({ routed: false, started: started.model, why: 'off' })
95      return started
96    }
97    const { value: said = [] } = await $.state.get(SAID)
98    // The problem the main agent is on goes along as background, with no hint (#41): the summary and the count, while the
99    // unresolved switch is on and there are any.
100    const { value: held } = isOn(UNRESOLVED_SWITCH) ? await $.state.get(COUNT) : { value: undefined }
101    const dispatch: Dispatch = {
102      user_message: said.join('\n'),
103      agent_type: e.subagentType,
104      description: e.description,
105      prompt: e.prompt,
106      requested_model: e.model ?? null,
107      ...(held?.summary === undefined ? {} : { problem_summary: renderSummary(held.summary, ctx.ask.language) }),
108      ...(held === undefined || held.count === 0 ? {} : { unresolved_count: held.count }),
109    }
110    const part = dispatchPart(dispatch, settings)
111    const request = mergeParts(dispatchState(dispatch, ctx.config.contextByKind.agent), [part])
112    const io: BackendIo = {
113      fetch: (url: string, init: HttpInit) => $.http.fetch(url, init),
114      sleep: (ms: number, signal: AbortSignal) => $.clock.sleep(ms, { signal }),
115      pace: { now: () => $.clock.now(), sentAt: { get: () => $.state.get(PPLX_RATE), set: (value, options) => $.state.set(PPLX_RATE, value, options) } },
116    }
117    const reporting: ReportIo = {
118      board: { get: () => $.state.get(BOARD), set: (value, options) => $.state.set(BOARD, value, options) },
119      decisions: { get: () => $.state.get(DECISIONS), set: (value, options) => $.state.set(DECISIONS, value, options) },
120      debug: (line) => $.ui.log(line, { to: 'debug' }),
121      now: () => $.clock.now(),
122      toast: (text) => $.ui.toast(text),
123    }
124    const about = `${quoteStart(e.description)}(${e.subagentType})`
125    /** What every report of this agent says of it; its id is known once it has started (a spawn is not yet an agent). */
126    const reportOf = (agent: string) => ({ feature: SWITCH, agent, subject: about, node: { kind: 'agent' as const, name: e.name ?? (e.description === '' ? e.subagentType : e.description), type: e.subagentType } })
127    const startedAt = await $.clock.now()
128    const asked = await ctx.backend.ask(io, request, ctx.config.timeoutMs)
129    const ms = (await $.clock.now()) - startedAt
130    $.ui.log(`request [${Object.keys(request.questions).join(', ')}] to ${ctx.backend.name} for agent ${about}: ${describeAsked(asked, ms)}`, { to: 'debug' })
131    // The agent starts as the main agent asked; the board says why it was not routed, once it has an id to say it of.
132    // A spawn refused beneath started no agent: there is nothing to say of it.
133    const notRouted = async (failure: Failure) => {
134      const started = await next(e)
135      if (started.deny !== undefined) return started
136      await report(reporting, { decision: { ...reportOf(started.agentId ?? e.tool_use_id), routed: false, failure: { backend: ctx.backend.name, ...failure } } })
137      leaveNote({ routed: false, started: started.model, why: { failure, backend: ctx.backend.name } })
138      return started
139    }
140    if (!asked.ok) return notRouted(asked.failure)
141    const decision = decideDispatch(answersFor(part, asked.answers), dispatch, settings)
142    if (!decision.answered) return notRouted({ kind: 'parse', detail: 'no answer about the agent' })
143    // The main agent's pick, when it stands, goes on as the main agent wrote it.
144    const spawn = decision.model !== null && decision.model !== modelFamily(e.model) ? { ...e, model: decision.model } : e
145    const result = await next(spawn)
146    if (result.deny !== undefined) return result
147    leaveNote({ routed: true, decision, started: result.model, requested: modelFamily(e.model), thetaOverride: settings.thetaOverride })
148    // The agent has started: nothing after this may fail its spawn.
149    try {
150      // The effort and the person's terms go into the plan (what changes the agent later keeps to the
151      // terms); the model is set on the spawn, and a planned model would also pin every step against the
152      // engine's overload fallback.
153      const terms = termsOf(decision)
154      if (result.agentId !== undefined && (decision.effort !== null || terms !== null)) {
155        await $.state.set({ ...AGENTS, id: result.agentId }, { effort: decision.effort, floor: null, model: null, terms })
156      }
157      const family = decision.model ?? modelFamily(result.model)
158      const model = family ?? result.model
159      const outcome = decision.effort === null ? model : `${model} ${decision.effort}`
160      await report(reporting, {
161        decision: {
162          ...reportOf(result.agentId ?? e.tool_use_id),
163          routed: true,
164          outcome,
165          reason: dispatchReason(decision, modelFamily(e.model), settings.thetaOverride, '主 agent 指定的'),
166          ...dispatchEvidence(decision),
167          ...(family === null ? {} : { model: family }),
168          ...(decision.effort === null ? {} : { effort: decision.effort }),
169        },
170      })
171    } catch (error) {
172      $.ui.log(`agent ${about} started (${result.agentId ?? 'no id'}), but its plan was not recorded: ${String(error)}`, { to: 'debug' })
173    }
174    return result
175  })
176}
177
hooks/features/escalation.ts 663 lines
1// Feature: a loop whose tool calls keep failing goes up a level (#7).
2//
3// Every tool call of the main agent and of the agents it dispatches is told
4// apart as it ends (core/outcomes.ts), and each loop's failures are counted
5// here, in one record (`escalation`) that the mid-turn re-decision (#5) sends
6// in its counts too: a call the person refused is no failure, one a hook
7// blocked counts toward escalating only with `hook-block-failures` on, and
8// the counts start over whenever this feature deals with them (a forced
9// raise, failures found expected, or nothing left to raise).
10//
11// When a loop's counted failures reach `escalateAfter`, the decision model is
12// asked at once, as the call that reaches them ends (story 18): the mid-turn
13// effort question with the trouble, and whether the failures are what the
14// work expects (a test written to fail first, a search that finds nothing).
15// The loop's next step takes the answer, waiting `rejudgeWaitMs` at most, as
16// a mid-turn re-decision does; an answer later than that is taken at a later
17// step. Expected: the counts start over and the answer is an ordinary
18// re-decision. Otherwise (also when no answer comes) the loop goes up: one
19// level (at most to xhigh) or straight to max, at most `escalateLimit` times
20// per turn (per run, for an agent); a haiku agent, which takes no effort,
21// goes on as another model. The person's terms for an agent's work hold
22// (core/plans.ts `AgentPlan`).
23//
24// It sits above the mid-turn feature (the entry registers it first), so a
25// step's own mid-turn answer finds the raise already in the turn's plan and
26// cannot undercut it.
27
28import type { EngineInterface, HttpInit, On, TurnStepInput } from 'claude-code'
29import { describeAsked, errorText, failureLine, within, type BackendIo, type Failure } from '../decision/backend.ts'
30import { AGENT_MODELS, effortFloor, modelFamily, type AgentModel, type Terms } from '../decision/dispatched-agent.ts'
31import { briefOf, forcedTarget, readExpected, rowsFromTranscript, stepsFromRows, stuckRequest, traceRaise, troubleText, type RaiseMode, type TranscriptRow } from '../decision/escalation.ts'
32import { higherEffort, isEffort, probsOf, readEffort, readingText, type Effort, type EffortReading } from '../decision/effort.ts'
33import { contentLanguage, judgeMidturn, MIDTURN_LEVEL, midturnRecord, outcomeOf, verdictReason, type MidturnInput, type MidturnLimits, type MidturnPosition, type MidturnRules, type MidturnVerdict } from '../decision/midturn.ts'
34import { modelId, type ResolvedModel } from '../decision/model-ids.ts'
35import { quoteStart } from '../decision/redact.ts'
36import { answersFor } from '../decision/system-one.ts'
37import { startedIn } from '../decision/workflow-labels.ts'
38import { endedAs, noteEnded, wasBlocked } from '../core/outcomes.ts'
39import { floorHeld, forced, MAIN, newTurn, redecided, replace, turnKey, update, type AgentPlan, type Cell, type TurnRecord } from '../core/plans.ts'
40import { report, type Decided, type ReportIo } from '../core/report.ts'
41import type { Ctx } from '../core/setup.ts'
42import { defineSwitch, isOn, masterOn } from '../core/switches.ts'
43
44const ESCALATION = { plugin: 'dispatch-pilot', key: 'escalation' } as const
45const TURNS = { plugin: 'dispatch-pilot', key: 'turns' } as const
46const AGENTS = { plugin: 'dispatch-pilot', key: 'agents' } as const
47const LOCK = { plugin: 'dispatch-pilot', key: 'lock' } as const
48const RUNS = { plugin: 'dispatch-pilot', key: 'workflowRuns' } as const
49const BOARD = { plugin: 'dispatch-pilot', key: 'board' } as const
50const DECISIONS = { plugin: 'dispatch-pilot', key: 'decisionLog' } as const
51const PPLX_RATE = { plugin: 'dispatch-pilot', key: 'pplxRate' } as const
52
53/** At most this many Workflow runs' directories are kept. */
54const MAX_RUNS = 8
55
56/** The feature's switch (`/dp escalation on|off`). */
57const SWITCH = 'escalation'
58/** Whether a call one of the person's hooks blocked counts as a failure (`/dp hook-block-failures on|off`); off by default. */
59const BLOCKS_SWITCH = 'hook-block-failures'
60
61/** One loop's failed calls and forced raises (`escalation` in types/index.d.ts): the main agent's for one turn, an agent's for its run. */
62type LoopRecord = {
63  /** The main turn the counts belong to; '' for an agent. */
64  turnId: string
65  /** Calls that failed, and calls a hook refused, since the loop began. */
66  failures: number
67  hookBlocks: number
68  /** The counts when they last started over: what is counted is the rest. */
69  base: Counted
70  /** Forced raises so far (failures found expected are not one). */
71  raises: number
72  /** The loop's latest step, the engine's effort and model on it: what a re-decision asked as a call ends reads. */
73  step: number | null
74  engine: Effort | null
75  model: string | null
76  /** The step a stuck re-decision was last asked for; null before the first. */
77  askedFor: number | null
78  /** An agent's step at its latest raise (the main agent's is its turn's `raisedAt`); null before one. */
79  raisedAt: number | null
80  /** The escalation switch was off when the loop was last seen: what was counted meanwhile is written off once it is back on. */
81  paused: boolean
82}
83
84type Counted = { failures: number; hookBlocks: number }
85
86/** What the feature runs with: the shared ctx and the options it reads. */
87type Settings = {
88  ctx: Ctx
89  /** Counted failures that make a loop stuck. */
90  after: number
91  mode: RaiseMode
92  /** Forced raises a loop gets at most. */
93  limit: number
94  /** The failures count as expected when the answer reaches this probability. */
95  thetaExpected: number
96  /** The model a failing haiku agent is switched to, as a step names it (the engine takes no alias for a step); null for none. */
97  haikuTo: ResolvedModel | null
98  /** `escalateHaikuTo` as written, for the log when it names no model. */
99  haikuToWritten: string
100  /** The models agents may run on: a ruled-out escalateHaikuTo gives way to the next of them up. */
101  models: readonly AgentModel[]
102  /** The mid-turn rules: what a raise the decision model itself suggests needs, and how an ordinary re-decision moves the level. */
103  rules: MidturnRules
104  limits: MidturnLimits
105  /** How long a step waits for a stuck re-decision not back yet. */
106  waitMs: number
107}
108
109/** What the decision model made of a stuck loop. */
110type Stuck = {
111  /** The effort answer; null when it was not asked or not answered. */
112  reading: EffortReading | null
113  /** The probability that the failures were expected; null without an answer. */
114  expected: number | null
115  /** Why the question went unanswered, when it did. */
116  failure: Failure | null
117  /** No transcript of the loop could be read: nothing was asked. */
118  unread: boolean
119}
120
121/**
122 * A stuck re-decision on its way, or answered and not yet taken, by loop id:
123 * asked for step `forStep` (of turn `turnId`, for the main agent), about the
124 * counts in `covered` (what starts over once it is dealt with). `answer`
125 * never rejects; once it resolves, `settled` holds it and `ms` how long it took.
126 * Promises cannot live in $.state: a reload drops them, and the loop's next
127 * step asks again.
128 */
129type Asking = { forStep: number; turnId: string; covered: Counted; about: string; answer: Promise<Stuck>; settled: Stuck | null; ms: number }
130const asking = new Map<string, Asking>()
131
132export function registerEscalation(on: On, ctx: Ctx): void {
133  defineSwitch({ name: SWITCH, info: '工具调用接连失败的 agent,强制升高它的 effort', parts: ['counts'] })
134  defineSwitch({ name: BLOCKS_SWITCH, info: '判断要不要强制升档时,把被你的 hook 拦下的调用也算失败', default: false })
135  // Read at each use: which decision model's settings these are is settled at the session start (core/setup.ts, `Ctx`).
136  const settings: Settings = {
137    ctx,
138    get after() {
139      return ctx.config.escalation.after
140    },
141    get mode() {
142      return ctx.config.escalation.mode
143    },
144    get limit() {
145      return ctx.config.escalation.limit
146    },
147    get thetaExpected() {
148      return ctx.config.escalation.thetaExpected
149    },
150    get haikuTo() {
151      return ctx.config.escalation.haikuTo
152    },
153    get haikuToWritten() {
154      return ctx.config.escalation.haikuToWritten
155    },
156    get models() {
157      return ctx.config.agents.models
158    },
159    get rules() {
160      return ctx.config.midturn.rules
161    },
162    get limits() {
163      return ctx.config.midturn.limits
164    },
165    get waitMs() {
166      return ctx.config.midturn.waitMs
167    },
168  }
169
170  on('tool.call', { tool: /(?:)/ }, async ($, e, next) => {
171    const result = await next(e)
172    // The calls the model made, not another plugin's `$.tool.call`; counted while Dispatch Pilot is on, whatever this
173    // feature's switch says (the mid-turn re-decision sends the counts too).
174    if (next.origin.plugin !== 'engine' || !masterOn()) return result
175    try {
176      // A Workflow's run: its directory holds each of its agents' transcripts, which the engine keeps from the mod.
177      const run = e.tool === 'Workflow' ? launchedRun(result.result) : null
178      if (run !== null) {
179        const runs: Cell<{ runId: string; dir: string }[]> = { get: () => $.state.get(RUNS), set: (value, options) => $.state.set(RUNS, value, options) }
180        await update(runs, (list) => [...(list ?? []).filter((kept) => kept.runId !== run.runId), run].slice(-MAX_RUNS))
181      }
182      const ended = outcomeOf(e.tool, result, wasBlocked(e.tool_use_id))
183      noteEnded(e.tool_use_id, ended)
184      // A refusal by a plugin's tool.call hook (`{ deny }`) is nobody's failure and no settings hook's block: the
185      // Workflow feature's hand-back of a script is one. Only what the tool reported, and what the settings hooks refused, count.
186      if (typeof result.deny === 'string' || (ended !== 'failed' && ended !== 'blocked')) return result
187      const id = e.agentId ?? MAIN
188      const ref = { ...ESCALATION, id }
189      const cell: Cell<LoopRecord> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
190      const record = await update(cell, (r) => {
191        const loop = r ?? newLoop('')
192        return ended === 'failed' ? { ...loop, failures: loop.failures + 1 } : { ...loop, hookBlocks: loop.hookBlocks + 1 }
193      })
194      if (!isOn(SWITCH)) return result
195      await showCounts($, id, record, null)
196      // The call that reaches the threshold asks at once: the answer is for the loop's next step (once per step).
197      if (record.step !== null && record.askedFor !== record.step + 1) await launch($, settings, id, e.agentId, record, record.step + 1)
198    } catch (error) {
199      $.ui.log(`escalation: ${errorText(error)}`, { to: 'debug' })
200    }
201    return result
202  })
203
204  on('turn.step', { turnId: /(?:)/ }, async function* ($, e, next) {
205    if (masterOn()) {
206      try {
207        await atStep($, settings, e)
208      } catch (error) {
209        $.ui.log(`escalation: ${errorText(error)}`, { to: 'debug' })
210      }
211    }
212    return yield* next(e)
213  })
214}
215
216/**
217 * At a loop's step: keeps where the loop is (a new main turn starts its counts
218 * afresh), then takes the stuck re-decision meant for this step and applies
219 * it. With none on its way while the counted failures call for one (nothing
220 * to raise when the call ended, or a reload dropped the request), deals with
221 * them here: written off when nothing can be raised, else asked now.
222 */
223async function atStep($: EngineInterface, s: Settings, e: TurnStepInput): Promise<void> {
224  const main = e.agentId === undefined
225  const id = e.agentId ?? MAIN
226  const ref = { ...ESCALATION, id }
227  const cell: Cell<LoopRecord> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
228  const engine = isEffort(e.effort) ? e.effort : null
229  const on = isOn(SWITCH)
230  /** Whether a record belongs to an earlier main turn (or there is none): this step starts the turn's counts afresh. */
231  const startsTurn = (r: LoopRecord | undefined) => main && (r === undefined || r.turnId !== e.turnId)
232  const { before, after } = await replace(cell, (r) => {
233    const loop = startsTurn(r) || r === undefined ? newLoop(main ? e.turnId : '') : r
234    // Back on after being off: what was counted meanwhile is written off, so it is counted from here on.
235    const written = on && loop.paused ? { ...loop, base: { failures: loop.failures, hookBlocks: loop.hookBlocks }, paused: false } : loop
236    return { ...written, step: e.index, engine, model: e.model, paused: !on }
237  })
238  let record = after
239  if (startsTurn(before)) {
240    // A new main turn: what the one before counted, and asked, is over.
241    asking.delete(MAIN)
242    await showCounts($, MAIN, null, null)
243  }
244  if (!on) return
245
246  let pending = asking.get(id)
247  if (pending === undefined || (main && pending.turnId !== e.turnId) || pending.forStep > e.index) {
248    if (pending !== undefined) return
249    // No re-decision on its way, though the counted failures may call for one: nothing could be raised when
250    // the call ended, or a reload dropped the request. Dealt with here: written off, or asked now.
251    const counted = countedOf(record)
252    if (counted.failures + counted.hookBlocks < s.after || record.raises >= s.limit || record.askedFor === e.index || s.ctx.backend.configured === false) return
253    record = await update(cell, (r) => ({ ...(r ?? record), askedFor: e.index }))
254    const raise = await raiseOf($, s, id, e.agentId, record, e.index)
255    if (raise.kind === 'none') return
256    if (raise.kind === 'keep') {
257      const kept = await startOver($, cell, id, countedNow(record), false)
258      const brief = e.agentId === undefined ? '' : briefOf((await agentRows($, e.agentId)) ?? [])
259      await decide($, id, kept, { outcome: raise.outcome, about: aboutOf(e.agentId, brief, e.index, counted), reason: raise.reason, tone: 'info' })
260      return
261    }
262    await launch($, s, id, e.agentId, record, e.index)
263    pending = asking.get(id)
264    if (pending === undefined) return
265  }
266  const answer = pending.settled ?? (await within((ms, signal) => $.clock.sleep(ms, { signal }), pending.answer, s.waitMs, null))
267  if (answer === null) {
268    // Not back yet: the step goes as it is, and the answer is taken at a later step (as a mid-turn re-decision's).
269    await showCounts($, id, record, 'late')
270    return
271  }
272  asking.delete(id)
273  await apply($, s, e, cell, pending.covered, pending.about, answer)
274}
275
276/** What a stuck loop's re-decision would do, from where the loop stands. */
277type Raise =
278  /** Leave the loop be: the person's lock holds the main agent's effort, or the step takes no effort level. */
279  | { kind: 'none' }
280  /** Nothing to raise: the failures are written off, and the decision recorded. */
281  | { kind: 'keep'; outcome: string; reason: string }
282  /** A higher effort, from `current`. */
283  | { kind: 'level'; current: Effort; target: Effort }
284  /** Another model (a haiku agent: it takes no effort). */
285  | { kind: 'model'; from: string; to: ResolvedModel; note: string }
286
287/**
288 * What raising the loop would be at step `at`, by its plan: the main agent's
289 * turn (unless the person locked its effort), or an agent's effective model
290 * (the plan's, else the engine's) and the person's terms for its work.
291 */
292async function raiseOf($: EngineInterface, s: Settings, id: string, agentId: string | undefined, record: LoopRecord, at: number): Promise<Raise> {
293  const top = (current: Effort) => ({ kind: 'keep' as const, outcome: `effort ${current}(保持)`, reason: s.mode === 'max' ? '已经是 max' : '升一档最高到 xhigh' })
294  if (agentId === undefined) {
295    const { value: lock = null } = await $.state.get(LOCK)
296    if (lock !== null || record.engine === null) return { kind: 'none' }
297    const { value: turn } = await $.state.get({ ...TURNS, id: turnKey(record.turnId, undefined) })
298    const current = higherEffort(turn?.effort ?? record.engine, floorHeld(turn, at)) as Effort
299    const target = forcedTarget(current, s.mode)
300    return target === null ? top(current) : { kind: 'level', current, target }
301  }
302  const { value: planned } = await $.state.get({ ...AGENTS, id })
303  const model = planned?.model ?? record.model ?? ''
304  const family = modelFamily(model)
305  // A model the mod does not know, and that takes no effort level: nothing to raise.
306  if (family === null && record.engine === null) return { kind: 'none' }
307  if (family === 'haiku') {
308    const to = haikuSwitch(s, planned?.terms ?? null)
309    return 'why' in to ? { kind: 'keep', outcome: `model ${model}(保持)`, reason: to.why } : { kind: 'model', from: model, to: to.to, note: to.note }
310  }
311  const named = planned?.terms?.effort ?? null
312  if (named !== null) return { kind: 'keep', outcome: `effort ${named}(保持)`, reason: `${named} 是你给它点名的 effort` }
313  // Without a level of its own (moved off a model that takes none), its steps go at the engine's own for an agent: medium (measured on 2.1.289).
314  const current = higherEffort(planned?.effort ?? record.engine ?? 'medium', planned?.floor ?? null) as Effort
315  const target = forcedTarget(current, s.mode)
316  return target === null ? top(current) : { kind: 'level', current, target }
317}
318
319/**
320 * Sends the stuck re-decision for the loop's step `forStep`, when its counted
321 * failures call for one and there is something to raise; the answer waits in
322 * `asking` for that step. Once per step, one at a time per loop.
323 */
324async function launch($: EngineInterface, s: Settings, id: string, agentId: string | undefined, record: LoopRecord, forStep: number): Promise<void> {
325  const counted = countedOf(record)
326  if (counted.failures + counted.hookBlocks < s.after || record.raises >= s.limit || s.ctx.backend.configured === false || asking.has(id)) return
327  const raise = await raiseOf($, s, id, agentId, record, forStep)
328  if (raise.kind !== 'level' && raise.kind !== 'model') return
329  const ref = { ...ESCALATION, id }
330  const cell: Cell<LoopRecord> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
331  await update(cell, (r) => ({ ...(r ?? record), askedFor: forStep }))
332  // What the counts say since they last started over, and what they start over from once this is dealt with.
333  const since = { failures: record.failures - record.base.failures, hookBlocks: record.hookBlocks - record.base.hookBlocks }
334  const entry: Asking = { forStep, turnId: record.turnId, covered: countedNow(record), about: '', answer: Promise.resolve(UNREAD), settled: null, ms: 0 }
335
336  let input: MidturnInput
337  if (agentId === undefined) {
338    const { value: turn } = await $.state.get({ ...TURNS, id: turnKey(record.turnId, undefined) })
339    const plan = turn ?? newTurn('', null, false)
340    const rows = (await $.session.messages().catch(() => [])) as TranscriptRow[]
341    entry.about = aboutOf(undefined, '', forStep, counted)
342    input = {
343      message: plan.prompt,
344      step: forStep,
345      current_effort: raise.kind === 'level' ? raise.current : 'medium',
346      counts: { judgments: plan.decisions, changes: plan.changes, failures: since.failures, hook_blocks: since.hookBlocks },
347      recent_steps: stepsFromRows(rows, { language: contentLanguage(plan.prompt), ended: endedAs }),
348      trouble: troubleText(counted),
349    }
350  } else {
351    const rows = await agentRows($, agentId)
352    const brief = rows === null ? '' : briefOf(rows)
353    entry.about = aboutOf(agentId, brief, forStep, counted)
354    if (rows === null) {
355      // No transcript to read (a workflow agent whose run the mod did not record): raised without asking.
356      entry.settled = UNREAD
357      asking.set(id, entry)
358      return
359    }
360    input = {
361      message: brief,
362      step: forStep,
363      current_effort: raise.kind === 'level' ? raise.current : 'medium',
364      counts: { judgments: 1 + record.raises, changes: record.raises, failures: since.failures, hook_blocks: since.hookBlocks },
365      recent_steps: stepsFromRows(rows, { language: contentLanguage(brief), ended: endedAs }),
366      trouble: troubleText(counted),
367    }
368  }
369  // A haiku agent takes no effort: only whether its failures were expected is asked.
370  const { request, effortPart, expectedPart } = stuckRequest(input, { limits: s.limits, ask: s.ctx.ask, effort: raise.kind === 'level' })
371  const io: BackendIo = {
372    fetch: (url: string, init: HttpInit) => $.http.fetch(url, init),
373    sleep: (ms: number, signal: AbortSignal) => $.clock.sleep(ms, { signal }),
374    pace: { now: () => $.clock.now(), sentAt: { get: () => $.state.get(PPLX_RATE), set: (value, options) => $.state.set(PPLX_RATE, value, options) } },
375  }
376  const startedAt = await $.clock.now()
377  entry.answer = s.ctx.backend.ask(io, request, s.ctx.config.timeoutMs).then(async (asked): Promise<Stuck> => {
378    entry.ms = (await $.clock.now()) - startedAt
379    $.ui.log(`request [${Object.keys(request.questions).join(', ')}] for ${entry.about} to ${s.ctx.backend.name}: ${describeAsked(asked, entry.ms)}`, { to: 'debug' })
380    let stuck: Stuck = { reading: null, expected: null, failure: null, unread: false }
381    if (!asked.ok) stuck = { ...stuck, failure: asked.failure }
382    else {
383      const expected = readExpected(answersFor(expectedPart, asked.answers))
384      stuck = {
385        reading: effortPart === null ? null : readEffort(answersFor(effortPart, asked.answers)[MIDTURN_LEVEL]),
386        expected,
387        failure: expected === null ? { kind: 'parse', detail: 'no answer to the question' } : null,
388        unread: false,
389      }
390    }
391    entry.settled = stuck
392    return stuck
393  })
394  asking.set(id, entry)
395}
396
397/** The answer when nothing could be asked: no transcript of the agent to read. */
398const UNREAD: Stuck = { reading: null, expected: null, failure: null, unread: true }
399
400/**
401 * Applies a stuck re-decision's answer at step `e`, to the loop as it stands
402 * now (a re-decision may have moved it since the question went out): expected
403 * failures leave an ordinary re-decision; otherwise the loop goes up (or, when
404 * it can no longer, its failures are written off). Either way the counts the
405 * question covered start over.
406 */
407async function apply($: EngineInterface, s: Settings, e: TurnStepInput, cell: Cell<LoopRecord>, covered: Counted, about: string, answer: Stuck): Promise<void> {
408  const id = e.agentId ?? MAIN
409  const { value: record } = await cell.get()
410  if (record === undefined) return
411  const raise = await raiseOf($, s, id, e.agentId, record, e.index)
412  if (raise.kind === 'none') return
413  if (raise.kind === 'keep') {
414    const kept = await startOver($, cell, id, covered, false)
415    await decide($, id, kept, { outcome: raise.outcome, about, reason: raise.reason, tone: 'info' })
416    return
417  }
418  const p = answer.expected
419  if (p !== null && p >= s.thetaExpected) {
420    const why = `这些失败是预期内的(概率 ${p.toFixed(2)},预期内失败门槛 ${s.thetaExpected.toFixed(2)}),不强制升档`
421    if (raise.kind === 'model') {
422      const kept = await startOver($, cell, id, covered, false)
423      await decide($, id, kept, { outcome: `model ${raise.from}(保持)`, about, reason: why, tone: 'info' })
424      return
425    }
426    // Expected: nothing is forced; the answer's effort is an ordinary re-decision, as mid-turn.
427    const level = await redecide($, s, e, record, raise.current, answer.reading)
428    const kept = await startOver($, cell, id, covered, false)
429    await decide($, id, kept, {
430      outcome: `effort ${level.effort}${level.effort === raise.current ? '(保持)' : `(原 ${raise.current})`}`,
431      about,
432      reason: level.why === null ? why : `${why};${level.why}`,
433      tone: level.effort === raise.current ? 'info' : 'ok',
434      ...(level.working === undefined
435        ? {}
436        : {
437            probs: probsOf(level.working.reading),
438            conf: level.working.verdict.confidence,
439            trace: [...level.working.verdict.trace.pick.steps, ...level.working.verdict.trace.steps],
440            mid: midturnRecord(level.working.verdict, level.working.position, s.rules),
441          }),
442    })
443    return
444  }
445
446  if (raise.kind === 'model') {
447    const ref = { ...AGENTS, id }
448    const planCell: Cell<AgentPlan> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
449    await update(planCell, (r) => ({ ...(r ?? { effort: null, floor: null, model: null, terms: null }), model: raise.to.id }))
450    const raised = await startOver($, cell, id, covered, true, e.index)
451    await decide($, id, raised, {
452      outcome: `model ${raise.to.id}(原 ${raise.from})`,
453      about,
454      reason: `haiku 没有 effort 可升,改用 ${raise.to.id}${raise.note};${knownReason(s, answer)}`,
455      tone: 'warn',
456      forced: { kind: 'model', from: raise.from, to: raise.to.id },
457    })
458    return
459  }
460  const { level, steps } = traceRaise(answer.reading, { from: raise.current, target: raise.target, mode: s.mode }, s.rules)
461  if (e.agentId === undefined) {
462    // Raised from this step on; for holdSteps steps nothing lowers it below the forced level, then ordinary re-decisions take over.
463    const ref = { ...TURNS, id: turnKey(e.turnId, undefined) }
464    const turnCell: Cell<TurnRecord> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
465    await update(turnCell, (r) => forced(r ?? newTurn('', null, false), level, raise.target, e.index, s.rules.holdSteps))
466  } else {
467    // The raise holds for the rest of the agent's run: nothing re-decides an agent mid-run, so it goes into its effort.
468    const ref = { ...AGENTS, id }
469    const planCell: Cell<AgentPlan> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
470    await update(planCell, (r) => ({ ...(r ?? { effort: null, floor: null, model: null, terms: null }), effort: level }))
471  }
472  const raised = await startOver($, cell, id, covered, true, e.index)
473  await decide($, id, raised, {
474    outcome: `effort ${level}(原 ${raise.current})`,
475    about,
476    reason: raiseReason(s, answer, level, raise.target),
477    tone: 'warn',
478    forced: { kind: 'effort', from: raise.current, to: level, floor: raise.target },
479    trace: steps,
480    ...(answer.reading === null ? {} : { probs: probsOf(answer.reading), ...(answer.reading.confidence === null ? {} : { conf: answer.reading.confidence }) }),
481  })
482}
483
484/**
485 * An ordinary re-decision from a stuck loop's effort answer, its failures
486 * found expected: the main agent's turn by the mid-turn rules (a turn the
487 * person started, as mid-turn), an agent's plan the same way. The level it
488 * goes on at, and why (null when there was nothing to decide from).
489 */
490async function redecide(
491  $: EngineInterface,
492  s: Settings,
493  e: TurnStepInput,
494  record: LoopRecord,
495  current: Effort,
496  reading: EffortReading | null,
497): Promise<{ effort: Effort; why: string | null; working?: { reading: EffortReading; verdict: MidturnVerdict; position: MidturnPosition } }> {
498  if (reading === null) return { effort: current, why: null }
499  if (e.agentId === undefined) {
500    const ref = { ...TURNS, id: turnKey(e.turnId, undefined) }
501    const turnCell: Cell<TurnRecord> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
502    const { value: turn } = await turnCell.get()
503    if (turn?.person !== true) return { effort: current, why: null }
504    const position = { current, sinceRaise: turn.raisedAt == null ? null : e.index - turn.raisedAt, atLeast: floorHeld(turn, e.index) }
505    const verdict = judgeMidturn(reading, position, s.rules)
506    await update(turnCell, (r) => redecided(r ?? turn, current, verdict.effort, e.index))
507    return { effort: verdict.effort, why: `${readingText(reading)};${verdictReason(verdict, position, s.rules)}`, working: { reading, verdict, position } }
508  }
509  const ref = { ...AGENTS, id: e.agentId }
510  const planCell: Cell<AgentPlan> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
511  const { value: plan } = await planCell.get()
512  // A routed agent's re-decision goes no lower than its model's floor either (effortFloor): its effort was decided with it.
513  const floor = higherEffort(plan?.floor ?? null, plan?.effort == null ? null : effortFloor(modelFamily(plan.model ?? record.model ?? '')))
514  const position = { current, sinceRaise: record.raisedAt === null ? null : e.index - record.raisedAt, atLeast: floor }
515  const verdict = judgeMidturn(reading, position, s.rules)
516  if (verdict.effort !== current) await update(planCell, (r) => ({ ...(r ?? { effort: null, floor: null, model: null, terms: null }), effort: verdict.effort }))
517  return { effort: verdict.effort, why: `${readingText(reading)};${verdictReason(verdict, position, s.rules)}`, working: { reading, verdict, position } }
518}
519
520/**
521 * The model a failing haiku agent goes on as: `escalateHaikuTo`, unless the
522 * person ruled it out for the agent's work, then the next model up that agents
523 * may run on and the person did not rule out. None (with why) when the person
524 * named haiku, when escalateHaikuTo names no model, or when every model up is
525 * ruled out.
526 */
527function haikuSwitch(s: Settings, terms: Terms | null): { to: ResolvedModel; note: string } | { why: string } {
528  if (terms?.model === 'haiku') return { why: 'haiku 是你给它点名的模型' }
529  const to = s.haikuTo
530  if (to === null) return { why: s.haikuToWritten === '' ? '没有设置失败的 haiku 改用哪个模型' : `设置的改用模型 ${JSON.stringify(s.haikuToWritten)} 这个 mod 不认识` }
531  const banned = terms?.banned ?? []
532  if (!banned.includes(to.family)) return { to, note: '' }
533  const up = AGENT_MODELS.slice(AGENT_MODELS.indexOf(to.family) + 1).find((family) => s.models.includes(family) && !banned.includes(family))
534  if (up !== undefined) return { to: { family: up, id: modelId(up) }, note: `(${to.family} 被你排除了)` }
535  const above = AGENT_MODELS.filter((family) => family !== 'haiku' && banned.includes(family))
536  return { why: `haiku 以上的模型都被你排除了(${above.join('、')})` }
537}
538
539/** The run a Workflow tool's result says it launched: its id and its directory; null for none (refused, failed). */
540function launchedRun(result: unknown): { runId: string; dir: string } | null {
541  if (typeof result !== 'object' || result === null) return null
542  const { runId, transcriptDir } = result as { runId?: unknown; transcriptDir?: unknown }
543  return typeof runId === 'string' && typeof transcriptDir === 'string' ? { runId, dir: transcriptDir } : null
544}
545
546/**
547 * An agent's transcript: as the session gives it, or for a workflow's agent
548 * (the engine keeps it from the mod) from the run's directory on disk, found
549 * among the runs this feature noted; null when neither can be read.
550 */
551async function agentRows($: EngineInterface, agentId: string): Promise<TranscriptRow[] | null> {
552  const found: unknown = await $.session.messages({ agentId }).catch(() => null)
553  if (Array.isArray(found)) return found as TranscriptRow[]
554  const { value: runs = [] } = await $.state.get(RUNS)
555  for (const run of [...runs].reverse()) {
556    const journal = await $.fs.read(`${run.dir}/journal.jsonl`).catch(() => null)
557    if (journal === null || startedIn(journal, agentId) === null) continue
558    const text = await $.fs.read(`${run.dir}/agent-${agentId}.jsonl`).catch(() => null)
559    return text === null ? null : rowsFromTranscript(text)
560  }
561  return null
562}
563
564/** What a decision about the loop is about: its step and the failures, and for an agent its task (`brief`, '' when unknown). */
565function aboutOf(agentId: string | undefined, brief: string, at: number, counted: Counted): string {
566  const failed = `工具调用失败 ${counted.failures + counted.hookBlocks} 次`
567  if (agentId === undefined) return `第 ${at} 步(${failed})`
568  return `${brief === '' ? `agent ${agentId}` : `agent ${quoteStart(brief)}`},第 ${at} 步(${failed})`
569}
570
571/** What the decision report needs of the host: its two cells, the debug log, the clock and the toast. */
572function ioOf($: EngineInterface): ReportIo {
573  return {
574    board: { get: () => $.state.get(BOARD), set: (value, options) => $.state.set(BOARD, value, options) },
575    decisions: { get: () => $.state.get(DECISIONS), set: (value, options) => $.state.set(DECISIONS, value, options) },
576    debug: (line) => $.ui.log(line, { to: 'debug' }),
577    now: () => $.clock.now(),
578    toast: (text) => $.ui.toast(text),
579  }
580}
581
582/** What a decision of this feature says, besides who it is about and the loop's counts (`decide` adds them). */
583type Said = Pick<Decided, 'outcome' | 'reason' | 'tone' | 'probs' | 'conf' | 'trace' | 'floor' | 'mid' | 'forced'> & { about: string }
584
585/**
586 * Reports a decision of this feature (debug log, `/dp log`, the board data), about loop `id`
587 * (`main`, or an agent's id) and with its counts as they stand. Beside the loop's own route: not on its node.
588 */
589async function decide($: EngineInterface, id: string, record: LoopRecord, said: Said): Promise<void> {
590  const { about, ...rest } = said
591  await report(ioOf($), {
592    decision: {
593      ...rest,
594      feature: SWITCH,
595      agent: id,
596      aside: true,
597      subject: about,
598      counts: { failed: record.failures, blocked: record.hookBlocks, raised: record.raises },
599    },
600  })
601}
602
603/** Why a forced raise went where it did, for the decision log. */
604function raiseReason(s: Settings, answer: Stuck, level: Effort, target: Effort): string {
605  const how = s.mode === 'max' ? '强制升到 max' : '强制升一档'
606  const higher = level === target ? '' : ',回答自己的判断更高'
607  return `${how}${higher};${knownReason(s, answer)}${answer.reading === null ? '' : `;${readingText(answer.reading)}`}`
608}
609
610/** Whether the failures were known to be expected, for the decision log: why they were not. */
611function knownReason(s: Settings, answer: Stuck): string {
612  if (answer.unread) return '读不到这个 agent 的记录,没有问这些失败是不是预期内的'
613  if (answer.expected !== null) return `不是预期内的失败(概率 ${answer.expected.toFixed(2)},预期内失败门槛 ${s.thetaExpected.toFixed(2)})`
614  return `决策模型没有回答(${failureLine(s.ctx.backend.name, answer.failure ?? { kind: 'parse', detail: 'no answer to the question' })})`
615}
616
617function newLoop(turnId: string): LoopRecord {
618  return { turnId, failures: 0, hookBlocks: 0, base: { failures: 0, hookBlocks: 0 }, raises: 0, step: null, engine: null, model: null, askedFor: null, raisedAt: null, paused: false }
619}
620
621/** The failures counted toward escalating: since the counts last started over, hook blocks only when they count. */
622function countedOf(record: LoopRecord): Counted {
623  return {
624    failures: record.failures - record.base.failures,
625    hookBlocks: isOn(BLOCKS_SWITCH) ? record.hookBlocks - record.base.hookBlocks : 0,
626  }
627}
628
629/** The counts as they stand: where they start over from once the failures are dealt with. */
630function countedNow(record: LoopRecord): Counted {
631  return { failures: record.failures, hookBlocks: record.hookBlocks }
632}
633
634/** The loop's counts start over from `covered` (a raise counts as one), and are shown. */
635async function startOver($: EngineInterface, cell: Cell<LoopRecord>, id: string, covered: Counted, raised: boolean, at?: number): Promise<LoopRecord> {
636  const record = await update(cell, (r) => ({
637    ...(r ?? newLoop('')),
638    base: { failures: Math.max(covered.failures, r?.base.failures ?? 0), hookBlocks: Math.max(covered.hookBlocks, r?.base.hookBlocks ?? 0) },
639    raises: (r?.raises ?? 0) + (raised ? 1 : 0),
640    ...(raised && at !== undefined ? { raisedAt: at } : {}),
641  }))
642  await showCounts($, id, record, null)
643  return record
644}
645
646/**
647 * Reports a loop's counts (on the board's node of it): the main agent's, or an agent's (null record: the main
648 * agent's start over with a new turn). `note`: `late`, a note for the band's event stream too.
649 */
650async function showCounts($: EngineInterface, id: string, record: LoopRecord | null, note: 'late' | null): Promise<void> {
651  await report(ioOf($), {
652    tally: {
653      feature: 'escalation',
654      agent: id,
655      failed: record?.failures ?? 0,
656      blocked: record?.hookBlocks ?? 0,
657      raised: record?.raises ?? 0,
658      ...(note === null ? {} : { late: true as const }),
659      ...(record === null ? { turnStart: true as const } : {}),
660    },
661  })
662}
663
hooks/features/find-skill.ts 234 lines
1// Feature: find_skill (#12). Most skills are left out of the main agent's
2// listing (features/skills.ts, ADR 0002), so when, partway through a turn,
3// the work turns out to need a skill nobody suggested, the main agent can ask:
4// it calls the tool with a few words on the work, and the decision model rates
5// the session's skills against them, the same way it rates them beside a
6// message (the same ranker, `modRanker`, in the same two stages (#11), over the
7// same skills and their profiles, with the recent conversation read by the
8// same rules). Nothing is pushed: the skills come back only as the tool's
9// answer.
10//
11// Its switch is `find-skill` (`/dp find-skill off`), apart from `skills`
12// (spec #56). The tool stays registered whatever the switches say, with a
13// description that never changes (the tool list is part of the prompt cache);
14// switched off, it says so when called.
15
16import type { EngineInterface, HttpInit, On } from 'claude-code'
17import { type Asked, type BackendIo, describeAsked, errorText, failureText } from '../decision/backend.ts'
18import { turnStartState } from '../decision/context.ts'
19import { quoteStart } from '../decision/redact.ts'
20import { modRanker, pickSkills, skillOpening, type SkillPick, type SkillPolicy, type SkillRanking } from '../decision/skills.ts'
21import { answersFor, mergeParts, type DecisionRequest } from '../decision/system-one.ts'
22import { readSessionSkills } from '../core/profiles.ts'
23import { report, type ReportIo } from '../core/report.ts'
24import type { Ctx } from '../core/setup.ts'
25import { describeStages, rankingSettings, type CatalogSkill } from '../core/skills.ts'
26import { defineSwitch, isOn, masterOn } from '../core/switches.ts'
27
28const CATALOG = { plugin: 'dispatch-pilot', key: 'skillCatalog' } as const
29const BOARD = { plugin: 'dispatch-pilot', key: 'board' } as const
30const DECISIONS = { plugin: 'dispatch-pilot', key: 'decisionLog' } as const
31const PPLX_RATE = { plugin: 'dispatch-pilot', key: 'pplxRate' } as const
32
33/** The switch's name, in `/dp` and in the decision log. */
34const SWITCH = 'find-skill'
35/** The skills feature's switch for profiles (#11): off, skills are rated by their descriptions here too. */
36const PROFILES = 'skill-profiles'
37/** The tool's name; the model calls it as `mcp__dispatch-pilot__find_skill`. */
38const TOOL = 'find_skill'
39const TOOL_CALLED = 'mcp__dispatch-pilot__find_skill'
40
41/** What the model reads about the tool. Fixed: nothing of the session goes into it. */
42const DESCRIPTION =
43  "Searches this session's skills for the ones that fit a piece of work, and returns each one's exact name, its relevance (0 to 1) and what it is for. Use it when the work at hand might have a skill you have not been shown: a file format, a service or its tooling, or a way of working such as reviewing, planning, testing or releasing. Load a skill it returns with the Skill tool by that exact name. Skills only the user can start are never returned."
44const INPUT = {
45  type: 'object',
46  properties: {
47    query: {
48      type: 'string',
49      description: 'The kind of work you need a skill for, in a few words, such as "fill in a form in a PDF file" or "review a branch before merging".',
50    },
51  },
52  required: ['query'],
53  additionalProperties: false,
54}
55/** How an answer that brings no skill ends: the main agent's way on. */
56const CARRY_ON = 'Carry on without it, or load a skill you know with the Skill tool by its exact name.'
57
58/**
59 * The session's skills: as read earlier this session (the skills feature
60 * reads them at its start), else read now, with the profiles the store holds
61 * (#11), and kept for the session; null when the session cannot be read.
62 */
63async function sessionCatalog($: EngineInterface, model: string): Promise<CatalogSkill[] | null> {
64  const { value } = await $.state.get(CATALOG)
65  if (value) return value.skills
66  const found = await readSessionSkills(
67    {
68      commands: () => $.command.list(),
69      listed: async () => (await $.session.usage({ breakdown: 'summary' })).context.breakdown?.skills?.skillFrontmatter ?? [],
70      overrides: async (source) => (await $.settings.read({ source })).skillOverrides,
71      home: () => $.env.get('HOME'),
72      cwd: () => $.session.cwd(),
73      exists: (path) => $.fs.exists(path),
74      read: (path) => $.fs.read(path),
75      list: (path) => $.fs.list(path),
76      get: (key) => $.store.get(key),
77    },
78    model,
79  )
80  if (found !== null) await $.state.set(CATALOG, { skills: found.skills })
81  return found?.skills ?? null
82}
83
84/** One decision request for find_skill through the person's decision model, its outcome in the debug log. */
85async function askLogged($: EngineInterface, ctx: Ctx, what: string, about: string, request: DecisionRequest, timeoutMs: number): Promise<Asked> {
86  const io: BackendIo = {
87    fetch: (url: string, init: HttpInit) => $.http.fetch(url, init),
88    sleep: (ms: number, signal: AbortSignal) => $.clock.sleep(ms, { signal }),
89    pace: { now: () => $.clock.now(), sentAt: { get: () => $.state.get(PPLX_RATE), set: (value, options) => $.state.set(PPLX_RATE, value, options) } },
90  }
91  const startedAt = await $.clock.now()
92  const asked = await ctx.backend.ask(io, request, timeoutMs)
93  const ms = (await $.clock.now()) - startedAt
94  $.ui.log(`${what} [${Object.keys(request.questions).join(', ')}] to ${ctx.backend.name} ${about}: ${describeAsked(asked, ms)}`, { to: 'debug' })
95  return asked
96}
97
98/** The opening of a catalog skill's SKILL.md for the ranking's second stage; null when it has no file. */
99async function openingOf($: EngineInterface, catalog: readonly CatalogSkill[], name: string): Promise<string | null> {
100  const file = catalog.find((skill) => skill.name === name)?.file ?? null
101  return file === null ? null : skillOpening(await $.fs.read(file))
102}
103
104export function registerFindSkill(on: On, ctx: Ctx): void {
105  defineSwitch({ name: SWITCH, info: '回答主 agent 的 find_skill:挑出适合它说的那项工作的 skill' })
106
107  /** Skills never offered (the option the skills feature reads too). */
108  const neverSuggested = new Set(ctx.config.skills.neverSuggested)
109  const policy = (): SkillPolicy => ctx.config.skills.find
110  /** How the mod's ranker ranks: the settings it rates the skills beside each message with. */
111  const rankBy = rankingSettings(ctx)
112  /** The model whose profiles the skills are offered by (#11). */
113  const model = ctx.config.skills.profileModel
114
115  // Registered once every plugin is loaded, under a match-all matcher (other
116  // features set themselves up at session start too). Without a decision
117  // model nothing could rate the skills, and the main agent keeps its listing.
118  on('session.start', { cwd: /(?:)/ }, async ($, e, next) => {
119    const result = await next(e)
120    if (ctx.backend.configured === false) return result
121    try {
122      const { tool } = await $.tool.register({ name: TOOL, description: DESCRIPTION, inputSchema: INPUT })
123      if (tool !== TOOL_CALLED) $.ui.log(`find_skill is registered as ${tool}, but the hook answers ${TOOL_CALLED}: its calls will fail`, { to: 'debug' })
124    } catch (error) {
125      $.ui.log(`find_skill was not registered: ${errorText(error)}`, { to: 'debug' })
126    }
127    return result
128  })
129
130  // The model's call: answered here, never passed on. Whatever goes wrong,
131  // the answer says so at once (fail open: the turn goes on without a skill).
132  on('tool.call', { tool: TOOL_CALLED }, async ($, e) => {
133    if (!masterOn()) return { result: `Dispatch Pilot is switched off (/dp on turns it back on), so find_skill rated no skills. ${CARRY_ON}` }
134    if (!isOn(SWITCH)) return { result: `find_skill is switched off (/dp ${SWITCH} on turns it back on), so it rated no skills. ${CARRY_ON}` }
135    // A dispatched agent's listing is left whole (ADR 0002): it has every skill in view.
136    if (e.agentId !== undefined) {
137      return { result: "find_skill rates skills for the main agent, whose skill listing is cut short. Yours lists this session's skills: pick from it and load one with the Skill tool by its exact name." }
138    }
139    const query = typeof e.query === 'string' ? e.query.replace(/\s+/g, ' ').trim() : ''
140    if (query === '') return { result: 'find_skill needs a query: a few words on the kind of work you need a skill for.' }
141    const io: ReportIo = {
142      board: { get: () => $.state.get(BOARD), set: (value, options) => $.state.set(BOARD, value, options) },
143      decisions: { get: () => $.state.get(DECISIONS), set: (value, options) => $.state.set(DECISIONS, value, options) },
144      debug: (line) => $.ui.log(line, { to: 'debug' }),
145      now: () => $.clock.now(),
146      toast: (text) => $.ui.toast(text),
147    }
148    // The main agent's own call, beside its decision for the turn: in the log, not on its node.
149    const call = { feature: SWITCH, agent: 'main', aside: true as const }
150    try {
151      // The skills the main agent can load, as a message asks about them: their question of stage
152      // one is the very one beside a message (those only the person can start have one of their
153      // own, which is not asked: they never come back); each by its profile, as beside a message,
154      // unless profiles are switched off.
155      const known = await sessionCatalog($, model)
156      const candidates = known?.filter((skill) => skill.by === 'model' && !neverSuggested.has(skill.name)).map((skill) => (isOn(PROFILES) ? skill : { ...skill, profile: null })) ?? null
157      if (candidates === null) {
158        await report(io, { decision: { ...call, skipped: 'unread' as const } })
159        return { result: `find_skill could not read this session's skills. ${CARRY_ON}` }
160      }
161      const about = `for find_skill ${quoteStart(query)}`
162      const ranker = modRanker(
163        {
164          ask: (request, timeoutMs) => askLogged($, ctx, 'second request', about, request, timeoutMs),
165          opening: (option) => openingOf($, candidates, option.name),
166        },
167        rankBy,
168      )
169      // The first request offers each skill by its profile where the decision model can read them all in time
170      // (BACKEND_DEFAULTS findSkillProfiles), else by its description; the second re-reads by both.
171      const part = ranker.part(ctx.config.skills.findByProfile ? candidates : candidates.map((skill) => ({ ...skill, profile: null })))
172      if (part === null) {
173        await report(io, { decision: { ...call, skipped: 'none' as const } })
174        return { result: `This session has no skill that find_skill could return. ${CARRY_ON}` }
175      }
176
177      // The same state as beside a message, the work named in place of the message; the ranker's second
178      // request (#11) asks about the same. Both requests share one wait: the second gets what the first left
179      // of it. The wait is the decision model's (findWaitMs: a message's timeoutMs with Jev),
180      // within the hook's own 10 s.
181      const waitMs = ctx.config.skills.findWaitMs
182      const startedAt = await $.clock.now()
183      const messages = ctx.config.context.messages > 0 ? await $.session.messages().catch(() => []) : []
184      const request = mergeParts(turnStartState({ prompt: query, messages, limits: ctx.config.context }), [part])
185      const asked = await askLogged($, ctx, 'request', about, request, waitMs)
186      const left = waitMs - ((await $.clock.now()) - startedAt)
187      const ranked = asked.ok ? await ranker.rank(answersFor(part, asked.answers), candidates, { state: request.state, timeoutMs: left }) : null
188      const failure = !asked.ok ? asked.failure : ranked === null ? { kind: 'parse' as const, detail: 'no answer about the skills' } : ranked.failed
189      if (ranked === null || failure !== undefined) {
190        const lost = failure ?? { kind: 'parse' as const, detail: 'no answer about the skills' }
191        const why = failureText(ctx.backend.name, lost)
192        await report(io, { decision: { ...call, failure: { backend: ctx.backend.name, ...lost } } })
193        return { result: `find_skill could not rate the skills (${why}). ${CARRY_ON}` }
194      }
195      const ranking = ranked
196
197      // Only skills the main agent can load were asked about, and they alone can come back.
198      const { suggest } = pickSkills(ranking, candidates, policy())
199      const names = suggest.map((skill) => skill.name).join('、')
200      await report(io, {
201        decision: {
202          ...call,
203          subject: quoteStart(query),
204          outcome: suggest.length > 0 ? `查到 ${names}` : '没查到 skill',
205          reason: describeRanking(ranking, policy()),
206          tone: suggest.length > 0 ? 'ok' : 'info',
207          skills: { suggest: suggest.map(({ name, relevance }) => ({ name, relevance })), try: [] },
208        },
209      })
210      return { result: found(query, suggest, policy()) }
211    } catch (error) {
212      $.ui.log(`find_skill failed: ${errorText(error)}`, { to: 'debug' })
213      await report(io, { decision: { ...call, skipped: 'error' as const } })
214      return { result: `find_skill could not rate the skills (an error in Dispatch Pilot, written to the debug log). ${CARRY_ON}` }
215    }
216  })
217}
218
219/** Why: what each stage of the ranking said (`describeStages`), and the bar a skill had to reach. */
220function describeRanking(ranking: SkillRanking, policy: SkillPolicy): string {
221  return `${describeStages(ranking)};相关度 ${policy.minRelevance.toFixed(2)} 起返回,最多 ${policy.max} 个`
222}
223
224/** The answer: the skills that fit, each by name, relevance and description, most relevant first; or that none does. */
225function found(query: string, suggest: readonly SkillPick[], policy: SkillPolicy): string {
226  if (suggest.length === 0) {
227    return `No skill fits "${query}": none reached relevance ${policy.minRelevance.toFixed(2)}. Carry on without one, try other words for the work, or load a skill you know with the Skill tool by its exact name.`
228  }
229  return [
230    `Skills that fit "${query}", rated by Dispatch Pilot’s decision model (relevance 0 to 1), most relevant first. Load one with the Skill tool by its exact name if it fits the work:`,
231    ...suggest.map((skill) => `- ${skill.name} (relevance ${skill.relevance.toFixed(2)})${skill.description ? `: ${skill.description}` : ''}`),
232  ].join('\n')
233}
234
hooks/features/main-effort.ts 197 lines
1// Feature: the main agent's effort, decided each time the person sends a
2// message, and when a report starts a turn of its own: a dispatched agent's
3// hand-back or a background task's notice reaching the idle session (those
4// turns are usually a look at a result and the next dispatch; the session's own
5// effort would be a waste). It adds the effort question to the message's ballot
6// (the core sends it); the answer becomes the effort of the turn the message
7// starts, on every step of that turn (the core's turn.step writer). A report's
8// text is read as the message; its turn is not re-decided mid-turn.
9//
10// A command turn (#19: `/implement #19`, a skill or markdown command the
11// person typed) is decided like their message: the decision model reads the
12// command as typed and what the command is for (core/commands.ts), never the
13// prompt it expands to.
14//
15// A message typed while a turn runs (`e.turnId`) is delivered into that turn
16// at its next step (a `queued_command` attachment; measured on 2.1.289), so
17// its decision takes the running turn from then on. It also waits as pending,
18// in case the turn ends first and the message starts a turn of its own.
19
20import type { EngineInterface, On } from 'claude-code'
21import { EFFORTS, LEVEL, probsOf, readEffort, readingText, traceEffort, type Effort, type EffortReading } from '../decision/effort.ts'
22import { quoteStart } from '../decision/redact.ts'
23import { turnStartPart } from '../decision/turn-start.ts'
24import { givesHint, judgeUnresolved, readUnresolved, UNRESOLVED } from '../decision/unresolved.ts'
25import { contribute, type PartOutcome } from '../core/ballot.ts'
26import { commandOf, commandState } from '../core/commands.ts'
27import { hintDecision, keptCount, keptSummary, moveCount, unresolvedDecision, type CountCell } from '../core/unresolved.ts'
28import { addPending, revise, turnKey, update, type Cell, type PendingDecision, type TurnRecord } from '../core/plans.ts'
29import { isPersonsMessage, startsReportTurn } from '../core/prompts.ts'
30import { report, type HintRecord, type ReportIo, type UnresolvedRecord } from '../core/report.ts'
31import type { Ctx } from '../core/setup.ts'
32import { defineSwitch, isOn } from '../core/switches.ts'
33import { UNRESOLVED_SWITCH } from './unresolved.ts'
34
35const PENDING = { plugin: 'dispatch-pilot', key: 'pending' } as const
36const TURNS = { plugin: 'dispatch-pilot', key: 'turns' } as const
37const BOARD = { plugin: 'dispatch-pilot', key: 'board' } as const
38const DECISIONS = { plugin: 'dispatch-pilot', key: 'decisionLog' } as const
39const CATALOG = { plugin: 'dispatch-pilot', key: 'skillCatalog' } as const
40const COUNT = { plugin: 'dispatch-pilot', key: 'unresolved' } as const
41
42/**
43 * What a command turn's command is for: its skill in the session's catalog as
44 * the skills feature read it (by its profile while profiles are on), else the
45 * command as `$.command.list()` describes it.
46 */
47async function describeCommand($: EngineInterface, name: string): Promise<Readonly<Record<string, string>> | null> {
48  const { value: catalog } = await $.state.get(CATALOG)
49  const found = catalog?.skills.find((skill) => skill.name === name)
50  const skill = found && !isOn('skill-profiles') ? { ...found, profile: null } : found
51  const listed = skill?.description ? undefined : (await $.command.list().catch(() => [])).find((command) => command.name === name)?.description
52  return commandState(name, skill, listed)
53}
54
55/** The unresolved count and summary in `$.state`, as a cell. */
56function countCell($: EngineInterface): CountCell {
57  return { get: () => $.state.get(COUNT), set: (value, options) => $.state.set(COUNT, value, options) }
58}
59
60export function registerMainEffort(on: On, ctx: Ctx): void {
61  defineSwitch({ name: 'main-effort', info: '发消息时决定主 agent 的 effort' })
62
63  on('prompt.submit', { text: /(?:)/ }, async ($, e, next) => {
64    // The person's own message, or a report that starts a turn of its own (a dispatched agent's hand-back, a task notice).
65    const handBack = !isPersonsMessage(e) && startsReportTurn(e)
66    if ((!isPersonsMessage(e) && !handBack) || !isOn('main-effort')) return next(e)
67    const pending: Cell<PendingDecision[]> = { get: () => $.state.get(PENDING), set: (value, options) => $.state.set(PENDING, value, options) }
68    let added: PendingDecision | null = null
69
70    /** The message waits for its turn, decided or not: the turn it starts is the person's own (mid-turn re-decisions are for such turns). */
71    const wait = async (effort: Effort | null) => {
72      const entry: PendingDecision = { text: e.text, effort, at: await $.clock.now(), ...(handBack ? { report: true as const } : {}) }
73      await update(pending, (list) => addPending(list ?? [], entry))
74      added = entry
75    }
76
77    const ran = handBack ? null : commandOf(e.text)
78    const command = ran === null ? null : await describeCommand($, ran.command)
79
80    // The person's own message also asks whether it says the problem they are on is still not solved (the unresolved
81    // count); a report that starts a turn is no word of theirs, and the switch can leave the question out. It travels
82    // in this part, so it is in the effort request with the 24000-token state it reads (ADR 0005).
83    const counting = !handBack && isOn(UNRESOLVED_SWITCH)
84    // The summary there is: a write that is not done yet is no reason to wait, the decision uses the one before it.
85    const summary = counting ? await keptSummary(countCell($)).catch(() => null) : null
86    // The count before this message (its own answer comes in the same request), and with it the strong hint once it has
87    // reached the setting: a sentence in the effort question about the kind of work, the level still the decision model's (ADR 0005).
88    const count = counting ? await keptCount(countCell($)).catch(() => 0) : 0
89    const maxAfter = ctx.config.unresolved.maxAfter
90    const hint: HintRecord | undefined = givesHint(count, maxAfter) ? { count, maxAfter } : undefined
91    // About the main agent of the turn this message starts, or of the one running when it was typed into it.
92    const forTurn = e.turnId === undefined ? ('next' as const) : ('current' as const)
93
94    /** The effort this message's turn goes out at, and the board's word for it. */
95    const decideEffort = async (outcome: PartOutcome, io: ReportIo, unresolved?: UnresolvedRecord): Promise<Effort | null> => {
96      const about = { feature: handBack ? 'main-effort (agent report)' : 'main-effort', agent: 'main', forTurn, subject: quoteStart(e.text) }
97      if (!outcome.ok) {
98        await report(io, { decision: { ...about, routed: false, failure: { backend: ctx.backend.name, ...outcome.failure } } })
99        await wait(null)
100        return null
101      }
102      const reading = readEffort(outcome.answers[LEVEL])
103      if (reading === null) {
104        await report(io, { decision: { ...about, routed: false, failure: { backend: ctx.backend.name, kind: 'parse', detail: 'no effort answer' } } })
105        await wait(null)
106        return null
107      }
108      // pickEffort's rules with their working: the board shows the steps, never recomputes them (#23).
109      const { effort, steps } = traceEffort(reading, ctx.config)
110      await report(io, {
111        decision: {
112          ...about,
113          routed: true,
114          outcome: `effort ${effort}`,
115          // The level as data: what reads the log (the band, the pane) never reads it out of the words.
116          effort,
117          reason: describeReading(reading, effort, ctx.config.thetaMax),
118          probs: probsOf(reading),
119          ...(reading.confidence === null ? {} : { conf: reading.confidence }),
120          trace: steps,
121          // What this message did to the unresolved count: the card draws it with the effort's.
122          ...(unresolved === undefined ? {} : { unresolved }),
123          // The card says the hint was given.
124          ...(hint === undefined ? {} : { hint }),
125        },
126      })
127      await wait(effort)
128      const running = e.turnId
129      if (running !== undefined) {
130        const ref = { ...TURNS, id: turnKey(running, undefined) }
131        const turn: Cell<TurnRecord> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
132        await update(turn, (record) => revise(record, effort))
133      }
134      return effort
135    }
136
137    /**
138     * What the answer to the unresolved question does to the count, which is moved here: what the log and the card say
139     * of it; null when there is no answer or the count cannot be kept (the debug log says so, the count stays).
140     */
141    const countUnresolved = async (answered: Extract<PartOutcome, { ok: true }>) => {
142      const reading = readUnresolved(answered.answers[UNRESOLVED])
143      if (reading === null) {
144        $.ui.log(`unresolved for ${quoteStart(e.text)}: no answer, the count stays`, { to: 'debug' })
145        return null
146      }
147      const judged = judgeUnresolved(reading)
148      try {
149        return unresolvedDecision(judged, await moveCount(countCell($), judged.change))
150      } catch (error) {
151        $.ui.log(`unresolved count not kept: ${error instanceof Error ? error.message : String(error)}`, { to: 'debug' })
152        return null
153      }
154    }
155
156    contribute(e.text, {
157      // Asked as the decision model's table says for these questions (Chinese with Jev); the other questions keep ctx.ask's.
158      ...turnStartPart({ ask: ctx.config.ask.turnStart, unresolved: counting, command, summary, count, maxAfter }),
159      settle: async (outcome) => {
160        const io: ReportIo = {
161          board: { get: () => $.state.get(BOARD), set: (value, options) => $.state.set(BOARD, value, options) },
162          decisions: { get: () => $.state.get(DECISIONS), set: (value, options) => $.state.set(DECISIONS, value, options) },
163          debug: (line) => $.ui.log(line, { to: 'debug' }),
164          now: () => $.clock.now(),
165          toast: (text) => $.ui.toast(text),
166        }
167        // The count first, so the effort's decision can say what this message did to it; neither depends on the other
168        // (either answer can be missing, and the effort is decided and kept whatever becomes of the count).
169        const counted = counting && outcome.ok ? await countUnresolved(outcome) : null
170        const effort = await decideEffort(outcome, io, counted?.unresolved)
171        // A count that moved is a decision of its own in the log (an answer that left it as it was is on the effort's card only).
172        if (counted?.unresolved !== undefined && counted.unresolved.before !== counted.unresolved.count) {
173          await report(io, { decision: { feature: UNRESOLVED_SWITCH, agent: 'main', aside: true, forTurn, subject: quoteStart(e.text), ...counted } })
174        }
175        // Each hint given is a decision of its own in the log: the count, that it was given, the level that came of it.
176        if (hint !== undefined) await report(io, { decision: { feature: UNRESOLVED_SWITCH, agent: 'main', aside: true, forTurn, subject: quoteStart(e.text), ...hintDecision(hint, effort) } })
177      },
178    })
179
180    const result = await next(e)
181    // Refused beneath: no turn will take this decision.
182    const withdrawn = added as PendingDecision | null
183    if (result.drop !== undefined && withdrawn !== null) {
184      await update(pending, (list) => (list ?? []).filter((entry) => !(entry.text === withdrawn.text && entry.at === withdrawn.at)))
185    }
186    return result
187  })
188}
189
190/** Why a level was picked: every level's probability, `max` held back below thetaMax when it was the most likely, and the backend's confidence. */
191function describeReading(reading: EffortReading, picked: Effort, thetaMax: number): string {
192  const p = reading.probabilities
193  const max = p[EFFORTS.length - 1] ?? 0
194  const held = picked !== 'max' && p.every((other) => other <= max)
195  return readingText(reading, held ? `max 的概率没到 max 门槛 ${thetaMax.toFixed(2)}` : undefined)
196}
197
hooks/features/midturn-effort.ts 391 lines
1// Feature: the main agent's effort, decided again while a turn runs (#5).
2//
3// Every N steps of a turn, and when the main agent dispatches an agent,
4// starts a Workflow or loads a skill, the decision model is asked again how
5// much reasoning the rest of the work needs. The question goes out the moment
6// a tool call of the main agent starts (`tool.call`), so it is answered while
7// the tool runs; the next step (`turn.step`, a layer above the core) takes
8// the answer, waits briefly when it is not back yet, and writes the turn's
9// plan, which the core then sends. Raising needs thetaUp; lowering needs
10// thetaDown, goes one level at a time and waits holdSteps after a raise (a
11// forced one too: the escalation feature, #7, marks it in the turn's plan).
12//
13// The counts the question carries (failed calls, hook blocks) are the
14// escalation feature's: one counter for the turn, since the counts last
15// started over.
16//
17// Measured on 2.1.289: the engine runs a step's tool calls while the response
18// still streams, so `tool.call` fires inside the step, before the step's
19// stream ends. This layer keeps the text of the step as it streams.
20
21import type { EngineInterface, HttpInit, On } from 'claude-code'
22import { describeAsked, errorText, within, type Asked, type BackendIo } from '../decision/backend.ts'
23import { messageText } from '../decision/context.ts'
24import { higherEffort, isEffort, probsOf, readEffort, readingText, type Effort } from '../decision/effort.ts'
25import {
26  contentLanguage,
27  judgeMidturn,
28  midturnEffortPart,
29  midturnRecord,
30  midturnState,
31  MIDTURN_LEVEL,
32  outcomeOf,
33  resultLine,
34  toolDetail,
35  verdictReason,
36  type MidturnCounts,
37  type MidturnInput,
38  type MidturnLimits,
39  type MidturnRules,
40  type Outcome,
41} from '../decision/midturn.ts'
42import { renderSummary } from '../decision/summary.ts'
43import { answersFor, mergeParts } from '../decision/system-one.ts'
44import { givesHint } from '../decision/unresolved.ts'
45import { wasBlocked } from '../core/outcomes.ts'
46import { floorHeld, MAIN, redecided, turnKey, update, type Cell, type TurnRecord } from '../core/plans.ts'
47import { hintDecision } from '../core/unresolved.ts'
48import { report, type HintRecord, type NodeFailure, type ReportIo } from '../core/report.ts'
49import type { Ctx } from '../core/setup.ts'
50import { defineSwitch, isOn } from '../core/switches.ts'
51import { UNRESOLVED_SWITCH } from './unresolved.ts'
52
53const TURNS = { plugin: 'dispatch-pilot', key: 'turns' } as const
54const MIDTURN = { plugin: 'dispatch-pilot', key: 'midturn' } as const
55const MAIN_STEP = { plugin: 'dispatch-pilot', key: 'mainStep' } as const
56const FAILURES = { plugin: 'dispatch-pilot', key: 'escalation', id: MAIN } as const
57const LOCK = { plugin: 'dispatch-pilot', key: 'lock' } as const
58const BOARD = { plugin: 'dispatch-pilot', key: 'board' } as const
59const DECISIONS = { plugin: 'dispatch-pilot', key: 'decisionLog' } as const
60const COUNT = { plugin: 'dispatch-pilot', key: 'unresolved' } as const
61const PPLX_RATE = { plugin: 'dispatch-pilot', key: 'pplxRate' } as const
62
63/** The feature's switch (`/dp midturn-effort on|off`). */
64const SWITCH = 'midturn-effort'
65
66type ToolEnd = { name: string; detail: string; outcome: Exclude<Outcome, 'running'> }
67type StepRecord = { index: number; text: string; tools: ToolEnd[] }
68/** The feature's own record of a main turn (`midturn` in types/index.d.ts). */
69type MidturnRecord = {
70  steps: number
71  engine: Effort | null
72  askedFor: number | null
73  recent: StepRecord[]
74}
75
76/** Steps kept in a turn's record. */
77const MAX_RECENT = 16
78
79/** Calls that start a new phase of the work (dispatch an agent, start a Workflow, load a skill): each re-decides. */
80const PHASE_TOOLS: ReadonlySet<string> = new Set(['Agent', 'Task', 'Workflow', 'Skill'])
81
82/** What the feature runs with: the shared ctx and the options it reads. */
83type Settings = {
84  ctx: Ctx
85  /** Re-decide at every step whose index is a multiple of this; 0 for never. */
86  every: number
87  rules: MidturnRules
88  /** How long a step waits for a re-decision not back yet. */
89  waitMs: number
90  limits: MidturnLimits
91}
92
93/** A step of the main agent: its turn and its index. */
94type MainStep = { turnId: string; index: number }
95/** A tool call as it starts: its name and what it works on. */
96type Starting = { name: string; detail: string }
97
98/**
99 * A re-decision on its way: asked for step `forStep` (`reason` says why);
100 * `answer` never rejects, and once it resolves `settled` holds it and `ms`
101 * how long it took.
102 */
103type InFlight = { forStep: number; reason: string; answer: Promise<Asked>; settled: Asked | null; ms: number; hint?: HintRecord }
104/** Re-decisions on their way, by turn key. Promises cannot live in $.state: a reload drops them (the step then keeps its effort). */
105const inFlight = new Map<string, InFlight>()
106/** The text of the main step streaming now, by turn key. */
107const streamed = new Map<string, { index: number; block: number; text: string }>()
108
109export function registerMidturnEffort(on: On, ctx: Ctx): void {
110  defineSwitch({ name: SWITCH, info: '一轮进行中重新判断主 agent 的 effort', parts: ['midturn'] })
111  // Read at each use: which decision model's settings these are is settled at the session start (core/setup.ts, `Ctx`).
112  const settings: Settings = {
113    ctx,
114    get every() {
115      return ctx.config.midturn.every
116    },
117    get rules() {
118      return ctx.config.midturn.rules
119    },
120    get waitMs() {
121      return ctx.config.midturn.waitMs
122    },
123    get limits() {
124      return ctx.config.midturn.limits
125    },
126  }
127
128  on('tool.call', { tool: /(?:)/ }, async ($, e, next) => {
129    // The main agent's own calls only: not a dispatched agent's, nor another plugin's $.tool.call.
130    if (e.agentId !== undefined || next.origin.plugin !== 'engine' || !isOn(SWITCH)) return next(e)
131    const detail = toolDetail(e)
132    let step: MainStep | null = null
133    try {
134      step = (await $.state.get(MAIN_STEP)).value ?? null
135      if (step !== null) await launch($, settings, step, { name: e.tool, detail })
136    } catch (error) {
137      $.ui.log(`midturn: ${errorText(error)}`, { to: 'debug' })
138    }
139    const result = await next(e)
140    if (step !== null) {
141      try {
142        const ended: ToolEnd = { name: e.tool, detail, outcome: outcomeOf(e.tool, result, wasBlocked(e.tool_use_id)) }
143        const at = step.index
144        const ref = { ...MIDTURN, id: turnKey(step.turnId, undefined) }
145        const cell: Cell<MidturnRecord> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
146        await update(cell, (r) => withTool(r ?? newRecord(), at, ended))
147      } catch (error) {
148        $.ui.log(`midturn: ${errorText(error)}`, { to: 'debug' })
149      }
150    }
151    return result
152  })
153
154  on('turn.step', { turnId: /(?:)/ }, async function* ($, e, next) {
155    if (e.agentId !== undefined || !isOn(SWITCH)) return yield* next(e)
156    const key = turnKey(e.turnId, undefined)
157    if (e.index === 0) {
158      // A new turn: what is left of earlier turns (a text, an answer never taken) is dropped.
159      for (const old of [...streamed.keys()]) if (old !== key) streamed.delete(old)
160      for (const old of [...inFlight.keys()]) if (old !== key) inFlight.delete(old)
161    }
162    try {
163      const note = await takeAnswer($, settings, e, key)
164      const engine = isEffort(e.effort) ? e.effort : null
165      await $.state.set(MAIN_STEP, { turnId: e.turnId, index: e.index })
166      const ref = { ...MIDTURN, id: key }
167      const cell: Cell<MidturnRecord> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
168      const record = await update(cell, (r) => ({ ...(r === undefined || e.index === 0 ? newRecord() : r), steps: e.index + 1, engine }))
169      const { value: turn } = await $.state.get({ ...TURNS, id: key })
170      // Quiet until the turn has been re-decided once (so a short turn shows nothing).
171      await report(ioOf($), {
172        tally: {
173          feature: 'midturn-effort',
174          agent: 'main',
175          steps: record.steps,
176          judged: turn?.decisions ?? 0,
177          changed: turn?.changes ?? 0,
178          quiet: record.askedFor === null || turn === undefined,
179          ...(note === null ? {} : 'late' in note ? { late: true as const } : { failure: note.failure }),
180        },
181      })
182    } catch (error) {
183      $.ui.log(`midturn: ${errorText(error)}`, { to: 'debug' })
184    }
185    streamed.set(key, { index: e.index, block: -1, text: '' })
186    const stream = next(e)
187    for await (const chunk of stream) {
188      if (chunk.kind === 'text') {
189        const live = streamed.get(key)
190        if (live !== undefined && live.index === e.index) {
191          live.text += live.block !== -1 && live.block !== chunk.index ? `\n${chunk.text}` : chunk.text
192          live.block = chunk.index
193        }
194      }
195      yield chunk
196    }
197    const result = await stream.result
198    try {
199      const text = messageText(result.answer, settings.limits.tokens)
200      const ref = { ...MIDTURN, id: key }
201      const cell: Cell<MidturnRecord> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
202      await update(cell, (r) => withText(r ?? newRecord(), e.index, text))
203    } catch (error) {
204      $.ui.log(`midturn: ${errorText(error)}`, { to: 'debug' })
205    }
206    return result
207  })
208}
209
210/**
211 * Asks again about the turn's effort, for the step after `step`, when a call
212 * starting is a reason to: a phase tool, or a step index that is a multiple of
213 * `every`. Once per step; the answer is left for that step to take.
214 */
215async function launch($: EngineInterface, s: Settings, step: MainStep, starting: Starting): Promise<void> {
216  const upcoming = step.index + 1
217  const key = turnKey(step.turnId, undefined)
218  const [{ value: turn }, { value: record }, { value: lock = null }, { value: failures }, sentAt] = await Promise.all([
219    $.state.get({ ...TURNS, id: key }),
220    $.state.get({ ...MIDTURN, id: key }),
221    $.state.get(LOCK),
222    $.state.get(FAILURES),
223    $.clock.now(),
224  ])
225  // Only a turn the person's own message started is re-decided, whether its start was decided or not (that request
226  // may have failed: then from the session's own effort); never without a decision model set up (every request would
227  // fail at once).
228  if (turn === undefined || turn.person !== true || lock !== null || record === undefined || record.engine === null || s.ctx.backend.configured === false) return
229  const reason = PHASE_TOOLS.has(starting.name) ? starting.name : s.every > 0 && upcoming % s.every === 0 ? `每 ${s.every} 步` : null
230  if (reason === null || record.askedFor === upcoming || inFlight.get(key)?.forStep === upcoming) return
231  // The problem the person is on, as of now (this turn's own message has moved the count already): the summary and the
232  // count go along, and the strong hint once the count has reached the setting (#41). Nothing of it with the switch off.
233  const { value: held } = isOn(UNRESOLVED_SWITCH) ? await $.state.get(COUNT) : { value: undefined }
234  const count = held?.count ?? 0
235  const summary = held?.summary
236  const maxAfter = s.ctx.config.unresolved.maxAfter
237  const hint: HintRecord | undefined = givesHint(count, maxAfter) ? { count, maxAfter, where: 'mid' } : undefined
238  const input: MidturnInput = {
239    message: turn.prompt,
240    step: upcoming,
241    current_effort: higherEffort(turn.effort ?? record.engine, floorHeld(turn, upcoming)) as Effort,
242    counts: countsOf(turn, step.turnId, failures),
243    recent_steps: summaries(record, step.index, key, starting, contentLanguage(turn.prompt)),
244    ...(summary === undefined ? {} : { problem_summary: renderSummary(summary, s.ctx.ask.language) }),
245    ...(count > 0 ? { unresolved_count: count } : {}),
246  }
247  const request = mergeParts(midturnState(input, s.limits), [midturnEffortPart(s.ctx.ask, { hint: hint !== undefined })])
248  const io: BackendIo = {
249    fetch: (url: string, init: HttpInit) => $.http.fetch(url, init),
250    sleep: (ms: number, signal: AbortSignal) => $.clock.sleep(ms, { signal }),
251    pace: { now: () => $.clock.now(), sentAt: { get: () => $.state.get(PPLX_RATE), set: (value, options) => $.state.set(PPLX_RATE, value, options) } },
252  }
253  const asking = s.ctx.backend.ask(io, request, s.ctx.config.timeoutMs)
254  const entry: InFlight = { forStep: upcoming, reason, answer: asking, settled: null, ms: 0, ...(hint === undefined ? {} : { hint }) }
255  entry.answer = asking.then(async (answered) => {
256    entry.ms = (await $.clock.now()) - sentAt
257    entry.settled = answered
258    return answered
259  })
260  inFlight.set(key, entry)
261  const ref = { ...MIDTURN, id: key }
262  const cell: Cell<MidturnRecord> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
263  await update(cell, (r) => ({ ...(r ?? newRecord()), askedFor: upcoming }))
264}
265
266/**
267 * The turn's counts as the decision model reads them: its decisions and level
268 * changes, and the failed calls and hook blocks the escalation feature counted
269 * for it, since its counts last started over (a forced raise, or failures
270 * found expected).
271 */
272function countsOf(turn: TurnRecord, turnId: string, failures: { turnId: string; failures: number; hookBlocks: number; base: { failures: number; hookBlocks: number } } | undefined): MidturnCounts {
273  const counted = failures !== undefined && failures.turnId === turnId
274  return {
275    judgments: turn.decisions,
276    changes: turn.changes,
277    failures: counted ? failures.failures - failures.base.failures : 0,
278    hook_blocks: counted ? failures.hookBlocks - failures.base.hookBlocks : 0,
279  }
280}
281
282/**
283 * Takes the re-decision meant for this step, if any, and writes what it
284 * decides into the turn's plan. An answer not back yet gets `waitMs` more;
285 * still none, the step keeps the effort it had and the answer stays for a
286 * later step. Resolves to what the tally should add: `late`, or the failed
287 * request that is why there is no answer; null otherwise.
288 */
289async function takeAnswer($: EngineInterface, s: Settings, e: { index: number; effort?: unknown }, key: string): Promise<{ late: true } | { failure: NodeFailure } | null> {
290  const pending = inFlight.get(key)
291  if (pending === undefined || pending.forStep > e.index) return null
292  let asked = pending.settled
293  if (asked === null) {
294    asked = await within((ms, signal) => $.clock.sleep(ms, { signal }), pending.answer, s.waitMs, null)
295  }
296  if (asked === null) return { late: true }
297  inFlight.delete(key)
298  $.ui.log(`request [midturn.${MIDTURN_LEVEL}] for step ${pending.forStep} (${pending.reason}) to ${s.ctx.backend.name}: ${describeAsked(asked, pending.ms)}`, { to: 'debug' })
299  const engine = isEffort(e.effort) ? e.effort : null
300  const [{ value: turn }, { value: lock = null }] = await Promise.all([$.state.get({ ...TURNS, id: key }), $.state.get(LOCK)])
301  if (turn === undefined || lock !== null || engine === null) return null
302  const floor = floorHeld(turn, e.index)
303  const current = higherEffort(turn.effort ?? engine, floor) as Effort
304  const reading = asked.ok ? readEffort(answersFor(midturnEffortPart(s.ctx.ask), asked.answers)[MIDTURN_LEVEL]) : null
305  /** The hint given with this request is a decision of its own in the log: the level the decision model came to, when it did. */
306  const logHint = async (picked: Effort | null) => {
307    if (pending.hint !== undefined) await report(ioOf($), { decision: { feature: UNRESOLVED_SWITCH, agent: 'main', aside: true, subject: `第 ${e.index} 步(${pending.reason})`, ...hintDecision(pending.hint, picked) } })
308  }
309  if (reading === null) {
310    await logHint(null)
311    return { failure: { backend: s.ctx.backend.name, ...(asked.ok ? { kind: 'parse' as const, detail: 'no effort answer' } : asked.failure) } }
312  }
313
314  const ref = { ...TURNS, id: key }
315  const turnCell: Cell<TurnRecord> = { get: () => $.state.get(ref), set: (value, options) => $.state.set(ref, value, options) }
316  const sinceRaise = turn.raisedAt == null ? null : e.index - turn.raisedAt
317  const position = { current, sinceRaise, atLeast: floor }
318  const verdict = judgeMidturn(reading, position, s.rules)
319  await update(turnCell, (r) => redecided(r ?? turn, current, verdict.effort, e.index))
320  // Beside the turn's own decision (the node's link to it stays), with the rules' working: the answer's pick, then the move.
321  await report(ioOf($), {
322    decision: {
323      feature: SWITCH,
324      agent: 'main',
325      aside: true,
326      subject: `第 ${e.index} 步(${pending.reason})`,
327      outcome: `effort ${verdict.effort}${verdict.effort === current ? '(保持)' : `(原 ${current})`}`,
328      reason: `${readingText(reading)};${verdictReason(verdict, position, s.rules)}`,
329      tone: verdict.effort === current ? 'info' : 'ok',
330      probs: probsOf(reading),
331      conf: verdict.confidence,
332      trace: [...verdict.trace.pick.steps, ...verdict.trace.steps],
333      mid: midturnRecord(verdict, position, s.rules),
334    },
335  })
336  await logHint(verdict.picked)
337  return null
338}
339
340function newRecord(): MidturnRecord {
341  return { steps: 0, engine: null, askedFor: null, recent: [] }
342}
343
344/**
345 * The latest steps as the decision model reads them, oldest first: each
346 * step's text and its tool calls with how they ended. Step `index` comes
347 * last, its text as streamed so far, the call `starting` in it marked running.
348 */
349function summaries(record: MidturnRecord, index: number, key: string, starting: Starting, language: 'en' | 'zh'): MidturnInput['recent_steps'] {
350  const live = streamed.get(key)
351  const now = record.recent.find((step) => step.index === index)
352  const line = (t: { name: string; detail: string; outcome: Outcome }) => ({ name: t.name, result: resultLine(t.outcome, t.detail, language) })
353  return [
354    ...record.recent.filter((step) => step.index < index).map((step) => ({ assistant_text: step.text, tools: step.tools.map(line) })),
355    {
356      assistant_text: live?.index === index ? live.text : (now?.text ?? ''),
357      tools: [...(now?.tools ?? []).map(line), line({ ...starting, outcome: 'running' })],
358    },
359  ]
360}
361
362/** The record with a tool call's ending added to its step. */
363function withTool(record: MidturnRecord, index: number, ended: ToolEnd): MidturnRecord {
364  const at = record.recent.findIndex((step) => step.index === index)
365  const step = at >= 0 ? (record.recent[at] as StepRecord) : { index, text: '', tools: [] }
366  const recent = [...record.recent]
367  if (at >= 0) recent[at] = { ...step, tools: [...step.tools, ended] }
368  else recent.push({ ...step, tools: [ended] })
369  return { ...record, recent: recent.sort((a, b) => a.index - b.index).slice(-MAX_RECENT) }
370}
371
372/** The record with a step's text, once the step has streamed whole. */
373function withText(record: MidturnRecord, index: number, text: string): MidturnRecord {
374  const at = record.recent.findIndex((step) => step.index === index)
375  const recent = [...record.recent]
376  if (at >= 0) recent[at] = { ...(recent[at] as StepRecord), text }
377  else recent.push({ index, text, tools: [] })
378  return { ...record, recent: recent.sort((a, b) => a.index - b.index).slice(-MAX_RECENT) }
379}
380
381/** What the decision report needs of the host: its two cells, the debug log, the clock and the toast. */
382function ioOf($: EngineInterface): ReportIo {
383  return {
384    board: { get: () => $.state.get(BOARD), set: (value, options) => $.state.set(BOARD, value, options) },
385    decisions: { get: () => $.state.get(DECISIONS), set: (value, options) => $.state.set(DECISIONS, value, options) },
386    debug: (line) => $.ui.log(line, { to: 'debug' }),
387    now: () => $.clock.now(),
388    toast: (text) => $.ui.toast(text),
389  }
390}
391
hooks/features/unresolved.ts 161 lines
1// Feature: the unresolved count and the problem summary (#39, #40; GLOSSARY 未解决次数,
2// 问题摘要). Each of the person's messages is asked, in the effort request, whether it
3// says the problem they and the main agent are on is still not solved, is solved, or is
4// another one; the answer moves the count (core/unresolved.ts). The question and the
5// reading of its answer are in features/main-effort.ts, which owns the effort part of
6// the ballot they travel in (one contribution per part); this file owns the switch,
7// the count's start-over and the summary.
8//
9// The summary is a record of the problem for the decision model to read beside the
10// conversation. When a turn that the person's own message started ends, a cheap model
11// continues it, in the background: the hook does not wait, and a message that comes
12// before it is written is decided with the summary as it was (main-effort.ts reads it).
13// Writes go one after the other, each continuing what the one before wrote. A write
14// that fails (the model errs, times out, answers something that is no summary) or that
15// a reload of the mod lost leaves the summary as it was and is told to the decision
16// report; one the count's start-over overtook is dropped when it lands.
17//
18// The count is per session: `/clear` and a new session start it over (`session.end`
19// fires for both; `session.start` also fires on a hot reload, which must keep it),
20// a compaction keeps it. Switched off, the question is not asked, the count stays as it
21// is and no summary is written.
22
23import type { EngineInterface, ModelCompleteResult, On } from 'claude-code'
24import { errorText } from '../decision/backend.ts'
25import { clipToTokens } from '../decision/context.ts'
26import { quoteStart, redactSecrets } from '../decision/redact.ts'
27import { readSummary, SUMMARY_MAX_REPLY, SUMMARY_SYSTEM, SUMMARY_TIMEOUT_MS, summaryPrompt, turnTools, type TurnInput } from '../decision/summary.ts'
28import { turnKey } from '../core/plans.ts'
29import { report, type ReportIo } from '../core/report.ts'
30import type { Ctx } from '../core/setup.ts'
31import { defineSwitch, isOn } from '../core/switches.ts'
32import { clearCount, dropSummary, landSummary, lostSummaries, queueSummary, summaryFailure, summaryFor, type CountCell, type SummaryFailure } from '../core/unresolved.ts'
33
34const COUNT = { plugin: 'dispatch-pilot', key: 'unresolved' } as const
35const TURNS = { plugin: 'dispatch-pilot', key: 'turns' } as const
36const BOARD = { plugin: 'dispatch-pilot', key: 'board' } as const
37const DECISIONS = { plugin: 'dispatch-pilot', key: 'decisionLog' } as const
38
39/** The switch's name: `/dp unresolved on|off`. */
40export const UNRESOLVED_SWITCH = 'unresolved'
41
42/** The writes of this load of the mod: after the one before it (they are chained), and the turns they are for. A hot reload empties both. */
43let chain: Promise<void> = Promise.resolve()
44const running = new Set<string>()
45/** The engine refused the summary's model: no more is asked of it until the conversation ends (a reload asks again). */
46let refused = false
47
48/** What a write is given when the turn ends: the turn, and what the cheap model is shown of it. */
49type Write = { turn: string; subject: string; shown: Omit<TurnInput, 'previous'> }
50
51export function registerUnresolved(on: On, ctx: Ctx): void {
52  // Off until the person turns it on (#48, ADR 0006); what they set by hand is in $.store and wins.
53  defineSwitch({ name: UNRESOLVED_SWITCH, info: '每条消息判断同一个问题是否仍未解决,数未解决的次数;每轮结束后在后台写问题摘要(默认关)', default: false })
54
55  on('session.end', { reason: /(?:)/ }, async ($, e, next) => {
56    refused = false
57    try {
58      await clearCount(countCell($))
59    } catch (error) {
60      $.ui.log(`unresolved count not cleared: ${error instanceof Error ? error.message : String(error)}`, { to: 'debug' })
61    }
62    return next(e)
63  })
64
65  // A reload of the mod loses the writes in flight (the engine drops their calls with the old module): the state still lists them.
66  on('session.start', { cwd: /(?:)/ }, async ($, e, next) => {
67    const result = await next(e)
68    try {
69      for (const turn of await lostSummaries(countCell($), running)) await tell($, turn, 'lost', ctx.config.unresolved.summaryModel)
70    } catch (error) {
71      $.ui.log(`summaries lost to a reload not told: ${errorText(error)}`, { to: 'debug' })
72    }
73    return result
74  })
75
76  // A turn the person's own message started ends: its summary is written, in the background.
77  on('turn.complete', { turnId: /(?:)/ }, async ($, e, next) => {
78    const result = await next(e)
79    // The summary is a paraphrase of the conversation: none is written when the person sends the decision model none (`contextMessages` 0).
80    if (e.agentId !== undefined || refused || !isOn(UNRESOLVED_SWITCH) || ctx.backend.configured === false || ctx.config.context.messages <= 0) return result
81    try {
82      const { value: turn } = await $.state.get({ ...TURNS, id: turnKey(e.turnId, undefined) })
83      // Only the person's own message: a hand-back or a task notice that started the turn is no word of theirs.
84      if (turn === undefined || !turn.person) return result
85      const messages = await $.session.messages().catch(() => [])
86      const write: Write = { turn: e.turnId, subject: quoteStart(turn.prompt), shown: { person: turn.prompt, reply: e.answer, tools: turnTools(messages) } }
87      await queueSummary(countCell($), e.turnId)
88      running.add(e.turnId)
89      chain = chain.then(() => writeSummary($, ctx, write)).catch(() => undefined)
90    } catch (error) {
91      $.ui.log(`summary not queued: ${errorText(error)}`, { to: 'debug' })
92    }
93    return result
94  })
95}
96
97function countCell($: EngineInterface): CountCell {
98  return { get: () => $.state.get(COUNT), set: (value, options) => $.state.set(COUNT, value, options) }
99}
100
101/** The decision report's closures over this hook's `$`. */
102function reportIo($: EngineInterface): ReportIo {
103  return {
104    board: { get: () => $.state.get(BOARD), set: (value, options) => $.state.set(BOARD, value, options) },
105    decisions: { get: () => $.state.get(DECISIONS), set: (value, options) => $.state.set(DECISIONS, value, options) },
106    debug: (line) => $.ui.log(line, { to: 'debug' }),
107    now: () => $.clock.now(),
108    toast: (text) => $.ui.toast(text),
109  }
110}
111
112/** A write that did not make a summary, told to the decision report (the one that was there stays). `turn` only names the write in the debug log. */
113async function tell($: EngineInterface, turn: string, failure: SummaryFailure, model: string, subject = ''): Promise<void> {
114  $.ui.log(`summary for turn ${turn}: ${failure === 'lost' ? 'lost to a reload' : failure}`, { to: 'debug' })
115  await report(reportIo($), { decision: { feature: UNRESOLVED_SWITCH, agent: 'main', aside: true, subject, ...summaryFailure(failure, model) } })
116}
117
118/**
119 * Continues the summary with one turn, with the cheap model, once the writes before it are done. The summary it
120 * continues is read when it starts, so it holds what those wrote. Never throws.
121 */
122async function writeSummary($: EngineInterface, ctx: Ctx, write: Write): Promise<void> {
123  const cell = countCell($)
124  const model = ctx.config.unresolved.summaryModel
125  try {
126    const { wanted, previous } = await summaryFor(cell, write.turn)
127    // Cleared since the turn ended (solved, another problem, /clear), or switched off meanwhile: nothing to write.
128    if (!wanted || !isOn(UNRESOLVED_SWITCH) || refused) {
129      await dropSummary(cell, write.turn)
130      return
131    }
132    let reply: ModelCompleteResult
133    try {
134      reply = await $.model.complete({ model, system: SUMMARY_SYSTEM, prompt: summaryPrompt({ ...write.shown, previous }), maxTokens: SUMMARY_MAX_REPLY, timeoutMs: SUMMARY_TIMEOUT_MS })
135    } catch (error) {
136      refused = true
137      await dropSummary(cell, write.turn)
138      await tell($, write.turn, { reason: 'refused', detail: errorText(error) }, model, write.subject)
139      return
140    }
141    if (!reply.isAnswered) {
142      await dropSummary(cell, write.turn)
143      await tell($, write.turn, reply.reason === 'api-error' ? { reason: 'api-error', status: reply.status, error: reply.error } : { reason: reply.reason }, model, write.subject)
144      return
145    }
146    const summary = readSummary(reply.text)
147    if (summary === null) {
148      await dropSummary(cell, write.turn)
149      await tell($, write.turn, { reason: 'unfit', text: clipToTokens(redactSecrets(reply.text).replace(/\s+/g, ' ').trim(), 60) }, model, write.subject)
150      return
151    }
152    const kept = await landSummary(cell, write.turn, summary)
153    $.ui.log(`summary for turn ${write.turn}: ${kept ? 'written' : 'dropped, it was cleared meanwhile'} (${reply.usage.output_tokens} tokens from ${model})`, { to: 'debug' })
154  } catch (error) {
155    // The session went away under it, or a bug: either way the session is not held up.
156    $.ui.log(`summary for turn ${write.turn} not written: ${errorText(error)}`, { to: 'debug' })
157  } finally {
158    running.delete(write.turn)
159  }
160}
161