SLOPSHOPPER

codemode

One codemode tool: the model writes a script that calls the session's tools, and only the script's output returns.

newrowsguardprompttoolprocess
★ 1v0.6.0MITupdated 2026-10-09ejklock/claude-code-mode
A shopper browsing a rack in a slop shop
README

codemode for Claude Code: a bridge to Pi's code mode

Code mode for Claude Code: one tool that lets the model write a JavaScript script that calls the session's own tools (Read, Bash, Write, Edit, the MCP resource tools, and the session's MCP tools), and returns only the script's output. Many tool calls become one, and large intermediate results stay out of the context window.

Inspired by Pi's codemode, by Earendil. Earendil's post "You Said No MCP!" explains the idea and why it works, and Armin Ronacher's "What is Codemode" describes it from the harness side. This project does not reimplement it. It is a bridge: it runs Pi's own script runtime, @earendil-works/pi-codemode (QuickJS in WebAssembly), and connects each tools.* call the script makes to Claude Code's own tool call. Your permission rules, prompts and hooks still apply to every one of them, and a script written for Pi's codemode reads the same here.

Status: pilot (v0.6.0). It exposes Read, Bash, Write, Edit, the three MCP resource tools, and every MCP tool connected in the session. The design is in ADR 0001, which is still Proposed. The bridge overhead and the task-level savings are measured — see Benchmarks and issue 0005.

Claude Code picks the codemode tool on its own: one script lists the TypeScript files with git, reads all five in parallel, filters the TODO lines, and only the filtered output returns

The prompt never mentions codemode. Claude writes one script that runs git ls-files, reads five files in parallel with Promise.allSettled, and keeps only the TODO lines. The nested calls are listed live as they run, and only the script's output returns to the model.

What it is and where it shines

What it is: one tool, codemode. The model sends a short JavaScript script. The script runs in a sandbox and calls the session's tools as tools.Read(...), tools.Bash(...), tools.Write(...), tools.Edit(...), the MCP resource tools (tools.ListMcpResourcesTool(...), tools.ReadMcpResourceTool(...), tools.ReadMcpResourceDirTool(...)) and tools.mcp__server__tool(...). Only what the script prints returns to the model. Every nested call still meets your permission rules, prompts and hooks.

What it is for: work that takes many tool calls, where each call needs a model turn and each result fills the context.

Where it shines:

  • Batching: read or search many files at once with Promise.allSettled, instead of one call per file.
  • Filtering: keep the matching lines and drop the rest, so a large result never reaches the context window.
  • Chaining: use the result of one call as the input of the next, such as a git listing followed by a read of each file, in one turn.
  • MCP tools: call several tools of a server, or of different servers, from one script.

Where it does not: a single command, such as one git grep. Claude still calls Bash directly, which is the right choice.

Benchmarks

Measured with no model in the loop (node scripts/overhead.ts, 15 runs per size): the bridge costs ~98 ms to open (child start, sandbox, socket, close) and ~0.1–0.3 ms per nested call — 100 sequential calls add ~3 ms, within run-to-run noise. The bridge is expensive to open and nearly free to use.

The same tasks, with and without the plugin (node scripts/savings.ts, 2026-10-06, Claude Code 2.1.292, claude-opus-5-5, 3 runs per side after one discarded warm-up, medians over all-correct runs, the prompt never naming codemode):

TaskTurns with / withoutCache-read tokens with / withoutOutput tokens with / withoutCost with / without
read 5 tracked files, report their TODOs2 / 738,188 / 54,286574 / 966$0.023 / $0.039
git log → read each changed file2 / 438,222 / 55,041443 / 414$0.021 / $0.024
write 3 files6 / 8111,935 / 93,5811,229 / 1,321$0.062 / $0.059
one git grep2 / 238,329 / 32,780164 / 259$0.011 / $0.029

Why turns matter more than output tokens: every turn sends the whole context to the model again. With the prompt cache, that reread is billed as cache-read tokens, not as input, which is why the input column is near zero and the cache-read column carries the reread. Fewer turns mean fewer rereads, and the nested results a script handles never enter the context, so the turns that remain are smaller too.

Where many reads batch into one script, codemode cut the turns to less than a third, the cache-read tokens by ~30%, the output tokens to ~60% and the cost to ~60%, and the wall time fell with the turns (7.7 s against 11.9 s). On the write task the turns fell a quarter, but the cache reads rose (likely because the tool's schema and the script's confirmation output ride every remaining turn), so tokens and cost were level. On the single git grep — where codemode should not help — the model never used it (0 of 3 runs) and called Bash directly; that row's differences are run-to-run noise, not a saving or a cost of the tool. Input tokens are near zero on both sides because the context rides the prompt cache; the codemode tool's own schema (~430 tokens) is inside the with-side numbers. Ranges, method and the honest caveats are in issue 0005; with 3 runs per side, only the large differences above are claimed.

When a script fails, the result also lists the nested calls that already ran (#<id> <tool> <args> — <state>: <detail>), so a retry redoes only what did not. A call that threw, or never answered, is marked unknown rather than failed: it may have taken effect, so check before redoing it. In headless checks (node scripts/partial.ts, 3 runs, a store whose create tool fails once, and with --task ambiguous one whose fourth create is stored but never answered) the model duplicated no write, before or after the change, so no saving is claimed; see issue 0006.

Read more: Earendil's Pi codemode and "You Said No MCP!"; Armin Ronacher's "What is Codemode"; Cloudflare's Code Mode; Anthropic's Code execution with MCP.

Why code mode

Calling tools one at a time costs a model turn per call, and every result, however large, lands in the context. In code mode the model writes a short program instead:

const root = (await tools.Bash({ command: 'pwd' })).trim()
const files = (await tools.Bash({ command: 'git diff --name-only main' })).split('\n').filter(Boolean)
for (const file of files) {
  const source = await tools.Read({ file_path: `${root}/${file}` })
  if (source.includes('TODO')) text(file)
}

One tool call runs the whole loop, and only the matching file names come back.

The pattern is known as code mode: see Cloudflare's Code Mode and Anthropic's Code execution with MCP. This plugin makes it native to Claude Code.

How it works

sequenceDiagram
    participant M as Model
    participant Mod as codemode mod (hooks module)
    participant C as Node child (pi-codemode)
    participant CC as Claude Code tool call
    M->>Mod: codemode({ code })
    Mod->>C: spawn node child/main.ts, send the script
    C-->>Mod: listening on a Unix socket
    loop each tools.X(args) in the script
        C-->>Mod: call { id, tool, input } (stdout line)
        Mod->>CC: $.tool.call (permission check, hooks)
        CC-->>Mod: result or denial
        Mod->>C: POST the answer over the socket
    end
    C-->>Mod: done { output }
    Mod-->>M: only the script's output
  • The tool: the mod registers the tool, which the model sees as mcp__codemode__codemode.
  • Separate process: a mod's hooks module has no WebAssembly and no eval. So each script runs in a short-lived Node child process that hosts pi-codemode, pinned to an exact version.
  • Permissions: every nested call runs through $.tool.call, never $.mcp.call. Your allow and deny rules, permission prompts and PreToolUse hooks see each one.
  • Denials: a denied or failed call throws an Error inside the script, and the script can catch it and continue.

Requirements

  • Claude Code with mods support (tested on 2.1.291 and 2.1.292; the mods API is early access).
  • Node.js 22.19 or newer on your PATH (tested on Node 26).

Install

Run one command:

curl -fsSL https://raw.githubusercontent.com/ejklock/claude-code-mode/main/install.sh | sh

The script adds the marketplace, installs the plugin, and installs its dependency. It is safe to run again; a second run updates the plugin.

To choose where the plugin is installed, set SCOPE to user (the default), project or local:

curl -fsSL https://raw.githubusercontent.com/ejklock/claude-code-mode/main/install.sh | SCOPE=project sh

If you prefer not to pipe a script, install by hand. From Claude Code:

/plugin install codemode --marketplace ejklock/claude-code-mode

Answer y to add the marketplace, then pick a scope.

Known gap: the child needs @earendil-works/pi-codemode, and node_modules is not in the repository. For the manual route, run npm ci --omit=dev in the installed plugin's folder after you install. The install script does this for you.

To develop or try it from a clone:

git clone https://github.com/ejklock/claude-code-mode
cd claude-code-mode
npm ci
claude --plugin-dir .

Use

Give Claude a task with many steps. Like Pi, the plugin keeps the tool declared up front, describes the script API to the model, and adds one line to the system prompt and one note to the end of Bash's description. With that, Claude picks codemode on its own when batching, chaining or filtering helps: in a measured run of a "read every file and report its TODOs" task, it chose codemode in 5 of 5 runs, against 0 of 5 without these hints (issue 0004). For a single command, such as one git grep, it still calls Bash directly.

The transcript draws each call as two boxes. The first holds the highlighted script and its nested calls, each with its state and duration. The second holds a summary and the output:

The codemode tool in the Claude Code transcript: the highlighted script, six nested calls with durations, then a green summary box with the filtered output

In a script:

APIWhat it does
await tools.Read({ file_path, offset?, limit? })Resolves to the file's text.
await tools.Bash({ command, timeout? })Resolves to the command's output.
await tools.Write({ file_path, content })Writes the file; resolves to a confirmation.
await tools.Edit({ file_path, old_string, new_string, replace_all? })Replaces text in the file; resolves to a confirmation.
await tools.ListMcpResourcesTool({ server? })Lists the resources MCP servers offer; resolves to the list.
await tools.ReadMcpResourceTool({ server, uri })Reads an MCP resource; resolves to its text.
await tools.ReadMcpResourceDirTool({ server, uri })Lists the resources under an MCP resource directory; resolves to the list.
await tools.mcp__server__tool(args) / await tools"mcp__server__tool"Calls a connected MCP tool by its full name; resolves to its text. A tool that is not connected is absent, and ALL_TOOLS lists those that are.
text(value) / console.log(value)Adds to the output that returns to the model.

The script is the body of an async function, so top-level await works. return and exit() end the script; only what it prints with text() or console.log() comes back. Scripts time out after 120 seconds. Every tools.* call meets the session's permission mode and rules as a direct call does, so Write and Edit follow acceptEdits, allow rules and deny rules, and a refusal reaches the script as a rejection.

MCP exposure modes

Each MCP server, or each of its tools, takes one of Pi's four exposure modes. The modes set what the model sees and what a script can do:

ModeWhat the model seesIn a script
codemodeThe tool is not declared up front. Its description tells the model to use it only through the codemode tool, and the codemode description lists it.Callable.
deferredThe tool is not declared until tool search loads it.Callable.
directThe tool is declared with its full schema every turn, like a built-in tool.Callable.
hiddenNothing. A direct call is refused.Refused, and left out of ALL_TOOLS.

A server named in no list keeps Claude Code's own placement and stays callable from scripts. Installing the plugin changes no server until you name it. This differs from Pi, where the default is codemode.

Four settings hold the entries, one list for each mode: mcpCodemode, mcpDeferred, mcpDirect and mcpHidden. An entry is a server name (codegraph), one tool (claude_ai_Gmail__trash_message), or a server and a tool pattern, where * matches any characters (claude_ai_Gmail__trash_*). Write the name without the mcp__ prefix.

/plugin configure codemode draws each setting as one text line and stores what you type as one comma-separated string, such as codegraph, claude_ai_Gmail. To set them by hand, put the lists in settings.json, under the plugin's name:

{
  "pluginConfigs": {
    "codemode": {
      "options": {
        "mcpCodemode": ["claude_ai_Gmail", "codegraph"],
        "mcpHidden": ["claude_ai_Gmail__trash_*"]
      }
    }
  }
}

When more than one entry matches a tool, an exact tool name wins over a pattern, and a pattern wins over a server name. When two patterns match, the lists are read in the order mcpHidden, mcpCodemode, mcpDeferred, mcpDirect, and the first match wins. The same entry in two lists, or twice in one, fails the load with a message that names it. In the example, every Gmail tool is in codemode mode except the trash_* tools, which are hidden.

Limits:

  • deferred takes effect only for a tool the engine is willing to defer. With ENABLE_TOOL_SEARCH=auto, the engine can keep a tool's schema in the request, and the model can call it directly.
  • A direct call to a codemode tool still runs, as in Pi. The mode keeps the model from seeing the tool; it does not refuse the call.
  • A hidden tool is refused for the model and missing from scripts.
  • A codemode tool stays in the codemode description, unlike Pi, because scripts here have no searchTools() or describeTool().

The measurements, including the adoption runs and the subagent view, are in issue 0011.

Develop

npm ci
claude plugin validate .   # manifest and hooks module
claude plugin test .       # the mod's side, with a stand-in child
npm test                   # the real child on a real socket, and invariants
npm run typecheck
node scripts/e2e.ts        # headless end to end with claude -p (spends model tokens)
sh scripts/install-check.sh # runs install.sh for real into a throwaway config (slow)
node scripts/partial.ts --runs 3  # duplicated writes after a failed script, headless (spends model tokens)

npm test and the end-to-end run open a Unix socket, so run them outside a sandbox that blocks socket listen().

Decisions and open work live in docs/: the constitution, the ADRs, the issues and the research.

Roadmap

  • Every built-in tool and the Agent tool as typed tools.*.
  • store() / load(), tool search, and parity with Pi's lower-case names.
  • Re-measure the task savings at five or more runs per side once the rate window allows (issue 0005).

Credits

This project is not affiliated with Anthropic or Earendil.

License

MIT © 2026 Evaldo Klock

Source 9 files
hooks/register.ts 109 lines
1import { atom, update } from 'claude-code'
2import type { Register, ToolInfo } from 'claude-code'
3
4import { CodemodeBridge } from './bridge.ts'
5import { BASH_NOTE, GUIDELINE, codeDescription, describeCodemode } from './describe.ts'
6import { exposureSync, registerExposure, withoutHidden } from './expose.ts'
7import { readExposure } from './exposure.ts'
8import { registerRender } from './render.tsx'
9import { CODEMODE_TOOL_ID } from '../shared/protocol.ts'
10
11// The state scan reads the reference from this file, so render.tsx spells its
12// own; an invariant spec fails when the two differ.
13const RUNS = atom({ plugin: 'codemode', key: 'runs' } as const, [])
14
15const TOOL_NAME = 'codemode'
16const SCRIPT_TIMEOUT_MS = 120_000
17
18const TOOL_ID = CODEMODE_TOOL_ID
19
20/** The `code` property's description carries the tool sections, which the engine sends whole. */
21const inputSchemaWith = (codeText: string) => ({
22  type: 'object',
23  properties: { code: { type: 'string', description: codeText } },
24  required: ['code'],
25})
26
27type ListTools = () => Promise<ToolInfo[]>
28
29/** The `code` description for the tools connected now; `undefined` when the list cannot be read. */
30async function readCodeText(list: ListTools): Promise<string | undefined> {
31  try {
32    const tools = await list()
33    const mcp = tools.filter(tool => tool.mcp && tool.name !== TOOL_ID)
34    return codeDescription(mcp.map(({ name, description }) => ({ name, description })))
35  } catch {
36    return undefined
37  }
38}
39
40export const register: Register = (on, options) => {
41  // Read first so a bad setting fails the load.
42  const exposure = readExposure(options)
43  registerRender(on, options)
44  registerExposure(on, exposure)
45  const syncExposure = exposureSync(exposure)
46  const visibleTools = async (list: ListTools): Promise<ToolInfo[]> => withoutHidden(exposure, await list())
47  // Lost on a hot reload, which costs one more registration of the same text.
48  let registeredText: string | undefined
49
50  on('session.start', async ($, e, next) => {
51    const started = await next(e)
52    const codeText = (await readCodeText(() => visibleTools(() => $.tool.list()))) ?? codeDescription()
53    await $.tool.register({ name: TOOL_NAME, description: describeCodemode(), inputSchema: inputSchemaWith(codeText) })
54    registeredText = codeText
55    return started
56  })
57
58  // MCP servers may connect or change after the session starts; the prompt cache
59  // is spent only when the rendered sections differ from the last registered.
60  on('turn.start', async ($, e, next) => {
61    await syncExposure({ tool: { list: () => $.tool.list() }, ui: { invalidate: event => $.ui.invalidate(event) } })
62    const codeText = await readCodeText(() => visibleTools(() => $.tool.list()))
63    if (codeText !== undefined && codeText !== registeredText) {
64      await $.tool.register({ name: TOOL_NAME, description: describeCodemode(), inputSchema: inputSchemaWith(codeText) })
65      registeredText = codeText
66    }
67    return next(e)
68  }).catch((_$, e, next) => next(e))
69
70  on('tool.describe', { tool: TOOL_ID }, () => ({ description: describeCodemode(), isDeferred: false })).catch(
71    (_$, e, next) => next(e),
72  )
73
74  on('tool.describe', { tool: 'Bash' }, async (_$, e, next) => {
75    const answer = await next(e)
76    if (answer.description.endsWith(BASH_NOTE)) return answer
77    return { ...answer, description: `${answer.description}\n\n${BASH_NOTE}` }
78  }).catch((_$, e, next) => next(e))
79
80  on('prompt.compose', async (_$, e, next) => {
81    const { sections } = await next(e)
82    if (!e.tools.includes(TOOL_ID)) return { sections }
83    return { sections: [...sections, { id: 'codemode:guideline', text: GUIDELINE, scope: 'session' as const }] }
84  }).catch((_$, e, next) => next(e))
85
86  on('tool.call', { tool: TOOL_ID }, async ($, e) => {
87    if (typeof e.code !== 'string') return { deny: 'codemode needs a `code` string.' }
88    const bridge = new CodemodeBridge(
89      {
90        pluginRoot: $.plugin.root,
91        spawn: request => $.process.spawn(request),
92        callTool: input => $.tool.call(input),
93        listTools: () => visibleTools(() => $.tool.list()),
94        post: (url, init) => $.http.fetch(url, init),
95        publish: async change => {
96          await update($, RUNS, change)
97        },
98        now: () => $.clock.now(),
99        sleep: ms => $.clock.sleep(ms),
100      },
101      SCRIPT_TIMEOUT_MS,
102    )
103    const outcome = await bridge.run(e.code, e.tool_use_id)
104    return outcome.ok ? { result: outcome.output } : { deny: outcome.error }
105  }).catch((_$, _e, next) => ({
106    deny: `codemode failed unexpectedly: ${next.error.message ?? next.error.kind}`,
107  }))
108}
109
hooks/bridge.ts 440 lines
1import type {
2  HookStream,
3  HttpInit,
4  HttpResponse,
5  ProcessSpawnChunk,
6  ProcessSpawnRequest,
7  ProcessSpawnResult,
8  ToolCallArgs,
9  ToolCallResult,
10  ToolInfo,
11} from 'claude-code'
12
13import { ANSWER_PATH, CODEMODE_TOOL_ID, isExposedTool, parseChildMessage } from '../shared/protocol.ts'
14import type { CallAnswer, ChildMessage, McpTool, RunRequest } from '../shared/protocol.ts'
15import type { CodemodeCall, CodemodeCallState, CodemodeRun } from '../types/index.d.ts'
16
17/**
18 * What the bridge needs of the engine. The hooks loader refuses a module that
19 * passes `$` around, so the hook builds these closures where it spells `$`.
20 */
21export type BridgeHost = {
22  pluginRoot: string
23  spawn: (request: ProcessSpawnRequest) => HookStream<ProcessSpawnChunk, ProcessSpawnResult>
24  callTool: (input: ToolCallArgs) => Promise<ToolCallResult>
25  /** The tools the model has now, built-in and MCP alike. */
26  listTools: () => Promise<ToolInfo[]>
27  post: (url: string, init: HttpInit) => Promise<HttpResponse>
28  /** Applies a change to the runs the transcript draws from. */
29  publish: (change: (runs: CodemodeRun[]) => CodemodeRun[]) => Promise<void>
30  now: () => Promise<number>
31  /** Resolves after `ms` milliseconds; the hooks loader gives a module no timer of its own. */
32  sleep: (ms: number) => Promise<void>
33}
34
35/** How long a failed run waits for the child's closing line, so a script error is not lost to a late-answer failure. */
36export const CLOSING_GRACE_MS = 2000
37
38/** Runs kept in `$.state`: the transcript rarely draws more than the latest few. */
39export const RUN_LIMIT = 20
40/** Calls kept per run; a script looping over files would otherwise grow one value without end. */
41export const CALL_LIMIT = 100
42const LABEL_CHARS = 80
43const REASON_CHARS = 120
44const ARGS_CHARS = 80
45
46type Settled = {
47  answer: CallAnswer
48  state: Exclude<CodemodeCallState, 'running'>
49  reason?: string
50  /** Ledger only: the call may have taken effect though no answer says so; the transcript state stays as it is. */
51  unknown?: true
52  /** Ledger only: the tool held the input read-only, so redoing the call is safe; set on done calls. */
53  readOnly?: true
54}
55
56// Built from strings so the source holds no raw control character. CSI and OSC
57// sequences go whole; a lone ESC with its next character goes after them.
58const ESCAPE_SEQUENCES = new RegExp(
59  ['\\u001b\\[[0-9;?]*[ -/]*[@-~]', '\\u001b\\][^\\u0007\\u001b]*(?:\\u0007|\\u001b\\\\)?', '\\u001b.?'].join('|'),
60  'g',
61)
62const CONTROL_CHARACTERS = /[\u0000-\u001f\u007f]/g
63
64/** One printable line: what the state keeps must not carry a terminal's control codes. */
65function firstLine(text: string, limit: number): string {
66  const first = text.split(/[\r\n]/)[0] ?? ''
67  const line = first.replace(ESCAPE_SEQUENCES, '').replaceAll('\t', ' ').replace(CONTROL_CHARACTERS, '').trim()
68  return line.length > limit ? `${line.slice(0, limit - 1)}…` : line
69}
70
71function labelOf(input: Record<string, unknown>): string {
72  const target = input.file_path ?? input.command
73  return typeof target === 'string' ? firstLine(target, LABEL_CHARS) : ''
74}
75
76function mapRun(runs: CodemodeRun[], id: string, change: (run: CodemodeRun) => CodemodeRun): CodemodeRun[] {
77  return runs.map(run => (run.id === id ? change(run) : run))
78}
79
80function withCall(run: CodemodeRun, call: CodemodeCall): CodemodeRun {
81  const calls = [...run.calls, call]
82  const dropped = Math.max(0, calls.length - CALL_LIMIT)
83  return { ...run, calls: calls.slice(dropped), omitted: run.omitted + dropped }
84}
85
86/**
87 * Publishes one run's progress for the transcript to draw. Drawing is a side
88 * view: a failed publish never reaches the script or the model.
89 */
90class RunTracker {
91  private readonly host: BridgeHost
92  private readonly runId: string
93
94  constructor(host: BridgeHost, runId: string) {
95    this.host = host
96    this.runId = runId
97  }
98
99  begin(code: string): Promise<void> {
100    return this.safely(async () => {
101      const scriptWidth = Math.max(...code.split('\n').map(line => line.length))
102      const run: CodemodeRun = {
103        id: this.runId,
104        startedAt: await this.host.now(),
105        scriptWidth,
106        calls: [],
107        omitted: 0,
108      }
109      await this.host.publish(runs => [...runs, run].slice(-RUN_LIMIT))
110    })
111  }
112
113  startCall(call: CallMessage): Promise<void> {
114    return this.safely(async () => {
115      const entry: CodemodeCall = {
116        id: call.id,
117        tool: call.tool,
118        label: labelOf(call.input),
119        state: 'running',
120        startedAt: await this.host.now(),
121      }
122      await this.host.publish(runs => mapRun(runs, this.runId, run => withCall(run, entry)))
123    })
124  }
125
126  settleCall(id: number, settled: Settled): Promise<void> {
127    return this.safely(async () => {
128      const endedAt = await this.host.now()
129      const reason = settled.reason === undefined ? undefined : firstLine(settled.reason, REASON_CHARS)
130      const patch = { state: settled.state, endedAt, ...(reason === undefined ? {} : { reason }) }
131      await this.host.publish(runs =>
132        mapRun(runs, this.runId, run => ({
133          ...run,
134          calls: run.calls.map(call => (call.id === id ? { ...call, ...patch } : call)),
135        })),
136      )
137    })
138  }
139
140  finish(): Promise<void> {
141    return this.safely(async () => {
142      const endedAt = await this.host.now()
143      await this.host.publish(runs => mapRun(runs, this.runId, run => ({ ...run, endedAt })))
144    })
145  }
146
147  private async safely(work: () => Promise<void>): Promise<void> {
148    try {
149      await work()
150    } catch {
151      // The drawing is optional; the script's answers do not depend on it.
152    }
153  }
154}
155
156type LedgerEntry = {
157  id: number
158  tool: string
159  args: string
160  state: CodemodeCallState
161  detail: string
162  unknown: boolean
163  readOnly: boolean
164}
165
166const LEDGER_HEADING = 'Nested calls before the failure:'
167
168function detailOf(settled: Settled): string {
169  const text = settled.answer.ok ? settled.answer.text : (settled.reason ?? '')
170  return firstLine(text, REASON_CHARS)
171}
172
173const MAY_HAVE_RUN = '(it may have taken effect; check before redoing it)'
174
175function renderEntry(entry: LedgerEntry): string {
176  const unknown = entry.unknown || entry.state === 'running'
177  const detail = entry.state === 'running' ? 'no answer' : entry.detail
178  const marked = entry.readOnly && entry.state === 'done' && !unknown
179  const state = unknown ? 'unknown' : marked ? 'done (read-only)' : entry.state
180  const text = detail === '' ? state : `${state}: ${detail}`
181  return `#${entry.id} ${entry.tool} ${entry.args} — ${unknown ? `${text} ${MAY_HAVE_RUN}` : text}`
182}
183
184function renderOmitted(omitted: number): string[] {
185  if (omitted === 0) return []
186  return [`(${omitted} earlier ${omitted === 1 ? 'call' : 'calls'} left out)`]
187}
188
189/**
190 * What the model reads after a failure: the nested calls that already ran.
191 * It lives in the run state, apart from the published progress, which may fail.
192 */
193export class CallLedger {
194  private entries: LedgerEntry[] = []
195  private omitted = 0
196
197  begin(call: CallMessage): void {
198    const args = firstLine(JSON.stringify(call.input), ARGS_CHARS)
199    this.entries.push({ id: call.id, tool: call.tool, args, state: 'running', detail: '', unknown: false, readOnly: false })
200    const dropped = Math.max(0, this.entries.length - CALL_LIMIT)
201    this.entries = this.entries.slice(dropped)
202    this.omitted += dropped
203  }
204
205  settle(id: number, settled: Settled): void {
206    this.entries = this.entries.map(entry =>
207      entry.id === id ? { ...entry, state: settled.state, detail: detailOf(settled), unknown: settled.unknown === true, readOnly: settled.readOnly === true } : entry,
208    )
209  }
210
211  /** The section to append to a failure's text; empty when no call was made. */
212  render(): string {
213    if (this.entries.length === 0) return ''
214    return [LEDGER_HEADING, ...renderOmitted(this.omitted), ...this.entries.map(renderEntry)].join('\n')
215  }
216}
217
218export type CodemodeOutcome = { ok: true; output: string } | { ok: false; error: string }
219
220type CallMessage = Extract<ChildMessage, { type: 'call' }>
221type DoneMessage = Extract<ChildMessage, { type: 'done' }>
222
223type RunState = {
224  socketPath: string | undefined
225  closing: DoneMessage | undefined
226  problem: string | undefined
227  stderr: string
228  /** The MCP tool names this run's script may call, besides the built-ins. */
229  mcpNames: ReadonlySet<string>
230  answers: Promise<void>[]
231  /** Settles when a problem is recorded, so a read blocked on the child can stop. */
232  aborted: Promise<void>
233  abort: () => void
234  tracker: RunTracker
235  ledger: CallLedger
236}
237
238function newRunState(tracker: RunTracker, mcpNames: ReadonlySet<string>): RunState {
239  let abort = (): void => {}
240  const aborted = new Promise<void>(resolve => {
241    abort = resolve
242  })
243  return {
244    socketPath: undefined,
245    closing: undefined,
246    problem: undefined,
247    stderr: '',
248    mcpNames,
249    answers: [],
250    aborted,
251    abort,
252    tracker,
253    ledger: new CallLedger(),
254  }
255}
256
257function recordProblem(state: RunState, problem: string): void {
258  state.problem ??= problem
259  state.abort()
260}
261
262const STDERR_TAIL_CHARS = 2000
263const BAD_LINE_CHARS = 200
264
265/** Keys the engine reserves on a call; a script must never set them. */
266const RESERVED_KEYS = ['tool', 'tool_use_id', 'consent', 'agentId']
267
268/** Cuts a byte stream into whole lines; a line may span pieces or share one. */
269class LineSplitter {
270  private pending = ''
271
272  push(text: string): string[] {
273    const pieces = (this.pending + text).split('\n')
274    this.pending = pieces.pop() ?? ''
275    return pieces.filter(piece => piece.trim() !== '')
276  }
277}
278
279/**
280 * Runs one codemode script in a child process and serves the child's nested
281 * tool calls through `$.tool.call`, so each runs under the session's
282 * permission check and hooks.
283 */
284export class CodemodeBridge {
285  private readonly host: BridgeHost
286  private readonly timeoutMs: number
287
288  constructor(host: BridgeHost, timeoutMs: number) {
289    this.host = host
290    this.timeoutMs = timeoutMs
291  }
292
293  /** `runId` is the codemode call's tool_use_id, the key its transcript row draws from. */
294  async run(code: string, runId: string): Promise<CodemodeOutcome> {
295    const mcpTools = await this.connectedMcpTools()
296    const request: RunRequest = { code, timeoutMs: this.timeoutMs, mcpTools }
297    const tracker = new RunTracker(this.host, runId)
298    await tracker.begin(code)
299    const state = newRunState(tracker, new Set(mcpTools.map(tool => tool.name)))
300    const exit = await this.readChild(JSON.stringify(request), state)
301    await Promise.allSettled(state.answers)
302    await tracker.finish()
303    return this.outcome(state, exit)
304  }
305
306  /** A list that cannot be read leaves the script with the built-ins; the model is told nothing new. */
307  private async connectedMcpTools(): Promise<McpTool[]> {
308    try {
309      const listed = await this.host.listTools()
310      return listed
311        .filter(tool => tool.mcp && tool.name !== CODEMODE_TOOL_ID)
312        .map(({ name, description }) => ({ name, description }))
313    } catch {
314      return []
315    }
316  }
317
318  private async readChild(input: string, state: RunState): Promise<string> {
319    const argv = ['node', `${this.host.pluginRoot}/child/main.ts`]
320    const stream = this.host.spawn({ argv, input })
321    const lines = new LineSplitter()
322    try {
323      while (state.closing === undefined || state.problem === undefined) {
324        const next = await Promise.race([stream.next(), state.aborted])
325        if (next === undefined) break
326        if (next.done) return this.describeExit(next.value.code, next.value.signal)
327        this.takeChunk(next.value, lines, state)
328      }
329      // Not awaited: a return() waits behind a read still pending on the child.
330      stream.return({ code: null, signal: null }).catch(() => undefined)
331      return 'killed by the mod'
332    } catch (error) {
333      recordProblem(state, `the codemode child could not run: ${errorMessage(error)}`)
334      return 'did not start'
335    }
336  }
337
338  private takeChunk(chunk: ProcessSpawnChunk, lines: LineSplitter, state: RunState): void {
339    if (chunk.stream === 'stderr') state.stderr = (state.stderr + chunk.text).slice(-STDERR_TAIL_CHARS)
340    else for (const line of lines.push(chunk.text)) this.handleLine(line, state)
341  }
342
343  private describeExit(code: number | null, signal: string | null): string {
344    if (code !== null) return `exited with code ${code}`
345    return `was stopped by signal ${signal ?? 'unknown'}`
346  }
347
348  private handleLine(line: string, state: RunState): void {
349    const message = parseChildMessage(line)
350    if (message === undefined) {
351      recordProblem(state, `the codemode child sent a malformed line: ${line.slice(0, BAD_LINE_CHARS)}`)
352    } else if (message.type === 'listening') {
353      state.socketPath = message.socketPath
354    } else if (message.type === 'done') {
355      state.closing = message
356    } else if (state.socketPath === undefined) {
357      recordProblem(state, 'the codemode child asked for a tool before it announced its socket')
358    } else {
359      state.answers.push(this.serve(message, state.socketPath, state))
360    }
361  }
362
363  private async serve(call: CallMessage, socketPath: string, state: RunState): Promise<void> {
364    state.ledger.begin(call)
365    await state.tracker.startCall(call)
366    const settled = await this.execute(call, state.mcpNames)
367    state.ledger.settle(call.id, settled)
368    await state.tracker.settleCall(call.id, settled)
369    try {
370      await this.host.post(`http://bridge${ANSWER_PATH}`, {
371        method: 'POST',
372        body: JSON.stringify(settled.answer),
373        socketPath,
374      })
375    } catch (error) {
376      // The child may be ending on a script error: the read goes on for its closing line, but only for the grace.
377      state.problem ??= `the answer to a nested call could not reach the child: ${errorMessage(error)}`
378      void this.host.sleep(CLOSING_GRACE_MS).then(state.abort, state.abort)
379    }
380  }
381
382  private async execute(call: CallMessage, mcpNames: ReadonlySet<string>): Promise<Settled> {
383    const refused = (reason: string): Settled => ({
384      answer: { id: call.id, ok: false, error: reason },
385      state: 'denied',
386      reason,
387    })
388    const failed = (reason: string): Settled => ({
389      answer: { id: call.id, ok: false, error: reason },
390      state: 'failed',
391      reason,
392    })
393    if (!isExposedTool(call.tool) && !mcpNames.has(call.tool)) return refused(`tool ${call.tool} is not available to codemode scripts`)
394    const input = Object.fromEntries(
395      Object.entries(call.input).filter(([key]) => !RESERVED_KEYS.includes(key)),
396    )
397    try {
398      // The tool's own schema validates the arguments; the script chose them.
399      const result = await this.host.callTool({ ...input, tool: call.tool } as ToolCallArgs)
400      if (result.deny !== undefined) return refused(result.deny)
401      if (result.isError === true) return failed(result.text ?? `${call.tool} failed`)
402      const answer = { id: call.id, ok: true as const, text: result.text ?? '' }
403      return result.isReadOnly === true ? { answer, state: 'done', readOnly: true } : { answer, state: 'done' }
404    } catch (error) {
405      return thrownOutcome(call.id, error)
406    }
407  }
408
409  private outcome(state: RunState, exit: string): CodemodeOutcome {
410    if (state.problem === undefined && state.closing?.ok === true) return { ok: true, output: state.closing.output }
411    return { ok: false, error: withLedger(this.failure(state, exit), state.ledger) }
412  }
413
414  private failure(state: RunState, exit: string): string {
415    if (state.closing?.ok === false) {
416      const printed = state.closing.output === '' ? '' : `\n\nOutput before the failure:\n${state.closing.output}`
417      // A late answer that cannot reach the exited child is a symptom of the script's end, so it follows the cause.
418      const late = state.problem === undefined ? '' : `\n\n${state.problem}`
419      return `${state.closing.error}${printed}${late}`
420    }
421    if (state.problem !== undefined) return state.problem
422    const stderr = state.stderr.trim() === '' ? '' : `\n${state.stderr.trim()}`
423    return `the codemode child ${exit} without a closing line${stderr}`
424  }
425}
426
427function withLedger(text: string, ledger: CallLedger): string {
428  const section = ledger.render()
429  return section === '' ? text : `${text}\n\n${section}`
430}
431
432function errorMessage(error: unknown): string {
433  return error instanceof Error ? error.message : String(error)
434}
435
436export function thrownOutcome(id: number, error: unknown): Settled {
437  const reason = errorMessage(error)
438  return { answer: { id, ok: false, error: reason }, state: 'failed', reason, unknown: true }
439}
440
hooks/describe.ts 203 lines
1import { EXPOSED_TOOLS, TOOL_SPECS } from '../shared/protocol.ts'
2import type { McpTool, ToolArg, ToolSpec } from '../shared/protocol.ts'
3
4/** The line the system prompt carries so the model reaches for codemode unprompted. */
5export const GUIDELINE =
6  'Use codemode to batch independent tool calls (Promise.allSettled), chain them, or filter large output, instead of many separate calls.'
7
8/** The note Bash's description ends with, so data-processing scripts go to codemode instead of inline python or node. */
9export const BASH_NOTE =
10  'To process data, batch tool calls or filter large output with a script, use the codemode tool (JavaScript calling `tools.<name>(args)`) instead of inline python or node in Bash. Keep Bash for running commands.'
11
12/** How a script calls one tool, and what the call resolves to. */
13export type ToolDoc = {
14  name: string
15  summary: string
16  args: string
17  resolves: string
18}
19
20function describeArg([name, arg]: [string, ToolArg]): string {
21  return arg.note === undefined ? `\`${name}\`` : `\`${name}\` (${arg.note})`
22}
23
24function describeArgs(args: Record<string, ToolArg>): string {
25  const entries = Object.entries(args)
26  const required = entries.filter(([, arg]) => arg.isRequired).map(describeArg)
27  const optional = entries.filter(([, arg]) => !arg.isRequired).map(describeArg)
28  const optionalText = optional.length === 0 ? [] : [`optional ${optional.join(' and ')}`]
29  return [...required, ...optionalText].join(', ')
30}
31
32/** One doc per tool in `names`, read from `specs`, the source the child declares its tools from. */
33export function toolDocs(
34  specs: Record<string, ToolSpec> = TOOL_SPECS,
35  names: readonly string[] = EXPOSED_TOOLS,
36): ToolDoc[] {
37  return names.flatMap(name => {
38    const spec = specs[name]
39    return spec === undefined ? [] : [{ name, summary: spec.summary, args: describeArgs(spec.args), resolves: spec.resolves }]
40  })
41}
42
43export const EXPOSED_DOCS: readonly ToolDoc[] = toolDocs()
44
45const INTRO = [
46  'Runs JavaScript that calls other tools. The input is raw JavaScript (not JSON, no code fence), run as an async function body in a sandbox: top-level `await` works. No Node, file system, network, or timers.',
47  '- `await tools.<name>({ ...args })` resolves to the tool\'s text and rejects with an Error when the call fails or a permission rule refuses it; catch it to continue.',
48  '- Only what the script prints comes back, so filter and combine results in the script.',
49  '- A failed run lists the nested calls that already ran, so a retry redoes only what did not.',
50  '- A call marked unknown may have taken effect: read the current state before redoing it.',
51  '- A call marked read-only is safe to redo.',
52  "- A script with writes prints each step as it completes, catches each item's failure apart, and passes an idempotency key when a tool takes one, derived from the data, never at random.",
53  '- Prefer writes that are safe to repeat: overwrite, `mkdir -p`, upsert, check then act.',
54  '- To search, run `rg` or `git grep` through Bash, print only the matches, then read only the files that matter.',
55  '- Keep a handle a tool returns in a variable and pass it to the next call; never print it.',
56].join('\n')
57
58const GLOBALS = [
59  'Globals:',
60  '- `text(value)` and `console.log(...)` add output; non-strings are JSON-stringified. A top-level `return` ends the script, and its value is not sent back.',
61  '- `exit()` ends the script successfully, keeping its output.',
62  '- `ALL_TOOLS` lists `{ name, description }` for each tool a script can call.',
63  '- Connected MCP tools are callable too, as `tools.<name>(args)` by their full `mcp__server__tool` name, and listed in `ALL_TOOLS`.',
64  '- Each nested tool has a section in the description of the `code` parameter; one with no section there is still callable, and `ALL_TOOLS` is how to find it.',
65].join('\n')
66
67/** The first line of the `code` property's description. */
68const CODE_LEAD = 'The script to run.'
69
70function toolSection(doc: ToolDoc): string {
71  return [
72    `### \`${doc.name}\``,
73    `${doc.summary} \`tools.${doc.name}(args)\` takes ${doc.args}, and resolves to ${doc.resolves}.`,
74  ].join('\n')
75}
76
77/** The most characters of a tool's description the build sends to the model. */
78export const DESCRIPTION_CAP = 2048
79
80/** What the sections may cost together, in estimated tokens (characters divided by four). */
81export const SECTIONS_BUDGET = 3000
82const CHARS_PER_TOKEN = 4
83
84/** One nested tool's section; `server` is absent for the built-ins. */
85export type Section = {
86  name: string
87  server: string | undefined
88  text: string
89}
90
91/** A group's sections as the budget left them: `shown` in the group's order, out of `total`. */
92export type Group = {
93  server: string | undefined
94  shown: Section[]
95  total: number
96}
97
98const costOf = (section: Section): number => Math.ceil(section.text.length / CHARS_PER_TOKEN)
99
100/** The identifier a script uses for a tool: a character invalid in an identifier becomes `_`. */
101export function toIdentifier(name: string): string {
102  let identifier = ''
103  for (const char of name) {
104    const isValid = identifier === '' ? /^[A-Za-z_$]$/.test(char) : /^[A-Za-z0-9_$]$/.test(char)
105    identifier += isValid ? char : '_'
106  }
107  return identifier === '' ? '_' : identifier
108}
109
110/** The server of `mcp__server__tool`: the part between the first two `__`. */
111function serverOf(name: string): string {
112  return name.split('__')[1] ?? ''
113}
114
115export function builtinSection(doc: ToolDoc): Section {
116  return { name: doc.name, server: undefined, text: toolSection(doc) }
117}
118
119// A description is data, not markup: one line with no fence and no leading `#`
120// cannot end its section or open another.
121function inert(text: string): string {
122  return text.replace(/\s+/g, ' ').trim().replace(/`{3,}/g, "'''").replace(/^#/, '\\#')
123}
124
125/** An MCP tool's section: its heading, its own description, and how a script calls it. */
126export function mcpSection(tool: McpTool): Section {
127  const id = toIdentifier(tool.name)
128  const heading = id === tool.name ? `### \`${id}\`` : `### \`${id}\` (\`${inert(tool.name)}\`)`
129  const description = inert(tool.description)
130  const call = `\`tools.${id}(args)\` takes an open object of arguments, and resolves to the tool's text.`
131  const lines = [heading, ...(description === '' ? [] : [description]), call]
132  return { name: tool.name, server: serverOf(tool.name), text: lines.join('\n') }
133}
134
135function serversOf(sections: readonly Section[]): string[] {
136  const names = new Set(sections.flatMap(section => (section.server === undefined ? [] : [section.server])))
137  return [...names].sort((a, b) => a.localeCompare(b))
138}
139
140/**
141 * Picks the sections that fit `budget` tokens: in each round every group, the
142 * built-ins first and then the servers by name, places its cheapest remaining
143 * section; a group whose next one does not fit drops out while the others go on.
144 */
145export function selectSections(sections: readonly Section[], budget: number): Group[] {
146  const servers: (string | undefined)[] = [undefined, ...serversOf(sections)]
147  // Unlike Pi, which keeps the input order, a server's ties and shown order go by name,
148  // so the text depends on the set of tools and not on the order they arrive in;
149  // the built-ins are a fixed list and keep its order.
150  const inGroupOrder = (server: string | undefined, list: Section[]): Section[] =>
151    server === undefined ? list : list.sort((a, b) => a.name.localeCompare(b.name))
152  const groups = servers
153    .map(server => ({ server, all: inGroupOrder(server, sections.filter(section => section.server === server)) }))
154    .filter(group => group.all.length > 0)
155  const queues = groups.map(group => [...group.all].sort((a, b) => costOf(a) - costOf(b)))
156  const shown = new Set<Section>()
157  let remaining = budget
158  let active = queues
159  while (active.length > 0) {
160    active = active.filter(queue => {
161      const next = queue.shift()
162      if (next === undefined) return false
163      if (costOf(next) > remaining) return false
164      remaining -= costOf(next)
165      shown.add(next)
166      return queue.length > 0
167    })
168  }
169  return groups.map(group => ({
170    server: group.server,
171    shown: group.all.filter(section => shown.has(section)),
172    total: group.all.length,
173  }))
174}
175
176function serverHeading(group: Group): string[] {
177  if (group.server === undefined) return []
178  if (group.shown.length === group.total) return [`## ${group.server}`]
179  const listing = group.shown.length === 0 ? 'tools not listed' : 'some tools not listed'
180  return [`## ${group.server} (${listing})`]
181}
182
183/** The sections that fit the budget, under `Nested tools:`; empty when there is no tool at all. */
184export function renderSections(sections: readonly Section[], budget: number = SECTIONS_BUDGET): string {
185  if (sections.length === 0) return ''
186  const parts = selectSections(sections, budget).flatMap(group => [
187    ...serverHeading(group),
188    ...group.shown.map(section => section.text),
189  ])
190  return ['Nested tools:', ...parts].join('\n\n')
191}
192
193/** The `code` property's description: what the script is, then one section per nested tool. */
194export function codeDescription(mcpTools: readonly McpTool[] = [], docs: readonly ToolDoc[] = EXPOSED_DOCS): string {
195  const sections = renderSections([...docs.map(builtinSection), ...mcpTools.map(mcpSection)])
196  return sections === '' ? CODE_LEAD : `${CODE_LEAD}\n\n${sections}`
197}
198
199/** The tool's own description: the intro and the globals; the sections ride in the `code` property. */
200export function describeCodemode(): string {
201  return [INTRO, GLOBALS].join('\n\n')
202}
203
hooks/expose.ts 163 lines
1import type { InvalidatableEventName, On, ToolInfo } from 'claude-code'
2
3import { toScriptIdentifier } from '../shared/identifier.ts'
4import { modeOf } from './exposure.ts'
5import type { Exposure } from './exposure.ts'
6
7const MCP_PREFIX = 'mcp__'
8/** The part of the hook context this file reads. */
9type Session = {
10  readonly tool: { readonly list: () => Promise<ToolInfo[]> }
11  readonly ui: { readonly invalidate: (event: InvalidatableEventName) => void }
12}
13
14const FUNCTION_OPEN = '<function>'
15
16const serverOf = (tool: string): string => {
17  const rest = tool.slice(MCP_PREFIX.length)
18  return rest.slice(0, Math.max(rest.indexOf('__'), 0))
19}
20
21const callNote = (tool: string): string =>
22  `Call this tool inside the codemode tool as \`tools.${tool}(args)\`.`
23
24/** The tool a `<function>` block line declares, when the line is one and its JSON parses. */
25function declaredTool(line: string): string | undefined {
26  const trimmed = line.trim()
27  if (!trimmed.startsWith(FUNCTION_OPEN)) return undefined
28  const body = trimmed.slice(FUNCTION_OPEN.length).replace(/<\/function>$/, '')
29  try {
30    const parsed: unknown = JSON.parse(body)
31    const name = (parsed as { name?: unknown } | null)?.name
32    return typeof name === 'string' ? name : undefined
33  } catch {
34    return undefined
35  }
36}
37
38/** The text without the deferred-list lines and the `<function>` lines of the hidden tools; the rest is byte-identical. */
39function withoutTools(text: string, hidden: readonly string[]): string {
40  if (!hidden.some(name => text.includes(name))) return text
41  const gone = new Set(hidden)
42  return text
43    .split('\n')
44    .filter(line => !gone.has(line.trim()) && !gone.has(declaredTool(line) ?? ''))
45    .join('\n')
46}
47
48/** One line naming the tools, a server whose every listed tool is in the group by its wildcard. */
49function instructionsLine(group: readonly string[], connected: readonly string[]): string {
50  const names: string[] = []
51  const servers = new Set(group.map(serverOf))
52  for (const server of servers) {
53    const inGroup = group.filter(tool => serverOf(tool) === server)
54    const listed = connected.filter(tool => serverOf(tool) === server)
55    if (inGroup.length === listed.length) names.push(`${toScriptIdentifier(`${MCP_PREFIX}${server}__`)}*`)
56    else names.push(...inGroup.map(toScriptIdentifier))
57  }
58  return (
59    `The tools ${names.join(', ')} run inside the codemode tool as \`tools.<name>(args)\`; ` +
60    `where a server's instructions above name one of its tools, call it from a codemode script.`
61  )
62}
63
64const connected = async ($: Session): Promise<string[] | undefined> => {
65  try {
66    const tools = await $.tool.list()
67    return tools.filter(tool => tool.mcp).map(tool => tool.name)
68  } catch {
69    return undefined
70  }
71}
72
73const inCodemode = (exposure: Exposure, names: readonly string[]): string[] =>
74  names.filter(name => modeOf(exposure, name) === 'codemode')
75
76type Describe = { readonly description: string; readonly isDeferred?: boolean }
77
78/** The describe answer a tool's mode asks for; a mode the plugin does not set keeps the engine's answer. */
79function describeAnswer<A extends Describe>(exposure: Exposure, tool: string, answer: A): A {
80  const mode = modeOf(exposure, tool)
81  if (mode === 'deferred' || mode === 'hidden') return { ...answer, isDeferred: true }
82  if (mode === 'direct') return { ...answer, isDeferred: false }
83  if (mode !== 'codemode') return answer
84  return { ...answer, description: `${callNote(tool)}\n\n${answer.description}`, isDeferred: true }
85}
86
87const inHidden = (exposure: Exposure, names: readonly string[]): string[] =>
88  names.filter(name => modeOf(exposure, name) === 'hidden')
89
90/** The tools with the hidden ones left out; what the codemode description and the scripts may see. */
91export function withoutHidden<T extends { readonly name: string }>(exposure: Exposure, tools: readonly T[]): T[] {
92  return tools.filter(tool => modeOf(exposure, tool.name) !== 'hidden')
93}
94
95/**
96 * The attachment text without the codemode-mode and hidden tools; the instructions line
97 * is added on an instructions delta and names the codemode-mode tools only.
98 */
99function attachmentText(
100  text: string,
101  type: string,
102  groups: { readonly codemode: readonly string[]; readonly hidden: readonly string[] },
103  names: readonly string[],
104): string {
105  const kept = withoutTools(text, [...groups.codemode, ...groups.hidden])
106  if (type !== 'mcp_instructions_delta' || groups.codemode.length === 0) return kept
107  return `${kept}\n\n${instructionsLine(groups.codemode, names)}`
108}
109
110const hiddenReason = (tool: string): string =>
111  `${tool} is hidden by the codemode plugin's settings (mcpHidden); no call to it runs.`
112
113/** The refusal when the check itself failed: every tool is denied, and only a hidden one is told why. */
114const failedCheckVerdict = (exposure: Exposure, tool: string) => ({
115  decision: 'deny' as const,
116  reason:
117    modeOf(exposure, tool) === 'hidden'
118      ? hiddenReason(tool)
119      : `The codemode plugin's permission check failed, so the call to ${tool} is refused.`,
120})
121
122/** Hides the codemode-mode MCP tools from the model outside the codemode tool, and refuses the hidden ones everywhere. */
123export function registerExposure(on: On, exposure: Exposure): void {
124  on('tool.describe', async (_$, e, next) => describeAnswer(exposure, e.tool, await next(e))).catch((_$, e, next) =>
125    next(e),
126  )
127
128  on('tool.check', (_$, e, next) =>
129    modeOf(exposure, e.tool) === 'hidden' ? { decision: 'deny' as const, reason: hiddenReason(e.tool) } : next(e),
130  ).catch((_$, e) => failedCheckVerdict(exposure, e.tool))
131
132  on('prompt.attachment', async ($, e, next) => {
133    const answer = await next(e)
134    if (answer.text === null) return answer
135    const names = await connected($)
136    if (names === undefined) return answer
137    const groups = { codemode: inCodemode(exposure, names), hidden: inHidden(exposure, names) }
138    if (groups.codemode.length + groups.hidden.length === 0) return answer
139    return { ...answer, text: attachmentText(answer.text, e.type, groups, names) }
140  }).catch((_$, e, next) => next(e))
141}
142
143/** One string for the two sets; a blank line separates them, which no tool name holds. */
144const groupsKey = (codemode: readonly string[], hidden: readonly string[]): string =>
145  `${[...codemode].sort().join('\n')}\n\n${[...hidden].sort().join('\n')}`
146
147/**
148 * The turn-start check: the engine caches the two answers `registerExposure` rewrites for the session,
149 * so a codemode-mode tool that connects later needs them asked again.
150 */
151export function exposureSync(exposure: Exposure): (session: Session) => Promise<void> {
152  let seen = groupsKey([], [])
153  return async session => {
154    const names = await connected(session)
155    if (names === undefined) return
156    const now = groupsKey(inCodemode(exposure, names), inHidden(exposure, names))
157    if (now === seen) return
158    seen = now
159    session.ui.invalidate('prompt.attachment')
160    session.ui.invalidate('tool.describe')
161  }
162}
163
hooks/exposure.ts 93 lines
1import { CODEMODE_TOOL_ID } from '../shared/protocol.ts'
2
3export type ExposureMode = 'codemode' | 'deferred' | 'direct' | 'hidden'
4
5// Structural, so the node specs load this file without the kit's type package.
6type PluginOptions = Readonly<Record<string, unknown>>
7
8type Entry = { text: string; mode: ExposureMode; pattern: RegExp | undefined }
9
10type Table = {
11  readonly exact: ReadonlyMap<string, ExposureMode>
12  readonly patterns: readonly Entry[]
13  readonly servers: ReadonlyMap<string, ExposureMode>
14}
15
16declare const brand: unique symbol
17
18/** Opaque: built by `readExposure`, read by `modeOf`; the table behind it stays in this file. */
19export type Exposure = { readonly [brand]: 'exposure' }
20
21const tables = new WeakMap<Exposure, Table>()
22
23// The order patterns are tried in: the first list to match wins.
24const LISTS = [
25  ['mcpHidden', 'hidden'],
26  ['mcpCodemode', 'codemode'],
27  ['mcpDeferred', 'deferred'],
28  ['mcpDirect', 'direct'],
29] as const satisfies readonly (readonly [string, ExposureMode])[]
30
31const MCP_PREFIX = 'mcp__'
32const SEPARATOR = '__'
33
34const patternOf = (text: string): RegExp =>
35  new RegExp(`^${text.split('*').map(part => part.replace(/[.+?^${}()|[\]\\]/g, '\\$&')).join('.*')}$`)
36
37// `/plugin configure` stores a multiple-string field as one comma-separated string; a hand-written array also arrives.
38const listOf = (options: PluginOptions, key: string): readonly string[] => {
39  const value = options[key]
40  if (value === undefined) return []
41  if (Array.isArray(value)) {
42    if (value.every((item): item is string => typeof item === 'string')) return value
43  } else if (typeof value === 'string') {
44    const pieces = value.split(',')
45    // A trailing comma is a typing slip, and an empty string splits to one empty piece.
46    if (pieces.at(-1)?.trim() === '') pieces.pop()
47    return pieces
48  }
49  throw new Error(`The ${key} setting takes a list of strings or one comma-separated string.`)
50}
51
52/** Reads the four lists; throws on a bad value, an empty entry or an entry named twice. */
53export function readExposure(options: PluginOptions): Exposure {
54  const seen = new Map<string, string>()
55  const exact = new Map<string, ExposureMode>()
56  const servers = new Map<string, ExposureMode>()
57  const patterns: Entry[] = []
58
59  for (const [key, mode] of LISTS) {
60    for (const written of listOf(options, key)) {
61      const trimmed = written.trim()
62      const text = trimmed.startsWith(MCP_PREFIX) ? trimmed.slice(MCP_PREFIX.length).trim() : trimmed
63      if (text === '') throw new Error(`The ${key} setting holds an empty entry.`)
64      const earlier = seen.get(text)
65      if (earlier !== undefined) {
66        const where = earlier === key ? `twice in ${key}` : `in ${earlier} and ${key}`
67        throw new Error(`The entry "${text}" is named ${where}; each entry takes one exposure mode.`)
68      }
69      seen.set(text, key)
70      if (text.includes('*')) patterns.push({ text, mode, pattern: patternOf(text) })
71      else if (text.includes(SEPARATOR)) exact.set(text, mode)
72      else servers.set(text, mode)
73    }
74  }
75  const handle = Object.freeze({}) as Exposure
76  tables.set(handle, { exact, patterns, servers })
77  return handle
78}
79
80/** The mode the settings give an MCP tool name; `undefined` leaves the tool as the host has it. */
81export function modeOf(exposure: Exposure, tool: string): ExposureMode | undefined {
82  if (!tool.startsWith(MCP_PREFIX) || tool === CODEMODE_TOOL_ID) return undefined
83  const table = tables.get(exposure)
84  if (table === undefined) return undefined
85  const name = tool.slice(MCP_PREFIX.length)
86  const exactMode = table.exact.get(name)
87  if (exactMode !== undefined) return exactMode
88  const hit = table.patterns.find(entry => entry.pattern?.test(name))
89  if (hit !== undefined) return hit.mode
90  const server = name.slice(0, Math.max(name.indexOf(SEPARATOR), 0))
91  return table.servers.get(server)
92}
93
hooks/render.tsx 294 lines
1import { atom, read } from 'claude-code'
2import type { Register, StateDollar } from 'claude-code'
3
4import type { CodemodeCall, CodemodeRun } from '../types/index.d.ts'
5
6// The state scan reads the reference from this file, so register.ts spells its
7// own; an invariant spec fails when the two differ.
8const RUNS = atom({ plugin: 'codemode', key: 'runs' } as const, [])
9
10const CODEMODE_TOOL = 'mcp__codemode__codemode'
11const TITLE = 'codemode · script.js'
12/** The columns the engine's line-number gutter takes: up to 999 lines plus a separator; its exact width is undocumented, so the engine's cut absorbs any miscount. */
13const GUTTER = 5
14
15const GLYPHS = { running: '…', done: '✓', denied: '✗', failed: '✗' } as const
16
17const LONG_PATH_CHARS = 30
18const KEPT_SEGMENTS = 2
19const LABEL_MAX = 30
20/** The widest a box's content grows without a measured surface, so one long script line never stretches the row. */
21const WIDTH_CAP = 100
22/** Border and padding on both sides of a box's content. */
23const FRAME = 4
24/** The narrowest content a measured surface gets; a box on a viewport under FRAME + this overflows it. */
25const MIN_CONTENT = 4
26/** The most output lines a result box draws; the model still receives the whole output. */
27const RESULT_LINES = 10
28const ZERO_WIDTH: [number, number][] = [
29  [0x300, 0x36f],
30  [0x200b, 0x200f],
31  [0xfe00, 0xfe0f],
32]
33const DOUBLE_WIDTH: [number, number][] = [
34  [0x1100, 0x115f],
35  [0x2e80, 0xa4cf],
36  [0xac00, 0xd7a3],
37  [0xf900, 0xfaff],
38  [0xfe30, 0xfe6f],
39  [0xff00, 0xff60],
40  [0xffe0, 0xffe6],
41  [0x1f300, 0x1f64f],
42  [0x1f900, 0x1f9ff],
43  [0x20000, 0x3fffd],
44]
45const GAP = '  '
46/** The title's width at its widest plausible line count; the result box cannot know the count, so both boxes floor here and stay equal. */
47const TITLE_FLOOR = columnsWide(`${TITLE} · 999 lines`)
48
49type Columns = { tool: number; label: number; tail: number }
50type Row = { call: CodemodeCall; label: string; tail: string }
51
52function scriptOf(input: unknown): string | undefined {
53  if (typeof input !== 'object' || input === null || !('code' in input)) return undefined
54  return typeof input.code === 'string' ? input.code : undefined
55}
56
57function duration(ms: number): string {
58  return ms < 1000 ? `${ms} ms` : `${(ms / 1000).toFixed(1)} s`
59}
60
61/** What the row's last column says: the duration, or the verdict of a call that did not run. */
62function tailOf(call: CodemodeCall): string {
63  if (call.state === 'running') return ''
64  if (call.state === 'done') return duration((call.endedAt ?? call.startedAt) - call.startedAt)
65  return call.state
66}
67
68function summaryOf(run: CodemodeRun): string {
69  const count = run.calls.length + run.omitted
70  const parts = [`${count} ${count === 1 ? 'call' : 'calls'}`]
71  if (run.endedAt !== undefined) parts.push(duration(run.endedAt - run.startedAt))
72  for (const state of ['denied', 'failed'] as const) {
73    const total = run.calls.filter(call => call.state === state).length
74    if (total > 0) parts.push(`${total} ${state}`)
75  }
76  return parts.join(' · ')
77}
78
79/** Shortens each long path in a label to its last segments; the state keeps the full label. */
80function shorten(label: string): string {
81  return label
82    .split(' ')
83    .map(word => {
84      const segments = word.split('/')
85      const isLongPath = segments.length > KEPT_SEGMENTS + 1 && word.length > LONG_PATH_CHARS
86      return isLongPath ? `…/${segments.slice(-KEPT_SEGMENTS).join('/')}` : word
87    })
88    .join(' ')
89}
90
91function clip(label: string): string {
92  return label.length > LABEL_MAX ? `${label.slice(0, LABEL_MAX - 1)}…` : label
93}
94
95function columnsOf(rows: Row[]): Columns {
96  const widest = (pick: (row: Row) => string): number => Math.max(0, ...rows.map(row => pick(row).length))
97  return { tool: widest(row => row.call.tool), label: widest(row => row.label), tail: widest(row => row.tail) }
98}
99
100function rowWidth(columns: Columns): number {
101  return 2 + columns.tool + 1 + columns.label + GAP.length + columns.tail
102}
103
104function rowsOf(run: CodemodeRun | undefined): Row[] {
105  return (run?.calls ?? []).map(call => ({ call, label: clip(shorten(call.label)), tail: tailOf(call) }))
106}
107
108/** The width a measured surface leaves a box's content, never under a few columns so the `…` cut stays legible; the transcript takes no margin, as the engine's own rules reach the same edge. */
109function surfaceWidth(surfaceColumns: number): number {
110  return Math.max(MIN_CONTENT, surfaceColumns - FRAME)
111}
112
113/** The one width both boxes of a call share: the surface's room when measured, else the script's and the rows'. */
114function contentWidth(longest: number, columns: Columns, surfaceColumns: number | undefined): number {
115  if (surfaceColumns !== undefined) return surfaceWidth(surfaceColumns)
116  return Math.min(WIDTH_CAP, Math.max(TITLE_FLOOR, longest + GUTTER, rowWidth(columns)))
117}
118
119/** The result box's width: the surface's room when measured, else the run's stored script width, else none. */
120function resultWidth(run: CodemodeRun | undefined, surfaceColumns: number | undefined): number | undefined {
121  if (surfaceColumns !== undefined) return surfaceWidth(surfaceColumns)
122  return run?.scriptWidth === undefined ? undefined : contentWidth(run.scriptWidth, columnsOf(rowsOf(run)), undefined)
123}
124
125function within(point: number, ranges: [number, number][]): boolean {
126  return ranges.some(([low, high]) => point >= low && point <= high)
127}
128
129/** A character's terminal columns from short range lists; no full Unicode width table, no grapheme clusters. */
130function charWidth(point: number): number {
131  if (within(point, ZERO_WIDTH)) return 0
132  return within(point, DOUBLE_WIDTH) ? 2 : 1
133}
134
135function columnsWide(text: string): number {
136  return [...text].reduce((total, char) => total + charWidth(char.codePointAt(0) ?? 0), 0)
137}
138
139/** The longest start of the line that fits `room` columns, never splitting a code point. */
140function fitting(line: string, room: number): string {
141  let used = 0
142  let kept = ''
143  for (const char of line) {
144    used += charWidth(char.codePointAt(0) ?? 0)
145    if (used > room) break
146    kept += char
147  }
148  return kept
149}
150
151/** Cuts every line wider than the box, in terminal columns, with `…`, so none wraps inside it. */
152function cutLines(text: string, width: number): string {
153  return text
154    .split('\n')
155    .map(line => (columnsWide(line) > width ? `${fitting(line, width - 1)}…` : line))
156    .join('\n')
157}
158
159/** The text to draw and how many lines it leaves out; one trailing newline is not a line, and a text within the cap is returned whole. */
160function capLines(text: string): { shown: string; hidden: number } {
161  const lines = text.split('\n')
162  const count = lines.length > 1 && lines[lines.length - 1] === '' ? lines.length - 1 : lines.length
163  if (count <= RESULT_LINES) return { shown: text, hidden: 0 }
164  return { shown: lines.slice(0, RESULT_LINES).join('\n'), hidden: count - RESULT_LINES }
165}
166
167/** The title of a script's row; lines are counted as `capLines` counts them. */
168function titleOf(script: string): string {
169  const lines = script.split('\n')
170  const count = lines.length > 1 && lines[lines.length - 1] === '' ? lines.length - 1 : lines.length
171  return `${TITLE} · ${count} ${count === 1 ? 'line' : 'lines'}`
172}
173
174function moreLine(hidden: number): string {
175  return `… ${hidden} more ${hidden === 1 ? 'line' : 'lines'}`
176}
177
178async function runOf($: StateDollar, id: string): Promise<CodemodeRun | undefined> {
179  return (await read($, RUNS)).find(run => run.id === id)
180}
181
182/**
183 * Draws a codemode call in the transcript as two bordered boxes, each as wide
184 * as its content: the script with its nested calls in columns, then the
185 * summary and the output. Every other row, and any codemode row it cannot
186 * read, stays the engine's own. A refused call's reason stays in the state and
187 * in the model's result; the row says only that it was denied or failed.
188 */
189export const registerRender: Register = on => {
190  on('ui.render', { component: 'ToolUse' }, async ($, e, next) => {
191    const script = e.props.tool === CODEMODE_TOOL ? scriptOf(e.props.input) : undefined
192    if (script === undefined) return next(e)
193    const run = await runOf($, e.requestId)
194    const { Box, Text, Code } = $.ui.resolve(e)
195    const rows = rowsOf(run)
196    const columns = columnsOf(rows)
197    const longest = Math.max(...script.split('\n').map(line => line.length))
198    const width = contentWidth(longest, columns, e.viewport?.columns)
199
200    return (
201      <Box
202        key="codemode-row"
203        flexDirection="column"
204        alignSelf="flex-start"
205        width={width + FRAME}
206        borderStyle="round"
207        borderDimColor
208        paddingX={1}
209      >
210        <Box key="title">
211          <Text dimColor>{cutLines(titleOf(script), width)}</Text>
212        </Box>
213        <Code source={script} language="javascript" startLine={1} wrap="truncate-end" />
214        {rows.length > 0 ? (
215          <Box key="divider">
216            <Text dimColor>{'─'.repeat(width)}</Text>
217          </Box>
218        ) : null}
219        {run !== undefined && run.omitted > 0 ? (
220          <Text dimColor>{`… ${run.omitted} earlier calls not shown`}</Text>
221        ) : null}
222        {rows.map(({ call, label, tail }) => (
223          <Box key={`call-${call.id}`}>
224            <Text>{`${GLYPHS[call.state]} ${call.tool.padEnd(columns.tool)} ${label.padEnd(columns.label)}`}</Text>
225            <Text
226              dimColor={call.state === 'running' || call.state === 'done'}
227              color={call.state === 'denied' || call.state === 'failed' ? 'error' : undefined}
228            >
229              {`${GAP}${tail.padStart(columns.tail)}`}
230            </Text>
231          </Box>
232        ))}
233      </Box>
234    )
235  }).catch(($, e, next) => next(e))
236
237  on('ui.render', { component: 'ToolResult' }, async ($, e, next) => {
238    const output = e.props.output
239    if (e.props.tool !== CODEMODE_TOOL || typeof output !== 'string') return next(e)
240    const { Box, Text } = $.ui.resolve(e)
241    const run = await runOf($, e.requestId)
242    // Without a viewport or the run's script width the box shrink-wraps its output.
243    const width = resultWidth(run, e.viewport?.columns)
244    const fit = (text: string): string => (width === undefined ? text : cutLines(text, width))
245    const boxWidth = width === undefined ? {} : { width: width + FRAME }
246    const { shown, hidden } = capLines(output)
247    const more = hidden > 0 ? <Box key="more"><Text dimColor>{fit(moreLine(hidden))}</Text></Box> : null
248
249    if (e.props.isErrored) {
250      return (
251        <Box
252          key="codemode-result"
253          flexDirection="column"
254          alignSelf="flex-start"
255          {...boxWidth}
256          borderStyle="round"
257          borderColor="error"
258          borderDimColor
259          paddingX={1}
260        >
261          <Box key="error">
262            <Text color="error">{fit(`✗ ${shown}`)}</Text>
263          </Box>
264          {more}
265        </Box>
266      )
267    }
268
269    return (
270      <Box
271        key="codemode-result"
272        flexDirection="column"
273        alignSelf="flex-start"
274        {...boxWidth}
275        borderStyle="round"
276        borderColor="success"
277        borderDimColor
278        paddingX={1}
279      >
280        <Box key="summary">
281          <Text bold color="success">
282            ✓
283          </Text>
284          <Text>{` ${run === undefined ? 'done' : summaryOf(run)}`}</Text>
285        </Box>
286        <Box key="output">
287          <Text>{fit(shown)}</Text>
288        </Box>
289        {more}
290      </Box>
291    )
292  }).catch(($, e, next) => next(e))
293}
294
shared/protocol.ts 244 lines
1/**
2 * The wire format between the mod and the codemode child. Both import this
3 * file, and it holds types and pure parsers only, so either side may load it.
4 */
5
6/** The tools a script may call; the child declares them and the mod refuses others. */
7export const EXPOSED_TOOLS = [
8  'Read',
9  'Bash',
10  'Write',
11  'Edit',
12  'ListMcpResourcesTool',
13  'ReadMcpResourceTool',
14  'ReadMcpResourceDirTool',
15] as const
16
17export type ExposedTool = (typeof EXPOSED_TOOLS)[number]
18
19/** One argument of an exposed tool: the child declares it, the description names it. */
20export type ToolArg = {
21  type: 'string' | 'number' | 'boolean'
22  isRequired: boolean
23  /** How the model-facing description words the argument. */
24  note?: string
25  /** The argument's `description` in the schema a script can inspect. */
26  sandboxNote?: string
27}
28
29/** What a script may pass an exposed tool and what the call resolves to. */
30export type ToolSpec = {
31  /** The tool's `description` in the sandbox's `ALL_TOOLS`. */
32  sandboxDescription: string
33  summary: string
34  resolves: string
35  args: Record<string, ToolArg>
36}
37
38export const TOOL_SPECS: Record<ExposedTool, ToolSpec> = {
39  Read: {
40    sandboxDescription: 'Reads a file; resolves to its text.',
41    summary: 'Reads a file.',
42    resolves: 'the file text',
43    args: {
44      file_path: { type: 'string', isRequired: true, note: 'absolute path', sandboxNote: 'Absolute path of the file.' },
45      offset: { type: 'number', isRequired: false, note: 'first line, from 1', sandboxNote: 'First line to read, from 1.' },
46      limit: { type: 'number', isRequired: false, note: 'number of lines', sandboxNote: 'Number of lines to read.' },
47    },
48  },
49  Bash: {
50    sandboxDescription: 'Runs a shell command; resolves to its output.',
51    summary: 'Runs a shell command.',
52    resolves: 'the command output',
53    args: {
54      command: { type: 'string', isRequired: true },
55      timeout: { type: 'number', isRequired: false, note: 'milliseconds', sandboxNote: 'Milliseconds.' },
56    },
57  },
58  Write: {
59    sandboxDescription: 'Writes a file, replacing it; resolves to a confirmation.',
60    summary: 'Writes a file.',
61    resolves: 'a confirmation',
62    args: {
63      file_path: { type: 'string', isRequired: true, note: 'absolute path', sandboxNote: 'Absolute path of the file.' },
64      content: { type: 'string', isRequired: true, sandboxNote: 'The content to write.' },
65    },
66  },
67  Edit: {
68    sandboxDescription: 'Replaces text in a file; resolves to a confirmation.',
69    summary: 'Replaces text in a file.',
70    resolves: 'a confirmation',
71    args: {
72      file_path: { type: 'string', isRequired: true, note: 'absolute path', sandboxNote: 'Absolute path of the file.' },
73      old_string: { type: 'string', isRequired: true, sandboxNote: 'The text to replace.' },
74      new_string: { type: 'string', isRequired: true, sandboxNote: 'The replacement text.' },
75      replace_all: { type: 'boolean', isRequired: false, note: 'every match', sandboxNote: 'Replace every match, not only the first.' },
76    },
77  },
78  ListMcpResourcesTool: {
79    sandboxDescription: 'Lists the resources MCP servers offer; resolves to the list.',
80    summary: 'Lists the resources MCP servers offer.',
81    resolves: 'the resources, each with its uri, name and server',
82    args: {
83      server: { type: 'string', isRequired: false, note: 'server name', sandboxNote: 'Only this MCP server.' },
84    },
85  },
86  ReadMcpResourceTool: {
87    sandboxDescription: 'Reads an MCP resource; resolves to its text.',
88    summary: 'Reads an MCP resource.',
89    resolves: 'the resource text',
90    args: {
91      server: { type: 'string', isRequired: true, note: 'server name', sandboxNote: 'Name of the MCP server.' },
92      uri: { type: 'string', isRequired: true, sandboxNote: 'The resource uri.' },
93    },
94  },
95  ReadMcpResourceDirTool: {
96    sandboxDescription: 'Lists the resources under an MCP resource directory; resolves to the list.',
97    summary: 'Lists the resources under an MCP resource directory.',
98    resolves: 'the resources under it',
99    args: {
100      server: { type: 'string', isRequired: true, note: 'server name', sandboxNote: 'Name of the MCP server.' },
101      uri: { type: 'string', isRequired: true, sandboxNote: 'The directory uri.' },
102    },
103  },
104}
105
106/** The JSON schema of a tool's input, as the child declares it to the sandbox. */
107export function inputSchemaOf(spec: ToolSpec): Record<string, unknown> {
108  const entries = Object.entries(spec.args)
109  return {
110    type: 'object',
111    properties: Object.fromEntries(
112      entries.map(([name, arg]) => [name, { type: arg.type, ...(arg.sandboxNote === undefined ? {} : { description: arg.sandboxNote }) }]),
113    ),
114    required: entries.filter(([, arg]) => arg.isRequired).map(([name]) => name),
115  }
116}
117
118/** What the child declares for an exposed tool: its description, input schema and output schema. */
119export function declarationOf(name: ExposedTool): {
120  description: string
121  inputSchema: Record<string, unknown>
122  outputSchema: Record<string, unknown>
123} {
124  const spec = TOOL_SPECS[name]
125  return {
126    description: spec.sandboxDescription,
127    inputSchema: inputSchemaOf(spec),
128    outputSchema: { type: 'string' },
129  }
130}
131
132/** Path the mod POSTs each answer to, over the child's Unix socket. */
133export const ANSWER_PATH = '/answer'
134
135/** The session's own tool; a script never calls it, or a codemode call could nest without end. */
136export const CODEMODE_TOOL_ID = 'mcp__codemode__codemode'
137
138/** A connected MCP tool a script may call, by its full `mcp__server__tool` name. */
139export type McpTool = {
140  name: string
141  description: string
142}
143
144/** What the child declares for an MCP tool: no input schema is known at run time, so it is open. */
145export function mcpDeclarationOf(tool: McpTool): {
146  description: string
147  inputSchema: Record<string, unknown>
148  outputSchema: Record<string, unknown>
149} {
150  return { description: tool.description, inputSchema: { type: 'object' }, outputSchema: { type: 'string' } }
151}
152
153/** What the mod writes to the child's standard input. */
154export type RunRequest = {
155  code: string
156  timeoutMs: number
157  /** Absent means none; the parser always fills it. */
158  mcpTools?: McpTool[]
159}
160
161/** One JSON line the child writes to standard output. */
162export type ChildMessage =
163  | { type: 'listening'; socketPath: string }
164  | { type: 'call'; id: number; tool: string; input: Record<string, unknown> }
165  | { type: 'done'; ok: true; output: string }
166  | { type: 'done'; ok: false; error: string; output: string }
167
168/** The body of a POST from the mod: how the nested call `id` ended. */
169export type CallAnswer =
170  | { id: number; ok: true; text: string }
171  | { id: number; ok: false; error: string }
172
173type Json = Record<string, unknown>
174
175function parseJson(text: string): Json | undefined {
176  try {
177    const value: unknown = JSON.parse(text)
178    const isObject = typeof value === 'object' && value !== null && !Array.isArray(value)
179    return isObject ? (value as Json) : undefined
180  } catch {
181    return undefined
182  }
183}
184
185export function isExposedTool(name: string): name is ExposedTool {
186  return (EXPOSED_TOOLS as readonly string[]).includes(name)
187}
188
189export function parseRunRequest(text: string): Required<RunRequest> | undefined {
190  const json = parseJson(text)
191  const isValid = json !== undefined && typeof json.code === 'string' && typeof json.timeoutMs === 'number'
192  if (!isValid) return undefined
193  const mcpTools = parseMcpTools(json.mcpTools)
194  return mcpTools === undefined ? undefined : { code: json.code as string, timeoutMs: json.timeoutMs as number, mcpTools }
195}
196
197/** An absent list is no tools; a list with any malformed entry is rejected whole. */
198function parseMcpTools(value: unknown): McpTool[] | undefined {
199  if (value === undefined) return []
200  if (!Array.isArray(value)) return undefined
201  const tools: McpTool[] = []
202  for (const entry of value as unknown[]) {
203    const isObject = typeof entry === 'object' && entry !== null && !Array.isArray(entry)
204    const { name, description } = isObject ? (entry as Json) : {}
205    if (typeof name !== 'string' || name === '' || typeof description !== 'string') return undefined
206    tools.push({ name, description })
207  }
208  return tools
209}
210
211export function parseChildMessage(line: string): ChildMessage | undefined {
212  const json = parseJson(line)
213  if (json === undefined) return undefined
214  if (json.type === 'listening' && typeof json.socketPath === 'string') {
215    return { type: 'listening', socketPath: json.socketPath }
216  }
217  if (json.type === 'call') return parseCall(json)
218  if (json.type === 'done') return parseDone(json)
219  return undefined
220}
221
222function parseCall(json: Json): ChildMessage | undefined {
223  const { id, tool, input } = json
224  const isInput = typeof input === 'object' && input !== null && !Array.isArray(input)
225  if (typeof id !== 'number' || typeof tool !== 'string' || !isInput) return undefined
226  return { type: 'call', id, tool, input: input as Record<string, unknown> }
227}
228
229function parseDone(json: Json): ChildMessage | undefined {
230  const { ok, output, error } = json
231  if (typeof output !== 'string') return undefined
232  if (ok === true) return { type: 'done', ok, output }
233  if (ok === false && typeof error === 'string') return { type: 'done', ok, error, output }
234  return undefined
235}
236
237export function parseCallAnswer(text: string): CallAnswer | undefined {
238  const json = parseJson(text)
239  if (json === undefined || typeof json.id !== 'number') return undefined
240  if (json.ok === true && typeof json.text === 'string') return { id: json.id, ok: true, text: json.text }
241  if (json.ok === false && typeof json.error === 'string') return { id: json.id, ok: false, error: json.error }
242  return undefined
243}
244
types/index.d.ts 36 lines
1export type CodemodeCallState = 'running' | 'done' | 'denied' | 'failed'
2
3/** One nested call a codemode script made through the session's tools. */
4export type CodemodeCall = {
5  /** The id the script's child gave the call; unique within its run. */
6  id: number
7  tool: string
8  /** One short line naming the target: a file path, the first line of a command. */
9  label: string
10  state: CodemodeCallState
11  /** Epoch milliseconds. */
12  startedAt: number
13  endedAt?: number
14  /** Why the call was denied or failed, as one short line. */
15  reason?: string
16}
17
18/** One codemode call as the transcript draws it, keyed by the call's tool_use_id. */
19export type CodemodeRun = {
20  id: string
21  startedAt: number
22  endedAt?: number
23  /** Length of the script's longest line, so the result row sizes its box like the script row. */
24  scriptWidth?: number
25  /** The most recent calls, oldest first, capped. */
26  calls: CodemodeCall[]
27  /** How many older calls the cap dropped from `calls`. */
28  omitted: number
29}
30
31declare module 'claude-code' {
32  interface PluginState {
33    codemode: { runs: CodemodeRun[] }
34  }
35}
36
shared/identifier.ts 15 lines
1/**
2 * The identifier a codemode script uses for a tool name: each character outside
3 * A-Za-z0-9_$ becomes `_`, a leading digit too, and an empty name becomes `_`.
4 * A copy of `toCodemodeIdentifier` from @earendil-works/pi-codemode, because the
5 * hooks loader allows only relative imports; test/node/identifier.spec.ts checks it.
6 */
7export function toScriptIdentifier(name: string): string {
8  let identifier = ''
9  for (const char of name) {
10    const valid = identifier === '' ? /^[A-Za-z_$]$/.test(char) : /^[A-Za-z0-9_$]$/.test(char)
11    identifier += valid ? char : '_'
12  }
13  return identifier === '' ? '_' : identifier
14}
15