SLOPSHOPPER

codex

OpenAI Codex as native Claude Code subagents: the Agent tool starts one and the mod drives codex app-server in its place, no model in between

newbandspinnerrowscommandprocess
★ 1v0.4.5Apache-2.0updated 2026-10-04alex2481kobe/claude-mods/plugins/codex
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · codex
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM › /codex-model ⎿ codex: Open a codex agent's view to use /codex-model. ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts
README

codex

OpenAI Codex as a native Claude Code subagent.

Two Codex agents on different models in Claude Code's agent list

Claude starts Codex with the Agent tool, the same way it starts any other subagent, and Codex behaves like one:

  • it shows in the agent list with its model and effort (· Sol 6.1 (xhigh)), its running time and token count, and clears when it finishes; the count is Codex's current context plus what it has written, the way Claude Code counts a Claude subagent, not the session's running total
  • its row's activity line updates as Codex works, and Enter opens its view, where Codex's steps appear live and you can message it
  • it runs in the background, and its report comes back to Claude
  • you can message it, while it runs or after it finishes; each message resumes the same Codex session. A message Claude sends while Codex works joins Codex's running turn, so Codex reads it then and its answer covers it; one you type in the agent's view waits for the turn to end (· 1 queued)
  • when Codex asks for an approval or an answer, the agent reports the question and its next message answers it, in the same paused Codex turn
  • in its view, /codex-* commands change its model, effort, sandbox and approvals for its next Codex turns and show its status
  • stopping the agent stops Codex, and so does Claude Code exiting or crashing; like a stopped Claude subagent, its next message carries on in the same Codex session rather than starting the task again

No Claude model runs inside the agent: the mod drives codex app-server in its place and shows what Codex does in the agent's view as it happens. (A small Claude model stands in only if the mod itself fails, to report that failure.)

Agent typeShown asCodex sandbox and approvalsUse it for
codex:readCodex read-onlyread-only, approvals offsecond opinions, review, research, scoping
codex:writeCodex workspace-writeworkspace-write, approvals offbounded implementation in the working directory
codex:runCodexyour Codex config, changed by flagsanything else Codex can be set up to do

Requirements

  • Claude Code with mods (tested on 2.1.287 and 2.1.288)
  • macOS or Linux (the mod talks to Codex through a named pipe)
  • The Codex CLI with codex app-server (tested on 0.159), installed and logged in:
  npm install -g @openai/codex
  codex login

codex must be on the PATH Claude Code starts with.

Install

/plugin marketplace add alex2481kobe/claude-mods
/plugin install codex@claude-mods

Use

Ask Claude for it by name:

Have codex:read review the changes in src/auth and report anything risky.

Use codex:write to add input validation to parseConfig in src/config.ts.

Codex cannot see your conversation with Claude, so Claude passes it a self-contained prompt. Codex uses your own Codex login, model and config.

Choosing a model and settings

The Agent tool's own model option names Claude models, so Codex is set up in the prompt instead. The prompt may open with Codex CLI flags, one per line, spelled as codex exec --help spells them (model: gpt-6-astra, --sandbox workspace-write, or a flag alone on its line). The mod passes them to Codex and removes them from the task. Model: and Effort: may be written in any case:

model: gpt-6-astra
effort: high
Review app.js for bugs and report back.

The agent types' descriptions list the models your Codex knows, so you can ask in plain words: "have Codex review this on gpt-6-astra and gpt-6.1-sol".

Only an exact option line is an option: a flag's name with one plain value (a path, or a config key=value, may hold spaces), or a flag that takes no value alone on its line. The first line that is not one starts the task, so a prompt that opens with prose such as search: every call to fetch or color: change the header color is passed to Codex whole.

codex:read and codex:write take model, effort and the flags that leave their sandbox alone; sandbox, approvals, add-dir and cd are refused there. codex:run takes every flag that applies to a session the mod drives:

FlagWhat Codex gets
model, effortmodel, model_reasoning_effort
sandbox (s)sandbox_mode
ask-for-approval (a)approval_policy
approve-for-methe automatic reviewer, in workspace-write
dangerously-bypass-approvals-and-sandboxdanger-full-access with approvals off
add-diran extra writable root, beside your config's
search, local-provider, cd (C), image (i)live web search, the model provider, the folder, images
config (c), enable, disable, strict-configpassed as given
ephemeral, output-schemaan unsaved session, a JSON Schema for the answer

Everything left out comes from your Codex config. An option line for a flag codex app-server has no use for (profile: fast, worktree, json, ...) stops the run with the reason rather than being dropped.

Answering Codex

When Codex asks for something, the agent hands the question back:

Codex asks to run:
  printf 'hi' > note.txt
in /path/to/project
Reason: requires approval by policy
Reply "approve", "approve for session", "decline", or "cancel" (decline and stop the turn).

Claude answers it itself or asks you, then sends the reply to the agent as a message, and Codex carries on in the same turn. Questions for the user and MCP servers' forms work the same way.

A question lives as long as the Codex that asked it. If that Codex is gone (it exited, the session was resumed, or the mod reloaded), a reply such as "approve" is not sent as a new task: the agent reports that the question has expired, and the task has to be sent again.

Who Codex asks is set by its config. With approvals_reviewer set to Codex's automatic reviewer, Codex never asks: its reviewer decides, and the transcript shows what it decided. To be asked instead, set approvals to come to you, in your config or for one agent:

ask-for-approval: on-request
config: approvals_reviewer="user"
Create note.txt containing hi.

Commands in the agent's view

Open a codex agent's view (select it in the agent list, press Enter) and run a command; the / menu there lists them. The reply shows above the prompt in that view, and neither Codex nor Claude is sent it:

CommandWhat it does
/codex-model <id>the Codex model, from the agent's next Codex turn
/codex-effort <level>the reasoning effort, from the next turn
`/codex-sandbox <read-only\workspace-write\danger-full-access>`the sandbox, from the next turn
`/codex-approvals <untrusted\on-request\never>`when Codex asks for approval, from the next turn
/codex-statuswhat Codex said the session ran with at its last turn, what changes next turn, the Codex session and its running token total
/codex-helpthe list
model     gpt-6-astra from the next Codex turn (now gpt-6-luna)
effort    low
sandbox   workspace-write
approvals on-request
session   01a0f9f0-0000-7000-8000-000000000000
tokens    141,551 in (127,488 cached), 296 out, running total

A setting is kept for that agent and goes with each of its later Codex turns, and the header of each turn names the model and effort Codex reports for it. Values follow the option rules above: codex:read and codex:write refuse /codex-sandbox and /codex-approvals, since they pin their sandbox, and a value that is not one plain word is refused. An unknown /codex- command, or one without its value, answers with the list. Claude can send one to the agent with SendMessage; then the reply is the agent's report, and a message sent together with it goes to Codex on its own.

While a codex agent's view is open, the footer and the / menu list these commands alone: Claude Code's own commands act on the main session, not on the agent, so they are hidden there (typed in full, they still run). Elsewhere the /codex- commands are hidden, and one typed in full says to open a codex agent's view.

How it works

  • The mod registers the three agent types.
  • When a codex:* agent's loop asks its model for a response, the mod answers instead: it starts codex app-server with the agent's flags as config overrides, starts or resumes the Codex session in the session's working directory with the settings its commands chose, shows Codex's messages and commands in the agent's view as they happen (each is also appended as a notice, which is what refreshes the agent list's activity line; the detailed transcript, ctrl+o, shows both), and reports Codex's token usage on the agent's row: each turn hands Claude Code the input of Codex's last request (its current context, cached tokens apart) and the output the turn generated. Claude Code keeps the latest input and adds up the outputs, as it does for a Claude subagent. The running total, which counts every cached re-read of the session and soon reaches millions, is in /codex-status.
  • A mod's process takes its input once, so the mod writes to codex app-server through a named pipe in a private temporary folder, which goes when the process does. The mod writes only to that pipe, and never once the server has gone.
  • Codex runs in a process group of its own under a small shell. Stopping the agent ends the shell and Codex with it; if Claude Code exits without stopping it (a crash, kill -9), the shell sees its parent gone within a second or two and ends Codex and the folder.
  • Codex's final message, or its question, goes back as the agent's report: through the SubagentHandback tool in an interactive session, or as the final text where that tool does not exist (headless, SDK). A handback you interrupt (Esc in the agent's view) does not change that. A report that never reached the caller (its handback interrupted, or failed for want of the tool) is given again the next time the agent's loop runs, once: ahead of the answer to a new message, or alone when there is none.
  • A step interrupted before Codex finished (Esc, or stopping the agent) stops Codex, and the agent hands back only codex: stopped before Codex finished.: what Codex said on the way is never handed back, and that line is never given again as an undelivered report. (Claude Code's own notice that the agent was stopped still quotes what the agent had shown so far.) A message counts as passed on once Codex has taken it (its turn started, or its question answered), finished or not: the next message resumes the same Codex session with that message alone, and the stopped task is not sent again. Only a step cut off before Codex took its messages (Esc, or the session moving host, while the session was starting) passes nothing on, and the next time the agent's loop runs, those messages are given to Codex again. The agent's options come from the prompt it was spawned with, which the mod records at spawn, so a first task run again keeps them however Claude Code places the messages sent since; the task goes first, then those messages. Claude Code's interruption marker ([Request interrupted by user]) never reaches Codex. When the interruption is the session moving to the background (the session list opening while the agent's first turn runs), Claude Code 2.1.287 continues the agent in a forked session whose conversation, as the mod reads it, no longer holds the task, so the agent reports codex: nothing new to send to Codex. and Claude has to send the task again.
  • Codex runs only while it works or waits on a question. Every message sent to the agent is passed to Codex once, in the sender's own words: as the answer to a waiting question, added to Codex's running turn (turn/steer, for a message Claude sends with SendMessage while Codex works), or as a new turn of the same Codex session. Claude Code places a message sent to a running agent twice, wrapped in its own instructions and as typed; the mod counts it once and drops the wrapping, so the same words sent twice are asked twice. A /codex- message is the mod's own and never reaches Codex.

Limits

  • The Agent tool's model and cwd options are ignored: Codex uses your Codex config, in the session's working directory.
  • What Codex may do is decided by its sandbox, approvals and requirements, not by Claude Code's permission prompts. codex:run takes every flag Codex accepts, dangerously-bypass-approvals-and-sandbox included. Codex's own requirements (allowed_sandbox_modes, allowed_approval_policies) are the place to forbid one; the mod has not been tested against them. In auto mode, Claude may note that its safety check could not review the agent's output, since no Claude model wrote it.
  • codex app-server takes no profiles: set those values with config: lines.
  • codex:read and codex:write run with approvals off, so the sandbox is the limit. Codex's workspace-write sandbox keeps .git read-only, so codex:write cannot stage or commit (git add fails on .git/index.lock); commit its work yourself, or use codex:run with approvals that let Codex ask.
  • A question left unanswered keeps that Codex waiting until it is answered or the session ends.
  • An agent started with ephemeral cannot take follow-ups: Codex does not keep its session, so there is nothing to resume.
  • Claude Code 2.1.287 shows a mod a message typed in an agent's view only once the agent's turn has ended, so such a message cannot join Codex's running turn; it waits for Codex's current task, then runs as its next turn.
  • A command's reply shows above the prompt only while that view stays open, and goes when the view closes or the mod reloads.
  • What the agent answers also reaches Claude as the agent's report, and Claude reads a message typed in the view as one the agent got.
  • Claude Code may run the agent's loop again for the copy of a message it places later. Codex is not asked again, but the loop has to report, so Claude gets the one line codex: nothing new to send to Codex., and the view shows codex: nothing to run. (Ending without a report would have Claude Code tell Claude that no report came and to message the agent for one.)
  • Opening the session list (← from the prompt) moves the conversation into a background process. On macOS Claude Code 2.1.287 sometimes starts that process as the Claude Code app itself, which macOS checks on its own for access to Documents, Desktop and Downloads. If the plugin's folder is under one of those and Claude Code has not been given access, that process cannot read the plugin, so the conversation continues there without it: the codex agent types are gone until you restart. A plugin installed from the marketplace lives under ~/.claude and is not affected; for a --plugin-dir or a local marketplace, keep the folder outside those three.
  • Claude Code's task list (/tasks) names the stand-in's model, Haiku, for a codex agent; the agent's row and header show Codex's. The agent keeps a Claude model so that a run the mod does not answer (the mod not loaded, or the session resumed without it) reaches the stand-in, which reports that Codex did not run, rather than failing on a Codex model id.
  • Tested on macOS with codex-cli 0.159 and Claude Code 2.1.287 and 2.1.288, in an interactive terminal session (agent list, agent view and its commands, footer and / menu, background agents, messages and queued messages, approvals, stop) and headless (claude -p). Linux should behave the same; Windows is not supported (the mod needs sh and a named pipe).
  • The mods API is early access and may change between Claude Code releases.

Develop

claude plugin validate plugins/codex
claude plugin test plugins/codex
claude --plugin-dir plugins/codex
Source 14 files
hooks/register.ts 97 lines
1import { atom, update } from 'claude-code'
2import type { EngineInterface, Register } from 'claude-code'
3
4import { specsOf, TYPES } from './agents'
5import { replyBand } from './band'
6import { HINT, isCommand, SPECS } from './commands'
7import { flagsOf } from './flags'
8import { codexConfig, codexModels, labelOf, type Files } from './model'
9import { closeHeld, command, steer, step } from './step'
10import { openView, replyIn, setOpenView } from './view'
11
12// Codex as native subagent types. The Agent tool starts one like any other
13// subagent (task list, background, SendMessage); a turn.step hook answers its
14// loop's model request by driving `codex app-server`, streaming what Codex
15// does into the agent's transcript, then hands Codex's answer back. What
16// Codex asks on the way (an approval, a question) is handed back the same
17// way, and the agent's next message answers it. No Claude model runs.
18
19// Codex's files as model.ts reads them; `$.env.get` takes literal names.
20async function files($: EngineInterface): Promise<Files> {
21  const [CODEX_HOME, HOME, USERPROFILE] = [await $.env.get('CODEX_HOME'), await $.env.get('HOME'), await $.env.get('USERPROFILE')]
22  return { env: { CODEX_HOME, HOME, USERPROFILE }, read: path => $.fs.read(path) }
23}
24
25// The prompt each codex agent was spawned with, which carries its options:
26// the engine may later place another message ahead of it in the agent's
27// conversation, so step.ts reads the options from here.
28const openings = atom({ plugin: 'codex', key: 'openings' } as const, {})
29
30export const register: Register = on => {
31  on('session.start', async ($, e, next) => {
32    for (const spec of specsOf(await codexModels(await files($)))) await $.agent.register(spec)
33    for (const spec of SPECS) await $.command.register(spec)
34    return next(e)
35  })
36
37  on('session.end', async ($, e, next) => {
38    closeHeld()
39    return next(e)
40  })
41
42  on('agent.spawn', async ($, e, next) => {
43    const type = TYPES[e.subagentType]
44    if (!type) return next(e)
45    const flags = flagsOf(e.prompt, type.pin)
46    // The agent keeps the stand-in's Claude model: a run the mod does not
47    // answer (the mod not loaded, a resume elsewhere) is sent to it, so it
48    // must be one the Anthropic API serves, never the Codex model.
49    const label = labelOf(await codexConfig(await files($)), 'error' in flags ? {} : flags)
50    const started = await next({ ...e, description: `${e.description} · ${label}` })
51    const agentId = started.agentId
52    if (agentId) await update($, openings, all => ({ ...all, [agentId]: e.prompt }))
53    return started
54  })
55
56  // The transcript's Agent row names the type in words, not `codex:read`.
57  on('ui.render', { component: 'ToolUse' }, async ($, e, next) => {
58    const input = e.props.input as { subagent_type?: unknown } | undefined
59    const type = e.props.tool === 'Agent' && typeof input?.subagent_type === 'string' ? TYPES[input.subagent_type] : undefined
60    if (!type) return next(e)
61    return next({ ...e, props: { ...e.props, input: { ...input, subagent_type: type.shown } } })
62  })
63
64  on('turn.step', step)
65
66  // A message sent to a codex agent while Codex works joins Codex's running
67  // turn, so Codex reads it now and its answer covers it; the agent is not
68  // also sent it, which would start another Codex turn once this one ends.
69  on('session.send', async ($, e, next) => {
70    const agent = (await $.agent.list()).find(a => a.id === e.to || a.name === e.to)
71    if (!agent || !TYPES[agent.type] || isCommand(e.text)) return next(e)
72    return (await steer(agent.id, e.text)) ? { isDelivered: true } : next(e)
73  })
74
75  // The codex agent whose view is open: the band above the prompt is drawn
76  // for the view on screen, and shows the reply to the last /codex- command
77  // run there.
78  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
79    const id = e.props.view.agentId
80    const agent = id === undefined ? undefined : (await $.agent.list()).find(a => a.id === id)
81    if (setOpenView(agent && TYPES[agent.type] ? id : undefined)) {
82      $.ui.invalidate('ui.render')
83      $.ui.invalidate('command.describe')
84    }
85    const view = openView()
86    const reply = view === undefined ? undefined : replyIn(view)
87    return reply ? replyBand($.ui.resolve(e), reply, await next(e)) : next(e)
88  })
89
90  // There the footer names the agent's commands, and the "/" menu lists them
91  // alone: Claude Code's act on the main session rather than the agent.
92  // Elsewhere the agent's commands are left out.
93  on('ui.render', { component: 'PromptHint' }, async ($, e, next) => (openView() ? next({ ...e, props: { ...e.props, tail: HINT } }) : next(e)))
94  on('command.describe', async ($, e, next) => (Boolean(openView()) !== isCommand(`/${e.command}`) ? next({ ...e, isHidden: true }) : next(e)))
95  on('command.run', command)
96}
97
hooks/agents.ts 50 lines
1import type { AgentSpec } from 'claude-code'
2
3import { FLAG_NAMES, type Pin } from './flags'
4
5// The three codex agent types: what each pins, how it is shown, and how it
6// is offered to the model.
7
8export const TYPES: Record<string, { pin?: Pin; shown: string }> = {
9  'codex:read': { pin: { sandbox: 'read-only' }, shown: 'Codex read-only' },
10  'codex:write': { pin: { sandbox: 'workspace-write' }, shown: 'Codex workspace-write' },
11  'codex:run': { shown: 'Codex' },
12}
13
14// The agent's system prompt, which no model reads while the mod answers its
15// turns. The agent list still summarises a running agent from it, so it opens
16// with what the agent does; a stand-in model, reached only if the turn.step
17// hook fails, reports that instead of answering.
18const FALLBACK = `This agent passes its task to OpenAI Codex, which does the work in its place: the codex mod runs Codex and reports Codex's answer. A model reading this means the mod did not run, so do nothing else and give as your final report that Codex did not run and that this agent's transcript and the debug log say why.`
19
20const CHOOSING = (models: string[]) =>
21  ` Codex CLI flags may open the prompt, one per line as \`codex exec --help\` names them without dashes (\`model: <id>\`, \`effort: <level>\`` +
22  (models.length > 0 ? `; models: ${models.join(', ')}` : '') +
23  `); left out, Codex uses its own config. When Codex asks for an approval or an answer, the agent reports the question; send it the reply as a message.`
24
25// The agent types as `$.agent.register` takes them, offering these models.
26export function specsOf(models: string[]): AgentSpec[] {
27  const common = { prompt: FALLBACK, tools: ['Read'], model: 'haiku', omitClaudeMd: true } as const
28  const choosing = CHOOSING(models)
29  return [
30    {
31      ...common,
32      name: 'read',
33      description:
34        'OpenAI Codex in a read-only sandbox: a second opinion, review, research or scoping by a different model. Give it a self-contained prompt; it cannot see this conversation.' + choosing,
35    },
36    {
37      ...common,
38      name: 'write',
39      description:
40        'OpenAI Codex with workspace-write in the working directory: bounded implementation by a different model. Give it a self-contained prompt naming the files it owns; it cannot see this conversation.' + choosing,
41    },
42    {
43      ...common,
44      name: 'run',
45      description:
46        `OpenAI Codex as your Codex config sets it up (sandbox, approvals, reviewer), changed per call by any of its flags: ${FLAG_NAMES.join(', ')}. Give it a self-contained prompt; it cannot see this conversation.` + choosing,
47    },
48  ]
49}
50
hooks/band.tsx 14 lines
1import type { RenderElement } from 'claude-code'
2
3// The band above the prompt in a codex agent's view: the reply to the last
4// /codex- command run there, over whatever the band holds beneath.
5export function replyBand(ui: { Box: any; Text: any }, reply: string, below: RenderElement | null) {
6  const { Box, Text } = ui
7  return (
8    <Box flexDirection="column">
9      <Text>{reply}</Text>
10      {below}
11    </Box>
12  )
13}
14
hooks/commands.ts 112 lines
1import type { CodexOptionName, CodexRun } from '../types'
2import { APPROVALS, flagsOf, SANDBOXES, type Pin } from './flags'
3
4// `/codex-*` messages, which a codex agent answers itself and never passes to
5// Codex: settings for its next Codex turns, its status, and help. They are
6// typed in the agent's view, where Claude Code hands a command it does not
7// know to the agent as a message. A setting is the Codex option line it
8// names, so it is checked by the option rules and the agent type's pin.
9
10export type Options = Partial<Record<CodexOptionName, string>>
11
12// The reply, and the agent's settings when the command changed them.
13export type Answer = { reply: string; options?: Options }
14
15const SETTINGS: Record<string, { option: CodexOptionName; value: string; what: string }> = {
16  model: { option: 'model', value: '<id>', what: 'The Codex model' },
17  effort: { option: 'effort', value: '<level>', what: 'The reasoning effort' },
18  sandbox: { option: 'sandbox', value: `<${SANDBOXES.join('|')}>`, what: 'The sandbox' },
19  approvals: { option: 'ask-for-approval', value: `<${APPROVALS.join('|')}>`, what: 'When Codex asks for approval' },
20}
21
22// The commands as Claude Code registers them, listed in a codex agent's view.
23export const SPECS: { name: string; description: string; argumentHint?: string; immediate: true }[] = [
24  ...Object.entries(SETTINGS).map(([name, { value, what }]) => ({
25    name: `codex-${name}`,
26    description: `${what} for this codex agent, from its next Codex turn`,
27    argumentHint: value,
28    immediate: true as const,
29  })),
30  { name: 'codex-status', description: 'What Codex runs this agent with, its session and its tokens', immediate: true },
31  { name: 'codex-help', description: 'The codex agent commands', immediate: true },
32]
33
34const NAMES = SPECS.map(spec => `/${spec.name}`)
35
36const HELP = [
37  'Codex commands for this agent (a setting applies from its next Codex turn):',
38  ...Object.entries(SETTINGS).map(([name, { value }]) => `  /codex-${name} ${value}`),
39  '  /codex-status   what Codex runs with, the session and its tokens',
40  '  /codex-help     this list',
41].join('\n')
42
43// The footer hint while a codex agent's view is open.
44export const HINT = NAMES.join(' ')
45
46const COMMAND = /^\/codex-(\S*)(?:\s+([\s\S]*))?$/
47
48export function isCommand(text: string): boolean {
49  return text.trimStart().startsWith('/codex-')
50}
51
52export function answerOf(text: string, options: Options, pin: Pin | undefined, run: CodexRun | undefined): Answer {
53  const match = COMMAND.exec(text.trim())
54  const name = match?.[1] ?? ''
55  const value = (match?.[2] ?? '').trim()
56  if (name === 'status') return { reply: statusOf(options, run) }
57  const setting = SETTINGS[name]
58  if (!setting || value === '') {
59    const why = name === 'help' ? '' : setting ? `codex: /codex-${name} needs a value.\n\n` : `codex: no /codex-${name}.\n\n`
60    return { reply: `${why}${HELP}` }
61  }
62  // One plain word: a value that spans lines would carry option lines of its own.
63  if (/\s/.test(value)) return { reply: `codex: /codex-${name} takes one plain value, not "${value}".` }
64  const parsed = flagsOf(`${setting.option}: ${value}`, pin)
65  if ('error' in parsed) return { reply: `codex: ${parsed.error}` }
66  if (parsed.prompt !== '') return { reply: `codex: /codex-${name} takes one plain value, not "${value}".` }
67  return { reply: `codex: ${name} ${value} from the next Codex turn.`, options: { ...options, [setting.option]: value } }
68}
69
70// Commands answered in order, each seeing the settings the ones before it
71// left; the replies, and the agent's settings after them.
72export function answersOf(texts: readonly string[], options: Options, pin: Pin | undefined, run: CodexRun | undefined): { replies: string[]; options: Options } {
73  let mine = options
74  const replies = texts.map(text => {
75    const answer = answerOf(text, mine, pin, run)
76    mine = answer.options ?? mine
77    return answer.reply
78  })
79  return { replies, options: mine }
80}
81
82// The settings as thread/start and thread/resume take them: a resumed
83// session keeps the model, sandbox and approvals it ran with unless the call
84// names others, so they go with every call.
85export function threadParamsOf(options: Options): { model?: string; sandbox?: string; approvalPolicy?: string; config?: Record<string, string> } {
86  return {
87    ...(options.model ? { model: options.model } : {}),
88    ...(options.sandbox ? { sandbox: options.sandbox } : {}),
89    ...(options['ask-for-approval'] ? { approvalPolicy: options['ask-for-approval'] } : {}),
90    ...(options.effort ? { config: { model_reasoning_effort: options.effort } } : {}),
91  }
92}
93
94const count = (n: number) => String(n).replace(/\B(?=(\d{3})+(?!\d))/g, ',')
95
96function statusOf(options: Options, run: CodexRun | undefined): string {
97  const now: Record<CodexOptionName, string | undefined> = {
98    model: run?.model,
99    effort: run ? (run.effort ?? 'default') : undefined,
100    sandbox: run?.sandbox,
101    'ask-for-approval': run?.approvals,
102  }
103  const rows = Object.entries(SETTINGS).flatMap(([name, { option }]) => {
104    const next = options[option]
105    if (next !== undefined && next !== now[option]) return [`${name.padEnd(10)}${next} from the next Codex turn${now[option] ? ` (now ${now[option]})` : ''}`]
106    return now[option] ? [`${name.padEnd(10)}${now[option]}`] : []
107  })
108  if (!run) return ['codex: no Codex turn yet.', ...rows].join('\n')
109  const { input, cached, output } = run.tokens
110  return [...rows, `${'session'.padEnd(10)}${run.threadId}`, `${'tokens'.padEnd(10)}${count(input)} in (${count(cached)} cached), ${count(output)} out, running total`].join('\n')
111}
112
hooks/flags.ts 158 lines
1// A codex agent's prompt may open with Codex CLI flags, one per line, spelled
2// as `codex exec --help` spells them without the dashes: `model: gpt-6-astra`,
3// `sandbox: workspace-write`, `approve-for-me`, `config: key=value`. They
4// choose how Codex runs and are not part of the task. Codex's own config fills
5// in everything not chosen.
6
7export type Flags = {
8  // Arguments for `codex app-server`: config overrides and feature switches.
9  args: string[]
10  model?: string
11  effort?: string
12  cwd?: string
13  ephemeral?: boolean
14  images: string[]
15  outputSchema?: string
16  addDirs: string[]
17  prompt: string
18}
19
20export type Parsed = Flags | { error: string }
21
22// What the read and write types pin: a flag that changes it is refused there.
23export type Pin = { sandbox: string }
24
25const WORD = /^[A-Za-z0-9._/:@+-]+$/
26const CONFIG = /^[A-Za-z0-9_.-]+=.+$/
27const PINNED = /^(sandbox_mode|approval_policy|approvals_reviewer|sandbox_workspace_write)\b/
28
29export const SANDBOXES = ['read-only', 'workspace-write', 'danger-full-access']
30export const APPROVALS = ['untrusted', 'on-request', 'never']
31
32// `codex exec` flags that have no meaning for a session the mod drives.
33const EXEC_ONLY: Record<string, string> = {
34  profile: 'codex app-server takes no --profile; set the values with config: lines',
35  json: 'the mod already reads Codex events',
36  'output-last-message': 'the report goes back to Claude',
37  color: 'there is no terminal',
38  'skip-git-repo-check': 'the mod never needs a git repo',
39  worktree: 'codex app-server takes no --worktree',
40  'ignore-user-config': 'codex app-server takes no --ignore-user-config',
41  'ignore-rules': 'codex app-server takes no --ignore-rules',
42  'dangerously-bypass-hook-trust': 'codex app-server takes no --dangerously-bypass-hook-trust',
43  oss: 'name the provider with local-provider: lmstudio or ollama',
44}
45
46const ALIASES: Record<string, string> = {
47  m: 'model',
48  s: 'sandbox',
49  a: 'ask-for-approval',
50  approval: 'ask-for-approval',
51  c: 'config',
52  C: 'cd',
53  i: 'image',
54  auto: 'approve-for-me',
55  'reasoning-effort': 'effort',
56}
57
58// `--sandbox read-only`, `-s=read-only`, `sandbox: read-only`, or a flag that
59// takes no value alone on its line (`approve-for-me`). "search the code" has
60// neither the dashes nor the colon; see isOptionLine for prose that has one.
61const DASHED = /^--?([A-Za-z][A-Za-z-]*)(?:[ =]\s*(.*))?$/
62const PLAIN = /^([A-Za-z][A-Za-z-]*)(?::\s*(.*))?$/
63
64type Apply = (flags: Flags, value: string) => string | undefined
65
66function quoted(key: string, value: string, flags: Flags): string | undefined {
67  if (!WORD.test(value)) return `${key} needs a plain value, not "${value}"`
68  flags.args.push('-c', `${key}="${value}"`)
69}
70
71function oneOf(key: string, allowed: string[]): Apply {
72  return (flags, value) => (allowed.includes(value) ? quoted(key, value, flags) : `${key} is one of ${allowed.join(', ')}`)
73}
74
75// `path`: the value is a path, which may hold spaces when it starts like one.
76const FLAGS: Record<string, { apply: Apply; pinned?: true; bare?: true; path?: true }> = {
77  model: { apply: (f, v) => quoted('model', v, f) ?? void (f.model = v) },
78  effort: { apply: (f, v) => quoted('model_reasoning_effort', v, f) ?? void (f.effort = v) },
79  sandbox: { apply: oneOf('sandbox_mode', SANDBOXES), pinned: true },
80  'ask-for-approval': { apply: oneOf('approval_policy', APPROVALS), pinned: true },
81  // As `codex exec --approve-for-me`: Codex's automatic reviewer decides, in
82  // the workspace-write sandbox.
83  'approve-for-me': {
84    apply: f => void f.args.push('-c', 'approvals_reviewer="auto_review"', '-c', 'sandbox_mode="workspace-write"'),
85    pinned: true,
86    bare: true,
87  },
88  'dangerously-bypass-approvals-and-sandbox': {
89    apply: f => void f.args.push('-c', 'sandbox_mode="danger-full-access"', '-c', 'approval_policy="never"'),
90    pinned: true,
91    bare: true,
92  },
93  'add-dir': { apply: (f, v) => void f.addDirs.push(v), pinned: true, path: true },
94  search: { apply: f => void f.args.push('-c', 'web_search="live"'), bare: true },
95  'local-provider': { apply: (f, v) => quoted('model_provider', v, f) },
96  config: {
97    apply: (f, v) => (CONFIG.test(v) ? void f.args.push('-c', v) : `config takes key=value, not "${v}"`),
98  },
99  enable: { apply: (f, v) => (WORD.test(v) ? void f.args.push('--enable', v) : `enable takes a feature name`) },
100  disable: { apply: (f, v) => (WORD.test(v) ? void f.args.push('--disable', v) : `disable takes a feature name`) },
101  'strict-config': { apply: f => void f.args.push('--strict-config'), bare: true },
102  ephemeral: { apply: f => void (f.ephemeral = true), bare: true },
103  // Pinned like add-dir: moving the folder moves the writable root.
104  cd: { apply: (f, v) => void (f.cwd = v), pinned: true, path: true },
105  image: { apply: (f, v) => void f.images.push(v), path: true },
106  'output-schema': { apply: (f, v) => void (f.outputSchema = v), path: true },
107}
108
109export const FLAG_NAMES = Object.keys(FLAGS)
110
111// Whether a line is shaped as an option line for this flag: a bare flag
112// alone (or `true`), a config `key=value`, a path, or one plain value. A line
113// of prose that only starts with a flag's name ("search: every call") is not,
114// and is the task's.
115function isOptionLine(name: string, value: string): boolean {
116  const flag = FLAGS[name]
117  if (flag?.bare) return value === '' || /^(true|yes|on)$/i.test(value)
118  if (name === 'config') return CONFIG.test(value)
119  if (flag?.path && /^[/~.]/.test(value)) return true
120  return value !== '' && !/\s/.test(value)
121}
122
123// `Model:` and `Effort:` read in any case, as 0.2.5 read them; every other
124// name as `codex exec --help` spells it, since `-C` and `-c` differ.
125function nameOf(raw: string): string {
126  const lower = raw.toLowerCase()
127  if (lower === 'model' || lower === 'effort') return lower
128  return ALIASES[raw] ?? raw
129}
130
131// Reads the leading option lines of a prompt. The first line that is not one
132// ends them, and it and the rest are the task. An option line used wrongly,
133// or one a pinned type refuses, is an error for the caller rather than a
134// silent change of what Codex runs with.
135export function flagsOf(text: string, pin?: Pin): Parsed {
136  const flags: Flags = { args: [], images: [], addDirs: [], prompt: text }
137  if (pin) flags.args.push('-c', `sandbox_mode="${pin.sandbox}"`, '-c', 'approval_policy="never"')
138  const lines = text.split('\n')
139  let used = 0
140  for (const line of lines) {
141    const match = DASHED.exec(line.trim()) ?? PLAIN.exec(line.trim())
142    if (!match) break
143    const name = nameOf(match[1]!)
144    const value = (match[2] ?? '').trim()
145    if (!(name in FLAGS || name in EXEC_ONLY) || !isOptionLine(name, value)) break
146    if (name in EXEC_ONLY) return { error: `${name}: ${EXEC_ONLY[name]}` }
147    const flag = FLAGS[name]!
148    if (pin && (flag.pinned || (name === 'config' && PINNED.test(value)))) {
149      return { error: `${name}: this agent type pins the ${pin.sandbox} sandbox; use codex:run to choose` }
150    }
151    const error = flag.apply(flags, value)
152    if (error) return { error }
153    used++
154  }
155  flags.prompt = lines.slice(used).join('\n').trim()
156  return flags
157}
158
hooks/model.ts 68 lines
1import { nameOf } from './names'
2
3// What Codex runs with: its config.toml's top-level model and reasoning
4// effort, under what a codex agent's prompt chose.
5
6export type Choice = { model?: string; effort?: string }
7
8export function configOf(toml: string): Choice {
9  const choice: Choice = {}
10  for (const line of toml.split('\n')) {
11    const text = line.trim()
12    if (text.startsWith('[')) break
13    const pair = /^([A-Za-z_]+)\s*=\s*["']([^"']+)["']/.exec(text)
14    if (pair?.[1] === 'model') choice.model = pair[2]
15    if (pair?.[1] === 'model_reasoning_effort') choice.effort = pair[2]
16  }
17  return choice
18}
19
20// `Sol 6.1 (high)`: the chosen model and effort over the config's. With
21// neither naming a model, Codex picks its own; which one is known only once
22// Codex runs (each turn's header names it), so the label says so.
23export function labelOf(config: Choice, options: Choice): string {
24  const id = options.model ?? config.model
25  const model = id ? nameOf(id) : 'Codex default'
26  const effort = options.effort ?? config.effort
27  return effort ? `${model} (${effort})` : model
28}
29
30// The model ids Codex's model cache lists, newest first as it keeps them;
31// empty when the cache is missing or its shape is not the one read here.
32export function modelsOf(json: string): string[] {
33  try {
34    const cache = JSON.parse(json)
35    const list = Array.isArray(cache) ? cache : (cache?.models ?? cache?.data)
36    if (!Array.isArray(list)) return []
37    return list.map((m: any) => m?.slug ?? m?.id).filter((s: unknown): s is string => typeof s === 'string')
38  } catch {
39    return []
40  }
41}
42
43// What reading Codex's own files needs of the engine: the environment
44// variables that place them, and a file read.
45export type Files = {
46  env: { CODEX_HOME?: string; HOME?: string; USERPROFILE?: string }
47  read: (path: string) => Promise<string>
48}
49
50async function readOr(files: Files, name: string): Promise<string> {
51  const { CODEX_HOME, HOME, USERPROFILE } = files.env
52  try {
53    return await files.read(`${CODEX_HOME ?? `${HOME ?? USERPROFILE}/.codex`}/${name}`)
54  } catch {
55    return ''
56  }
57}
58
59// The model and effort Codex's config.toml names; empty when it names none.
60export async function codexConfig(files: Files): Promise<Choice> {
61  return configOf(await readOr(files, 'config.toml'))
62}
63
64// The model ids Codex's model cache lists.
65export async function codexModels(files: Files): Promise<string[]> {
66  return modelsOf(await readOr(files, 'models_cache.json'))
67}
68
hooks/step.ts 282 lines
1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, Hook, TurnStepChunk, TurnUsage } from 'claude-code'
3
4import type { CodexRun } from '../types'
5import { TYPES } from './agents'
6import { answersOf, isCommand, threadParamsOf } from './commands'
7import { answerOf, apply, reportOf, STOPPED, type Run } from './events'
8import { flagsOf, type Flags, type Pin } from './flags'
9import { labelOf } from './model'
10import { expiredAnswer, questionOf, replyOf, type Asked } from './questions'
11import { HANDBACK, handsBack, lastAnswer, requestOf, rowsOf, undeliveredReport } from './request'
12import { NO_CODEX, open, type Server } from './server'
13import { openView, setReply } from './view'
14
15// One step of a codex agent's loop: its model request answered by driving
16// `codex app-server`, Codex's steps streamed as the agent's text, and Codex's
17// answer or question handed back.
18
19// A Codex turn paused on a question, by agent id: the server stays up until
20// the agent's next message answers it.
21const held = new Map<string, { server: Server; asked: Asked; threadId: string }>()
22
23// The Codex turn each codex agent is running now, by agent id, so a message
24// sent to the agent meanwhile can join it rather than wait for it to end.
25const working = new Map<string, { server: Server; threadId: string; turnId: string }>()
26
27// Adds a message to the agent's running Codex turn; false when no turn runs
28// or Codex refused (the turn had just ended), so the message takes the usual
29// way: the agent's next turn.
30export async function steer(agentId: string, text: string): Promise<boolean> {
31  const turn = working.get(agentId)
32  if (!turn) return false
33  try {
34    await turn.server.call('turn/steer', { threadId: turn.threadId, expectedTurnId: turn.turnId, input: [{ type: 'text', text }] })
35    return true
36  } catch {
37    return false
38  }
39}
40
41// Ends every Codex turn left waiting on a question.
42export function closeHeld(): void {
43  for (const { server } of held.values()) server.close()
44  held.clear()
45}
46
47// What each codex agent has passed to Codex, so a message is sent once; the
48// prompt it was spawned with (register.ts records it); what its commands
49// set; and what Codex last reported running it with.
50const sent = atom({ plugin: 'codex', key: 'sent' } as const, {})
51const openings = atom({ plugin: 'codex', key: 'openings' } as const, {})
52const options = atom({ plugin: 'codex', key: 'options' } as const, {})
53const runs = atom({ plugin: 'codex', key: 'runs' } as const, {})
54
55// Records what Codex reported the agent's session runs with (or keeps the
56// last), with its token total once Codex has counted it.
57async function record($: EngineInterface, agentId: string, report: Omit<CodexRun, 'tokens'> | undefined, tokens?: CodexRun['tokens']): Promise<void> {
58  await update($, runs, all => {
59    const base = report ?? all[agentId]
60    return base ? { ...all, [agentId]: { ...base, tokens: tokens ?? all[agentId]?.tokens ?? { input: 0, cached: 0, output: 0 } } } : all
61  })
62}
63
64const shown = (text: string): TurnStepChunk => ({ kind: 'text', index: 0, text })
65
66// Codex's progress also as notices in the agent's conversation, appended as
67// it happens: a notice is what refreshes the agent list's activity line, while
68// the agent's view shows the step's text. The model never reads a notice, and
69// only the detailed transcript (ctrl+o) shows both. Display only, so a refused
70// append changes nothing else.
71async function note($: EngineInterface, agentId: string, text: string): Promise<void> {
72  await $.session
73    .append({ agentId, message: { type: 'system', content: [{ type: 'text', text: text.trimEnd() }] } })
74    .catch(() => undefined)
75}
76
77// Resolves undefined when the step is aborted first; the listener goes once
78// the wait settles, so a long turn does not pile them up on the signal.
79export function unlessAborted<T>(signal: AbortSignal, promise: Promise<T>): Promise<T | undefined> {
80  if (signal.aborted) return Promise.resolve(undefined)
81  let stop = () => {}
82  const aborted = new Promise<undefined>(resolve => {
83    stop = () => resolve(undefined)
84    signal.addEventListener('abort', stop, { once: true })
85  })
86  return Promise.race([promise, aborted]).finally(() => signal.removeEventListener('abort', stop))
87}
88
89const EXPIRED =
90  'codex: the question Codex asked has expired: the Codex turn that asked it is gone (Codex exited, the session was resumed, or the mod reloaded), so nothing was answered. Send the task again to start a new turn.'
91
92// The agent's `/codex-*` commands, answered in order; a setting is kept for
93// its next Codex turn.
94async function answered($: EngineInterface, agentId: string, commands: string[], pin?: Pin): Promise<string[]> {
95  const last = (await read($, runs))[agentId]
96  let replies: string[] = []
97  await update($, options, all => {
98    const answers = answersOf(commands, all[agentId] ?? {}, pin, last)
99    replies = answers.replies
100    return { ...all, [agentId]: answers.options }
101  })
102  return replies
103}
104
105// A registered /codex- command acts on the agent whose view is open and
106// answers in that view's band; the main conversation is not sent the reply.
107export const command: Hook<'command.run'> = async ($, e, next) => {
108  if (!isCommand(`/${e.command}`)) return next(e)
109  const id = openView()
110  const agent = id === undefined ? undefined : (await $.agent.list()).find(a => a.id === id)
111  const type = agent && TYPES[agent.type]
112  if (!agent || !type) return { text: `Open a codex agent's view to use /${e.command}.` }
113  const [reply] = await answered($, agent.id, [`/${e.command} ${e.args}`.trim()], type.pin)
114  setReply(agent.id, reply!)
115  $.ui.invalidate('ui.render')
116  return {}
117}
118
119export const step: Hook<'turn.step'> = async function* ($, e, next) {
120  const agentId = e.agentId
121  if (!agentId) return yield* next(e)
122  const agent = (await $.agent.list()).find(a => a.id === agentId)
123  const type = agent && TYPES[agent.type]
124  if (!type) return yield* next(e)
125
126  const api = await $.session.messages({ as: 'api', agentId })
127  if ('deny' in api) throw new Error(api.deny)
128  const rows = rowsOf(api)
129  const given = (await read($, sent))[agentId] ?? []
130  const request = requestOf(rows, given, (await read($, openings))[agentId])
131  // Codex's progress is the transcript's text; the report goes back with a
132  // handback call, or as the final text where the loop has no such tool.
133  const handback = handsBack(rows)
134
135  const run: Run = {}
136  let progress = ''
137  let question: string | undefined
138  let model = 'codex'
139  let server: Server | undefined
140  const show = (text: string) => {
141    progress += text
142    return shown(text)
143  }
144  // A `/codex-*` message is the mod's to answer; the rest is Codex's.
145  const asked = request?.texts.filter(text => !isCommand(text)) ?? []
146  const replies = request ? await answered($, agentId, request.texts.filter(isCommand), type.pin) : []
147  for (const reply of replies) {
148    yield show(`${reply}\n\n`)
149    await note($, agentId, reply)
150  }
151  let report: Omit<CodexRun, 'tokens'> | undefined
152  // Whether Codex took this step's messages: it started their turn, or its
153  // question was answered. From then they are passed on, finished or not.
154  let isTaken = false
155  if (request) {
156    const prompt = asked.join('\n\n')
157    const pending = held.get(agentId)
158    const reply = pending && replyOf(pending.asked, prompt)
159    const cwd = await $.session.cwd()
160    try {
161      if (asked.length === 0) {
162        // Commands alone: Codex is not asked, and a question it asked still waits.
163      } else if (pending?.server.isEnded() || (!pending && expiredAnswer(lastAnswer(rows), prompt))) {
164        held.delete(agentId)
165        question = EXPIRED
166      } else if (pending && !reply) {
167        // Not an answer: ask again, Codex still waiting.
168        question = `That does not answer Codex.\n\n${questionOf(pending.asked)}`
169      } else if (pending && reply) {
170        held.delete(agentId)
171        server = pending.server
172        run.threadId = pending.threadId
173        await server.respond(pending.asked.id, reply)
174        isTaken = true
175        yield show(`answered Codex\n`)
176        await note($, agentId, 'answered Codex')
177      } else {
178        // Once Codex has taken a message the agent's session goes on, even
179        // where the conversation keeps no session line (a step stopped first).
180        const sessionId = given.length > 0 ? (request.sessionId ?? (await read($, runs))[agentId]?.threadId) : undefined
181        const first = sessionId === undefined
182        // The spawn prompt's flags hold for every run of the agent; they are
183        // not part of the task.
184        const flags = flagsOf(request.opening, type.pin)
185        if ('error' in flags) throw new Error(flags.error)
186        const set = threadParamsOf((await read($, options))[agentId] ?? {})
187        server = await open({ spawn: r => $.process.spawn(r), write: (p, t) => $.fs.write(p, t), stat: p => $.fs.stat(p) }, flags.args, cwd)
188        const where = flags.cwd ?? cwd
189        const config = { ...set.config, ...(await rootsOf(server, flags.addDirs)) }
190        const overrides = { ...set, ...(Object.keys(config).length > 0 ? { config } : {}) }
191        const thread = first
192          ? await server.call('thread/start', { cwd: where, ...(flags.ephemeral ? { ephemeral: true } : {}), ...overrides })
193          : await server.call('thread/resume', { threadId: sessionId, cwd: where, ...overrides })
194        // What the session runs on, as Codex reports it.
195        report = reportOf(thread)
196        run.threadId = report.threadId
197        // Known now, for /codex-status while the turn runs.
198        await record($, agentId, report)
199        model = report.model
200        yield show(`codex ${labelOf({}, { model: report.model, effort: report.effort ?? undefined })} · ${type.shown}\ncodex session ${run.threadId}\n\n`)
201        const images = flags.images.map(path => ({ type: 'localImage', path }))
202        const schema = flags.outputSchema ? { outputSchema: JSON.parse(await $.fs.read(flags.outputSchema)) } : {}
203        const opens = asked[0] === request.opening
204        const text = opens ? (flagsOf(prompt, type.pin) as Flags).prompt : prompt
205        const started = await server.call('turn/start', { threadId: run.threadId, input: [{ type: 'text', text }, ...images], ...schema })
206        isTaken = true
207        if (started?.turn?.id) working.set(agentId, { server, threadId: run.threadId!, turnId: started.turn.id })
208      }
209      while (server && !question) {
210        const message = await unlessAborted(next.signal, server.next())
211        if (!message) {
212          if (!next.signal.aborted) run.error = `codex app-server exited: ${server.stderr().trim().split('\n').at(-1) || 'no reason given'}`
213          break
214        }
215        if (message.id !== undefined) {
216          const asked = { id: message.id, method: message.method, params: message.params }
217          question = questionOf(asked)
218          if (question) held.set(agentId, { server, asked, threadId: run.threadId! })
219          else await server.respond(message.id, replyOf(asked, '')!)
220          continue
221        }
222        const step = apply(run, message.method, message.params)
223        if (step !== undefined) {
224          yield show(step)
225          await note($, agentId, step)
226        }
227        if (run.isDone) break
228      }
229    } catch (err) {
230      run.error = (err as Error).message
231      // Only a missing CLI, or one too old for app-server, earns the hint.
232      if (`${run.error}\n${server?.stderr() ?? ''}`.includes(NO_CODEX) || /unrecognized subcommand '?app-server/.test(`${run.error}\n${server?.stderr() ?? ''}`)) {
233        run.error += '. The codex mod needs the Codex CLI with `codex app-server` (0.159 or newer): `npm install -g @openai/codex`, then `codex login`.'
234      }
235    } finally {
236      working.delete(agentId)
237      if (server && !held.has(agentId)) server.close()
238      // A step cut off (interrupted, or closed) before Codex took its messages,
239      // asked or failed has passed nothing on: the next step sends them again.
240      const isCutOff = asked.length > 0 && !isTaken && !question && !run.error
241      if (!isCutOff) await update($, sent, all => ({ ...all, [agentId]: [...(all[agentId] ?? []), ...request.texts] }))
242      // What Codex reported this session runs with, and its token total.
243      await record($, agentId, report, run.tokens)
244    }
245  }
246
247  // Codex's own token counts, so the agent's row shows what the run cost.
248  const usage: TurnUsage | null = run.usage ? { ...run.usage, model } : null
249  const codexSaid = question || (asked.length > 0 ? answerOf(run) : undefined)
250  // A report that never reached the caller goes ahead of this step's answer;
251  // with nothing new (the engine ran the loop again) the loop says only that.
252  // A stopped step says only that it stopped, which is never such a report.
253  const undelivered = undeliveredReport(rows)
254  const message =
255    codexSaid === STOPPED
256      ? STOPPED
257      : request
258        ? [...(undelivered ? [undelivered] : []), ...replies, ...(codexSaid ? [codexSaid] : [])].join('\n\n')
259        : (undelivered ?? 'codex: nothing new to send to Codex.')
260  if (!handback) {
261    // The session line lets a follow-up resume this run (see requestOf).
262    const text = run.threadId ? `${message}\n\ncodex session ${run.threadId}` : message
263    yield { kind: 'text', index: 1, text }
264    yield { kind: 'stop', stopReason: 'end_turn', usage }
265    return { turnId: e.turnId, index: e.index, answer: text, toolUses: [], stopReason: 'end_turn', usage }
266  }
267  if (progress === '' && !question) yield show(run.error ? `codex: ${run.error}\n` : 'codex: nothing to run.\n')
268  const input = { message }
269  yield { kind: 'tool', index: 1, id: `toolu_codex_${crypto.randomUUID().replaceAll('-', '')}`, name: HANDBACK }
270  yield { kind: 'input', index: 1, json: JSON.stringify(input) }
271  yield { kind: 'stop', stopReason: 'tool_use', usage }
272  return { turnId: e.turnId, index: e.index, answer: progress, toolUses: [{ name: HANDBACK, input }], stopReason: 'tool_use', usage }
273}
274
275// `add-dir`: the config's writable roots plus the ones asked for, since a
276// thread's setting replaces the config's list rather than adding to it.
277async function rootsOf(server: Server, dirs: readonly string[]): Promise<Record<string, unknown>> {
278  if (dirs.length === 0) return {}
279  const { config } = await server.call('config/read', {})
280  return { 'sandbox_workspace_write.writable_roots': [...(config?.sandbox_workspace_write?.writable_roots ?? []), ...dirs] }
281}
282
hooks/view.ts 29 lines
1// The codex agent whose view is open, if any, and the reply to the last
2// /codex- command run in each such view. The band above the prompt sees the
3// view (a render hook may not write $.state, so it is kept here) and draws the
4// reply; a command reads the view. Display only: a reload drops both, and
5// the next draw sets the view again.
6
7let open: string | undefined
8const replies = new Map<string, string>()
9
10export function openView(): string | undefined {
11  return open
12}
13
14// Says whether the open view changed; leaving a view drops its reply.
15export function setOpenView(agentId: string | undefined): boolean {
16  if (agentId === open) return false
17  if (open !== undefined) replies.delete(open)
18  open = agentId
19  return true
20}
21
22export function replyIn(agentId: string): string | undefined {
23  return replies.get(agentId)
24}
25
26export function setReply(agentId: string, reply: string): void {
27  replies.set(agentId, reply)
28}
29
hooks/names.ts 9 lines
1// Codex model ids as people say them: `gpt-6.1-sol` reads `Sol 6.1`. An id
2// of another shape is shown as it is.
3export function nameOf(id: string): string {
4  const match = /^gpt-([\d.]+)-([a-z]+)$/i.exec(id)
5  if (!match) return id
6  const name = match[2]!
7  return `${name[0]!.toUpperCase()}${name.slice(1)} ${match[1]}`
8}
9
hooks/events.ts 130 lines
1// Reads what `codex app-server` reports while a turn runs. Each notification
2// becomes a line of progress for the agent's transcript, the run's answer is
3// its last agent message, and its usage is what the agent's row counts.
4
5import type { ModelUsage } from 'claude-code'
6
7import type { CodexRun } from '../types'
8
9export type Run = {
10  threadId?: string
11  answer?: string
12  error?: string
13  usage?: ModelUsage
14  isDone?: boolean
15  // The thread's token total before this turn's first request, which the
16  // turn's output is counted from.
17  before?: ModelUsage
18  // The thread's token total as Codex counts it, cached tokens inside input.
19  tokens?: CodexRun['tokens']
20}
21
22// What a run stopped before Codex finished answers: what Codex said on the
23// way is shown in the agent's view, never handed back as its report.
24export const STOPPED = 'codex: stopped before Codex finished.'
25
26// The run's answer: Codex's last message once the turn ended, its failure, or
27// that it was stopped first.
28export function answerOf(run: Run): string {
29  if (run.error) return `${run.answer ? `${run.answer}\n\n` : ''}codex failed: ${run.error}`
30  return run.isDone ? (run.answer ?? 'codex: the turn ended without a message.') : STOPPED
31}
32
33// What a thread/start or thread/resume answer says the session runs with.
34export function reportOf(answer: any): Omit<CodexRun, 'tokens'> {
35  const policy = answer?.approvalPolicy
36  return {
37    threadId: String(answer?.thread?.id),
38    model: String(answer?.model),
39    effort: answer?.reasoningEffort ?? null,
40    sandbox: String(answer?.sandbox?.type ?? 'unknown').replace(/[A-Z]/g, c => `-${c.toLowerCase()}`),
41    approvals: typeof policy === 'string' ? policy : JSON.stringify(policy),
42  }
43}
44
45// Codex counts cached tokens inside `inputTokens`; the engine's shape counts
46// them apart, as the Messages API does.
47export function usageOf(usage: any): ModelUsage {
48  const cached = Number(usage?.cachedInputTokens) || 0
49  return {
50    input_tokens: Math.max(0, (Number(usage?.inputTokens) || 0) - cached),
51    output_tokens: Number(usage?.outputTokens) || 0,
52    cache_read_input_tokens: cached,
53    cache_creation_input_tokens: Number(usage?.cacheWriteInputTokens) || 0,
54  }
55}
56
57function minus(a: ModelUsage, b: ModelUsage): ModelUsage {
58  return {
59    input_tokens: Math.max(0, a.input_tokens - b.input_tokens),
60    output_tokens: Math.max(0, a.output_tokens - b.output_tokens),
61    cache_read_input_tokens: Math.max(0, (a.cache_read_input_tokens ?? 0) - (b.cache_read_input_tokens ?? 0)),
62    cache_creation_input_tokens: Math.max(0, (a.cache_creation_input_tokens ?? 0) - (b.cache_creation_input_tokens ?? 0)),
63  }
64}
65
66// Splits streamed text into complete lines, keeping the unfinished tail.
67export function lines(buffer: string, text: string): { done: string[]; rest: string } {
68  const parts = (buffer + text).split('\n')
69  return { done: parts.slice(0, -1), rest: parts.at(-1) ?? '' }
70}
71
72// `/bin/zsh -lc 'cat go.mod'` reads `cat go.mod`.
73export function commandOf(command: string): string {
74  return /^\S*sh -lc (['"])(.*)\1$/s.exec(command)?.[2] ?? command
75}
76
77// Folds one notification into the run and returns the progress it shows.
78// Notifications of another thread (none, while the mod runs one) are not
79// this run's.
80export function apply(run: Run, method: string, params: any): string | undefined {
81  if (params?.threadId && run.threadId && params.threadId !== run.threadId) return undefined
82  const item = params?.item
83  switch (method) {
84    case 'item/started':
85      return item?.type === 'commandExecution' ? `$ ${commandOf(item.command)}\n` : undefined
86    case 'item/completed':
87      if (item?.type === 'agentMessage') {
88        run.answer = item.text
89        return `${item.text.trimEnd()}\n\n`
90      }
91      if (item?.type === 'commandExecution') return item.exitCode ? `  exit ${item.exitCode}\n` : undefined
92      if (item?.type === 'fileChange') return `edited ${(item.changes ?? []).map((c: any) => c.path).join(', ')}\n`
93      if (item?.type === 'webSearch') return `searched the web: ${item.query}\n`
94      if (item?.type === 'mcpToolCall') return `${item.server}.${item.tool}${item.error ? ' failed' : ''}\n`
95      return undefined
96    case 'guardianWarning':
97      return `auto review: ${params.message}\n`
98    case 'thread/tokenUsage/updated': {
99      // The agent's row counts as Claude Code counts a Claude subagent: the
100      // latest step's input (fresh, cache read and cache write) replaces the
101      // last, and each step's output adds to the outputs before it. So the
102      // turn reports its last request's input, which is Codex's current
103      // context, and the output the whole turn generated. The thread's
104      // running total, which counts every cached re-read, goes to
105      // /codex-status instead.
106      const total = usageOf(params.tokenUsage?.total)
107      const last = usageOf(params.tokenUsage?.last)
108      const raw = params.tokenUsage?.total
109      run.tokens = { input: Number(raw?.inputTokens) || 0, cached: Number(raw?.cachedInputTokens) || 0, output: Number(raw?.outputTokens) || 0 }
110      run.before ??= minus(total, last)
111      run.usage = { ...last, output_tokens: minus(total, run.before).output_tokens }
112      return undefined
113    }
114    case 'error':
115      // A retried error is one Codex recovers from; the turn says how it ended.
116      if (params.willRetry) return undefined
117      run.error = params.error?.message
118      return `error: ${run.error}\n`
119    case 'turn/completed': {
120      run.isDone = true
121      const turn = params.turn
122      if (turn?.status === 'completed') run.error = undefined
123      else run.error = turn?.error?.message ?? run.error ?? `the turn ended ${turn?.status ?? 'without a status'}`
124      return undefined
125    }
126    default:
127      return undefined
128  }
129}
130
hooks/questions.ts 109 lines
1// What Codex asks while a turn runs (an approval, a question for the user, an
2// MCP server's form) goes back to Claude as the agent's report, and Claude's
3// next message to the agent is the answer. These turn one into the other.
4
5import { commandOf } from './events'
6
7export type Asked = { id: number | string; method: string; params: any }
8
9export type Reply = { result: unknown } | { error: { code: number; message: string } }
10
11const APPROVAL = 'Reply "approve", "approve for session", "decline", or "cancel" (decline and stop the turn).'
12
13// The question as Claude reads it, or undefined for a request no person can
14// answer, which the mod refuses for Codex.
15export function questionOf(asked: Asked): string | undefined {
16  const p = asked.params ?? {}
17  const why = p.reason ? `\nReason: ${p.reason}` : ''
18  switch (asked.method) {
19    case 'item/commandExecution/requestApproval':
20      return `Codex asks to run:\n  ${commandOf(p.command ?? '(no command given)')}\nin ${p.cwd ?? 'its working directory'}${why}\n${APPROVAL}`
21    case 'item/fileChange/requestApproval':
22      return `Codex asks to change files${p.grantRoot ? ` under ${p.grantRoot}` : ''}.${why}\n${APPROVAL}`
23    case 'item/permissions/requestApproval':
24      return `Codex asks for more permissions: ${JSON.stringify(p.permissions)}${why}\nReply "approve", "approve for session", or "decline".`
25    case 'item/tool/requestUserInput': {
26      const questions = (p.questions ?? []) as any[]
27      const listed = questions.map((q, i) => {
28        const options = (q.options ?? []).map((o: any) => o.label ?? o.value ?? String(o)).join(' / ')
29        return `${i + 1}. ${q.question}${options ? ` (${options})` : ''}`
30      })
31      return `Codex asks:\n${listed.join('\n')}\nReply with the answer${questions.length > 1 ? 's, one line each, in order' : ''}.`
32    }
33    case 'mcpServer/elicitation/request':
34      return `The ${p.serverName} MCP server asks: ${p.message ?? JSON.stringify(p)}\nReply "accept" (with the JSON it asks for on the next lines, if any), "decline", or "cancel".`
35    default:
36      return undefined
37  }
38}
39
40const SAID = (reply: string) => reply.trim().toLowerCase().replace(/[.!]+$/, '')
41
42function decisionOf(reply: string): 'accept' | 'acceptForSession' | 'decline' | 'cancel' | undefined {
43  const said = SAID(reply)
44  if (/^(approve|accept|yes|y|ok|allow)( it)?$/.test(said)) return 'accept'
45  if (/^(approve|accept|allow) (for (the )?session|always)$/.test(said)) return 'acceptForSession'
46  if (/^(decline|deny|no|n|reject)$/.test(said)) return 'decline'
47  if (/^(cancel|stop|abort)$/.test(said)) return 'cancel'
48  return undefined
49}
50
51// Codex's response to what it asked, from Claude's reply; undefined when the
52// reply does not answer it, so the question is asked again.
53export function replyOf(asked: Asked, reply: string): Reply | undefined {
54  const p = asked.params ?? {}
55  switch (asked.method) {
56    case 'item/commandExecution/requestApproval':
57    case 'item/fileChange/requestApproval': {
58      const decision = decisionOf(reply)
59      return decision && { result: { decision } }
60    }
61    case 'item/permissions/requestApproval': {
62      const decision = decisionOf(reply)
63      if (!decision) return undefined
64      const isGranted = decision === 'accept' || decision === 'acceptForSession'
65      return {
66        result: {
67          permissions: isGranted ? p.permissions : {},
68          scope: decision === 'acceptForSession' ? 'session' : 'turn',
69        },
70      }
71    }
72    case 'item/tool/requestUserInput': {
73      const questions = (p.questions ?? []) as any[]
74      const said = reply.trim().split('\n').map(l => l.replace(/^\d+[.)]\s*/, '').trim()).filter(Boolean)
75      if (said.length === 0) return undefined
76      const answers: Record<string, { answers: string[] }> = {}
77      questions.forEach((q, i) => {
78        answers[q.id] = { answers: [questions.length === 1 ? reply.trim() : (said[i] ?? '')] }
79      })
80      return { result: { answers } }
81    }
82    case 'mcpServer/elicitation/request': {
83      const [first = '', ...rest] = reply.trim().split('\n')
84      const action = SAID(first)
85      if (action === 'decline' || action === 'cancel') return { result: { action, content: null } }
86      if (action !== 'accept') return undefined
87      if (rest.join('\n').trim() === '') return { result: { action, content: null } }
88      try {
89        return { result: { action, content: JSON.parse(rest.join('\n')) } }
90      } catch {
91        return undefined
92      }
93    }
94    default:
95      return { error: { code: -32601, message: `the codex mod cannot answer ${asked.method}` } }
96  }
97}
98
99// A report that is Codex's question, as questionOf words it.
100const ASKS = /^(?:That does not answer Codex\.\s+)?(?:Codex asks|The \S+ MCP server asks)/
101
102// Whether a message is a decision sent to a question whose Codex is gone (the
103// session was resumed, the mod reloaded, or Codex exited while it waited):
104// the agent's last report was the question and the message decides it. Such
105// a message is not passed to Codex as a new turn.
106export function expiredAnswer(lastReport: string | undefined, message: string): boolean {
107  return lastReport !== undefined && ASKS.test(lastReport) && decisionOf(message) !== undefined
108}
109
hooks/request.ts 168 lines
1import { STOPPED } from './events'
2
3// What a codex agent's loop is being asked: every message in its conversation
4// not yet passed on. Position is no guide, since the engine moves messages
5// (a message sent while the agent runs is placed twice: wrapped, in its
6// first turn even ahead of the spawn prompt, and as typed, maybe a turn
7// later), so the caller keeps the words already passed on (`sent`) and
8// messages are counted: the same words count as often as they stand in the
9// form, wrapped or as typed, that holds them more often. Sent twice, they are
10// asked twice. The copies not yet passed on go in the order they stand.
11
12export const HANDBACK = 'SubagentHandback'
13
14const SESSION = /^codex session ([0-9a-f-]{36})$/m
15
16// Text the engine adds to a subagent's loop that is not the caller's words:
17// reminders, handback nudges, and the marker an interruption leaves.
18function isEngineText(text: string): boolean {
19  return text.startsWith('<system-reminder>') || text.startsWith('[handback') || text.startsWith('[Request interrupted by user')
20}
21
22// A message sent to an agent may arrive wrapped in the engine's words: a
23// first line naming who sent it ("The user sent a new message while you were
24// working:", "The coordinator sent a message while you were working:",
25// "Another Claude session sent a message …"), the message, a blank line and
26// one paragraph of instructions to a Claude subagent. Codex gets the message
27// as it was sent.
28const WRAPPED = /^[^\n]* sent a (?:new )?message while you were working:\n([\s\S]*)\n\n[^\n]+$/
29
30// The sender's words, and whether the engine wrapped them.
31export function formOf(text: string): { words: string; isWrapped: boolean } {
32  const wrapped = WRAPPED.exec(text.trim())?.[1]
33  return wrapped === undefined ? { words: text.trim(), isWrapped: false } : { words: wrapped.trim(), isWrapped: true }
34}
35
36// One turn of the conversation: its words, and its tool calls with the text
37// of each one's result, once there is one.
38export type Row = {
39  role: 'user' | 'assistant'
40  text: string
41  texts: readonly string[]
42  toolUses: readonly { tool: string; input: Record<string, unknown>; result?: string }[]
43}
44
45type ApiBlock = { type: string; [field: string]: unknown }
46export type ApiTurn = { role: 'user' | 'assistant'; content: readonly ApiBlock[] }
47
48// The conversation as the model would be sent it. The engine's own rows
49// (`$.session.messages()`) leave out meta rows, and a message queued while the
50// agent ran is one, so the API form is the one that holds every request.
51export function rowsOf(turns: readonly ApiTurn[]): Row[] {
52  const results = new Map<unknown, string>()
53  for (const turn of turns) {
54    for (const block of turn.content) {
55      if (block.type === 'tool_result') results.set(block.tool_use_id, resultText(block.content))
56    }
57  }
58  return turns.map(turn => {
59    const texts = turn.content
60      .filter(b => b.type === 'text' && typeof b.text === 'string' && !isEngineText(b.text))
61      .map(b => b.text as string)
62    return {
63      role: turn.role,
64      text: texts.join('\n\n'),
65      texts,
66    toolUses: turn.content
67      .filter(b => b.type === 'tool_use')
68      .map(b => ({
69        tool: String(b.name),
70        input: (b.input ?? {}) as Record<string, unknown>,
71        ...(results.has(b.id) ? { result: results.get(b.id) } : {}),
72      })),
73    }
74  })
75}
76
77// `opening` is the spawn prompt, which carries the agent's options; `texts`
78// are the words this request passes on, for the caller to add to `sent`.
79export type Request = { prompt: string; opening: string; texts: string[]; sessionId?: string }
80
81// `spawned` is the prompt the agent was spawned with, as recorded at spawn.
82export function requestOf(rows: readonly Row[], sent: readonly string[], spawned?: string): Request | undefined {
83  const forms = rows.filter(r => r.role === 'user').flatMap(r => r.texts).map(formOf)
84  // The opening carries the agent's options: the spawn prompt as recorded,
85  // since the engine may place a later message ahead of it, even before it
86  // is first passed on (a first run cut off). Unrecorded (an agent spawned
87  // before the mod loaded), the first text passed to Codex, else the first.
88  const opening = spawned?.trim() ?? sent[0] ?? forms[0]?.words
89  if (opening === undefined) return undefined
90
91  // Where each copy of the same words stands, wrapped and as typed.
92  const places = new Map<string, { wrapped: number[]; typed: number[] }>()
93  forms.forEach(({ words, isWrapped }, at) => {
94    const place = places.get(words) ?? { wrapped: [], typed: [] }
95    place[isWrapped ? 'wrapped' : 'typed'].push(at)
96    places.set(words, place)
97  })
98  // The copies not yet passed on are the latest ones, in the form that holds
99  // more; they go in the order they stand in.
100  const left: { at: number; words: string }[] = []
101  for (const [words, { wrapped, typed }] of places) {
102    const copies = wrapped.length >= typed.length ? wrapped : typed
103    const count = Math.max(0, copies.length - sent.filter(text => text === words).length)
104    for (const at of copies.slice(copies.length - count)) left.push({ at, words })
105  }
106  const texts = left.sort((a, b) => a.at - b.at).map(copy => copy.words)
107  if (texts.length === 0) return undefined
108  // Until the opening is passed on, it is what Codex reads first.
109  const at = texts.indexOf(opening)
110  if (at > 0 && !sent.includes(opening)) texts.unshift(...texts.splice(at, 1))
111
112  let sessionId: string | undefined
113  for (const row of rows) {
114    if (row.role === 'assistant') sessionId = SESSION.exec(row.text)?.[1] ?? sessionId
115  }
116  return { prompt: texts.join('\n\n'), opening, texts, sessionId: sent.length > 0 ? sessionId : undefined }
117}
118
119// A tool result's content: a string, or text blocks.
120function resultText(content: unknown): string {
121  if (typeof content === 'string') return content
122  if (!Array.isArray(content)) return ''
123  return content.map(b => (b?.type === 'text' && typeof b.text === 'string' ? b.text : '')).join('\n')
124}
125
126// Claude Code's error for a call to a tool the loop does not have.
127const NO_SUCH_TOOL = 'No such tool available'
128
129// Whether this loop reports through a SubagentHandback call. An interactive
130// session gives subagents the tool and insists on it; a headless or SDK run
131// has none and takes the final text as the report. Nothing the loop can read
132// says which ahead of time, so it hands back until a handback has failed for
133// want of the tool. Any other failure (the person interrupted it) says the
134// tool is there.
135export function handsBack(rows: readonly Row[]): boolean {
136  return !rows.some(r => r.toolUses.some(u => u.tool === HANDBACK && u.result?.includes(NO_SUCH_TOOL)))
137}
138
139// Claude Code's result for a handback that reached the caller (2.1.287
140// records it inside a JSON object).
141const DELIVERED = 'Report delivered to your caller.'
142
143// The report of the last handback the loop sent, unless it was delivered: a
144// handback that failed (no such tool) or was interrupted never reached the
145// caller, so it is given again, once: a later turn that carries it as text
146// (the loop then reports as text) has given it. A stopped run's line is no
147// report: the one before it is the last.
148export function undeliveredReport(rows: readonly Row[]): string | undefined {
149  let last: { use: Row['toolUses'][number]; at: number } | undefined
150  rows.forEach((row, at) => {
151    for (const use of row.toolUses) {
152      if (use.tool === HANDBACK && typeof use.input.message === 'string' && use.input.message !== STOPPED) last = { use, at }
153    }
154  })
155  if (!last || last.use.result?.includes(DELIVERED)) return undefined
156  const report = last.use.input.message as string
157  return rows.slice(last.at + 1).some(r => r.role === 'assistant' && r.text.includes(report)) ? undefined : report
158}
159
160// What the agent last reported: its last turn's handback, or its text where
161// the loop reports as text (headless).
162export function lastAnswer(rows: readonly Row[]): string | undefined {
163  const row = rows.findLast(r => r.role === 'assistant')
164  if (!row) return undefined
165  const use = row.toolUses.findLast(u => u.tool === HANDBACK && typeof u.input.message === 'string')
166  return use ? (use.input.message as string) : row.text
167}
168