Folds a finished turn's work under a "Worked for" row, leaving the final answer, like Codex

Xenopus, like the lab frog it's named after, is where I experiment with how I work with AI coding agents. It's my personal setup for Claude Code and Codex, with the agents, global instructions and skills I use every day. Everything here is something I actually use, tested and revised as I go.
Out of the box, coding agents tend to do too much and claim too much. They add safeguards, options and abstractions nobody asked for, report work as tested when it wasn't, and drift from what you asked toward what they think you might want. Most of this setup exists to push back on that.
This setup puts quality first. When quality pulls against speed or token cost, it picks quality. It checks facts instead of answering from memory, proves changes on the real path, gets an independent review before finishing work where a mistake would be costly, and runs most specialist agents on the strongest models at high effort. Expect more tokens and longer turns than a stock setup, especially on larger or riskier tasks.
It doesn't spend for nothing, though. Simple tasks are done directly, and checks that can't change the result are skipped. If you'd rather trade some quality for speed or cost, lower the agents' models and effort levels (see Make it yours) or relax the review rule under "Subagent delegation" in the global instructions.
coordinate-specialists decides when delegating is worth it, which is less often than you'd think, since simple work is done directly.CLAUDE.md, AGENTS.md) load into every session. They set the working rules above and point to the skills that own each kind of work.Pieces refer to each other by name. Several agents work by the engineering skill, and some load a skill directly (device-runner loads the device skill, prompt-evaluator the prompt skill). Install the whole set rather than picking pieces out.
The setup keeps itself current as you use it, through improve-personal-customizations.
Every change is logged in ~/.agents/CUSTOMIZATIONS.md, along with checks that are still pending and ideas that were tried and rejected, so you can see what changed and why.
create-images to generate images.codex-reviewer agent.create-images scripts.The two folders mirror where each runtime looks for its files.
claude/ copy into ~/.claude
agents/ specialist subagents
CLAUDE.md global instructions
hooks/ queue-until-done.py
mods/ Claude Code mods: collapse-work, queue
skills/ skills
codex/ copy into ~/.codex, except skills/
AGENTS.md global instructions
agents/ specialist subagents (.toml)
skills/ copy into ~/.agents/skills
The Claude and Codex copies of each skill and agent say the same thing and differ only in how each runtime names things (for example the coordinate-specialists skill in Claude and $coordinate-specialists in Codex). Two pieces exist for one runtime only, the codex-reviewer agent (Claude) and the handoff-task skill (Codex).
Each block backs up the files it's about to replace into a dated folder in your home directory, then copies the setup in. Your own skills and agents with other names are left alone. You can install the Claude or Codex half on its own. Start a new session afterwards, since instructions and skills load when a session starts.
if (Test-Path "$env:TEMP\Xenopus") {
Remove-Item "$env:TEMP\Xenopus" -Recurse -Force
}
git clone https://github.com/FrogAi/Xenopus.git "$env:TEMP\Xenopus"
$backup = "$HOME\xenopus-backup-claude-$(Get-Date -Format yyyyMMdd-HHmmss)"
New-Item -ItemType Directory $backup | Out-Null
foreach ($item in "agents", "CLAUDE.md", "hooks", "mods", "skills") {
if (Test-Path "$HOME\.claude\$item") {
Copy-Item "$HOME\.claude\$item" $backup -Recurse
}
}
New-Item -ItemType Directory "$HOME\.claude" -Force | Out-Null
Copy-Item "$env:TEMP\Xenopus\claude\*" "$HOME\.claude" -Recurse -Force
if (Test-Path "$env:TEMP\Xenopus") {
Remove-Item "$env:TEMP\Xenopus" -Recurse -Force
}
git clone https://github.com/FrogAi/Xenopus.git "$env:TEMP\Xenopus"
$backup = "$HOME\xenopus-backup-codex-$(Get-Date -Format yyyyMMdd-HHmmss)"
New-Item -ItemType Directory $backup | Out-Null
foreach ($item in ".agents\skills", ".codex\agents", ".codex\AGENTS.md") {
if (Test-Path "$HOME\$item") {
Copy-Item "$HOME\$item" (Join-Path $backup ($item -replace "\\", "-")) -Recurse
}
}
New-Item -ItemType Directory "$HOME\.agents\skills", "$HOME\.codex" -Force | Out-Null
Copy-Item "$env:TEMP\Xenopus\codex\agents", "$env:TEMP\Xenopus\codex\AGENTS.md" "$HOME\.codex" -Recurse -Force
Copy-Item "$env:TEMP\Xenopus\codex\skills\*" "$HOME\.agents\skills" -Recurse -Force
rm -rf /tmp/Xenopus
git clone https://github.com/FrogAi/Xenopus.git /tmp/Xenopus
backup=~/xenopus-backup-claude-$(date +%Y%m%d-%H%M%S)
mkdir -p "$backup"
for item in agents CLAUDE.md hooks mods skills; do
if [ -e ~/.claude/$item ]; then
cp -R ~/.claude/$item "$backup"/
fi
done
mkdir -p ~/.claude
cp -R /tmp/Xenopus/claude/. ~/.claude/
rm -rf /tmp/Xenopus
git clone https://github.com/FrogAi/Xenopus.git /tmp/Xenopus
backup=~/xenopus-backup-codex-$(date +%Y%m%d-%H%M%S)
mkdir -p "$backup"
for item in .agents/skills .codex/agents .codex/AGENTS.md; do
if [ -e ~/$item ]; then
cp -R ~/$item "$backup"/$(echo $item | tr / -)
fi
done
mkdir -p ~/.agents/skills ~/.codex
cp -R /tmp/Xenopus/codex/agents /tmp/Xenopus/codex/AGENTS.md ~/.codex/
cp -R /tmp/Xenopus/codex/skills/. ~/.agents/skills/
The hook and mods need one more step each, covered in Hook and mods.
improve-personal-customizations. It's an empty ~/.agents/CUSTOMIZATIONS.md with four headings, Decided against or reverted, Pending, Recent changes and Twins and intentional differences.write-in-my-voice/references/voice-profile.md ships as a blank template with prompts to replace.develop-on-comma-device. Say in your global instructions which SSH alias is your development device and which is your driving device..toml (Codex). Change them to models you have access to.| Skill | What it's for |
|---|---|
coordinate-specialists | When to delegate to subagents and when not to, how to brief them, and how to maintain the agent library. |
create-images | Making, editing and repairing AI-generated images, from composing several real subjects into one scene to removing artifacts and seams. |
develop-on-comma-device | openpilot development on a comma three or 3X, covering device roles, safe bench testing, measurements and crash investigation. |
engineer-production-changes | The engineering method for all software work. Understand the system first, fix causes where they start, build the smallest complete design, prove it on the real path and report honestly. |
engineer-prompts | Writing, revising and evaluating prompts, skills, agent definitions and instruction files. |
explore-frontend-designs | Exploring and comparing UI design directions before one is chosen. |
handoff-task (Codex) | Handing a task to a fresh conversation, or recovering context from an earlier one. |
implement-frontend-designs | Building an approved design in a real project with exact visual and interaction fidelity. |
improve-personal-customizations | Keeps your agents, instructions and skills improving as you work. It saves your stated preferences (not task notes), fixes customizations that prove wrong, keeps the Claude and Codex copies in step, and tests changes in proportion to their risk. Keeps a record at ~/.agents/CUSTOMIZATIONS.md (see Make it yours). |
translate-content | Translation and localization, from prose to UI string catalogs. |
write-in-my-voice | Drafting messages that sound like you. Ships with a blank voice profile to fill in. |
write-release-notes | Public release notes and update posts that are accurate and fun to read. |
Specialists that the main session delegates to through coordinate-specialists. Each has one bounded job and reports back; none decides what to ship.
| Agent | What it does |
|---|---|
audience-reviewer | Reviews prose and visuals for what the intended reader would actually understand. |
behavior-tracer | Traces an execution or data path through code and configuration. |
check-runner | Runs assigned existing checks and reports what actually happened. |
cloud-auditor | Read-only audit of cloud provider state, usage and cost. |
codex-reviewer | Claude only. Gets a second opinion from Codex through its CLI. It sends your material to OpenAI, so it runs only when you ask for it. |
design-exploration-reviewer | Reviews frontend design alternatives for fit, quality and fair comparison. |
device-runner | Runs checks or retrieves logs on a comma device over SSH, with develop-on-comma-device. |
engineering-reviewer | Reviews a design or implementation for correctness, regressions and unnecessary complexity. |
evidence-auditor | Checks whether the evidence actually supports a claim of correctness, testing or completion. |
frontend-reviewer | Reviews an implemented interface against its accepted design and user flows. |
image-artifact-reviewer | Inspects a generated or edited image at full zoom for AI artifacts, wrong details and compositing seams, and reports them without editing the image. |
log-triager | Organizes logs or telemetry into event groups, counts and a timeline. |
performance-investigator | Measures latency, throughput and resource use with controlled comparisons. |
privacy-reviewer | Reviews data that leaves a device or system for personal information. |
prompt-evaluator | Assesses a prompt, skill or agent definition against its task and target model. |
reliability-reviewer | Reviews rollouts, recovery and operational behavior across versions and dependencies. |
repository-locator | Finds files, symbols, references and configuration in a repository. |
root-cause-investigator | Establishes a defect's trigger, mechanism and responsible boundary. |
security-reviewer | Reviews a design or change for reachable security weaknesses. |
source-extractor | Extracts specified facts or passages from files and web sources, with provenance. |
targeted-implementer | Implements one small, understood change and verifies it. |
translation-reviewer | Reviews translations for natural language, fidelity and complete coverage. |
verification-designer | Designs the checks that would actually prove a change works. |
These rely on Claude Code internals that can change between releases, so check them after updating.
hooks/queue-until-done.py. In the desktop app, it holds a message you send while Claude is working and delivers it when the current task ends. It depends on undocumented behavior of "continue": false; last confirmed on 2.1.284. Register it for both UserPromptSubmit and SessionEnd in ~/.claude/settings.json, using the script's full path. "hooks": {
"UserPromptSubmit": [{ "hooks": [{ "type": "command", "command": "python /path/to/.claude/hooks/queue-until-done.py" }] }],
"SessionEnd": [{ "hooks": [{ "type": "command", "command": "python /path/to/.claude/hooks/queue-until-done.py" }] }]
}
Use python3 instead of python where that's your interpreter's name (macOS and most Linux).
mods/collapse-work. Folds a finished turn's work under a "Worked for" line.mods/queue. A Codex-style queue for messages sent while Claude works, sent one per turn.The mods need Claude Code 2.1.287 or later and were tested in the terminal on 2.1.288. Load one with claude --plugin-dir ~/.claude/mods/queue. Use either the hook or mods/queue, not both, since they act on the same message. Each mod has tests in its tests folder, which you can run from the mod's folder with claude plugin test ..
Public domain (Unlicense).
hooks/register.js 386 lines1const WORK_ROWS = ['AssistantMessage', 'ToolUse', 'ToolResult', 'ToolGroup', 'TurnDuration']
2
3const KEPT_SESSIONS = 30
4
5function span(durationMs) {
6 const seconds = Math.floor(durationMs / 1000)
7 const units = [
8 [Math.floor(seconds / 3600), 'h'],
9 [Math.floor((seconds % 3600) / 60), 'm'],
10 [seconds % 60, 's'],
11 ]
12
13 const shown = units.filter(([amount]) => amount > 0)
14
15 if (shown.length === 0) {
16 return '0s'
17 }
18
19 return shown.map(([amount, unit]) => amount + unit).join(' ')
20}
21
22function textKey(text) {
23 let hash = 2166136261
24
25 for (let index = 0; index < text.length; index += 1) {
26 hash = Math.imul(hash ^ text.charCodeAt(index), 16777619)
27 }
28
29 return text.length + ':' + (hash >>> 0)
30}
31
32function rowId(e) {
33 if (e.component === 'ToolGroup') {
34 return e.props.calls[0]?.tool_use_id ?? e.requestId.split('collapsed-').pop()
35 }
36
37 if (e.component === 'ToolUse' || e.component === 'ToolResult') {
38 return e.props.tool_use_id
39 }
40
41 return e.requestId
42}
43
44function isFolded(turn) {
45 return !turn.isStopped && turn.answer.length > 0 && turn.rows[0] !== turn.answer[0]
46}
47
48function isFoldedOnDesktop(table, turn) {
49 if (!isFolded(turn)) {
50 return false
51 }
52
53 const answerRow = table.rowOfText.get(turn.textKeys[turn.answer[0]])
54 return answerRow?.turn === turn
55}
56
57function remember(table, turn) {
58 table.turns.push(turn)
59
60 for (const id of turn.rows) {
61 table.turnOfRow.set(id, turn)
62 }
63
64 if (turn.durationRow !== undefined) {
65 table.turnOfRow.set(turn.durationRow, turn)
66 }
67
68 for (const [id, key] of Object.entries(turn.textKeys)) {
69 let owner = { turn, id }
70
71 if (table.rowOfText.has(key)) {
72 owner = undefined
73 }
74
75 table.rowOfText.set(key, owner)
76 }
77}
78
79async function load($, table, startedSessionId) {
80 const sessionId = startedSessionId ?? (await $.session.id())
81
82 const saved = (await $.store.get('turns:' + sessionId)) ?? []
83
84 table.sessionId = sessionId
85
86 table.turns = []
87
88 table.turnOfRow = new Map()
89
90 table.rowOfText = new Map()
91
92 table.running = undefined
93
94 for (const turn of saved) {
95 remember(table, { textKeys: {}, ...turn, isOpen: false, host: Infinity })
96 }
97}
98
99async function save($, table) {
100 const savedTurns = []
101
102 for (const turn of table.turns) {
103 savedTurns.push({
104 rows: turn.rows,
105 answer: turn.answer,
106 textKeys: turn.textKeys,
107 durationRow: turn.durationRow,
108 durationMs: turn.durationMs,
109 isStopped: turn.isStopped,
110 })
111 }
112
113 const storedSessions = (await $.store.get('sessions')) ?? []
114 const sessions = storedSessions.filter((id) => id !== table.sessionId)
115
116 sessions.push(table.sessionId)
117
118 const dropped = sessions.splice(0, sessions.length - KEPT_SESSIONS)
119
120 for (const id of dropped) {
121 await $.store.delete('turns:' + id)
122 }
123
124 await $.store.set('sessions', sessions)
125
126 await $.store.set('turns:' + table.sessionId, savedTurns)
127}
128
129function foldToggle($, e, turn) {
130 const { Box, Button } = $.ui.resolve(e)
131
132 let chevron = ' ▸'
133
134 if (turn.isOpen) {
135 chevron = ' ▾'
136 }
137
138 const toggle = Button({
139 key: 'fold',
140 label: 'Worked for ' + span(turn.durationMs) + chevron,
141 plain: true,
142 dimColor: true,
143 onPress: () => {
144 turn.isOpen = !turn.isOpen
145
146 $.ui.invalidate('ui.render')
147 },
148 })
149
150 return Box({ marginTop: 1, children: [toggle] })
151}
152
153function stoppedLabel($, e, turn) {
154 const { Box, Text } = $.ui.resolve(e)
155
156 const label = Text({ dimColor: true, children: ['You stopped after ' + span(turn.durationMs)] })
157 return Box({ marginTop: 1, children: [label] })
158}
159
160// The desktop app draws tool calls itself and gives no thinking to draw, so there the fold only holds the text written before the
161// answer; a turn with none has nothing to open.
162function desktopHeader($, e, turn) {
163 const hasTextBeforeAnswer = Object.keys(turn.textKeys).some((id) => !turn.answer.includes(id))
164
165 if (hasTextBeforeAnswer) {
166 return foldToggle($, e, turn)
167 }
168
169 const { Box, Text } = $.ui.resolve(e)
170
171 const label = Text({ dimColor: true, children: ['Worked for ' + span(turn.durationMs)] })
172 return Box({ marginTop: 1, children: [label] })
173}
174
175export function register(on) {
176 const table = {
177 sessionId: undefined,
178 turns: [],
179 turnOfRow: new Map(),
180 rowOfText: new Map(),
181 running: undefined,
182 loaded: undefined,
183 isDetailed: false,
184 }
185
186 on('classic.SessionStart', { source: ['clear', 'resume', 'fork'] }, async ($, e, next) => {
187 table.loaded = load($, table, e.session_id)
188
189 await table.loaded
190
191 $.ui.invalidate('ui.render')
192
193 return next(e)
194 })
195
196 on('turn.start', async ($, e, next) => {
197 await (table.loaded ??= load($, table))
198
199 table.running = {
200 rows: [],
201 answer: [],
202 textKeys: {},
203 durationRow: undefined,
204 durationMs: 0,
205 isStopped: false,
206 isOpen: false,
207 host: Infinity,
208 }
209
210 return next(e)
211 })
212
213 on('session.append', async ($, e, next) => {
214 const turn = table.running
215
216 if (turn === undefined || e.agentId !== undefined) {
217 return next(e)
218 }
219
220 if (e.message.name === 'turn_duration') {
221 turn.durationRow = e.uuid
222 }
223
224 if (e.door !== 'response') {
225 return next(e)
226 }
227
228 for (const block of e.message.content) {
229 if (block.type === 'text') {
230 turn.rows.push(e.uuid)
231
232 turn.answer.push(e.uuid)
233
234 turn.textKeys[e.uuid] = textKey(block.text)
235 }
236
237 if (block.type === 'tool_use') {
238 turn.rows.push(block.id)
239
240 turn.answer = []
241 }
242 }
243
244 return next(e)
245 })
246
247 on('turn.complete', async ($, e, next) => {
248 const turn = table.running
249
250 if (turn === undefined || e.agentId !== undefined) {
251 return next(e)
252 }
253
254 turn.durationMs = e.durationMs
255
256 turn.isStopped = e.isAborted
257
258 const result = await next(e)
259
260 table.running = undefined
261
262 remember(table, turn)
263
264 $.ui.invalidate('ui.render')
265
266 await save($, table)
267
268 return result
269 })
270
271 on('ui.render', { component: 'UserMessage', surface: 'terminal' }, async ($, e, next) => {
272 if (
273 e.viewport?.isFullscreen !== false &&
274 e.props.origin.kind === 'composer' &&
275 e.props.isExpanded !== table.isDetailed
276 ) {
277 table.isDetailed = e.props.isExpanded
278
279 $.ui.invalidate('ui.render')
280 }
281
282 return next(e)
283 })
284
285 on('ui.render', { component: WORK_ROWS, surface: 'terminal' }, async ($, e, next) => {
286 if (e.viewport?.isFullscreen === false) {
287 return next(e)
288 }
289
290 await (table.loaded ??= load($, table))
291
292 const id = rowId(e)
293
294 const turn = table.turnOfRow.get(id)
295
296 if (turn === undefined || table.isDetailed) {
297 return next(e)
298 }
299
300 if (!turn.isStopped && (!isFolded(turn) || turn.answer.includes(id))) {
301 return next(e)
302 }
303
304 const index = turn.rows.indexOf(id)
305
306 if (index !== -1 && index < turn.host && e.component !== 'ToolResult') {
307 if (turn.host !== Infinity) {
308 $.ui.invalidate('ui.render')
309 }
310
311 turn.host = index
312 }
313
314 const isHost = index === turn.host && e.component !== 'ToolResult'
315
316 const { Box } = $.ui.resolve(e)
317
318 if (turn.isStopped) {
319 if (!isHost) {
320 return next(e)
321 }
322
323 const label = stoppedLabel($, e, turn)
324
325 const content = await next(e)
326 return Box({ flexDirection: 'column', children: [label, content] })
327 }
328
329 if (!isHost) {
330 if (turn.isOpen) {
331 return next(e)
332 }
333
334 return Box({ children: [] })
335 }
336
337 const children = [foldToggle($, e, turn)]
338
339 if (turn.isOpen) {
340 children.push(await next(e))
341 }
342
343 return Box({ flexDirection: 'column', children })
344 })
345
346 on('ui.render', { component: 'AssistantMessage', surface: 'desktop' }, async ($, e, next) => {
347 await (table.loaded ??= load($, table))
348
349 const row = table.rowOfText.get(textKey(e.props.text))
350
351 if (row === undefined) {
352 return next(e)
353 }
354
355 const { turn, id } = row
356
357 const { Box } = $.ui.resolve(e)
358
359 if (turn.isStopped) {
360 if (id !== Object.keys(turn.textKeys)[0]) {
361 return next(e)
362 }
363
364 const label = stoppedLabel($, e, turn)
365
366 const content = await next(e)
367 return Box({ flexDirection: 'column', children: [label, content] })
368 }
369
370 if (!isFoldedOnDesktop(table, turn)) {
371 return next(e)
372 }
373
374 if (id === turn.answer[0]) {
375 const header = desktopHeader($, e, turn)
376 return Box({ flexDirection: 'column', children: [header, await next(e)] })
377 }
378
379 if (turn.answer.includes(id) || turn.isOpen) {
380 return next(e)
381 }
382
383 return Box({ children: [] })
384 })
385}
386