SLOPSHOPPER

harness

Skills and guard hooks: ai-code-cleanup, interrogate, shape-task, verify-change, setup, a shared-worktree git guard, and a guard against reading API-key config…

newtoastpromptprocesstimeragents
v0.4.12MITupdated 2026-10-04Paradox07127/claude-utopia/plugins/harness
A shopper browsing a rack in a slop shop
README

claude-utopia

Under active development. This repo is updated continuously, and interfaces, names and defaults can change between versions. Star or watch the repo to follow the changes.

简体中文

Four Claude Code plugins built on mods (function hooks), plus two agent templates.

PluginWhat it does
dashboardA status band above the prompt for running subagents and mmrun reviews, a line under the prompt with context use and the tightest rate-limit window, and a seven-page workbench (Overview / Agents / Reviews / GPU / Timeline / Usage / Progress) opened with /dashboard, /subagents, /mmrun, /gpu, /timeline. The Reviews, GPU and Progress tabs show once they have data: an ~/.claude/mmruns folder, a GPU host, a progress board. The Timeline page draws the main loop's turns as a waterfall of model requests and tool calls, with a hotspot view of the slowest tools and turns. Renders subagent cards, test summaries and blocked-command notices in the transcript.
harnessSkills: ai-code-cleanup, interrogate, shape-task, verify-change, setup. Guards: refuse subagents on blocked models, refuse git commands that discard uncommitted work in the shared main worktree, refuse tool calls that would print API-key config files (~/.claude.json and its backups, Claude settings, Codex config and auth, grok auth, OpenViking config) or an environment dump (env, printenv, export -p, set, …) into the transcript, run /compact when the main thread idles until the prompt cache is about to expire.
mmCross-model code review and delegation: /mm:review runs codex / grok / agy in parallel as read-only reviewers, /mm:run hands a task to another model in its own worktree. Ships the mmrun CLI. A guard refuses reading mmrun's *.raw event streams (except a lone tail of 50 lines or fewer) and moves a foreground mmrun wait to the background.
progressA per-project progress board. After each turn you typed that changed files or made a commit, the plugin asks Sonnet to record the work as 1–3 nodes (title, summary, status, kind, links to earlier nodes); the main model spends no tokens on it. The board is kept under ~/.claude/progress/ or committed with the project, as you choose once per project. The plugin mirrors the board to an optional Artifact canvas. Works without dashboard; with it, the workbench's Progress page lists the board.
agents/worker and researcher subagent templates to copy into ~/.claude/agents/.

Screenshots

The status band above the prompt while a worker subagent runs, with the transcript cards for the dispatched agent and the test summary. Context use and the rate-limit window show only in the line under the prompt:

Status band and transcript cards in the terminal

The workbench's Timeline page: each model request of the last turn, its tool calls and the critical path.

Workbench Timeline page in the terminal

Both screenshots come from a real terminal session with language set to English. With zh-CN, the same UI is drawn in Chinese; see 简体中文.

Requirements

  • Claude Code 2.1.287 or later (mods are on by default from that version). Drawing works in the terminal and the Desktop Code tab; the VS Code panel and claude -p run the hooks without drawing.
  • harness: python3.
  • mm: at least one of the codex, grok or agy CLIs, plus bash, git, jq, python3. The grok and agy read-only fence uses sandbox-exec, so it is macOS only; codex uses its own sandbox.
  • dashboard GPU page: key-based ssh to the hosts and nvidia-smi (or tegrastats) on them.
  • Optional: /mm:run suggests the /impeccable skill for frontend tasks, and verify-change suggests the codegraph_impact MCP tool. Neither ships here; they are used when installed, and the plugins work without them.

Install

Let Claude do it: paste this into a Claude Code session.

Fetch and follow the instructions in https://raw.githubusercontent.com/Paradox07127/claude-utopia/main/INSTALL.md

Or by hand:

claude plugin marketplace add Paradox07127/claude-utopia
claude plugin install dashboard@claude-utopia
claude plugin install harness@claude-utopia
claude plugin install mm@claude-utopia
claude plugin install progress@claude-utopia

Then start a new Claude Code session. Run the setup skill (/harness:setup) at any time to review and change the options below.

Options

Set with /plugin configure <plugin>@claude-utopia, or claude plugin configure <plugin>@claude-utopia --values-stdin with a JSON object of strings. Changes take effect in the next session.

PluginKeyDefaultMeaning
dashboardlanguageautoUI language: auto, zh-CN or en. auto follows the settings language, then LC_ALL / LANG, else English.
dashboardgpuHostsemptyComma-separated ssh hosts; ssh <host> … in Bash opens that host's GPU page.
dashboardcacheTtlMinutes60The fallback prompt cache TTL: after this many idle minutes the band warns that the next message rewrites the prompt cache. Used only until the real TTL is known; a TTL read from the transcript or reported by a model switch overrides it.
dashboardtoastPeerAskstrueToast when another Claude Code session asks a permission or its turn fails. Other sessions' toasts show only in the session you typed in last within two minutes, or in every session when you typed in none.
dashboardtoastPeerRepliestrueToast when another session replies after a turn of two minutes or more. Shown by the same rule.
dashboardaskSoundfalsePlay a short chime with the toast of another session asking a permission or failing. Needs toastPeerAsks.
dashboardtoastRunstrueToast when an mmrun model returns, fails or goes stale.
harnesslanguageautoSame as above, for the harness toast.
harnessblockedSubagentModelsemptyComma-separated; a subagent whose model name contains any entry is refused. Empty allows all.
harnesssharedTreeGitGuardtrueIn the main worktree, refuse git commands that discard uncommitted changes or rewrite HEAD: checkout <path>, checkout --force / -f, restore (except --staged alone), stash (except list, show, create), clean (except -n / --dry-run), switch --discard-changes / --force / -f, reset --hard, commit --amend. Linked worktrees are exempt.
harnessidleCompacttrueWhen the main thread sits idle until just before the prompt cache expires, run /compact automatically: with a 1h cache 10 minutes before, at 100k tokens or more; with a 5m cache 1 minute before, at 200k or more. The TTL is read from the transcript.
mmreviewModelscodex,grokModels /mm:review uses when no --models is given.

progress has no options.

Skills, command docs and every instruction the plugins send to the model are in English; skill descriptions also carry Chinese trigger words so Chinese prompts match them. Claude replies in your language. Only the drawn UI follows language.

Privacy and trust

Mods run in-process with your user permissions, like any plugin hook. mm sends the code under review to the external model CLIs you have installed, under their own accounts and terms. Nothing here phones home.

harness hides the built-in general-purpose agent when worker.md or researcher.md exists in ~/.claude/agents/ or the project's .claude/agents/.

License

MIT

Source 3 files
hooks/register.tsx 225 lines
1import { atom, read, update } from 'claude-code'
2import type { EngineInterface, HookFailure, Register, Timer } from 'claude-code'
3import type { HarnessIdle } from '../types'
4import { pickLang, strings } from './i18n'
5import type { Lang } from './i18n'
6
7const AGENT_DEFS = ['worker.md', 'researcher.md']
8
9const MINUTE_MS = 60_000
10const CACHE_TTL_MS = { '5m': 5 * MINUTE_MS, '1h': 60 * MINUTE_MS } as const
11// leadMs: the compaction request itself reads the cached prefix, so it has to land this long before the cache expires.
12// minTokens: below it a compaction saves little on the next request yet loses the conversation's detail.
13const COMPACT_RULE = {
14  '1h': { leadMs: 10 * MINUTE_MS, minTokens: 100_000 },
15  '5m': { leadMs: MINUTE_MS, minTokens: 200_000 },
16} as const
17// The last main-thread reply sits within this tail; $.fs.read refuses files over 4 MiB and transcripts grow past that.
18const TRANSCRIPT_TAIL_BYTES = 256 * 1024
19const TAIL_TIMEOUT_MS = 5_000
20// Sits in the prompt: any change to it spends the prompt cache.
21const ASK_GUIDANCE =
22  'Ask only what blocks you. Prefer one question per dialog, at most two. Give 2–3 distinct options, each with a one-line consequence, and mark one "(Recommended)" with its reason. Say which default you will assume for anything you do not ask. For an open-ended question, ask in plain prose instead of options.'
23
24type CacheTtl = keyof typeof CACHE_TTL_MS
25type TranscriptLine = {
26  type?: string
27  isSidechain?: boolean
28  message?: { usage?: { cache_creation?: { ephemeral_1h_input_tokens?: number; ephemeral_5m_input_tokens?: number } } }
29}
30
31const idle = atom({ plugin: 'harness', key: 'idle' } as const, { lastReplyAt: null, lastStepAt: null, ttl: null } as HarnessIdle)
32
33// Module timers die with the module on reload; session.start fires again then and re-arms it from $.state.
34let idleTimer: Timer | null = null
35
36// Lives here, not in i18n.ts: validate only follows $ into functions declared in the same file.
37async function resolveLang($: EngineInterface, option: unknown): Promise<Lang> {
38  const settings = await $.settings.read()
39  return pickLang(option, settings.language, (await $.env.get('LC_ALL')) || (await $.env.get('LANG')))
40}
41
42function logFailure($: EngineInterface, event: string, error: HookFailure): void {
43  $.ui.log(`harness: ${event} hook failed (${error.kind}): ${error.message ?? 'no message'}`, { to: 'debug' })
44}
45
46/** The rule for the stored TTL; state saved before the field existed lacks it, and unknown means 1h. */
47async function compactRule($: EngineInterface): Promise<{ ttlMs: number; leadMs: number; minTokens: number }> {
48  const ttl = (await read($, idle)).ttl ?? '1h'
49  return { ttlMs: CACHE_TTL_MS[ttl], ...COMPACT_RULE[ttl] }
50}
51
52/** The cache TTL of the last main-thread reply that wrote the cache, or null when the tail holds none. */
53function transcriptTtl(tail: string): CacheTtl | null {
54  const lines = tail.split('\n')
55  for (let i = lines.length - 1; i >= 0; i -= 1) {
56    let line: TranscriptLine | null
57    try {
58      line = JSON.parse(lines[i] ?? '') as TranscriptLine | null
59    } catch {
60      // The first line of a `tail -c` read is usually cut, the last may still be being written.
61      continue
62    }
63    if (line?.type !== 'assistant' || line.isSidechain === true) continue
64    const written = line.message?.usage?.cache_creation
65    if ((written?.ephemeral_1h_input_tokens ?? 0) > 0) return '1h'
66    if ((written?.ephemeral_5m_input_tokens ?? 0) > 0) return '5m'
67  }
68  return null
69}
70
71/** Compacts a context of at least the TTL's minTokens; a compaction refused mid-turn is not retried. */
72async function compactIdle($: EngineInterface, language: unknown): Promise<void> {
73  idleTimer = null
74  if (((await $.session.usage()).context.tokens ?? 0) < (await compactRule($)).minTokens) return
75  const done = await $.session.compact().catch(() => undefined)
76  if (done === undefined || done.skip !== undefined) return
77  await update($, idle, was => ({ ...was, lastReplyAt: null }))
78  $.ui.toast(strings[await resolveLang($, language)].idleCompacted)
79}
80
81/** Sets the one timer for the TTL's leadMs before the cache expires. */
82async function armIdle($: EngineInterface, language: unknown): Promise<void> {
83  idleTimer?.cancel()
84  idleTimer = null
85  const { lastReplyAt } = await read($, idle)
86  if (lastReplyAt === null) return
87  const { ttlMs, leadMs } = await compactRule($)
88  const left = Math.max(0, lastReplyAt + ttlMs - leadMs - (await $.clock.now()))
89  idleTimer = $.clock.after(left, () => void compactIdle($, language))
90}
91
92export const register: Register = (on, options) => {
93  const blocked = String(options.blockedSubagentModels).split(',').map(name => name.trim()).filter(name => name !== '')
94  on('agent.spawn', ($, e, next) => {
95    const model = e.model?.toLowerCase()
96    const hit = model === undefined ? undefined : blocked.find(name => model.includes(name.toLowerCase()))
97    return hit !== undefined
98      ? { deny: `harness: subagents never use ${hit}. Use opus, or omit model and let the agent definition decide.` }
99      : next(e)
100  }).catch(($, e, next) => {
101    // It cannot tell whether the model was blocked, so it fails closed: undefined would leave the hook absent and let the spawn through.
102    logFailure($, 'agent.spawn', next.error)
103    return { deny: 'harness: the subagent model check failed, so the spawn is refused. Retry, or omit model and let the agent definition decide.' }
104  })
105
106  // Stats per offer instead of caching at session.start: the root moves with /cd and worktree moves.
107  on('agent.offer', { agent: 'general-purpose' }, async ($, e, next) => {
108    const root = await $.session.root().catch(() => undefined)
109    const home = await $.env.get('HOME')
110    const dirs = [root, home].filter(dir => dir !== undefined).map(dir => `${dir}/.claude/agents`)
111    for (const dir of dirs) {
112      for (const name of AGENT_DEFS) {
113        const st = await $.fs.stat(`${dir}/${name}`).catch(() => undefined)
114        if (st?.kind === 'file') return { isOffered: false }
115      }
116    }
117    return next(e)
118  }).catch(($, e, next) => {
119    logFailure($, 'agent.offer', next.error)
120    return next(e)
121  })
122
123  on('tool.describe', { tool: 'AskUserQuestion' }, async ($, e, next) => {
124    const beneath = await next(e)
125    return { ...beneath, description: `${beneath.description}\n\n${ASK_GUIDANCE}` }
126  }).catch(($, e, next) => {
127    logFailure($, 'tool.describe', next.error)
128    return next(e)
129  })
130
131  if (!options.idleCompact) return
132
133  on('session.start', async ($, e, next) => {
134    await armIdle($, options.language)
135    return next(e)
136  }).catch(($, e, next) => {
137    logFailure($, 'session.start', next.error)
138    return next(e)
139  })
140
141  // A request that reads the cache refreshes it, so the cache expires counting from when the last main request was sent.
142  on('turn.step', async function* ($, e, next) {
143    if (e.agentId === undefined) {
144      const lastStepAt = await $.clock.now()
145      await update($, idle, was => ({ ...was, lastStepAt }))
146    }
147    return yield* next(e)
148  }).catch(async function* ($, e, next) {
149    logFailure($, 'turn.step', next.error)
150    return yield* next(e)
151  })
152
153  on('turn.complete', async ($, e, next) => {
154    if (e.agentId === undefined) {
155      const now = await $.clock.now()
156      await update($, idle, was => ({ ...was, lastReplyAt: was.lastStepAt ?? now }))
157      await armIdle($, options.language)
158    }
159    return next(e)
160  }).catch(($, e, next) => {
161    logFailure($, 'turn.complete', next.error)
162    return next(e)
163  })
164
165  on('prompt.submit', async ($, e, next) => {
166    idleTimer?.cancel()
167    idleTimer = null
168    await update($, idle, was => ({ ...was, lastReplyAt: null, lastStepAt: null }))
169    return next(e)
170  }).catch(($, e, next) => {
171    logFailure($, 'prompt.submit', next.error)
172    return next(e)
173  })
174
175  // A /clear, resume or exit ends the conversation the timer was armed for; session.start does not fire for the next one.
176  on('session.end', async ($, e, next) => {
177    idleTimer?.cancel()
178    idleTimer = null
179    await update($, idle, was => ({ ...was, lastReplyAt: null, lastStepAt: null }))
180    return next(e)
181  }).catch(($, e, next) => {
182    logFailure($, 'session.end', next.error)
183    return next(e)
184  })
185
186  // A compaction keeps the conversation (compactIdle already leaves no timer); startup is left to session.start.
187  on('classic.SessionStart', async ($, e, next) => {
188    if (e.agent_id !== undefined || e.source === 'compact' || e.source === 'startup') return next(e)
189    idleTimer?.cancel()
190    idleTimer = null
191    await update($, idle, was => ({ ...was, lastReplyAt: null, lastStepAt: null }))
192    return next(e)
193  }).catch(($, e, next) => {
194    logFailure($, 'classic.SessionStart', next.error)
195    return next(e)
196  })
197
198  on('classic.PostModelSwitch', async ($, e, next) => {
199    if (e.agent_id !== undefined) return next(e)
200    await update($, idle, was => ({ ...was, ttl: e.cache_ttl }))
201    await armIdle($, options.language)
202    return next(e)
203  }).catch(($, e, next) => {
204    logFailure($, 'classic.PostModelSwitch', next.error)
205    return next(e)
206  })
207
208  // The transcript tells the TTL the main thread really gets (5m on an API key or past the plan limit) without a model switch.
209  on('classic.Stop', async ($, e, next) => {
210    if (e.agent_id !== undefined) return next(e)
211    if (e.transcript_path !== '') {
212      const tail = await $.process.run(['tail', '-c', String(TRANSCRIPT_TAIL_BYTES), e.transcript_path], { timeoutMs: TAIL_TIMEOUT_MS })
213      const ttl = tail.exitCode === 0 ? transcriptTtl(tail.stdout) : null
214      if (ttl !== null) {
215        await update($, idle, was => ({ ...was, ttl }))
216        await armIdle($, options.language)
217      }
218    }
219    return next(e)
220  }).catch(($, e, next) => {
221    logFailure($, 'classic.Stop', next.error)
222    return next(e)
223  })
224}
225
hooks/i18n.ts 21 lines
1export type Lang = 'zh-CN' | 'en'
2
3const CHINESE = /^zh\b|chinese|中文|汉语|漢語|简体|繁體|mandarin/i
4
5/** `language` userConfig → UI language; `auto` follows settings `language`, then the locale, else English. */
6export function pickLang(option: unknown, settingsLanguage: unknown, locale: string | undefined): Lang {
7  if (option === 'zh-CN' || option === 'en') return option
8  if (typeof settingsLanguage === 'string' && settingsLanguage.trim() !== '') return CHINESE.test(settingsLanguage.trim()) ? 'zh-CN' : 'en'
9  return locale !== undefined && /^zh/i.test(locale) ? 'zh-CN' : 'en'
10}
11
12const zh = {
13  idleCompacted: '空闲将满缓存时效,已自动压缩上下文',
14}
15
16const en: typeof zh = {
17  idleCompacted: 'Idle near cache expiry: context compacted',
18}
19
20export const strings: Record<Lang, typeof zh> = { 'zh-CN': zh, en }
21
types/index.d.ts 18 lines
1/** The main thread's idle time, for compacting before the prompt cache expires. */
2export type HarnessIdle = {
3  /** Epoch ms the cache expiry counts from (the main thread's last request); null when no idle waits to be compacted. */
4  lastReplyAt: number | null
5  /** Epoch ms the main thread's last model request was sent; missing in state saved before it existed. */
6  lastStepAt: number | null
7  /** The cache TTL the transcript or a PostModelSwitch reported; null (or missing) keeps the default 1h. */
8  ttl: '5m' | '1h' | null
9}
10
11declare module 'claude-code' {
12  interface PluginState {
13    harness: {
14      idle: HarnessIdle
15    }
16  }
17}
18