SLOPSHOPPER

compost

One software-engineering workflow composted from Matt Pocock's skills, superpowers, pstack, Imbue's blueprint, and autonomous-sdlc: specs with BDD acceptance…

newcommandpromptprocessnetwork
★ 7v0.4.0MITupdated 2026-10-07JoshuaOliphant/claude-plugins/plugins/compost
A shopper browsing a rack in a slop shop
README

compost

One software-engineering workflow for Claude Code, made by breaking down several skill collections and rebuilding the best of each: Matt Pocock's skills, obra's superpowers, Lauren Tan's pstack, Imbue's blueprint and code guardian, and the autonomous-sdlc and stick-shift plugins it replaces. NOTICE lists the sources compost takes text from; pile.toml records which upstream files fed which skill.

How it works

There is no state machine. Each skill ends with Next moves, naming the skills that usually follow and when, so Claude takes the obvious path or steps sideways into diagnosis or design. Once issues exist, compost:implement runs build, verify, and review for each one on its own, with review done by subagents every time, and records its rulings on the issue.

SkillWhen
setupOnce per repo: issue tracker, domain docs, test gates, AGENTS.md migration
specA fuzzy idea becomes user stories with Given/When/Then acceptance criteria
sliceA spec becomes tracer-bullet issues in the tracker
implementThe entry point once issues exist; picks the flow from there
buildOne issue, with its tests fitted into the existing suite
verifyBefore any claim of done: gates, AC-to-test map, the proof ladder
reviewAfter verify, always: /compost:review-changes, then weigh the findings
finishIntegrate: re-run on the merge target, PR recap, clean up
diagnoseAny bug or failing test, before a fix
deepenModule and interface design, refactoring for testability
pauseStop at a safe point and pick the work back up later
fan-outCompeting designs or hypotheses, tried in parallel and judged
turnCheck the upstream sources for changes worth working in
skill-driftDaily: turn repeated skill drift, judged by the skill-drift mod, into proposed skill fixes

canon/ holds short essays on the ideas the skills lean on: tracer bullets, ubiquitous language, deep modules, seams, expand-contract, proving it works, and more. Skills link to the ones they use.

Work is tracked in the repo's issue tracker (GitHub by default), recorded by compost:setup in docs/agents/issue-tracker.md. Domain vocabulary lives in CONTEXT.md, decisions in docs/adr/.

Jev tools

Many steps in the workflow are narrow judgments over lists that code can already produce: which existing test an acceptance criterion belongs in, whether a review finding is worth a skeptic subagent, whether an upstream change needs reading. compost hands those to TypeSafe Jev, which answers typed questions (Choice, Noul, Score) in a second or two for a fraction of a cent, so the frontier model spends its tokens on the rest.

Skills call uv run ${CLAUDE_PLUGIN_ROOT}/scripts/jev.py <tool> --input <file.json>; jev.py list names them all. Code gathers the candidates, Jev picks or checks, and thresholds in code decide. Results are cached by content hash in ~/.cache/compost/. The key comes from TYPESAFE_API_KEY or the macOS Keychain item typesafe. Without one, a tool exits 3 and the skill makes the call itself, so Jev speeds compost up and never blocks it. jev.py status checks for a key without calling the API; compost:setup runs it and tells you how to store one.

ToolPrimitiveUsed byLive eval (cases)
find-testChoice + Noulbuild, diagnoseright test 33/34, extend-or-add 33/34 (34)
duplicate-testNoulbuild23/23 (23 pairs)
ac-exercisedNoulverify20/22 (22)
test-valueScoreverify24/24 (24 uncovered blocks)
claim-backedNoulverify22/22 (22 claims)
triage-findingChoicereviewer, review-changesserious findings sent to a skeptic 15/15, minor spared 8/11 (26)
review-riskScorereview-changes, review24–25/28 structural calls (28 files)
canon-pickChoicereviewerprecision and recall ~0.6 (22)
locateChoice + Noulreviewer, specanchored 17/18, absent 8/8 (26)
spec-classChoicespec21/22, never lighter than labeled (22)
question-valueScorespec24–25/25 (25 questions)
ac-qualityScore ×3spec24/24 (24)
adr-worthyNoul ×3spec, deepen, implement24/24 (24 rulings)
classify-changeChoiceturnevery adopt/adapt read 4/4, no ignore recorded wrongly, 15/26 ignores skipped (30)
routeChoiceturn, description changes27/30 (30 real prompts)
rulings-lintNoul ×8turn, compost's own text10/13 violations, 2–4 false flags (36 passages)
stop-guardNoul ×2Stop hook in implement runs23/23 (23 final messages)
skill-drift modNoul ×3every skill-using turndeviated recall 10/11, precision 0.77–0.83; missing guidance 2/2 and corrected 1/1 flagged, thresholds provisional (60 turns)

compost also ships the skill-drift mod (hooks/skill-drift.ts), the collector behind compost:skill-drift. After each turn that loaded a skill it sends one Jev request per skill loaded, asking whether the turn deviated from the skill or lacked its guidance; when your next prompt follows such a turn it sends one more per skill, asking whether you are correcting that work. A request averages about 5,000 input tokens, a fraction of a cent. Hits go to ~/.claude/skill-drift/, and /drift lists them by skill. It never shows anything or holds a prompt back: your prompt enters before the correction check runs, and a TypeSafe outage is skipped. Without a key it asks and writes nothing. It is an in-process hooks module, so it loads only where Claude Code runs plugin hooks modules (CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 while they roll out); without them compost works as before and the log stays empty.

Thresholds were chosen from the same labeled cases they are measured on, so treat the numbers as upper bounds until the tools have run on real work. The labeled cases are drawn from private projects and sessions, so they live in a separate private repo, compost-evals. With a checkout, COMPOST_EVALS=<checkout>/evals uv run --group dev pytest -m jev re-runs every eval against the live API for a few cents; without one, those tests skip.

Install

/plugin marketplace add joshuaoliphant/claude-plugins
/plugin install compost@oliphant-plugins

Then run /compost:setup in each repo.

compost replaces several collections that cover the same stages, listed in the [replaces] table of pile.toml. Leaving them on gives Claude two answers to every request, so setup finds the ones still active on the machine and, with your yes, turns them off: it disables their plugins and sets their personal and skills-CLI skills to "off" under skillOverrides in ~/.claude/settings.json. Plugins synced from claude.ai have to be turned off there. Nothing is deleted, so turning one back on is a settings change. To check a machine directly: uv run scripts/pile.py replaced [--apply].

Development

From plugins/compost (the 100% coverage gate applies when pytest runs from there):

uv run --group dev pytest        # tooling tests (100% coverage) and checks on the shipped files
uv run scripts/pile.py status    # upstream changes since each source's pin (Python 3.11+ via uv)
Source 2 files
hooks/skill-drift.ts 194 lines
1// ABOUTME: Watches each skill the model loads and asks TypeSafe Jev whether the turn deviated from it, lacked its guidance, or drew a correction.
2// ABOUTME: Hits go quietly to ~/.claude/skill-drift/<date>-<session>.jsonl for compost:skill-drift to review; /drift lists them.
3import type { EngineInterface, Register, SessionMessage } from 'claude-code'
4
5import judgments from './skill-drift-judgments.ts'
6
7const ENDPOINT = 'https://api.typesafe.ai/v1/systemone'
8const TYPED_BY_A_PERSON = new Set(['composer', 'bridge'])
9const SKILL_CHARS = 8000
10const TURN_CHARS = 12000
11const EVIDENCE_CHARS = 4000
12const { thresholds: THRESHOLDS, turn: TURN_QUESTIONS, next_message: NEXT_MESSAGE_QUESTIONS } = judgments
13
14export type DriftKind = keyof typeof TURN_QUESTIONS | 'corrected'
15export type Hit = {
16  skill: string
17  kind: DriftKind
18  p: number
19  at: number
20  session: string
21  turn: string
22  skill_hash: string
23  prompt: string
24  evidence: string
25}
26type Watched = { skill: string; text: string; trace: string; prompt: string; turn: string }
27
28async function apiKey($: EngineInterface): Promise<string | undefined> {
29  const fromEnv = await $.env.get('TYPESAFE_API_KEY')
30  if (fromEnv) return fromEnv
31  const keychain = await $.process.run(['security', 'find-generic-password', '-s', 'typesafe', '-w'])
32  return keychain.exitCode === 0 ? keychain.stdout.trim() || undefined : undefined
33}
34
35async function nouls($: EngineInterface, key: string, state: unknown, questions: Record<string, string>) {
36  const response = await $.http.fetch(ENDPOINT, {
37    method: 'POST',
38    headers: { Authorization: `Bearer ${key}`, 'Content-Type': 'application/json' },
39    body: JSON.stringify({
40      model: 'jev-latest',
41      state,
42      questions: Object.fromEntries(Object.entries(questions).map(([id, text]) => [id, { type: 'noul', instructions: text }])),
43    }),
44  })
45  if (!response.ok) throw new Error(`TypeSafe ${response.status}: ${response.text.slice(0, 200)}`)
46  const answers = JSON.parse(response.text).answers as Record<string, { noul: number }>
47  return Object.fromEntries(Object.entries(answers).map(([id, answer]) => [id, answer.noul]))
48}
49
50export function skillHash(text: string): string {
51  let hash = 0x811c9dc5
52  for (let i = 0; i < text.length; i += 1) {
53    hash ^= text.charCodeAt(i)
54    hash = Math.imul(hash, 0x01000193) >>> 0
55  }
56  return hash.toString(16).padStart(8, '0')
57}
58
59function isTypedPrompt(message: SessionMessage): boolean {
60  return message.role === 'user' && message.text.trim() !== '' && !message.toolResults?.length
61}
62
63export function turnTrace(messages: SessionMessage[]): { prompt: string; trace: string } {
64  let start = messages.length - 1
65  while (start > 0 && !isTypedPrompt(messages[start])) start -= 1
66  const lines: string[] = []
67  for (const message of messages.slice(start)) {
68    if (message.role === 'user' && isTypedPrompt(message)) lines.push(`USER: ${message.text}`)
69    if (message.role === 'assistant') {
70      if (message.text.trim()) lines.push(`AGENT: ${message.text}`)
71      for (const use of message.toolUses) {
72        const outcome = use.isError ? `ERROR ${use.text ?? ''}` : (use.text ?? '')
73        lines.push(`TOOL ${use.tool} ${JSON.stringify(use.input).slice(0, 300)} -> ${outcome.slice(0, 400)}`)
74      }
75    }
76  }
77  const [opening = '', ...rest] = lines
78  const body = rest.join('\n')
79  const trace = body.length > TURN_CHARS ? `${opening}\n…\n${body.slice(-TURN_CHARS)}` : [opening, body].join('\n')
80  return { prompt: messages[start]?.text ?? '', trace }
81}
82
83function evidence(trace: string): string {
84  return trace.length > EVIDENCE_CHARS ? `${trace.slice(0, EVIDENCE_CHARS / 2)}\n…\n${trace.slice(-EVIDENCE_CHARS / 2)}` : trace
85}
86
87async function logDir($: EngineInterface): Promise<string> {
88  return `${await $.env.get('HOME')}/.claude/skill-drift`
89}
90
91async function record($: EngineInterface, found: Hit[]) {
92  if (found.length === 0) return
93  const day = new Date(found[0].at).toISOString().slice(0, 10)
94  const path = `${await logDir($)}/${day}-${found[0].session}.jsonl`
95  const before = (await $.fs.exists(path)) ? String(await $.fs.read(path)) : ''
96  await $.fs.write(path, before + found.map(hit => JSON.stringify(hit) + '\n').join(''))
97}
98
99export function summarize(log: Hit[]): string {
100  if (log.length === 0) return 'skill-drift: no drift recorded yet'
101  const bySkill = new Map<string, Hit[]>()
102  for (const hit of log) bySkill.set(hit.skill, [...(bySkill.get(hit.skill) ?? []), hit])
103  const ranked = [...bySkill.entries()].sort((a, b) => b[1].length - a[1].length)
104  const lines: string[] = []
105  for (const [skill, hits] of ranked) {
106    const kinds = new Map<string, Set<string>>()
107    for (const hit of hits) kinds.set(hit.kind, (kinds.get(hit.kind) ?? new Set()).add(hit.session))
108    lines.push(`${skill}  ${[...kinds.entries()].map(([kind, sessions]) => `${kind}×${sessions.size} sessions`).join('  ')}`)
109    const latest = hits.reduce((a, b) => (b.at > a.at ? b : a))
110    lines.push(`  latest: ${latest.kind} p=${latest.p.toFixed(2)} on "${latest.prompt.slice(0, 90)}"`)
111  }
112  return lines.join('\n')
113}
114
115async function readLog($: EngineInterface): Promise<Hit[]> {
116  const dir = await logDir($)
117  if (!(await $.fs.exists(dir))) return []
118  const files = (await $.fs.list(dir)).filter(entry => entry.kind === 'file' && entry.name.endsWith('.jsonl'))
119  const texts = await Promise.all(files.map(entry => $.fs.read(`${dir}/${entry.name}`)))
120  return texts.flatMap(text => String(text).split('\n').filter(Boolean).map(line => JSON.parse(line) as Hit))
121}
122
123export const register: Register = on => {
124  const loaded = new Map<string, string>()
125  let pending: Watched[] = []
126
127  on('session.start', async ($, e, next) => {
128    await $.command.register({ name: 'drift', description: 'List skill drift the mod has recorded, by skill' })
129    return next(e)
130  })
131
132  on('skill.prompt', async ($, e, next) => {
133    const result = await next(e)
134    loaded.set(e.skill, result.text.slice(0, SKILL_CHARS))
135    return result
136  })
137
138  on('turn.complete', async ($, e, next) => {
139    const result = await next(e)
140    if (e.agentId || e.isAborted || loaded.size === 0) return result
141    const key = await apiKey($)
142    if (!key) return result
143
144    const { prompt, trace } = turnTrace(await $.session.messages())
145    const watched = [...loaded.entries()].map(([skill, text]) => ({ skill, text, trace, prompt, turn: e.turnId }))
146    loaded.clear()
147    pending = watched
148
149    const answered = await Promise.all(
150      watched.map(w => nouls($, key, { skill: { name: w.skill, text: w.text }, turn: w.trace }, TURN_QUESTIONS)),
151    )
152    const [session, at] = await Promise.all([$.session.id(), $.clock.now()])
153    await record(
154      $,
155      watched.flatMap((w, i) =>
156        Object.entries(answered[i])
157          .filter(([kind, p]) => p >= THRESHOLDS[kind as keyof typeof THRESHOLDS])
158          .map(([kind, p]) => hitOf(w, kind as DriftKind, p, session, at, w.prompt)),
159      ),
160    )
161    return result
162  })
163
164  on('prompt.submit', async ($, e, next) => {
165    if (!TYPED_BY_A_PERSON.has(e.origin?.kind)) return next(e)
166    const watched = pending
167    pending = []
168    const result = await next(e)
169    if (watched.length === 0 || e.text.trim().startsWith('/')) return result
170    const key = await apiKey($)
171    if (!key) return result
172
173    const answered = await Promise.all(
174      watched.map(w =>
175        nouls($, key, { skill: { name: w.skill, text: w.text }, turn: w.trace, next_user_message: e.text }, NEXT_MESSAGE_QUESTIONS),
176      ),
177    )
178    const [session, at] = await Promise.all([$.session.id(), $.clock.now()])
179    await record(
180      $,
181      watched
182        .map((w, i) => hitOf(w, 'corrected', answered[i].corrected, session, at, e.text))
183        .filter(hit => hit.p >= THRESHOLDS.corrected),
184    )
185    return result
186  })
187
188  on('command.run', { name: 'drift' }, async $ => ({ text: summarize(await readLog($)) }))
189}
190
191function hitOf(w: Watched, kind: DriftKind, p: number, session: string, at: number, prompt: string): Hit {
192  return { skill: w.skill, kind, p, at, session, turn: w.turn, skill_hash: skillHash(w.text), prompt, evidence: evidence(w.trace) }
193}
194
hooks/skill-drift-judgments.ts 13 lines
1// ABOUTME: The Jev questions and thresholds the skill-drift mod asks; tests/test_evals_skill_drift.py replays them live.
2// ABOUTME: Everything after `export default` must stay plain JSON, because the Python eval parses it.
3export default {
4  "thresholds": { "deviated": 0.7, "missing_guidance": 0.6, "corrected": 0.55 },
5  "turn": {
6    "deviated": "The agent loaded the skill in `skill.text` during the turn in `turn`. Did the agent skip, reorder, or contradict a step or rule that the skill states, in a way that matters for this turn?",
7    "missing_guidance": "During `turn`, did the agent have to work out, guess, or look up something that the skill in `skill.text` is about and should have told it, but does not?"
8  },
9  "next_message": {
10    "corrected": "The user's message in `next_user_message` follows the agent's turn in `turn`, where the agent used the skill in `skill.text`. Is the user correcting or pushing back on how the agent handled something that skill covers?"
11  }
12}
13