One software-engineering workflow composted from Matt Pocock's skills, superpowers, pstack, Imbue's blueprint, and autonomous-sdlc: specs with BDD acceptance…

One software-engineering workflow for Claude Code, made by breaking down several skill collections and rebuilding the best of each: Matt Pocock's skills, obra's superpowers, Lauren Tan's pstack, Imbue's blueprint and code guardian, and the autonomous-sdlc and stick-shift plugins it replaces. NOTICE lists the sources compost takes text from; pile.toml records which upstream files fed which skill.
There is no state machine. Each skill ends with Next moves, naming the skills that usually follow and when, so Claude takes the obvious path or steps sideways into diagnosis or design. Once issues exist, compost:implement runs build, verify, and review for each one on its own, with review done by subagents every time, and records its rulings on the issue.
| Skill | When |
|---|---|
setup | Once per repo: issue tracker, domain docs, test gates, AGENTS.md migration |
spec | A fuzzy idea becomes user stories with Given/When/Then acceptance criteria |
slice | A spec becomes tracer-bullet issues in the tracker |
implement | The entry point once issues exist; picks the flow from there |
build | One issue, with its tests fitted into the existing suite |
verify | Before any claim of done: gates, AC-to-test map, the proof ladder |
review | After verify, always: /compost:review-changes, then weigh the findings |
finish | Integrate: re-run on the merge target, PR recap, clean up |
diagnose | Any bug or failing test, before a fix |
deepen | Module and interface design, refactoring for testability |
pause | Stop at a safe point and pick the work back up later |
fan-out | Competing designs or hypotheses, tried in parallel and judged |
turn | Check the upstream sources for changes worth working in |
skill-drift | Daily: turn repeated skill drift, judged by the skill-drift mod, into proposed skill fixes |
canon/ holds short essays on the ideas the skills lean on: tracer bullets, ubiquitous language, deep modules, seams, expand-contract, proving it works, and more. Skills link to the ones they use.
Work is tracked in the repo's issue tracker (GitHub by default), recorded by compost:setup in docs/agents/issue-tracker.md. Domain vocabulary lives in CONTEXT.md, decisions in docs/adr/.
Many steps in the workflow are narrow judgments over lists that code can already produce: which existing test an acceptance criterion belongs in, whether a review finding is worth a skeptic subagent, whether an upstream change needs reading. compost hands those to TypeSafe Jev, which answers typed questions (Choice, Noul, Score) in a second or two for a fraction of a cent, so the frontier model spends its tokens on the rest.
Skills call uv run ${CLAUDE_PLUGIN_ROOT}/scripts/jev.py <tool> --input <file.json>; jev.py list names them all. Code gathers the candidates, Jev picks or checks, and thresholds in code decide. Results are cached by content hash in ~/.cache/compost/. The key comes from TYPESAFE_API_KEY or the macOS Keychain item typesafe. Without one, a tool exits 3 and the skill makes the call itself, so Jev speeds compost up and never blocks it. jev.py status checks for a key without calling the API; compost:setup runs it and tells you how to store one.
| Tool | Primitive | Used by | Live eval (cases) |
|---|---|---|---|
find-test | Choice + Noul | build, diagnose | right test 33/34, extend-or-add 33/34 (34) |
duplicate-test | Noul | build | 23/23 (23 pairs) |
ac-exercised | Noul | verify | 20/22 (22) |
test-value | Score | verify | 24/24 (24 uncovered blocks) |
claim-backed | Noul | verify | 22/22 (22 claims) |
triage-finding | Choice | reviewer, review-changes | serious findings sent to a skeptic 15/15, minor spared 8/11 (26) |
review-risk | Score | review-changes, review | 24–25/28 structural calls (28 files) |
canon-pick | Choice | reviewer | precision and recall ~0.6 (22) |
locate | Choice + Noul | reviewer, spec | anchored 17/18, absent 8/8 (26) |
spec-class | Choice | spec | 21/22, never lighter than labeled (22) |
question-value | Score | spec | 24–25/25 (25 questions) |
ac-quality | Score ×3 | spec | 24/24 (24) |
adr-worthy | Noul ×3 | spec, deepen, implement | 24/24 (24 rulings) |
classify-change | Choice | turn | every adopt/adapt read 4/4, no ignore recorded wrongly, 15/26 ignores skipped (30) |
route | Choice | turn, description changes | 27/30 (30 real prompts) |
rulings-lint | Noul ×8 | turn, compost's own text | 10/13 violations, 2–4 false flags (36 passages) |
stop-guard | Noul ×2 | Stop hook in implement runs | 23/23 (23 final messages) |
skill-drift mod | Noul ×3 | every skill-using turn | deviated recall 10/11, precision 0.77–0.83; missing guidance 2/2 and corrected 1/1 flagged, thresholds provisional (60 turns) |
compost also ships the skill-drift mod (hooks/skill-drift.ts), the collector behind compost:skill-drift. After each turn that loaded a skill it sends one Jev request per skill loaded, asking whether the turn deviated from the skill or lacked its guidance; when your next prompt follows such a turn it sends one more per skill, asking whether you are correcting that work. A request averages about 5,000 input tokens, a fraction of a cent. Hits go to ~/.claude/skill-drift/, and /drift lists them by skill. It never shows anything or holds a prompt back: your prompt enters before the correction check runs, and a TypeSafe outage is skipped. Without a key it asks and writes nothing. It is an in-process hooks module, so it loads only where Claude Code runs plugin hooks modules (CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 while they roll out); without them compost works as before and the log stays empty.
Thresholds were chosen from the same labeled cases they are measured on, so treat the numbers as upper bounds until the tools have run on real work. The labeled cases are drawn from private projects and sessions, so they live in a separate private repo, compost-evals. With a checkout, COMPOST_EVALS=<checkout>/evals uv run --group dev pytest -m jev re-runs every eval against the live API for a few cents; without one, those tests skip.
/plugin marketplace add joshuaoliphant/claude-plugins
/plugin install compost@oliphant-plugins
Then run /compost:setup in each repo.
compost replaces several collections that cover the same stages, listed in the [replaces] table of pile.toml. Leaving them on gives Claude two answers to every request, so setup finds the ones still active on the machine and, with your yes, turns them off: it disables their plugins and sets their personal and skills-CLI skills to "off" under skillOverrides in ~/.claude/settings.json. Plugins synced from claude.ai have to be turned off there. Nothing is deleted, so turning one back on is a settings change. To check a machine directly: uv run scripts/pile.py replaced [--apply].
From plugins/compost (the 100% coverage gate applies when pytest runs from there):
uv run --group dev pytest # tooling tests (100% coverage) and checks on the shipped files
uv run scripts/pile.py status # upstream changes since each source's pin (Python 3.11+ via uv)hooks/skill-drift.ts 194 lines1// ABOUTME: Watches each skill the model loads and asks TypeSafe Jev whether the turn deviated from it, lacked its guidance, or drew a correction.
2// ABOUTME: Hits go quietly to ~/.claude/skill-drift/<date>-<session>.jsonl for compost:skill-drift to review; /drift lists them.
3import type { EngineInterface, Register, SessionMessage } from 'claude-code'
4
5import judgments from './skill-drift-judgments.ts'
6
7const ENDPOINT = 'https://api.typesafe.ai/v1/systemone'
8const TYPED_BY_A_PERSON = new Set(['composer', 'bridge'])
9const SKILL_CHARS = 8000
10const TURN_CHARS = 12000
11const EVIDENCE_CHARS = 4000
12const { thresholds: THRESHOLDS, turn: TURN_QUESTIONS, next_message: NEXT_MESSAGE_QUESTIONS } = judgments
13
14export type DriftKind = keyof typeof TURN_QUESTIONS | 'corrected'
15export type Hit = {
16 skill: string
17 kind: DriftKind
18 p: number
19 at: number
20 session: string
21 turn: string
22 skill_hash: string
23 prompt: string
24 evidence: string
25}
26type Watched = { skill: string; text: string; trace: string; prompt: string; turn: string }
27
28async function apiKey($: EngineInterface): Promise<string | undefined> {
29 const fromEnv = await $.env.get('TYPESAFE_API_KEY')
30 if (fromEnv) return fromEnv
31 const keychain = await $.process.run(['security', 'find-generic-password', '-s', 'typesafe', '-w'])
32 return keychain.exitCode === 0 ? keychain.stdout.trim() || undefined : undefined
33}
34
35async function nouls($: EngineInterface, key: string, state: unknown, questions: Record<string, string>) {
36 const response = await $.http.fetch(ENDPOINT, {
37 method: 'POST',
38 headers: { Authorization: `Bearer ${key}`, 'Content-Type': 'application/json' },
39 body: JSON.stringify({
40 model: 'jev-latest',
41 state,
42 questions: Object.fromEntries(Object.entries(questions).map(([id, text]) => [id, { type: 'noul', instructions: text }])),
43 }),
44 })
45 if (!response.ok) throw new Error(`TypeSafe ${response.status}: ${response.text.slice(0, 200)}`)
46 const answers = JSON.parse(response.text).answers as Record<string, { noul: number }>
47 return Object.fromEntries(Object.entries(answers).map(([id, answer]) => [id, answer.noul]))
48}
49
50export function skillHash(text: string): string {
51 let hash = 0x811c9dc5
52 for (let i = 0; i < text.length; i += 1) {
53 hash ^= text.charCodeAt(i)
54 hash = Math.imul(hash, 0x01000193) >>> 0
55 }
56 return hash.toString(16).padStart(8, '0')
57}
58
59function isTypedPrompt(message: SessionMessage): boolean {
60 return message.role === 'user' && message.text.trim() !== '' && !message.toolResults?.length
61}
62
63export function turnTrace(messages: SessionMessage[]): { prompt: string; trace: string } {
64 let start = messages.length - 1
65 while (start > 0 && !isTypedPrompt(messages[start])) start -= 1
66 const lines: string[] = []
67 for (const message of messages.slice(start)) {
68 if (message.role === 'user' && isTypedPrompt(message)) lines.push(`USER: ${message.text}`)
69 if (message.role === 'assistant') {
70 if (message.text.trim()) lines.push(`AGENT: ${message.text}`)
71 for (const use of message.toolUses) {
72 const outcome = use.isError ? `ERROR ${use.text ?? ''}` : (use.text ?? '')
73 lines.push(`TOOL ${use.tool} ${JSON.stringify(use.input).slice(0, 300)} -> ${outcome.slice(0, 400)}`)
74 }
75 }
76 }
77 const [opening = '', ...rest] = lines
78 const body = rest.join('\n')
79 const trace = body.length > TURN_CHARS ? `${opening}\n…\n${body.slice(-TURN_CHARS)}` : [opening, body].join('\n')
80 return { prompt: messages[start]?.text ?? '', trace }
81}
82
83function evidence(trace: string): string {
84 return trace.length > EVIDENCE_CHARS ? `${trace.slice(0, EVIDENCE_CHARS / 2)}\n…\n${trace.slice(-EVIDENCE_CHARS / 2)}` : trace
85}
86
87async function logDir($: EngineInterface): Promise<string> {
88 return `${await $.env.get('HOME')}/.claude/skill-drift`
89}
90
91async function record($: EngineInterface, found: Hit[]) {
92 if (found.length === 0) return
93 const day = new Date(found[0].at).toISOString().slice(0, 10)
94 const path = `${await logDir($)}/${day}-${found[0].session}.jsonl`
95 const before = (await $.fs.exists(path)) ? String(await $.fs.read(path)) : ''
96 await $.fs.write(path, before + found.map(hit => JSON.stringify(hit) + '\n').join(''))
97}
98
99export function summarize(log: Hit[]): string {
100 if (log.length === 0) return 'skill-drift: no drift recorded yet'
101 const bySkill = new Map<string, Hit[]>()
102 for (const hit of log) bySkill.set(hit.skill, [...(bySkill.get(hit.skill) ?? []), hit])
103 const ranked = [...bySkill.entries()].sort((a, b) => b[1].length - a[1].length)
104 const lines: string[] = []
105 for (const [skill, hits] of ranked) {
106 const kinds = new Map<string, Set<string>>()
107 for (const hit of hits) kinds.set(hit.kind, (kinds.get(hit.kind) ?? new Set()).add(hit.session))
108 lines.push(`${skill} ${[...kinds.entries()].map(([kind, sessions]) => `${kind}×${sessions.size} sessions`).join(' ')}`)
109 const latest = hits.reduce((a, b) => (b.at > a.at ? b : a))
110 lines.push(` latest: ${latest.kind} p=${latest.p.toFixed(2)} on "${latest.prompt.slice(0, 90)}"`)
111 }
112 return lines.join('\n')
113}
114
115async function readLog($: EngineInterface): Promise<Hit[]> {
116 const dir = await logDir($)
117 if (!(await $.fs.exists(dir))) return []
118 const files = (await $.fs.list(dir)).filter(entry => entry.kind === 'file' && entry.name.endsWith('.jsonl'))
119 const texts = await Promise.all(files.map(entry => $.fs.read(`${dir}/${entry.name}`)))
120 return texts.flatMap(text => String(text).split('\n').filter(Boolean).map(line => JSON.parse(line) as Hit))
121}
122
123export const register: Register = on => {
124 const loaded = new Map<string, string>()
125 let pending: Watched[] = []
126
127 on('session.start', async ($, e, next) => {
128 await $.command.register({ name: 'drift', description: 'List skill drift the mod has recorded, by skill' })
129 return next(e)
130 })
131
132 on('skill.prompt', async ($, e, next) => {
133 const result = await next(e)
134 loaded.set(e.skill, result.text.slice(0, SKILL_CHARS))
135 return result
136 })
137
138 on('turn.complete', async ($, e, next) => {
139 const result = await next(e)
140 if (e.agentId || e.isAborted || loaded.size === 0) return result
141 const key = await apiKey($)
142 if (!key) return result
143
144 const { prompt, trace } = turnTrace(await $.session.messages())
145 const watched = [...loaded.entries()].map(([skill, text]) => ({ skill, text, trace, prompt, turn: e.turnId }))
146 loaded.clear()
147 pending = watched
148
149 const answered = await Promise.all(
150 watched.map(w => nouls($, key, { skill: { name: w.skill, text: w.text }, turn: w.trace }, TURN_QUESTIONS)),
151 )
152 const [session, at] = await Promise.all([$.session.id(), $.clock.now()])
153 await record(
154 $,
155 watched.flatMap((w, i) =>
156 Object.entries(answered[i])
157 .filter(([kind, p]) => p >= THRESHOLDS[kind as keyof typeof THRESHOLDS])
158 .map(([kind, p]) => hitOf(w, kind as DriftKind, p, session, at, w.prompt)),
159 ),
160 )
161 return result
162 })
163
164 on('prompt.submit', async ($, e, next) => {
165 if (!TYPED_BY_A_PERSON.has(e.origin?.kind)) return next(e)
166 const watched = pending
167 pending = []
168 const result = await next(e)
169 if (watched.length === 0 || e.text.trim().startsWith('/')) return result
170 const key = await apiKey($)
171 if (!key) return result
172
173 const answered = await Promise.all(
174 watched.map(w =>
175 nouls($, key, { skill: { name: w.skill, text: w.text }, turn: w.trace, next_user_message: e.text }, NEXT_MESSAGE_QUESTIONS),
176 ),
177 )
178 const [session, at] = await Promise.all([$.session.id(), $.clock.now()])
179 await record(
180 $,
181 watched
182 .map((w, i) => hitOf(w, 'corrected', answered[i].corrected, session, at, e.text))
183 .filter(hit => hit.p >= THRESHOLDS.corrected),
184 )
185 return result
186 })
187
188 on('command.run', { name: 'drift' }, async $ => ({ text: summarize(await readLog($)) }))
189}
190
191function hitOf(w: Watched, kind: DriftKind, p: number, session: string, at: number, prompt: string): Hit {
192 return { skill: w.skill, kind, p, at, session, turn: w.turn, skill_hash: skillHash(w.text), prompt, evidence: evidence(w.trace) }
193}
194hooks/skill-drift-judgments.ts 13 lines1// ABOUTME: The Jev questions and thresholds the skill-drift mod asks; tests/test_evals_skill_drift.py replays them live.
2// ABOUTME: Everything after `export default` must stay plain JSON, because the Python eval parses it.
3export default {
4 "thresholds": { "deviated": 0.7, "missing_guidance": 0.6, "corrected": 0.55 },
5 "turn": {
6 "deviated": "The agent loaded the skill in `skill.text` during the turn in `turn`. Did the agent skip, reorder, or contradict a step or rule that the skill states, in a way that matters for this turn?",
7 "missing_guidance": "During `turn`, did the agent have to work out, guess, or look up something that the skill in `skill.text` is about and should have told it, but does not?"
8 },
9 "next_message": {
10 "corrected": "The user's message in `next_user_message` follows the agent's turn in `turn`, where the agent used the skill in `skill.text`. Is the user correcting or pushing back on how the agent handled something that skill covers?"
11 }
12}
13