SLOPSHOPPER

craftsman

Stack-agnostic doc-first command loop plus a deterministic quality engine — regression-only lint/type gates, acceptance-criteria enforcement, routed review…

newstatusprocess
v?MITupdated 2026-10-09plantgreytrees/craftsman
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · craftsman
› fix the failing auth test and add an audit log call ⏺ Read(src/auth.ts) ⎿ Read 6 lines ⏺ Update(src/auth.ts) ⎿ Added 2 lines, removed 1 line ⏺ Bash(bun test) ⎿ 3 pass, 1 fail ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts ⚠ craftsman: craftsman · ctx 49%
README

craftsman

v2.1.0 — see CHANGELOG.md for what changed.

A Claude Code plugin that makes code better automatically — in any language. It runs your formatters, linters, and type-checkers on every edit (flagging only the issues you introduce), refuses to call a task "done" if it broke the tests or left plan criteria unmet, and adds a clean plan → build → review workflow.

Drop it into a Python, Go, Rust, Java, Ruby, JS/TS, C#, Swift, Kotlin, PHP, C++, Terraform, or Kubernetes repo and run /craftsman:init — it detects your stack and configures itself. It's the generalized evolution of the craftsman v0.2 starter.


Quick start

# 1 — install  (from GitHub once published; see PUBLISH.md)
/plugin marketplace add plantgreytrees/craftsman
/plugin install craftsman@craftsman-marketplace

# 2 — restart Claude Code so the hooks load, then in your repo:
/craftsman:init        # audit, upgrade older setup, and repair required structure
/craftsman:baseline    # snapshot existing lint issues so only NEW ones get flagged

That's the whole setup. Installing from a local folder instead of GitHub? Use /plugin marketplace add /absolute/path/to/craftsman. Full walkthrough and troubleshooting is in INSTALL.md.


What it does

1. A quality engine that runs itself. Three layers, cheapest first — each only does what the layer below can't:

LayerWhat happensCost
Deterministicon every edit: format → lint → type-check; at "done": secrets scan + test-regression check (skipped when nothing changed since the last check). Only newly-introduced issues are reported.~free
Feedbackwhen a check fails, its output is fed straight back to Claude mid-turn to fix~free
JudgmentLLM review — correctness, security, house-style conformance, plan soundness — runs only where linters are structurally blind, and only when a cheap router says a given diff is worth itgated

2. A doc-first workflow: IDEA → ARCHITECT → PLAN → ORCHESTRATE → SCRUTINISE → SYNC-DOCS → MERGE. /idea vets a new idea, /architect turns it into enforced decisions, and /instruction packs the rest into one /goal prompt; /understand and /investigate enter at PLAN for existing code and bugs. One slug names the work from idea doc to plan, fix round and sync. /craftsman:auto is the single entry for the whole loop: it asks every open decision up front, runs the units unattended, asks once more about what came back, and lands it. Planners write a short plan doc; /orchestrate executes it, reading each unit's actual diff through gate-select.mjs to deterministically route it to the specialists that apply — ui-ux-reviewer, migration-reviewer, api-reviewer, dependency-auditor, performance-reviewer, observability-reviewer — instead of trusting a file-pattern rule to be remembered correctly; /scrutinise reviews the result; /sync-docs keeps the docs honest; /craftsman:merge lands what's left on the base branch and cleans up. Acceptance criteria written at plan-time are enforced before a task can finish. /craftsman:digest reads the same tracker state at any point for a done/decisions/next/%-complete summary.

Language-agnostic by design: every check is a config entry that silently skips if its tool isn't installed — so the same plugin lints Python with ruff, Go with staticcheck, Rust with clippy, and so on, wherever those tools exist. Adding a language is a one-block config edit, never a code change.

3. Pre-existing issues don't just vanish. The gate only ever reports issues you introduce — but /craftsman:baseline also writes what it skipped to docs/errors/KNOWN_ISSUES.md (worst file first, regenerated on every run), so a legacy repo's debt stays visible and addressable instead of living only in a gitignored local snapshot.


Commands

CommandWhat it's for
/craftsman:initaudit or upgrade the setup, detect the stack, and repair required project structure
/craftsman:mergeland a finished branch or worktree when you say "merge": tests, merge, push, then remove the worktree and the local and remote branch
/craftsman:upgradeupdate every Craftsman install (user and per-project) to the latest commit, then tell you what to reload and re-init
/craftsman:workspace-initexplicitly register selected existing Git projects without scanning the workspace
/ideascrutinise an idea before building it: overlap scan, whole-system fit, research, one isolated critic, scored verdict (--deep for a stronger critic)
/architectturn a vetted idea into confirmed engineering/data/systems decisions, written as enforced architecture docs (--deep, --init, --update; --backfill derives them for a project with existing plans and legacy docs)
/craftsman:autoone entry for the whole loop: ask every open decision first, run the units unattended in fresh agent contexts, ask about parked rows, land; repeat until every row is COMPLETE
/instructiongenerate one paste-ready /goal prompt that drives the whole loop to verified completion, plus an effort estimate
/understandbuild a cited understanding of a feature/area before touching it
/planturn a request into a build-ready plan doc (the only command that writes plans)
/orchestrateimplement a plan doc across the repo (per-unit review is routed too, same reasoning as /scrutinise)
/investigateroot-cause a bug into a fix-ready plan
/scrutinisereview what was built (routed, so trivial diffs stay cheap)
/sync-docsreconcile the docs with the code that actually shipped
/fix-testsget a red test suite green
/craftsman:digestADHD-friendly progress digest: done, decisions needed, next, % complete per plan
/craftsman:baselinesnapshot pre-existing lint issues (run once per repo)
/craftsman:statssee which gates actually fire — delete the ones that don't earn their keep
/craftsman:toggleturn the gates on/off for this repo
/craftsman:builtin-checkafter claude update: verify the wrapped Claude Code built-ins still match craftsman's assumptions, and propose new ones to wrap

Configure it for your stack

/craftsman:init writes a starter craftsman.config.json in your repo root; edit it any time. It deep-merges over the plugin's defaults, so you only write what you're changing. Common tweaks (full guide in EXTENDING.md):

{
  "stopGate": {
    "commands": { "package.json": "pnpm test" },   // your real test command
    "extraChecks": ["make lint"]                    // repo guards run before "done"
  },
  "languages": {
    "elixir": { "extensions": [".ex"], "format": ["mix format {file}"], "check": ["mix credo {file}"] }
  }
}

Turn everything off with CRAFTSMAN=off (env) or /craftsman:toggle off.


Good to know

  • Protected base branches (GitHub, GitLab, Azure DevOps) — by default a finished unit is merged locally and the base branch is pushed. If your host rejects direct pushes, set repoExec.land: "pr": each unit's branch is pushed and lands through an auto-merging pull/merge request opened with gh, glab or az, so your branch protection, CI and approvals apply. Worktree isolation is the same either way — see EXTENDING.md.
  • Doc-write authority — the guarded paths (docs/plans/, docs/ideas/, docs/architecture/**) are blocked until a writer command grants the session. /plan writes plans; /orchestrate, /scrutinise and /sync-docs --tracker update tracker rows; /idea writes idea docs (/architect marks them architected); /architect writes paired architecture areas (<area>.rules.md + <area>.md). Unpaired legacy prose under docs/architecture/ stays /sync-docs' until /architect --backfill pairs it. /understand and /investigate write nothing and hand off to /plan. The rest of docs/ is ordinary prose, freely editable; widen the guard in config if you want it stricter. The block is a hook; the grant covers every guarded path for that session, so which command writes which path is the commands' contract.
  • Architecture decisions are enforced, not just written down — /architect writes each area as a pair: <area>.md for people (plain English, mermaid, rationale), which the loop never loads, and a terse <area>.rules.md of numbered ARCH-… rules plus the paths it governs. /plan cites the rules per unit, plan-reviewer and the scrutineer check them, and scope.mjs refuses to activate a unit that writes a governed path without loading and citing its rules — so /orchestrate cannot edit there until it does. Breaking a decided rule takes /architect, never a workaround. Switch the gate off with architecture.enforce: false.
  • Safe with concurrent sessions — each session's state (test baseline, plan criteria, doc authority) is isolated under .craftsman/sessions/<id>/; one session finishing never blocks another. Each execution unit must activate a separate Git worktree with worktree_path; mutating Git commands from another checkout are blocked until the binding is released for the locked merge. A claim left behind by a crashed session can only be recovered once it's confirmed stale (24h+ by default) — see EXTENDING.md.
  • Structured tracker state — the readable tracker stays concise and link-first, while scripts/tracker.mjs records validated unit identities, legal status transitions, and terse evidence in the local .craftsman/tracker/events.jsonl ledger. Claims write IN_PROGRESS there automatically. Both the ledger and TRACKER.md live in the main checkout even when a session works in a worktree, so every parallel session sees one tracker. Each transition re-renders the file's generated ledger block, and the tracker-sync hook does the same at SessionStart, after edits and at Stop. pre-guard blocks writes to a worktree's copy and hand edits of the block.
  • Fast — session start never blocks on a build (the baseline runs in the background); tool detection is scoped to your detected stack and cached; lint results are content-hash cached; the Stop-gate's test-regression check only re-runs when something was actually edited since the last check, so an answer-only turn doesn't re-run the suite for nothing. At Stop, every test command, extraChecks entry, and secrets scan run concurrently against one shared wall-clock budget (stopGate.totalBudgetMs) rather than summed — see EXTENDING.md.
  • Secrets are a hard block, never silently skipped — a scan that gets cut off (budget or its own timeout) fails the Stop gate with a clear message instead of passing quietly; a clean working tree still gets scanned at least once per session, and gitignored .env files are checked too, not just tracked changes. If the configured scanner (gitleaks by default) isn't on PATH, the gate falls back to a dependency-free built-in scanner instead of just blocking on a missing tool (opt out with security.builtinFallback: false).
  • Root-only by default, mechanically enforced — /orchestrate and the planners do the work themselves, sequentially, instead of fanning out subagents; execution.agentMode (default root-only) is enforced by a PreToolUse hook that hard-blocks Task calls, not just a convention. The standing exceptions are /plan's one decomposition step and one isolated fresh-context agent per run of /scrutinise (scrutineer, which reviews merged work without having seen it written), /idea (idea-critic, which judges the idea without hearing the pitch) and /architect --deep (architect-analyst, on Fable). The first three run on opus; everything else in /plan, and most agents, run cheaper — implementer/security-auditor stay on sonnet. Each run is granted exactly one such agent, so a second or parallel dispatch is blocked. Claude Code's read-only built-in agents listed in execution.builtinAgents (default ["Explore"]) also pass, since they're read-only by design. Flip execution.agentMode to "subagents" to restore real fan-out.
  • Autonomous runs keep the root context light — root-only mode was chosen for lower total tokens. For autonomous runs that rationale is superseded: root context comes first (ARCH-ENGINE-10), because a long unattended run fails when root fills, not when the bill grows. /craftsman:auto therefore runs each unit in a fresh agent context, chosen by execution.engine: workflow (default) runs the plugin workflow workflows/run.js, subagent dispatches one unit-runner per unit, and root keeps the old in-root behaviour. Measured, root grew ~20k tokens per workflow unit against ~47k–108k per root-only unit (docs/plans/autonomous-e2e-loop-evidence.md). /plan refuses a unit whose scoped files and task text exceed execution.unitContextBytes (default 122,880), so every unit fits one fresh agent.
  • Built-in skills, craftsman's rules — /orchestrate runs Claude Code's own maintained simplify (step 6) and code-review (step 8) skills inside the loop. A PreToolUse hook on Skill hands them the governing standards, ARCH rules and open acceptance criteria. Their output is only a candidate list: code-reviewer still owns the verdict, and the scope guard and quality gate cover every edit they make.
  • Context stays bounded without stopping — every unit ends in a hand-off on disk, and the next unit re-reads only the tracker and its own scope. Claude Code's auto-compact reclaims context when it fills, and SessionStart restores the hand-off afterwards, so runs don't halt waiting for a manual /compact. Worktrees left behind after a merge, oversized shipped docs, and unbounded plan-memory/tracker growth are all caught the same way — by a hook or a CI check, not by remembering to do it.

Layout

.claude-plugin/{plugin.json, marketplace.json}   plugin + marketplace manifests
craftsman.config.json                            default tunables (a project can override)
hooks/hooks.json                                 SessionStart / PreToolUse / PostToolUse / Stop
scripts/                                         the Node engine (no dependencies)
commands/                                        the workflow + engine commands
agents/                                          the specialist review/implement agents
skills/language-aware-planning/                  19 per-language idiom checklists + plan template
skills/design-review/                            the UI/UX review checklist ui-ux-reviewer loads
skills/deep-research/                            forked read-only repo research (/deep-research)
output-styles/craftsman-terse.md                 an optional terse response style
.github/workflows/ci.yml                         syntax/test/JSON checks on every push and PR

Docs

Upgrading from craftsman v0.2

This replaces the v0.2 starter (it includes that engine plus the full workflow). Don't run both — they share .craftsman/ state and the craftsman plugin name. Uninstall v0.2 first (see INSTALL.md, step 0).

Contributing

Issues and PRs welcome — see CONTRIBUTING.md. The plugin is dependency-free; node --test scripts/*.test.mjs scripts/lib/*.test.mjs (every script's unit/subprocess suite, including claim/handoff/plan-graph/repo-exec/ scope/toggle/workspace coverage), node --check on every script, and claude plugin validate . are the whole test suite — CI runs all of it plus a diff-whitespace check on every push and PR. Run claude plugin validate . from this directory (the repo root is one level up and has no manifest); because plugin.json and marketplace.json share .claude-plugin/, the CLI validates the marketplace manifest. Keep it stack-agnostic — project-specific behaviour belongs in a project's own .claude/, not in the plugin.

License

MIT © 2026 Kieran. Requires Node (ships with Claude Code); no other dependencies.

Source 1 files
hooks/register.js 116 lines
1// Craftsman's mod: launch, observe, draw — nothing else (ARCH-MOD-01). It
2// registers no tool.call, tool.check or agent.spawn handler and reads no
3// settings. Every engine call is feature-detected by trying it (ARCH-MOD-04).
4//
5// On turn end:
6//   observe — the engine's own context size goes to scripts/telemetry.mjs (ARCH-MOD-03);
7//   draw    — a status band: plan · unit · tracker % · context %;
8//   launch  — a .craftsman/instructions/<slug>.goal.txt written since the session
9//             started starts once: command.run("goal"), else prompt.submit,
10//             else the band asks for /craftsman:auto-go (ARCH-MOD-02).
11// /craftsman:auto-go (commands/auto-go.md) answers with the newest goal file.
12
13const GOALS = ".craftsman/instructions";
14const LAUNCHED = "craftsman.launched-goals";
15
16// The newest *.goal.txt in a fs.list result, optionally only those newer than `since`.
17function newestGoal(entries, since = 0, launched = {}) {
18  const goals = (Array.isArray(entries) ? entries : [])
19    .filter((f) => f && f.kind === "file" && /\.goal\.txt$/.test(f.name) && f.mtimeMs > since && launched[f.name] !== f.mtimeMs)
20    .sort((a, b) => b.mtimeMs - a.mtimeMs);
21  return goals[0] || null;
22}
23
24// The newest scope file name in a session dir listing.
25function newestScope(entries) {
26  const scopes = (Array.isArray(entries) ? entries : [])
27    .filter((f) => f && f.kind === "file" && /^scope(?:@[A-Za-z0-9_-]+?)?(?:-[0-9a-f]{16})?\.json$/.test(f.name))
28    .sort((a, b) => b.mtimeMs - a.mtimeMs);
29  return scopes[0] ? scopes[0].name : null;
30}
31
32// Share of the plan's units at MERGED or COMPLETE, from the tracker ledger text.
33function trackerPercent(ledger, plan) {
34  if (typeof ledger !== "string" || !plan) return null;
35  const latest = new Map();
36  for (const line of ledger.split("\n")) {
37    let row;
38    try { row = JSON.parse(line); } catch { continue; }
39    if (row && row.plan === plan && row.unit) latest.set(row.unit, row.status);
40  }
41  if (!latest.size) return null;
42  const done = [...latest.values()].filter((s) => s === "MERGED" || s === "COMPLETE").length;
43  return Math.round((100 * done) / latest.size);
44}
45
46function band({ plan, unit, tracker, context }) {
47  const slug = plan ? String(plan).replace(/^.*\//, "").replace(/\.md$/, "") : null;
48  return ["craftsman", slug, unit, tracker === null || tracker === undefined ? null : `tracker ${tracker}%`,
49    context === null || context === undefined ? null : `ctx ${context}%`].filter(Boolean).join(" · ");
50}
51
52const goalArgs = (text) => String(text).trim().replace(/^\/goal\s+/, "");
53
54export function register(on) {
55  on("command.run", ($, e, next) => {
56    if (e.command !== "craftsman:auto-go" && e.command !== "auto-go") return next(e);
57    const none = { text: "No goal file under .craftsman/instructions — run /craftsman:auto <slug> first." };
58    return $.fs.list(GOALS)
59      .then((entries) => {
60        const goal = newestGoal(entries);
61        return goal ? $.fs.read(`${GOALS}/${goal.name}`).then((text) => ({ text: String(text).trim() })) : none;
62      })
63      .catch(() => none);
64  });
65
66  on("turn.complete", async ($, e, next) => {
67    const result = await next(e);
68    let sid = null;
69    let usage = null;
70    try { sid = await $.session.id({}); } catch { /* unsupported */ }
71    try { usage = await $.session.usage({}); } catch { /* unsupported */ }
72    const context = usage && usage.context ? usage.context : null;
73
74    let scope = null;
75    try {
76      const name = newestScope(await $.fs.list(`.craftsman/sessions/${sid}`));
77      if (name) scope = JSON.parse(await $.fs.read(`.craftsman/sessions/${sid}/${name}`));
78    } catch { /* no active scope */ }
79
80    if (sid && context && typeof context.tokens === "number") {
81      const argv = ["node", `${$.plugin.root}/scripts/telemetry.mjs`, "--sid", sid, "--tokens", String(context.tokens), "--percent", String(context.percent)];
82      if (usage.cost && typeof usage.cost.usd === "number") argv.push("--cost", String(usage.cost.usd));
83      if (scope && scope.unit) argv.push("--unit", scope.unit);
84      try { await $.process.run(argv); } catch { /* telemetry is best-effort */ }
85    }
86
87    let tracker = null;
88    try { tracker = trackerPercent(await $.fs.read(".craftsman/tracker/events.jsonl"), scope && scope.plan); } catch { /* no ledger */ }
89    let status = band({ plan: scope && scope.plan, unit: scope && scope.unit, tracker, context: context && context.percent });
90
91    let goal = null;
92    let launched = {};
93    try {
94      // Without the session's start time no goal file can be shown to be new, so none launches.
95      if (usage && typeof usage.startedAt === "number" && usage.startedAt > 0) {
96        launched = (await $.store.get(LAUNCHED)) || {};
97        goal = newestGoal(await $.fs.list(GOALS), usage.startedAt, launched);
98      }
99    } catch { /* no goal dir or no store: nothing to launch */ }
100    if (goal) {
101      await $.store.set(LAUNCHED, { ...launched, [goal.name]: goal.mtimeMs });
102      const text = String(await $.fs.read(`${GOALS}/${goal.name}`)).trim();
103      let started = false;
104      try { await $.command.run({ command: "goal", args: goalArgs(text) }); started = true; } catch { /* no /goal here */ }
105      if (!started) {
106        // prompt.submit refuses a leading "/", so this submits the directive alone.
107        try { await $.prompt.submit({ text: goalArgs(text) }); started = true; } catch { /* cannot submit as the user */ }
108      }
109      status = started ? `${status} · launched ${goal.name}` : `${status} · type /craftsman:auto-go to launch ${goal.name}`;
110    }
111
112    try { await $.ui.status(status); } catch { /* no status surface */ }
113    return result;
114  });
115}
116