SLOPSHOPPER

self-compact

Compaction at task boundaries: the model can ask to compact once its turn ends. Above a fill threshold it is asked whether to, after a turn decision-model…

newguardtoastprompttool
★ 1v0.1.0NOASSERTIONupdated 2026-10-09kzarzycki/claude-mods/plugins/self-compact
A shopper browsing a rack in a slop shop
README

claude-mods

Claude Code mods: plugins whose hooks run inside Claude Code's engine through the in-process $ API. Each one installs on its own from this repository's plugin marketplace. Mods that make decisions with a model (auto-effort, resume-on-stop, ttsr-rules) depend on decision-model, which provides that model.

oh-my-pi features

The first set rebuilds features of oh-my-pi (omp), a coding-agent harness, as mods. The research behind it rates every omp feature by how naturally it fits Claude Code: proposal (written before the build, where the project was called omc). The rest of this README covers that set: how to install it, how it works, and what building it found out about the platform.

ModWhat it does
decision-modelThe decision model, for other mods: $.decision.ask({ state, questions }), with typed questions (choice, noul = probability of yes, score) in the request and answer shape of TypeSafe's System One API. Any endpoint that speaks that protocol is a backend (presets openrouter and typesafe, or a URL; Jev is the default model). Without a key, a small Claude model answers the same questions as text. It does no deciding of its own. /decision shows the backend and every recent decision with the mod that asked; /decision ask <question> -- <state> tries one.
auto-effortAuto effort: asks the decision model how open-ended each prompt is and runs that turn at the chosen effort, up to a ceiling. /effort-rate <prompt>.
resume-on-stopUnexpected stops: a turn that ends on "Let me run the tests next." without acting is spotted by the decision model and gets one resume turn.
ttsr-rulesomp's TTSR rule files (condition, scope, globs, question) from ~/.omp/agent/rules, .omp/rules, ~/.claude/ttsr, .claude/ttsr. A condition rule denies the tool call that would break it and hands the model the rule as the reason. A question rule is put to the decision model after each turn; a yes leaves the rule for the model's next turn. /ttsr, /ttsr reload.
compact-methodsTwo of omp's compaction methods. /compact shake moves old tool output to files under ~/.claude/compact-methods/<session>/ and leaves a pointer the model can Read (no model call). /compact handoff [focus] replaces the context with a handoff document written by a fork of the main thread. autoMethod picks what runs when the context fills.
self-compactCompaction at task boundaries rather than at overflow. The model can call compact_after_turn (with a focus for what to keep) and the compaction runs as /compact once its turn ends. Above threshold (60% of the auto-compact window, not the model's), a turn decision-model judges a finished task gets a short follow-up turn asking the model whether to compact; without decision-model, each prompt carries a hidden line with the fill level instead. Optional idle compaction (idleSeconds).
agent-hub/hub pane: this session's subagents with status, model, turns, tokens and their latest answer, plus a box that sends a message to a running one. /hub list, /hub send <id> <message>.
eval-kernelOne eval tool backed by a long-lived Bun kernel. State persists between cells (top-level declarations, imports). Cells call Claude Code tools (await tool.Read({ file_path })), models (completion()), subagents (agent(), with a JSON Schema schema for a parsed answer) and the decision model (judge(), when decision-model is loaded), in parallel with Promise.all. A subagent's completion notice goes to the cell, not the conversation. Interrupting a cell resets the kernel. Every call goes back through Claude Code, so permissions and other mods' hooks still apply. /eval <code>, /eval vars, /eval reset.

Install

Built and tested on Claude Code 2.1.288. The mods use the in-process $ hook API, so older versions won't load them.

From GitHub (to use them)

claude plugin marketplace add kzarzycki/claude-mods
claude plugin install auto-effort@claude-mods     # pulls in decision-model
claude plugin install resume-on-stop@claude-mods
claude plugin install ttsr-rules@claude-mods
claude plugin install compact-methods@claude-mods
claude plugin install self-compact@claude-mods
claude plugin install agent-hub@claude-mods
claude plugin install eval-kernel@claude-mods

Pick any subset; each mod works alone, and the three that make decisions pull in decision-model. Inside a session, /plugin does the same from a menu. claude plugin marketplace update claude-mods fetches new versions.

From a clone (to change them)

git clone https://github.com/kzarzycki/claude-mods && cd claude-mods
claude --plugin-dir plugins/decision-model --plugin-dir plugins/auto-effort   # one session

To load them in every session, including ones another tool starts (omnigent, an IDE), list the folders in ~/.claude/settings.json; an interactive session reloads a mod when you save its files:

{ "env": { "CLAUDE_CODE_PLUGIN_DIRS": "/path/to/claude-mods/plugins/decision-model:/path/to/claude-mods/plugins/auto-effort" } }

Options of mods loaded this way are under <name>@inline (in /config, or claude plugin configure decision-model@inline).

Setup

  • Jev for decisions. Without a key, decision-model asks Claude Haiku. To use Jev, give it an OpenRouter key without echoing it: ``sh read -rs K && printf '{"apiKey":"%s"}' "$K" | claude plugin configure decision-model@claude-mods --values-stdin; unset K ` (decision-model@inline for a clone), or set OPENROUTER_API_KEY. A TypeSafe key needs "endpoint":"typesafe" too; endpoint also takes the URL of any other System One server, with model naming its model. /decision` shows which backend answers.
  • eval-kernel needs Bun on the PATH (or its bun option set to one).
  • TTSR rules are Markdown files in .claude/ttsr/ (project) or ~/.claude/ttsr/; existing omp rule folders are read too. See e2e/ttsr.sh for one of each kind.

Checks

  • scripts/check.sh: validate, unit-test (claude plugin test) and type-check every mod, plus the kernel's self-check. No model calls, under ten seconds.
  • e2e/run.sh [decision eval ttsr compact hub self-compact]: real claude sessions with the mods loaded, asserting on output and on files the mods write. They make real model calls; the whole run takes about four minutes. hub and the question-rule check drive an interactive session through expect, because headless claude -p differs there (see below).

How the pieces work

  • A decision model, and the places that use it. decision-model owns the protocol and the backend; each place a decision is made (effort, stops, question rules) is its own mod that asks through $.decision.ask and names its purpose for the log. Swapping Jev for another model, or for Claude, changes no consumer.
  • $.decision across mods. A mod adds a noun to $ in engine.create. The noun's methods are only placeholders: calling $.decision.ask(x) raises the event decision.ask, which decision-model's hook answers with a $ of its own. A $ captured at engine.create can't be used later; the validator refuses it.
  • Eval kernel transport. A mod can write to a child's stdin only once, so the kernel serves HTTP on a Unix socket instead ($.http.fetch with socketPath). A cell that needs the host parks the call and returns it as a call event; the mod runs it and answers with /resume.
  • Subagents from a cell. agent() spawns the subagent and the eval hook waits on /next. The subagent's hand-back (SubagentHandback, or its final turn.complete) answers the call through /answer. Every wait happens inside $.http.fetch, which doesn't count against the hook's 10-second budget. Waiting on an ordinary promise would count.
  • Kernel lifetime. The kernel is started from session.start and restarted when it exits (/eval reset simply exits it). It dies with the mod's module, and a parent-pid watchdog ends it if Claude Code dies.

What the spike found about the platform

Each item was observed in a run, not read from docs.

  • Settings rows are interactive-only. Plugin userConfig rows are in $.config.list() in an interactive session but not under claude -p, where only the engine's 40 rows come back.
  • Notes appended after the last turn are lost headless. A $.session.append made after the last turn of a claude -p run never reaches the transcript, because the process exits. Interactively the model reads it on its next turn.
  • Compaction can't start from a command. A command.run hook may not call $.session.compact; the engine refuses because the command holds the turn. Hence /compact shake, not /shake.
  • Fork can fail after a resume. $.model.fork has nothing to fork right after a session is resumed headless. Handoff falls back to $.model.complete over the messages session.compact passes in.
  • A missing dependency blocks the load. A mod whose dependencies aren't loaded doesn't load at all. A marketplace install pulls the dependency in (+ 1 dependency).
  • Jev returns real distributions; the Claude backend can't. Jev's choice answers carry probabilities (xhigh 0.99, high 0.01, confidence 0.98); the Claude text judge answers one-hot. Jev's noul can sit near the middle where a reader would say yes: "Is DROP TABLE users; in production irreversible?" came back 0.51, so thresholds need tuning per question.
  • A mod compacts only between turns. $.session.compact and $.command.run("compact") are refused inside tool.call and prompt.submit hooks ("would compact under the turn this hook is holding"), so a model can't compact mid-turn and a mod can't compact just before a new prompt runs. From turn.complete both work interactively; headless (-p) only $.command.run("compact") does.
  • Child sessions don't save transcripts. A session started from inside another Claude Code session inherits CLAUDE_CODE_CHILD_SESSION and stops saving transcripts; the e2e scripts unset it.
  • Haiku subagents can miss the task. A Haiku subagent sometimes answers its injected system context ("System initialization acknowledged…") instead of the task. The e2e checks use Sonnet for subagents.
  • Test-kit gaps (claude plugin test):
  • A test hook on session.append never sees $.session.append.
  • A test hook answering agent.spawn gets no agent id.
  • Test plugins run without their closures.

Those paths are covered by e2e instead.

Not done or not verified

  • Other System One servers. Jev through the OpenRouter preset is checked live (the decision and ttsr e2e suites pass against it, answering as typesafe/jev-1.13-20260917). The TypeSafe preset and custom URLs are checked only against faked replies.
  • Steering a running subagent (/hub send, the pane's input). No automated check: the kit can't intercept $.session.append, and an e2e needs a subagent that stays running long enough.
  • TTSR rules scoped to edit/write miss Bash. A model that's refused a Write can write the same file with printf … > file (seen in an e2e run). Scope a rule to tool:bash too if that matters; matching shell redirections to file globs isn't done.
  • Eval cell timeouts. A cell has no timeout. A synchronous infinite loop blocks the kernel until /eval reset (or reset: true). agent()'s schema is parsed, not validated: a wrong shape reaches the cell as is.
  • omp features left out of this MVP. Snapcompact, omp's soft compaction, idle compaction, TTSR rules on streamed text (Claude Code can't take back text it has already shown), and multi-vendor models.
Source 1 files
hooks/register.ts 140 lines
1import type { EngineInterface, Register } from "claude-code";
2
3// self-compact: compaction at the end of a finished task rather than when the context overflows.
4// The model decides whether and what to keep; this mod decides when to ask it. With decision-model,
5// a turn it judges a finished task (above the threshold) gets a short follow-up turn asking the
6// model whether to compact. Without it, each prompt above the threshold carries a hidden line with
7// the fill level instead, which costs no extra turns. A mod can't compact
8// inside a tool call or a prompt.submit (the host refuses), so everything happens at turn.complete:
9// the compact_after_turn tool only records the request, and the compaction runs as `/compact`
10// once the turn ends (`$.command.run` works headless too, unlike `$.session.compact`).
11
12const TOOL = "compact_after_turn";
13const TOOL_ID = "mcp__self-compact__compact_after_turn";
14const DESCRIPTION =
15  "Compact the conversation once this turn ends. Call it after finishing a task when the next task won't need this one's details. `focus` tells the summary what to keep for what comes next. Nothing changes until the turn ends, so finish your reply as usual.";
16
17const FINISHED_TASK = {
18  type: "noul",
19  instructions:
20    "Classify whether this assistant message ends a self-contained task: the work it was doing is done or reported, and it does not announce further steps it is about to take or ask the user to decide something.",
21} as const;
22
23type Ask = { decision: { ask: (r: object) => Promise<{ answers: Record<string, { type: string; noul?: number }> }> } };
24
25/** Whether this turn looks like a finished task, by the decision model. */
26async function looksFinished($: EngineInterface, answer: string): Promise<boolean> {
27  try {
28    const r = await ($ as unknown as Ask).decision.ask({ purpose: "self-compact", state: answer.slice(-2000), questions: { done: FINISHED_TASK } });
29    const a = r.answers.done;
30    return a?.type === "noul" && (a.noul ?? 0) >= 0.5;
31  } catch {
32    return false;
33  }
34}
35
36/**
37 * How full the context is, in percent of the window auto-compaction measures against. That window
38 * can be much smaller than the model's (CLAUDE_CODE_AUTO_COMPACT_WINDOW), and `context.percent`
39 * is of the model's, so a threshold on it could sit past the point where auto-compaction fires.
40 */
41async function fill($: EngineInterface): Promise<number> {
42  const { context } = await $.session.usage({ breakdown: "summary" });
43  const b = context.breakdown;
44  return b && b.rawMaxTokens > 0 ? (100 * b.totalTokens) / b.rawMaxTokens : (context.percent ?? 0);
45}
46
47async function compact($: EngineInterface, method: string, focus: string) {
48  const args = [method === "native" ? "" : method, focus].filter(Boolean).join(" ");
49  $.ui.toast(`self-compact: compacting${focus ? ` (keep: ${focus.slice(0, 60)})` : ""}`);
50  await $.command.run({ command: "compact", args }).catch(() => {});
51}
52
53export const register: Register = (on, options) => {
54  const threshold = Number(options.threshold ?? 60);
55  const toolFloor = Number(options.toolFloor ?? 20);
56  const nudge = options.nudge !== false;
57  const method = String(options.method || "native");
58  const idleSeconds = Number(options.idleSeconds ?? 0);
59
60  let requested: string | undefined; // the model's focus, once it called the tool this turn
61  let nudgePending = false; // the next main turn answers our nudge
62  let hasDecision = false; // whether decision-model is loaded, checked at session start
63  let turns = 0;
64  let compactedAt = -2; // the turn the last compaction was queued at
65  let generation = 0; // bumped by every prompt; an idle timer fires only if nothing came since
66
67  on("session.start", async ($, e, next) => {
68    hasDecision = await ($ as unknown as Ask).decision.ask({ purpose: "self-compact", state: "", questions: {} }).then(() => true, () => false);
69    await $.tool.register({
70      name: TOOL,
71      description: DESCRIPTION,
72      inputSchema: { type: "object", properties: { focus: { type: "string", description: "What the summary must keep for the next task" } } },
73    });
74    return next(e);
75  });
76
77  // ponytail: matched by hand because the generated types list only MCP tools connected when the mod was saved.
78  on("tool.call", async ($, e, next) => {
79    if ((e.tool as string) !== TOOL_ID) return next(e);
80    if (e.agentId) return { isError: true, result: "compact_after_turn is for the main conversation" };
81    requested = String((e as unknown as { focus?: unknown }).focus ?? "").trim();
82    return { result: "Compaction will run when this turn ends." };
83  });
84
85  // Without a decision model: a person's prompt above the threshold carries the fill level, so the
86  // model can compact on its own after the task it starts. Attached to the prompt rather than the
87  // system prompt, which would break the prompt cache each time the figure changed.
88  on("prompt.submit", async ($, e, next) => {
89    generation++;
90    if (!nudge || hasDecision || e.origin.kind === "plugin") return next(e);
91    const percent = await fill($);
92    if (percent < threshold) return next(e);
93    const line = `Context is ${Math.round(percent)}% full. When you finish the task this prompt starts, if the next task won't need its details, call the ${TOOL} tool with a focus saying what to keep.`;
94    return next({ ...e, context: [...(e.context ?? []), line] });
95  });
96
97  on("turn.complete", async ($, e, next) => {
98    const done = await next(e);
99    if (e.agentId || e.isAborted || e.reason !== "answer") return done;
100    turns++;
101    const answered = nudgePending;
102    nudgePending = false;
103    const percent = await fill($);
104
105    if (requested !== undefined) {
106      const focus = requested;
107      requested = undefined;
108      // A second request right after one (say from a resumed turn) would compact twice.
109      if (percent >= toolFloor && turns - compactedAt >= 2) {
110        compactedAt = turns;
111        void compact($, method, focus);
112      }
113      return done;
114    }
115    if (answered || percent < threshold) return done; // a declined nudge, or not full enough
116
117    const idle = generation;
118    if (idleSeconds > 0) {
119      void (async () => {
120        await $.clock.sleep(idleSeconds * 1000);
121        if (generation !== idle || (await $.prompt.read()).text.trim()) return;
122        if ((await fill($)) >= threshold) await compact($, method, "");
123      })().catch(() => {});
124    }
125
126    if (!nudge || !hasDecision) return done;
127    if (!(await looksFinished($, e.answer))) return done;
128    nudgePending = true;
129    $.ui.toast("self-compact: asking the model whether to compact");
130    void $.prompt
131      .submit({
132        text: `Context is ${Math.round(percent)}% full and this looks like the end of a task. If it is, call the ${TOOL} tool with a focus saying what the next task needs kept, then reply only "ok". If the work isn't finished, reply only "no".`,
133      })
134      .catch(() => {
135        nudgePending = false;
136      });
137    return done;
138  });
139};
140