SLOPSHOPPER

self-improvement-loop

Background reflection on Claude Code sessions that promotes recurring lessons into skills, hooks, rules and agents, tracks artifact usage feedback, and ships a…

new
v1.0.3PolyForm-Noncommercial-1.0.0updated 2026-10-02kolezka/self-improvement-loop
A shopper browsing a rack in a slop shop
Preview could not run: harness produced no result (56 | return head + [...names].map(n > `export const ${n} __stub;`).join('\n') 57 | } 58 | 59 | async function bundle(entry: string): Promise<{ code?: string; error?: st
README

self-improvement-loop

A Claude Code plugin that learns from your sessions. It reflects on them in the background, promotes recurring lessons into skills, hooks, rules and agents, and tracks whether those artifacts actually get used. Nothing reaches a shared remote without a human. A Codex agent can operate the same loop through the same skill.

CI Release License Runtime

Demo: how the loop works, then in the web console a staged skill is reviewed and accepted, its usage and votes are tracked, and a misfiring hook is flagged for a rewrite

About 30 seconds, no sound. Also as MP4.

Contents

How the loop works

flowchart LR
  A[Session hooks record events] --> B[Worker reflects out of session]
  B --> C[Reflections cluster by pattern]
  C -->|3 occurrences| D[Curriculum drafts an artifact on a branch]
  D --> E[Human accepts, rejects, rehomes or retires]
  E --> F[Lesson delivered to the next matching session]
  F --> G[Usage and feedback scorecards]
  G --> D

Hooks record what happened in a session. They are fast, one bundled script, never blocking, and they never call a model. A worker, running on a schedule outside any session, reflects on sessions that ended or went idle: it builds evidence from the transcript, makes one model call, and writes a reflection if there is a real lesson.

Reflections cluster by pattern. Three occurrences of the same pattern trigger curriculum, which drafts and stages an artifact on a branch. You review and accept it (or reject, rehome, retire it) in the web UI or the CLI. Once accepted, the next session that matches gets the lesson delivered as context, and usage of the new artifact starts feeding a scorecard that the next curriculum pass reads.

See docs/ARCHITECTURE.md for the full picture and docs/V1-PARITY.md for what carried over from the previous (dotfiles-next) version of this loop and what changed.

Hosts: Claude Code, Codex, OpenClaw

The engine is the sil CLI. A host is an agent whose sessions feed the loop, or whose agent operates it.

Claude CodeCodexOpenClaw
Sessions feed the loopyes, plugin hooksnoyes, plugin or sil openclaw scan
Lessons deliveredSessionStart and UserPromptSubmit contextno, sil lessons reads themmanaged block in AGENTS.md
Skill use countedyesnoyes, on a read of the SKILL.md
Agent skillships with the pluginsil codex installsil openclaw install
Slash commands/reflect, /loop, /curriculum, /feedbacknonenone

Claude Code and Codex load the same skill file, skills/self-improvement-loop/SKILL.md. OpenClaw has its own, because lessons reach it a different way. See docs/OPENCLAW.md.

Requirements

RequirementNotes
bun >= 1.4.2Runs every part of the plugin: hook fast path, CLI, worker, web UI.
gitArtifacts are staged on a branch in a target repo.
A model endpointAn OpenAI-compatible proxy (LiteLLM, Ollama) or the claude CLI on PATH.
Codex CLI (optional)Only to operate the loop from Codex.

Install

claude plugin marketplace add kolezka/marketplace
claude plugin install self-improvement-loop@kolezka

In any Claude Code session, /loop then prints status and confirms the plugin is wired up.

To operate the loop from Codex as well, run the plugin's own copy of the CLI once (no sil shim exists yet):

~/.claude/plugins/cache/kolezka/self-improvement-loop/<version>/scripts/sil codex install

It copies the skill to ~/.agents/skills/self-improvement-loop/SKILL.md, where Codex finds it, and writes the sil shim to ~/.local/bin/sil. Both are copies, so run sil codex install again after a plugin update. A symlink you put at either path is left alone.

See docs/INSTALL.md for local dev installs, the web UI on another machine, and running on a schedule.

Quick start

Until a shim exists, call the plugin's own copy of the CLI. Claude Code keeps it under its plugin cache:

SIL=~/.claude/plugins/cache/kolezka/self-improvement-loop/<version>/scripts/sil
"$SIL" init                     # config templates and a learned/ repo for the default world
$EDITOR ~/.config/self-improvement-loop/llm.yaml   # set models.critic, models.drafter, models.judge
"$SIL" status                   # worlds, queue, worker and model wiring
"$SIL" schedule install --web   # worker on a timer, web UI as a service, shim in ~/.local/bin
sil status                      # from here on, plain sil

The worker cannot call a model until llm.yaml names a model for each of the three roles. After sil schedule install, plain sil works from any shell that has ~/.local/bin on PATH, and the web UI answers at http://127.0.0.1:8766/ (the port is web.port in config.yaml).

Using the loop from an agent

The skill teaches an agent the real command surface: which commands need --world, how to read reflections and pending lessons, how to rate an artifact, and how to accept a staged one only after the human saw its diff.

Claude Code loads the skill with the plugin. These slash commands come with it:

CommandWhat it does
/reflectQueue the current session for a background reflection. Returns at once.
/loopPrint loop status. /loop web starts the web UI, /loop run triggers one worker pass.
/curriculumDry-run preview of what would be promoted. /curriculum apply drafts and stages branches, detached.
`/feedback <type>:<name> good\bad [note]`Record a vote on an artifact.

Codex gets the same skill from sil codex install. Mention it with $self-improvement-loop, or let Codex pick it from its description. A Codex agent can check status, read reflections and lessons, review and rate artifacts. It cannot queue its own session: nothing records Codex sessions yet, and the skill tells the agent to say so instead of running sil reflect.

CLI reference

CommandWhat it does
sil initWrite config templates and the default world.
sil statusShow worlds, queue, worker state and model wiring.
sil webServe the local review UI. It prints the URL and keeps running.
sil worker --onceRun one worker pass by hand: reflections, then curriculum when it is due. --no-curriculum skips curriculum. It calls a model.
sil reflect --session <id> or --cwd <path>Mark a pending Claude Code queue entry ended. --now also runs the worker.
sil reflections list, show <id>List reflections (--pattern, --limit) or print one.
sil lessonsList the lessons waiting in the inbox. Delivers nothing.
sil curriculum plan, run --applyPreview a promotion pass, or draft and stage branches.
sil review list, show, accept, reject, rehome, retireReview staged artifacts from the terminal.
sil artifactsList artifacts with their scorecards. sil artifacts rebuild --world <w> recomputes them.
`sil feedback add <type>:<name> good\bad`Rate an artifact. sil feedback list shows every vote.
sil worlds list, add, import-kbManage worlds.
sil llm list, use, set-modelInspect and switch model endpoints per role.
sil aliases list, set, rm, suggestFold near-duplicate pattern slugs together.
sil schedule install, uninstall, showRun the worker and web UI under systemd or launchd.
sil logs <name>Tail the hook, worker, web or curriculum log.
sil codex installInstall the skill and the sil shim for Codex.
sil openclaw install, sync, scan, statusRun the loop against an OpenClaw install.
sil export <path>Write a migration bundle: config, worlds, reflections, learned repos and history.
sil import bundle <path>Restore a migration bundle on another host.

Every command takes --help. Read commands resolve the world from the current directory. The commands that change a review, and reflections show, need --world <name>.

Reviewing a staged artifact

The web UI is the easiest way. From the terminal:

sil review list                                   # pattern, type, count, branch
sil review show verify-callsites --world work     # body, reviewed_state, any accept block
sil review show verify-callsites --world work --diff
sil review accept verify-callsites --world work --reviewed-state <hash from the diff>

Accept is bound to the reviewed_state digest of the exact diff you saw. If the branch moved since, accept refuses with reviewed state changed since preview; look at the new diff and accept that one. Accept merges into the world's target repo, and a world with remote: push or remote: pr also pushes or opens a PR.

sil review reject <pattern> --world <w> drops a staged branch. sil review retire <pattern> --world <w> --yes stages the removal of a live artifact, and sil review rehome <pattern> --type <type> --world <w> stages a move to another type. Both take effect only when their branch is accepted.

promotion.auto_merge: true in config.yaml skips the review: curriculum fast-forwards a staged branch into the target repo's default branch on its own. It is off by default, never applies to an llm: local world, and never pushes.

How lessons reach a session

In Claude Code, SessionStart injects the world's managed rules block (when rules_inject is on), up to 3 undelivered lessons and a one-line loop status. UserPromptSubmit delivers lessons that arrived since the session started. A lesson is delivered once, then marked so it never repeats in a later session.

OpenClaw gets the same lessons through a managed block in the workspace AGENTS.md. Codex gets nothing injected; sil lessons lists what is waiting without marking anything delivered.

Feedback and scorecards

/feedback skill:verify-callsites good "caught a real bug"

or

sil feedback add skill:verify-callsites good --note "caught a real bug"

Votes, tool and agent invocation counts, nudge fires and critic-noticed helpful/misfire signals all fold into a per-artifact scorecard (sil artifacts). The curriculum planner reads scorecards and proposes refine or retire-candidate; a human still decides.

Configuration

PathContents
~/.config/self-improvement-loop/config.yamlWorlds, promotion thresholds, worker cadence, web UI port.
~/.config/self-improvement-loop/llm.yamlModel endpoints and the three model roles (critic, drafter, judge).
~/.local/state/self-improvement-loop/Queue, usage events, feedback, inbox, logs, worker.lock.
~/.local/share/self-improvement-loop/Reflections and aliases per world, and the built-in learned/ repo.

SIL_CONFIG_DIR, SIL_STATE_DIR and SIL_DATA_DIR move each of them.

Each endpoint carries its own model names. sil llm list shows which endpoint serves each role, sil llm use <endpoint> switches all three, sil llm use <endpoint> --role critic switches one, and sil llm set-model <role> <model> changes the model a role asks for.

Worlds

A world owns its own reflections, ledger, target repo and model policy. Sessions map to a world by the longest repos prefix match on cwd; a world with no repos is the catch-all.

sil worlds add work --repos ~/Development/work --target ~/dotfiles --llm cloud

To keep a V1 (dotfiles-next) target repo's on-disk contract (claude/skills, claude/hooks/nudges, claude/agents, global.CLAUDE.md, claude/skills/promotions.json), add --layout v1:

sil worlds add legacy --target ~/dotfiles --layout v1

A V1 worlds.yaml manifest imports directly:

sil worlds import-kb ~/.config/kb/worlds.yaml

Model providers

llm.yaml supports three endpoint kinds:

KindUse it for
openaiAny OpenAI-compatible endpoint, including a local LiteLLM proxy or Ollama.
claude-cliShells out to claude -p --model <model>, no proxy needed.
system-oneA typed-decision API such as Jev or Laya. Only the judge role can run on it, see docs/OPERATIONS.md.

A role picks its endpoint from role_endpoints, else active, so the critic can run on claude -p while the drafter and judge stay on the proxy. Every chat call runs at temperature 0 for reproducibility. A world set to llm: local may only use models in llm.yaml's local_models allowlist; the loop refuses rather than silently falling back to a cloud model.

Patterns and aliases

Reflections cluster by their Pattern: slug, and the match is exact. Two reflections about the same mechanism under different slugs never reach the promotion threshold together. The alias map folds one into the other, one hop only:

sil aliases suggest          # near-duplicate slugs, by token overlap
sil aliases set stale-env stale-cached-env
sil aliases list

suggest is deterministic and calls no model. It proposes; you apply. set re-points any alias that pointed at the slug you just aliased, so a two hop chain (which would resolve to nothing) can never form.

Token overlap misses a pair that shares a mechanism but no words. Setting alias_semantic.enabled: true adds one typed question per candidate on the system-one endpoint that serves the judge, and prints the verdict under the candidate. It is off by default, it still only proposes, and a pair judged distinct stays on the list. See docs/OPERATIONS.md.

Moving to another host

sil export writes one bundle with everything the loop made: config.yaml, llm.yaml, and per world the reflections, aliases, scorecards, inbox and the built-in learned repo with its review branches, plus the usage and feedback history the scorecards are rebuilt from.

sil export ~/sil.tar.gz              # on the old host
sil import bundle ~/sil.tar.gz       # on the new one

A path that ends in .tar.gz or .tgz gives an archive; any other path gives a plain directory. Flags: --world <name...> picks the worlds, --no-history leaves usage and feedback behind, and --force overwrites.

On import the host always wins: a file this install already has is kept, and the count of kept files is printed. Pass --force to take the bundle's copy instead. Reflections are append-only and are never overwritten, not even with --force.

Two things do not travel. Credentials stay behind, because llm.yaml names an env var and never a key, so set the API key env vars again on the new host. A world with an external target repo is recorded and reported, but not copied: clone it on the new host yourself.

See docs/OPERATIONS.md.

Privacy

Reflections, the ledger, scorecards and the queue all live on disk under ~/.local/state/self-improvement-loop/ and ~/.local/share/self-improvement-loop/. Only the critic, drafter and judge calls leave the machine, to whatever endpoint you configured. Curriculum never pushes to a remote on its own; accept is a human action, and remote: push|pr is an explicit per-world opt-in on top of that.

sil codex install writes two files, the skill and the shim, and reads nothing from Codex.

A migration bundle carries config.yaml and llm.yaml only. No other file from the config directory is read, so an env file you keep next to them never enters a bundle.

Troubleshooting

SymptomCheck
ModelNotConfiguredA role has no model. sil llm list, then sil llm set-model <role> <model>. A scheduled worker also needs the endpoint's API key env var in its own environment.
Nothing ever gets reflectedsil schedule show. With no unit, the worker runs only when a Claude Code hook kicks it (worker.auto_kick, on by default) or on sil worker --once. Then sil logs worker.
Queue entry skipped: transcript not persistedThe session ran with --no-session-persistence. Nothing to fix.
Queue entry failedThe reflection ran and broke. sil logs worker has the error.
sil: command not foundRun sil schedule install or sil codex install once through the plugin's copy, see Quick start.

More in docs/OPERATIONS.md.

Development

The repo is a Bun workspace: apps/ (cli, hook, hook-module, server, web) and packages/ (core, critic, curriculum, feedback, nudges, openclaw, ops, providers, review, store, transcript, worker).

bun install
bun run lint:dashes   # fail on an em dash or en dash in tracked text
bun run typecheck
bun run check:web     # svelte-check
bun run build         # writes dist/, which the hooks run
bun test
make dev-install      # claude --plugin-dir .

CI runs the same steps in that order, plus a check that bun.lock did not change. Build before testing: dist/ is untracked, the hooks run dist/hook.js, and the drift tests skip without a build. bun run sil <args> runs the CLI from source.

The README demo is recorded by bun run build && bun run demo:record. It seeds a fake home in a temp dir (scripts/demo/seed.ts), drives sil web over it with headless Chromium, and writes docs/media/demo.mp4 and demo.gif. It needs ffmpeg and a Chromium binary (CHROMIUM_PATH, default /usr/bin/chromium). No model is called and no real install is read.

Documentation

DocumentContents
docs/ARCHITECTURE.mdDesign rules, runtime layout, full loop mechanics.
docs/INSTALL.mdMarketplace install, local dev install, Codex, scheduling.
docs/OPERATIONS.mdDaily loop, reviewing, host migration, key rotation, troubleshooting.
docs/OPENCLAW.mdRunning the loop on OpenClaw sessions.
docs/BENCHMARK.mdEngine benchmark: hook fast path, worker, curriculum, API timings.
docs/RELEASE.mdHow a version is cut and where the installed dist/ comes from.
docs/V1-PARITY.mdWhat carried over from V1 and what changed.

License

Source-available under the PolyForm Noncommercial License 1.0.0 (SPDX: PolyForm-Noncommercial-1.0.0). Any noncommercial use is free: run it, study it, change it, share it. That covers personal and hobby use, research, and use by schools, charities, public research bodies and government institutions.

Commercial use is not granted by that license. See COMMERCIAL-LICENSE.md or contact mariusz@raqz.pl.

This license is not OSI-approved, because it restricts a field of endeavour. Call it source-available, not open source.

Source 0 files

Not captured.