SLOPSHOPPER

ReasoningBank Lite

Saves a lesson when a failing check is made to pass, recalls it on similar work, and tests whether that helps.

newpanebandguardcommandtoast
v0.1.0MITupdated 2026-10-09ramankrishna/reasoning-bank-mod
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · reasoning-bank-mod
│ ┃ ReasoningBank Lite ✕ › fix the failing auth test and add an audit log call │ ┃ 1: Lessons (0) 2: Trial │ ┃ ⏺ Read(src/auth.ts) │ ┃ Mode: o: on f: off t: test ⎿ Read 6 lines │ ┃ ⏺ Update(src/auth.ts) │ ┃ No lessons yet. ⎿ Added 2 lines, removed 1 line │ ┃ One is saved when a failing check is made to ⏺ Bash(bun test) │ ┃ pass in a turn. ⎿ 3 pass, 1 fail │ ┃ Or type one: /bank add the tests hang => run │ ┃ them with -x first ● Done. refresh now rejects expired claims and logs an audit event. │ │ ✻ Worked for 42s · done 4:20 PM │ │ › /bank │ │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Pane · ReasoningBank Lite
1: Lessons (0) 2: Trial Mode: o: on f: off t: test No lessons yet. One is saved when a failing check is made to pass in a turn. Or type one: /bank add the tests hang => run them with -x first
README

ReasoningBank Lite

A mod for Claude Code that keeps the lessons from past failures in a project, hands them to Claude when similar work comes up, and measures whether doing so helps.

Read this first: it is unproven. The idea is plausible and easy to believe. The author's fuller implementation, reasoning-bank, ran a suite of experiments and reports in its README that it found no setting where accumulated memory produced a net gain. This mod does not assume a different outcome. It ships with an on/off trial so you can find out on your own work, and the rule for reading that trial was fixed before any data existed.

Where the idea comes from

The idea is from the ReasoningBank paper (Ouyang et al., arXiv:2509.25140): an agent distils lessons from past attempts and retrieves them on new tasks. This mod is a much smaller take on that idea, not an implementation of the paper. It learns only from one signal, a failing check made to pass, and it recalls by word overlap. It has no trajectory judging, no embeddings and no test-time scaling. It is not affiliated with the paper's authors.

How it works

Learn. The mod watches the test, build and lint commands Claude runs. When a turn contains a failing check that is later made to pass, the mod asks the model one question at the end of the turn: what is the one lesson here that would save time next time? The answer is saved as a short pair, a situation and what to do in it, with a few keywords. If the model says the failure was a one-off slip, nothing is saved.

Recall. When you send a prompt, the mod compares its words with the saved lessons. A lesson is handed to Claude only if it shares a keyword and at least one more word with the prompt. At most three are handed over, as context beside your prompt, with a note to use them only where they fit. Recall is plain word matching, not a model, so the same prompt always recalls the same lessons.

Measure. In test mode, turns that a lesson fits alternate between getting the lesson and not getting it. The mod counts failed checks in each group.

The rule for reading the trial

Set in advance, and written into the code as MIN_TURNS and MIN_DROP:

  • At least 30 turns in each group. Before that the trial says only "not enough turns yet".
  • At least 20% fewer failed checks per turn with lessons. Anything less is reported as "no effect seen".

If your trial reports no effect, the honest move is to turn the mod off.

Install

/plugin install reasoning-bank-mod --marketplace ramankrishna/reasoning-bank-mod

Answer y to add the marketplace, then pick a scope. Needs Claude Code v2.1.287 or later.

| Command | What it does | | :- | :- | | /bank | Opens the lessons and the trial | | /bank on | Learn and recall. This is the default | | /bank off | Fully off: nothing learned, nothing handed over, no model called | | /bank test | Runs the on/off trial | | /bank add <situation> => <what to do> | Saves a lesson you type yourself | | /bank forget <number> | Deletes the lesson with that number in /bank | | /bank reset | Sets the trial counts back to zero |

A line above the prompt tells you when lessons were handed to Claude, or held back by the trial.

What it runs, reads and stores

  • Calls the model at the end of a turn in which a failing check was made to pass, and only then. The call goes through your own Claude Code session, reads the session's conversation, and costs tokens on your account. It also adds a few seconds to the end of such a turn.
  • Adds text to your prompts. When lessons fit, each prompt carries up to three of them to Claude as hidden context.
  • Reads the shell commands Claude runs and whether each failed, to spot checks.
  • Stores on your machine, in Claude Code's plugin store, keyed by the project's path: the lessons (at most 200), the trial counts, and the mode.
  • Sends nothing anywhere else. No network calls of its own, no telemetry.

Like every mod, it runs with the same access to your machine as Claude Code itself. The full statement is in PRIVACY.md.

Limits

  • The trial is small and simple. Groups alternate; they are not randomised. The count is failed checks per turn, which depends heavily on what you happened to be working on. It can show a large effect and will miss a small one. It applies no significance test.
  • Only turns a lesson fits are in the trial. That is the right comparison, and it means the trial fills slowly.
  • A lesson is the model's opinion. Nothing checks that a saved lesson is true or still applies. Read the list now and then, and forget the ones that are wrong.
  • Recall is word matching. It misses a lesson phrased with different words, and can hand over one that only looks related.
  • Only recognised commands count as checks. It knows the common test, build, lint and type-check commands. A failure in a custom script is not seen.
  • Lessons are per project and stay on your machine. They are not shared with a team.

Develop

claude plugin validate .
claude plugin test .
claude --plugin-dir .

License

MIT

Source 3 files
hooks/register.tsx 393 lines
1// ReasoningBank Lite: lessons from the times a failing check was made to pass.
2//
3// Learn: after a turn in which a check failed and then passed, the model is
4// asked for the one lesson in it, and the lesson is saved for the project.
5// Recall: a later prompt that shares words with a lesson carries it to Claude.
6// Trial: an on/off comparison, built in, says whether any of this helps.
7
8import { atom, read, update } from 'claude-code'
9import type { EngineInterface, Register } from 'claude-code'
10
11import type { Lesson, Mode, Tab, Trial } from '../types'
12import {
13  addLesson,
14  armLine,
15  briefing,
16  clip,
17  isCheck,
18  markUsed,
19  MIN_DROP,
20  MIN_TURNS,
21  nextArm,
22  NO_TRIAL,
23  NO_TURN,
24  parseLesson,
25  parseTyped,
26  QUESTION,
27  recall,
28  recordCheck,
29  recordTurn,
30  verdictOf,
31} from './bank'
32
33const PANE = 'bank'
34
35const lessons = atom({ plugin: 'reasoning-bank-mod', key: 'lessons' } as const, [])
36const mode = atom({ plugin: 'reasoning-bank-mod', key: 'mode' } as const, 'on')
37const turn = atom({ plugin: 'reasoning-bank-mod', key: 'turn' } as const, NO_TURN)
38const trial = atom({ plugin: 'reasoning-bank-mod', key: 'trial' } as const, NO_TRIAL)
39const tab = atom({ plugin: 'reasoning-bank-mod', key: 'tab' } as const, 'lessons')
40const note = atom({ plugin: 'reasoning-bank-mod', key: 'note' } as const, '')
41
42let root = ''
43
44const key = (name: string): string => `${name}:${root}`
45
46const isMode = (value: unknown): value is Mode => value === 'on' || value === 'off' || value === 'test'
47
48/** Reads the project's lessons, its trial and its mode from the store. */
49async function load($: EngineInterface): Promise<void> {
50  root = await $.session.root()
51
52  const savedLessons = await $.store.get(key('lessons'))
53  const savedTrial = (await $.store.get(key('trial'))) as Trial | undefined
54  const savedMode = await $.store.get(key('mode'))
55
56  await update($, lessons, () => (Array.isArray(savedLessons) ? (savedLessons as Lesson[]) : []))
57  await update($, trial, () => (savedTrial?.with !== undefined && savedTrial.without !== undefined ? savedTrial : NO_TRIAL))
58  await update($, mode, () => (isMode(savedMode) ? savedMode : 'on'))
59}
60
61async function saveLessons($: EngineInterface, all: Lesson[]): Promise<void> {
62  await update($, lessons, () => all)
63  await $.store.set(key('lessons'), all)
64}
65
66async function setMode($: EngineInterface, value: Mode): Promise<void> {
67  await update($, mode, () => value)
68  await $.store.set(key('mode'), value)
69}
70
71async function forget($: EngineInterface, id: string): Promise<void> {
72  await saveLessons(
73    $,
74    (await read($, lessons)).filter(lesson => lesson.id !== id),
75  )
76}
77
78async function resetTrial($: EngineInterface): Promise<void> {
79  await update($, trial, () => NO_TRIAL)
80  await $.store.set(key('trial'), NO_TRIAL)
81}
82
83/** Asks the model for the lesson in the turn that just ended, and banks it. */
84async function learn($: EngineInterface): Promise<void> {
85  let said = ''
86
87  try {
88    const reply = await $.model.fork({ prompt: QUESTION })
89
90    if (!reply.isAnswered) {
91      await update($, note, () => `No lesson saved last time: the model call did not answer (${reply.reason}).`)
92
93      return
94    }
95
96    said = reply.text
97  } catch {
98    await update($, note, () => 'No lesson saved last time: the model could not be asked.')
99
100    return
101  }
102
103  const parts = parseLesson(said)
104
105  if (parts === null) {
106    await update($, note, () => 'No lesson saved last time: the model found nothing reusable.')
107
108    return
109  }
110
111  const before = await read($, lessons)
112  const all = addLesson(before, parts, await $.clock.now(), 'learned')
113
114  if (all.length === before.length && all.every((lesson, index) => lesson.id === before[index]?.id)) {
115    await update($, note, () => 'No lesson saved last time: the bank already holds it.')
116
117    return
118  }
119
120  await saveLessons($, all)
121  await update($, note, () => `Last saved: When ${parts.when}`)
122  $.ui.toast(`Lesson saved: when ${clip(parts.when, 80)}`)
123}
124
125const openPane = ($: EngineInterface) =>
126  $.ui.open({ id: PANE, title: 'ReasoningBank Lite', focus: true, closeOnEscape: true })
127
128const USAGE =
129  'Usage: /bank, /bank on, /bank off, /bank test, /bank add <situation> => <what to do>, /bank forget <number>, /bank reset'
130
131export const register: Register = on => {
132  on('session.start', async ($, e, next) => {
133    await $.command.register({
134      name: 'bank',
135      description: 'Lessons from past failures in this project, and whether they help',
136      argumentHint: '[on|off|test|add|forget|reset]',
137    })
138    await load($)
139
140    return next(e)
141  })
142
143  // /clear, /resume and /branch reset the session's state and raise no session.start.
144  on('classic.SessionStart', { source: ['clear', 'resume', 'fork'] }, async ($, e, next) => {
145    await load($)
146
147    return next(e)
148  })
149
150  on('command.run', { command: 'bank' }, async ($, e) => {
151    const args = e.args.trim()
152    const [verb = '', ...rest] = args.split(/\s+/).filter(Boolean)
153
154    await load($)
155
156    if (verb === '') {
157      await openPane($)
158
159      return {}
160    }
161    if (isMode(verb)) {
162      await setMode($, verb)
163
164      return {
165        text:
166          verb === 'on'
167            ? 'ReasoningBank Lite is on: lessons are learned and handed to Claude.'
168            : verb === 'off'
169              ? 'ReasoningBank Lite is off: nothing is learned or handed over, and no model is called.'
170              : `Trial started: turns a lesson fits alternate between getting it and not. It reports after ${MIN_TURNS} turns each way.`,
171      }
172    }
173    if (verb === 'add') {
174      const parts = parseTyped(args.slice(3))
175
176      if (parts === null) return { text: 'Write it as: /bank add <situation> => <what to do>' }
177
178      await saveLessons($, addLesson(await read($, lessons), parts, await $.clock.now(), 'typed'))
179
180      return { text: `Saved: When ${parts.when}: ${parts.do}` }
181    }
182    if (verb === 'forget') {
183      const shown = [...(await read($, lessons))].reverse()
184      const lesson = shown[Number(rest[0]) - 1]
185
186      if (lesson === undefined) return { text: 'Give the number shown beside the lesson in /bank.' }
187
188      await forget($, lesson.id)
189
190      return { text: `Forgot: When ${lesson.when}` }
191    }
192    if (verb === 'reset') {
193      await resetTrial($)
194
195      return { text: 'The trial counts are back to zero.' }
196    }
197
198    return { text: USAGE }
199  })
200
201  // --------------------------------------------------------------- recall
202
203  on('prompt.submit', async ($, e, next) => {
204    // A prompt folded into a running turn is part of that turn, not a new one.
205    if (e.turnId !== undefined) return next(e)
206
207    const now = await read($, mode)
208    const all = await read($, lessons)
209    const fit = now === 'off' ? [] : recall(all, e.text)
210    const arm = now === 'test' && fit.length > 0 ? nextArm(await read($, trial)) : null
211    const isHanded = fit.length > 0 && arm !== 'without'
212    const ids = fit.map(lesson => lesson.id)
213
214    await update($, turn, () => ({ ...NO_TURN, arm, applied: isHanded ? ids : [] }))
215
216    if (!isHanded) return next(e)
217
218    await saveLessons($, markUsed(all, ids))
219
220    return next({ ...e, context: [...(e.context ?? []), briefing(fit)] })
221  })
222
223  on('tool.call', { tool: 'Bash' }, async ($, e, next) => {
224    const ran = await next(e)
225
226    if (ran.deny !== undefined || !isCheck(e.command)) return ran
227
228    const record = ran.isError === true ? undefined : ran.result
229
230    if (record?.backgroundTaskId !== undefined) return ran
231
232    const ok = ran.isError !== true && record?.interrupted !== true
233
234    await update($, turn, was => recordCheck(was, ok))
235
236    return ran
237  })
238
239  // -------------------------------------------------------------- learning
240
241  on('turn.complete', async ($, e, next) => {
242    const done = await next(e)
243
244    if (e.agentId !== undefined) return done
245
246    const now = await read($, mode)
247    const ended = await read($, turn)
248
249    // Counted once: the turn's tally is cleared as it is recorded.
250    await update($, turn, () => NO_TURN)
251
252    if (now === 'test' && ended.arm !== null) {
253      const counted = recordTurn(await read($, trial), ended)
254
255      await update($, trial, () => counted)
256      await $.store.set(key('trial'), counted)
257    }
258    if (now !== 'off' && ended.hasRecovered && e.reason === 'answer') await learn($)
259
260    return done
261  })
262
263  // -------------------------------------------------------------- drawing
264
265  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
266    const now = await read($, turn)
267
268    if (e.props.hasSurvey || (now.applied.length === 0 && now.arm !== 'without')) return next(e)
269
270    const { Box, Button, Text } = $.ui.resolve(e)
271    const theirs = await next(e)
272    const count = now.applied.length
273
274    return (
275      <Box flexDirection="column">
276        {theirs}
277        <Box flexDirection="row" columnGap={2}>
278          <Text inverse> bank </Text>
279          {now.arm === 'without' ? (
280            <Text color="warning">trial: lessons held back this turn</Text>
281          ) : (
282            <Text dimColor>
283              {count} {count === 1 ? 'lesson' : 'lessons'} handed to Claude
284            </Text>
285          )}
286          <Button key="open" label="bank" hotkey="b" plain onPress={() => openPane($)} />
287        </Box>
288      </Box>
289    )
290  })
291
292  on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
293    const { Box, Button, Text } = $.ui.resolve(e)
294    const all = await read($, lessons)
295    const now = await read($, mode)
296    const counts = await read($, trial)
297    const shown = await read($, tab)
298    const last = await read($, note)
299    const width = Math.max(24, e.props.bodyColumns - 8)
300    const verdict = verdictOf(counts)
301
302    const tabButton = (to: Tab, label: string, hotkey: string) => (
303      <Button
304        key={`tab-${to}`}
305        label={label}
306        hotkey={hotkey}
307        plain
308        dimColor={shown !== to}
309        onPress={() => update($, tab, () => to)}
310      />
311    )
312    const modeButton = (to: Mode, hotkey: string) => (
313      <Button
314        key={`mode-${to}`}
315        label={to}
316        hotkey={hotkey}
317        plain
318        dimColor={now !== to}
319        onPress={() => setMode($, to)}
320      />
321    )
322
323    const list = (
324      <Box flexDirection="column" rowGap={1}>
325        {all.length === 0 && (
326          <Box flexDirection="column">
327            <Text>No lessons yet.</Text>
328            <Text dimColor>One is saved when a failing check is made to pass in a turn.</Text>
329            <Text dimColor>{'Or type one: /bank add the tests hang => run them with -x first'}</Text>
330          </Box>
331        )}
332        {[...all]
333          .reverse()
334          .slice(0, 12)
335          .map((lesson, index) => (
336            <Box flexDirection="column">
337              <Box flexDirection="row" columnGap={1}>
338                <Text dimColor>{index + 1}.</Text>
339                <Text>When {clip(lesson.when, width)}</Text>
340              </Box>
341              <Text dimColor>{clip(lesson.do, width + 4)}</Text>
342              <Box flexDirection="row" columnGap={2}>
343                <Text dimColor>
344                  used {lesson.uses} {lesson.uses === 1 ? 'time' : 'times'} {'·'} {lesson.source}
345                </Text>
346                <Button key={`forget-${lesson.id}`} label="forget" plain dimColor onPress={() => forget($, lesson.id)} />
347              </Box>
348            </Box>
349          ))}
350        {last !== '' && <Text dimColor>{clip(last, width + 4)}</Text>}
351      </Box>
352    )
353
354    const comparison = (
355      <Box flexDirection="column" rowGap={1}>
356        <Box flexDirection="column">
357          <Text dimColor>In test mode, turns a lesson fits alternate between getting it and not.</Text>
358          <Text dimColor>
359            The rule, set in advance: {MIN_TURNS} turns each way, and at least {Math.round(MIN_DROP * 100)}% fewer failed
360            checks a turn with lessons.
361          </Text>
362        </Box>
363        <Box flexDirection="column">
364          <Text>With lessons: {armLine(counts.with)}</Text>
365          <Text>Without: {armLine(counts.without)}</Text>
366        </Box>
367        {verdict.tone === null ? (
368          <Text dimColor>{verdict.text}</Text>
369        ) : (
370          <Text color={verdict.tone}>{verdict.text}</Text>
371        )}
372        <Button key="reset" label="r: reset the trial" hotkey="r" onPress={() => resetTrial($)} />
373      </Box>
374    )
375
376    return (
377      <Box flexDirection="column" rowGap={1}>
378        <Box flexDirection="row" columnGap={3}>
379          {tabButton('lessons', `Lessons (${all.length})`, '1')}
380          {tabButton('trial', 'Trial', '2')}
381        </Box>
382        <Box flexDirection="row" columnGap={2}>
383          <Text dimColor>Mode:</Text>
384          {modeButton('on', 'o')}
385          {modeButton('off', 'f')}
386          {modeButton('test', 't')}
387        </Box>
388        {shown === 'lessons' ? list : comparison}
389      </Box>
390    )
391  })
392}
393
hooks/bank.ts 268 lines
1// The bank's rules, with no engine in them: every function here is pure, so
2// the tests and the hooks read the same judgement. Recall is word overlap, not
3// a model: the same prompt always recalls the same lessons.
4
5import type { Arm, Lesson, Trial, TurnState } from '../types'
6
7export const MAX_LESSONS = 200
8export const RECALL = 3
9
10/** The turns each arm needs before the trial says anything. */
11export const MIN_TURNS = 30
12
13/** How much lower the failure rate must be with lessons to count as helping. */
14export const MIN_DROP = 0.2
15
16export const NO_TURN: TurnState = { arm: null, applied: [], checks: 0, failures: 0, hasRecovered: false }
17export const NO_TRIAL: Trial = {
18  with: { turns: 0, checks: 0, failures: 0 },
19  without: { turns: 0, checks: 0, failures: 0 },
20}
21
22// --------------------------------------------------------------- watching
23
24const TESTLIKE = new RegExp(
25  [
26    String.raw`\bpytest\b`,
27    String.raw`\bpy\.test\b`,
28    String.raw`\bpython3?\s+-m\s+(pytest|unittest)\b`,
29    String.raw`\b(tox|nox)\b`,
30    String.raw`\b(npm|pnpm|yarn|bun)\s+(run\s+)?(test|lint|typecheck|check|build)\b`,
31    String.raw`\bnpx\s+(jest|vitest|tsc|eslint|playwright)\b`,
32    String.raw`\b(jest|vitest|mocha)\b`,
33    String.raw`\bcargo\s+(test|check|clippy|build)\b`,
34    String.raw`\bgo\s+(test|vet|build)\b`,
35    String.raw`\bmake(\s+[\w-]+)?\s*$`,
36    String.raw`(\bmvn|\bgradle|\./gradlew)\s+(test|verify|check|build)\b`,
37    String.raw`\btsc\b`,
38    String.raw`\b(ruff|mypy|eslint|flake8)\b`,
39    String.raw`\b(rspec|phpunit|ctest)\b`,
40    String.raw`\b(dotnet|swift)\s+(test|build)\b`,
41    String.raw`\bclaude\s+plugin\s+(test|validate)\b`,
42  ].join('|'),
43)
44
45const NOT_A_RUN =
46  /^\s*(cat|grep|rg|ls|echo|which|head|tail|sed|awk|find|cd|export|git)\b|\b(install|uninstall|add|remove)\b/
47
48/** True when some part of the command runs tests, a build, a linter or a type check. */
49export function isCheck(command: string): boolean {
50  return command
51    .split(/&&|\|\||[;|\n]/)
52    .some(part => TESTLIKE.test(part) && !NOT_A_RUN.test(part))
53}
54
55/** Counts one check into the turn; a pass after a failure marks the turn as a recovery. */
56export function recordCheck(turn: TurnState, ok: boolean): TurnState {
57  return {
58    ...turn,
59    checks: turn.checks + 1,
60    failures: turn.failures + (ok ? 0 : 1),
61    hasRecovered: turn.hasRecovered || (ok && turn.failures > 0),
62  }
63}
64
65// ----------------------------------------------------------------- recall
66
67const STOP = new Set(
68  (
69    'the and for with that this from have has had not but are was were will would should could can ' +
70    'you your our its into out about when what which how why then than them they there here just ' +
71    'also any all some one two use using used get got make made need want please let lets now new ' +
72    'add fix run file files code does did done its it is in on at to of a an be as by or if so we i'
73  ).split(' '),
74)
75
76/** The words of a text that carry meaning, lowercased and lightly stemmed. */
77export function wordsOf(text: string): string[] {
78  const words = text.toLowerCase().match(/[a-z][a-z0-9_]{2,}/g) ?? []
79
80  return [
81    ...new Set(
82      words
83        .filter(word => !STOP.has(word))
84        .map(word =>
85          word.length > 5 && word.endsWith('ing')
86            ? word.slice(0, -3)
87            : word.length > 4 && word.endsWith('ed')
88              ? word.slice(0, -2)
89              : word.length > 3 && word.endsWith('s') && !word.endsWith('ss')
90                ? word.slice(0, -1)
91                : word,
92        ),
93    ),
94  ]
95}
96
97/** How well a lesson fits a prompt: a tag in common counts double a word in common. */
98export function scoreOf(lesson: Lesson, prompt: readonly string[]): number {
99  const asked = new Set(prompt)
100  const tags = new Set(wordsOf(lesson.tags.join(' ')))
101  const tagHits = [...tags].filter(tag => asked.has(tag)).length
102  const wordHits = wordsOf(lesson.when).filter(word => !tags.has(word) && asked.has(word)).length
103
104  // A lesson needs a tag in common and one thing more, or it is a coincidence.
105  return tagHits >= 1 && tagHits * 2 + wordHits >= 3 ? tagHits * 2 + wordHits : 0
106}
107
108/** The lessons that fit a prompt, best first, at most three. */
109export function recall(lessons: readonly Lesson[], prompt: string): Lesson[] {
110  const asked = wordsOf(prompt)
111
112  return lessons
113    .map(lesson => ({ lesson, score: scoreOf(lesson, asked) }))
114    .filter(one => one.score > 0)
115    .sort((a, b) => b.score - a.score || b.lesson.uses - a.lesson.uses || b.lesson.at - a.lesson.at)
116    .slice(0, RECALL)
117    .map(one => one.lesson)
118}
119
120/** What Claude reads beside a prompt the bank has lessons for. */
121export function briefing(lessons: readonly Lesson[]): string {
122  return [
123    'Lessons from earlier sessions in this project (ReasoningBank Lite). Each was saved when a failing check was made to pass.',
124    'Use one only where it fits this task, and say so when you do:',
125    ...lessons.map(lesson => `- When ${lesson.when}: ${lesson.do}`),
126  ].join('\n')
127}
128
129// ---------------------------------------------------------------- learning
130
131/** What the model is asked after a turn in which a failing check was made to pass. */
132export const QUESTION = [
133  'In the turn that just ended, a check failed and was then made to pass.',
134  'Write the one lesson from it that would save time in a future session on this project.',
135  'Reply with JSON only, on one line:',
136  '{"when": "the situation, under 120 characters", "do": "what to do, under 160 characters", "tags": ["three to six lowercase keywords someone would use when asking for similar work"]}.',
137  'If the failure was a one-off slip with nothing reusable in it, reply {"skip": true}.',
138  'Do not call any tool.',
139].join(' ')
140
141const tidy = (text: unknown, most: number): string =>
142  typeof text === 'string' ? text.replace(/\s+/g, ' ').trim().replace(/^when\s+/i, '').slice(0, most) : ''
143
144/** Reads the model's reply; answers the lesson's parts, or null when there is none. */
145export function parseLesson(reply: string): Pick<Lesson, 'when' | 'do' | 'tags'> | null {
146  const found = /\{[\s\S]*\}/.exec(reply)
147
148  if (found === null) return null
149
150  let data: unknown
151
152  try {
153    data = JSON.parse(found[0])
154  } catch {
155    return null
156  }
157
158  if (typeof data !== 'object' || data === null) return null
159
160  const { when, do: action, tags } = data as Record<string, unknown>
161  const lesson = {
162    when: tidy(when, 140),
163    do: tidy(action, 200),
164    tags: Array.isArray(tags)
165      ? [...new Set(tags.filter((tag): tag is string => typeof tag === 'string').map(tag => tag.toLowerCase().trim()))]
166          .filter(tag => /^[a-z0-9][a-z0-9 _.-]{1,30}$/.test(tag))
167          .slice(0, 6)
168      : [],
169  }
170
171  return lesson.when === '' || lesson.do === '' || lesson.tags.length === 0 ? null : lesson
172}
173
174/** Reads a lesson typed by hand: `when the tests hang => run them with -x first`. */
175export function parseTyped(text: string): Pick<Lesson, 'when' | 'do' | 'tags'> | null {
176  const [left, ...rest] = text.split('=>')
177  const when = tidy(left, 140)
178  const action = tidy(rest.join('=>'), 200)
179
180  if (when === '' || action === '') return null
181
182  return { when, do: action, tags: wordsOf(when).slice(0, 6) }
183}
184
185const plain = (text: string): string => text.toLowerCase().replace(/[^a-z0-9]+/g, ' ').trim()
186
187const same = (a: string, b: string): boolean => plain(a) === plain(b)
188
189/**
190 * Adds a lesson to the bank. One that says what a kept lesson says is not
191 * added twice; past the bank's size, the least used and oldest gives way.
192 */
193export function addLesson(
194  lessons: readonly Lesson[],
195  parts: Pick<Lesson, 'when' | 'do' | 'tags'>,
196  at: number,
197  source: Lesson['source'],
198): Lesson[] {
199  if (lessons.some(lesson => same(lesson.when, parts.when) && same(lesson.do, parts.do))) return [...lessons]
200
201  const lesson: Lesson = { id: `${at.toString(36)}-${lessons.length}`, at, uses: 0, source, ...parts }
202  const all = [...lessons, lesson]
203
204  if (all.length <= MAX_LESSONS) return all
205
206  const weakest = [...all.slice(0, -1)].sort((a, b) => a.uses - b.uses || a.at - b.at)[0]
207
208  return all.filter(one => one.id !== weakest?.id)
209}
210
211export const markUsed = (lessons: readonly Lesson[], ids: readonly string[]): Lesson[] =>
212  lessons.map(lesson => (ids.includes(lesson.id) ? { ...lesson, uses: lesson.uses + 1 } : lesson))
213
214// ------------------------------------------------------------------ trial
215
216/** Which arm the next eligible turn of a trial goes to: the two alternate. */
217export const nextArm = (trial: Trial): 'with' | 'without' =>
218  trial.with.turns <= trial.without.turns ? 'with' : 'without'
219
220export function recordTurn(trial: Trial, turn: TurnState): Trial {
221  if (turn.arm === null) return trial
222
223  const was: Arm = trial[turn.arm]
224
225  return {
226    ...trial,
227    [turn.arm]: { turns: was.turns + 1, checks: was.checks + turn.checks, failures: was.failures + turn.failures },
228  }
229}
230
231const rate = (arm: Arm): number => (arm.turns === 0 ? 0 : arm.failures / arm.turns)
232
233export type Verdict = { tone: 'success' | 'warning' | null; text: string }
234
235/**
236 * What the trial shows, by the rule set before any data came in: at least
237 * thirty turns in each arm, and a failure rate at least a fifth lower with
238 * lessons, or it is not a result.
239 */
240export function verdictOf(trial: Trial): Verdict {
241  const short = Math.min(trial.with.turns, trial.without.turns)
242
243  if (short < MIN_TURNS) {
244    return { tone: null, text: `Not enough turns yet: ${short} of ${MIN_TURNS} in the smaller arm.` }
245  }
246
247  const withRate = rate(trial.with)
248  const withoutRate = rate(trial.without)
249
250  if (withoutRate > 0 && withRate <= withoutRate * (1 - MIN_DROP)) {
251    return {
252      tone: 'success',
253      text: `Fewer failed checks with lessons: ${withRate.toFixed(2)} a turn against ${withoutRate.toFixed(2)}.`,
254    }
255  }
256
257  return {
258    tone: 'warning',
259    text: `No effect seen: ${withRate.toFixed(2)} failed checks a turn with lessons, ${withoutRate.toFixed(2)} without.`,
260  }
261}
262
263export const armLine = (arm: Arm): string =>
264  `${arm.turns} ${arm.turns === 1 ? 'turn' : 'turns'} · ${arm.failures} failed of ${arm.checks} checks · ${rate(arm).toFixed(2)} a turn`
265
266export const clip = (text: string, width: number): string =>
267  text.length > width ? `${text.slice(0, Math.max(1, width - 3))}...` : text
268
types/index.d.ts 50 lines
1/** One thing learned: a situation, and what to do in it. */
2export type Lesson = {
3  id: string
4  /** When it was saved, in milliseconds. */
5  at: number
6  /** The situation, completing "When ...". */
7  when: string
8  do: string
9  tags: string[]
10  /** How many prompts it has been handed to Claude beside. */
11  uses: number
12  /** `learned` from a turn by the model, or `typed` by the person. */
13  source: 'learned' | 'typed'
14}
15
16/** What one turn has seen so far. */
17export type TurnState = {
18  /** In a trial, which arm the turn is in; null when the turn is not part of one. */
19  arm: 'with' | 'without' | null
20  /** The lessons handed to Claude beside this turn's prompt. */
21  applied: string[]
22  checks: number
23  failures: number
24  /** True once a check passed after one failed. */
25  hasRecovered: boolean
26}
27
28export type Arm = { turns: number; checks: number; failures: number }
29
30/** The on/off comparison: turns with lessons against turns without. */
31export type Trial = { with: Arm; without: Arm }
32
33export type Mode = 'on' | 'off' | 'test'
34
35export type Tab = 'lessons' | 'trial'
36
37declare module 'claude-code' {
38  interface PluginState {
39    'reasoning-bank-mod': {
40      lessons: Lesson[]
41      mode: Mode
42      turn: TurnState
43      trial: Trial
44      tab: Tab
45      /** What the last attempt to learn came to, for the pane. */
46      note: string
47    }
48  }
49}
50