SLOPSHOPPER

knowledge-store

Query and build knowledge stores made with knowledge-store-builder. knowledge-store: ask a store questions (architecture, duplication, journeys, business…

newguard
v?MITupdated 2026-10-09hmcts/knowledge-store-builder
A shopper browsing a rack in a slop shop
README

knowledge-store-builder

Ask questions about a software estate and get answers that cite the code, the commits and the tickets behind them.

Which applications implement their own address formatting, and which tickets changed them?

<img src="docs/images/explorer-answering-a-question.png" alt="The explorer page answering 'how are addresses validated?'. A headline verdict reads: each application formats addresses with its own copy of AddressPipe; there is no shared implementation. Below it, sections headed How it works, Where it lives, and What this is NOT, then the business features in the area and the commits that changed them, each citing a repository, a file path or a ticket id." width="580">

That is explorer.html, one of the artefacts a build produces. It is a single file, it runs from file://, and it answers with no server, no network and no LLM. The screenshot is this repository's own test fixture, so you can produce it yourself: python3 tests/explorer/fixture.py.

Note the section headed What this is NOT. Two applications have a same-named AddressPipe and no edge connects them, so the store reports them as independent implementations rather than guessing they are shared. Absence of evidence is a finding here, not a silence.

What a build produces

Point the library at a GitHub organisation and choose the repositories. It drives graphify for the extraction itself, then enriches and indexes what comes back. The build writes static files you commit alongside the code:

ArtefactWhat it holds
graphify-out/graph.jsonthe estate graph, merged from graphify's per-repository extraction and enriched here with business features, package and deployment edges
graphify-out/explorer.htmlthe self-contained page above
knowledge/git-history/per-repository commit history as NDJSON
knowledge/intent/which tickets changed which files
knowledge/summaries/, docs/topics/, docs/deep-dives/prose an LLM wrote at build time from graph evidence, then reviewed

Everything is a committed file. Consumers clone and read; nothing is computed at query time.

The three ways to ask

In a browser — open explorer.html. No install, no Claude licence, no network.

In Claude Code — install the plugin and ask in English.

/plugin marketplace add hmcts/knowledge-store-builder
/plugin install knowledge-store@knowledge-store-builder
/reload-plugins

/reload-plugins is not optional; the skills do not load without it. Then ask:

> how are addresses validated across these services?

  Each application formats addresses with its own copy of AddressPipe.
  There is no shared implementation.

  Where it lives
    demo-app-a  src/pipes/address.pipe.ts         (graph)
    demo-app-b  src/pipes/address.pipe.ts         (graph)

  What this is NOT
    These two are not one component. They share a name and no edge
    connects them, so the store reports them separately rather than
    assuming they are the same.

  Answered from: the graph, and the commit history for both files.

> export that as a finding I can send to the platform team

The repositories and the file are the ones in tests/explorer/fixture.py, which builds the page in the screenshot. Every line names the layer it came from, and an answer the store cannot support says so instead of filling the gap.

The plugin also carries skills for building and refreshing a store, exporting a finding, and assessing a backlog of tickets against what the platform now does.

From the terminal — graphify query against the committed graph.

Building a store

The library ships the knowledgestore command, one stage per step. A build is that sequence run in order, and each stage writes files the next one reads:

knowledgestore discover        # list the estate's repositories
knowledgestore sync            # clone or update them
knowledgestore extract-ast     # the code layer, one repository at a time
knowledgestore export-history  # per-repository commit history
knowledgestore intent          # join files to the tickets that changed them
knowledgestore explorer        # build the page
knowledgestore status          # what is present, what is stale

knowledgestore with no arguments lists every stage with a line each. knowledgestore <stage> --help explains one. The full sequence, the extraction extras and the authoring steps are in Creating a knowledge store, which also carries the install command — the guides own install detail so there is one copy to keep correct.

Building needs Python 3.10+, Git, the GitHub CLI and graphify, which does the extraction. This library prepares its inputs and enriches its output; it does not re-implement it.

Start here

You want toGo to
ask questions about a store someone builtAsking questions
build a store for your estateCreating a knowledge store
refresh a store you maintainRefreshing a store
see every command with nothing around itCHEATSHEET.md

Asking needs the plugin and nothing else — no Python, no pip. Without a Claude licence, explorer.html answers in a browser.

How it is designed

  • The store is the product. Outputs are committed static files. Consumers clone and read; nothing is built at query time.
  • The browser has no query-time LLM. Whoever builds a store may have a licence; the people querying it may not, and explorer.html is committed for them. Everything an LLM writes during the build is committed as reviewed static text. Claude Code reads the same evidence when a question needs a new prose answer.
  • Deterministic where it can be. Extraction, indexing and page composition are pure functions of the sources; two runs on the same inputs produce byte-identical output.
  • Per-commit history stays out of the graph. It is exported alongside as NDJSON, because "what changed last sprint" is a dataset query, not a graph traversal — and it keeps the committed graph an order of magnitude smaller.
  • Absence of evidence is a finding. Same-named components with no connecting edge are independent implementations, and the tooling says that rather than guessing.

Reference documentation

DocumentFor
CHEATSHEET.mdthe commands, per surface, with nothing else around them
docs/asking-questions.mdasking questions with Claude Code, explorer.html or graphify query
docs/creating-a-store.mdcreating, building and publishing a new store
docs/refreshing-a-store.mdrefreshing an existing store and changing its pinned library version
docs/configuring-a-store.mdpipeline settings, BDD support and stage outputs
docs/building-a-knowledge-store.mdthe operator's judgement: defining an estate, what extraction yields, refresh economics, the traps
docs/grounding-and-verification.mdwhether a store's answers are fact-based, and how to verify subagent-authored content
docs/retrieval-architecture.mdhow this differs from vector RAG, and where each answer layer lives
docs/how-it-works.mdthe science: each mechanism, its constants, and where its behaviour is proven
CLAUDE.mdworking on this repository: the dev install, the checks, and what has bitten us

Licence

MIT. See LICENSE.

Source 3 files
hooks/register.ts 29 lines
1import type { Register } from "claude-code";
2// @ts-expect-error - a plain ES module beside this one, checked by its own harness
3import { decide } from "./guards.mjs";
4
5// A literal ref, because `claude plugin validate` holds every key to the contract.
6const MERGE_INPUTS_RAN = { plugin: "knowledge-store", key: "mergeInputsRan" } as const;
7
8export const register: Register = (on) => {
9  on("tool.call", { tool: "Bash" }, async ($, e, next) => {
10    try {
11      const command = String(e.command ?? "");
12      const held = await $.state.get(MERGE_INPUTS_RAN);
13      const verdict = decide({
14        command,
15        state: { mergeInputsRan: held.value === true },
16      });
17      if (!verdict.allow) return { deny: verdict.deny };
18      if (/\bknowledgestore\s+merge-inputs\b/.test(command)) {
19        await $.state.set(MERGE_INPUTS_RAN, true);
20      }
21    } catch {
22      // Fail open. These guards are an early warning, not a gate: the real
23      // checks are in CI and the library, so a fault here must never stop a
24      // command the operator is entitled to run.
25    }
26    return next(e);
27  });
28};
29
hooks/guards.mjs 233 lines
1// Refuses command shapes this library has documented as destructive. Pure:
2// it reads nothing, writes nothing and calls nothing outside its arguments.
3
4/** @param {unknown} value @returns {string} */
5function asText(value) {
6  return typeof value === "string" ? value : "";
7}
8
9/** Blank the contents of quoted spans, keeping the quotes and the length, so
10 *  a separator or a flag inside a string cannot manufacture a segment or a
11 *  token. Masking hides text, so on its own it can only stop a guard firing.
12 *  The one place that hiding would instead cause a refusal is the graphify-out
13 *  exclusion, whose value may be quoted: that is read from the unmasked text
14 *  (see chainLinks and unexcludedClean), so the invariant holds there too. */
15/** @param {unknown} command @returns {string} */
16export function maskQuoted(command) {
17  let out = "";
18  let quote = null;
19  const text = asText(command);
20  // Per UTF-16 unit, not per code point: an astral character is two units, and
21  // chainLinks slices the original at offsets found in this string.
22  for (const ch of text.split("")) {
23    if (quote) {
24      out += ch === quote ? ch : "X";
25      if (ch === quote) quote = null;
26    } else if (ch === "'" || ch === '"') {
27      quote = ch;
28      out += ch;
29    } else {
30      out += ch;
31    }
32  }
33  return out;
34}
35
36/** Segments with the separator that precedes each one. `text` is masked, `raw`
37 *  is the same slice of the original command, and `before` is "&&", "||", ";"
38 *  or "\n" (a bare line break), null for the first. */
39/** @param {unknown} command @returns {{ text: string, raw: string, before: string|null }[]} */
40export function chainLinks(command) {
41  const original = asText(command);
42  const masked = maskQuoted(original);
43  const links = [];
44  // An operator wins over a line break, so `a &&\nb` keeps its `&&`.
45  const separator = /(&&|\|\||;)[ \t\r\n]*|(\r?\n)/g;
46  let last = 0;
47  let before = null;
48  let match;
49  while ((match = separator.exec(masked)) !== null) {
50    links.push({
51      text: masked.slice(last, match.index).trim(),
52      raw: original.slice(last, match.index).trim(),
53      before,
54    });
55    before = match[1] ?? "\n";
56    last = match.index + match[0].length;
57  }
58  links.push({ text: masked.slice(last).trim(), raw: original.slice(last).trim(), before });
59  return links.filter((link) => link.text);
60}
61
62/** Split a command into its &&, ||, ; and line-break separated segments. */
63/** @param {unknown} command @returns {string[]} */
64export function chainSegments(command) {
65  return chainLinks(command).map((link) => link.text);
66}
67
68/** The first command of a segment's pipeline. */
69/** @param {unknown} segment @returns {string} */
70export function pipelineHead(segment) {
71  return asText(segment).split("|")[0].trim();
72}
73
74/** A segment's whitespace-separated tokens. Quoting is not honoured, which
75 *  can only cause a guard to stay silent, never to fire wrongly. */
76/** @param {unknown} segment @returns {string[]} */
77export function argsOf(segment) {
78  const t = asText(segment).trim();
79  return t ? t.split(/\s+/) : [];
80}
81
82const EXCLUDES_GRAPHIFY_OUT = /(?:--exclude|-e)[=\s]*["']?graphify-out["']?/;
83
84/** @param {string} argument @returns {boolean} */
85function isFlag(argument) {
86  return argument.startsWith("-");
87}
88
89/** True when the clean names a path to clean, which cannot reach
90 *  graphify-out at the repository root. The value after -e or --exclude is
91 *  that flag's, not a pathspec. */
92/** @param {string[]} args @returns {boolean} */
93function hasPathspec(args) {
94  const after = args.slice(args.indexOf("clean") + 1);
95  for (let i = 0; i < after.length; i++) {
96    if (after[i] === "-e" || after[i] === "--exclude") {
97      i++;
98      continue;
99    }
100    if (!isFlag(after[i])) return true;
101  }
102  return false;
103}
104
105/** @param {string} segment @param {{ mergeInputsRan?: boolean }} _state @param {string|undefined} command @param {string} [raw] the segment's unmasked text @returns {string|null} */
106function unexcludedClean(segment, _state, command, raw) {
107  const args = argsOf(segment);
108  if (args[0] !== "git" || !args.includes("clean")) return null;
109  const flags = args.filter((a) => /^-[^-]/.test(a)).join("");
110  if (!(flags.includes("f") && flags.includes("d"))) return null;
111  if (flags.includes("n") || args.includes("--dry-run")) return null;
112  if (EXCLUDES_GRAPHIFY_OUT.test(raw ?? segment)) return null;
113  if (hasPathspec(args)) return null;
114  // The context is read from the whole command, not this segment: the form an
115  // agent writes is `cd repositories/<name> && git clean -fd`, where the cd is
116  // a different segment. A clone and a store root that merely sits under a
117  // directory named repositories share a path shape, so cwd cannot tell them
118  // apart and is not consulted.
119  if (!/(^|\s)repositories\//.test(maskQuoted(command))) return null;
120  return (
121    "This clean would delete the per-repo graphs. They live untracked at " +
122    "repositories/<name>/graphify-out/, and a clean without the exclusion " +
123    "forces a full re-extraction; it has destroyed 61 of 81 graphs here once. " +
124    "Run: git clean -fd -e graphify-out"
125  );
126}
127
128/** @param {string} segment @param {{ mergeInputsRan?: boolean }} _state @param {string|undefined} _command @param {string} [_raw] @returns {string|null} */
129function indiscriminateStage(segment, _state, _command, _raw) {
130  const args = argsOf(segment);
131  if (args[0] !== "git" || args[1] !== "add") return null;
132  const rest = args.slice(2);
133  if (!rest.some((a) => a === "-A" || a === "--all" || a === ".")) return null;
134  return (
135    "Stage explicit paths. The working tree here routinely holds generated " +
136    "pipeline output, another session's edits and scratch files at once, and " +
137    "-A cannot tell them apart. Read git status --short, then name the paths."
138  );
139}
140
141const EXTRACTION_VERBS = new Set(["update", "extract"]);
142
143/** @param {string[]} args @returns {boolean} */
144function asksForHelp(args) {
145  return args.includes("-h") || args.includes("--help");
146}
147
148/** @param {string} segment @param {{ mergeInputsRan?: boolean }} _state @param {string|undefined} _command @param {string} [_raw] @returns {string|null} */
149function outsideExtraction(segment, _state, _command, _raw) {
150  const args = argsOf(segment);
151  if (asksForHelp(args)) return null;
152  // Only the extraction verbs. merge-graphs is given repositories/*/... by the
153  // build skill itself, so matching every graphify subcommand would refuse the
154  // documented merge.
155  if (args[0] !== "graphify" || !EXTRACTION_VERBS.has(args[1])) return null;
156  // A flag's value is that flag's, not a path to extract.
157  const rest = args.slice(2);
158  if (!rest.some((a, i) => a.startsWith("repositories/") && !(i > 0 && isFlag(rest[i - 1])))) return null;
159  return (
160    "Extract from inside the repository. A path like repositories/<name> " +
161    "prefixes every source_file with it, which breaks the file-to-ticket join " +
162    "silently - the only symptom is that nodes lose their tickets. " +
163    "Run: ( cd repositories/<name> && graphify update . )"
164  );
165}
166
167// Real checkers only. A general-purpose interpreter in a pipeline is ordinary
168// work, not a gate, so node, npm, npx and bare python are deliberately absent.
169const CHECKERS = /^(?:pytest|ruff|pyright|tsc|eslint|mypy)\b|^python3?\s+-m\s+(?:unittest|pytest)\b/;
170const PUBLISHES = /^git\s+(?:push|commit)\b/;
171
172/** Reads the whole command rather than one segment: the hazard is the
173 *  relationship between a piped checker and the commit or push straight after
174 *  it. Only an immediate `&&` gates it - `;` and a line break run the next
175 *  command whatever happened, and `||` runs it only on failure, so none of
176 *  them can be masked by the pipe. */
177/** @param {string|undefined} command @returns {string|null} */
178function pipedGate(command) {
179  const links = chainLinks(command);
180  for (let i = 0; i < links.length; i++) {
181    if (!links[i].text.includes("|")) continue;
182    if (!CHECKERS.test(pipelineHead(links[i].text))) continue;
183    // Only a publish immediately `&&`-chained is ungated. Any link between the
184    // checker and the publish is reading the saved output, which is the remedy.
185    const next = links[i + 1];
186    if (next?.before === "&&" && PUBLISHES.test(next.text)) {
187      return (
188        "A pipeline's exit status is its last command's, not the checker's, " +
189        "so this commits or pushes whatever the checker did. Redirect the " +
190        "checker to a file and read it, or test ${PIPESTATUS[0]}."
191      );
192    }
193  }
194  return null;
195}
196
197/** @param {string} segment @param {{ mergeInputsRan?: boolean }} state @param {string|undefined} _command @param {string} [_raw] @returns {string|null} */
198function unreconciledMerge(segment, state, _command, _raw) {
199  const args = argsOf(segment);
200  if (asksForHelp(args)) return null;
201  if (args[0] !== "graphify" || args[1] !== "merge-graphs") return null;
202  if (state?.mergeInputsRan) return null;
203  return (
204    "Run knowledgestore merge-inputs first and read its output. The merge is " +
205    "driven by a shell glob, and a glob has picked up a previous run's outputs " +
206    "here before - the merge reported a healthy count over the wrong inputs. " +
207    "This guard does not inspect any graph; merge-inputs makes that judgement."
208  );
209}
210
211const GUARDS = [unexcludedClean, indiscriminateStage, outsideExtraction, unreconciledMerge];
212
213/** @typedef {{ allow: true } | { allow: false, deny: string }} Verdict */
214
215const RAN_MERGE_INPUTS = /\bknowledgestore\s+merge-inputs\b/;
216
217/** @param {{ command?: string, state?: { mergeInputsRan?: boolean } }} [input] @returns {Verdict} */
218export function decide({ command, state = {} } = {}) {
219  const whole = pipedGate(command);
220  if (whole) return { allow: false, deny: whole };
221  const masked = maskQuoted(command);
222  let mergeInputsRan = state.mergeInputsRan === true;
223  for (const link of chainLinks(command)) {
224    for (const guard of GUARDS) {
225      const deny = guard(link.text, { mergeInputsRan }, masked, link.raw);
226      if (deny) return { allow: false, deny };
227    }
228    // Checked before marking, so merge-inputs after the merge does not count.
229    if (RAN_MERGE_INPUTS.test(link.text)) mergeInputsRan = true;
230  }
231  return { allow: true };
232}
233
types/index.d.ts 9 lines
1declare module "claude-code" {
2  interface PluginState {
3    "knowledge-store": {
4      /** Set once `knowledgestore merge-inputs` has run in this session. */
5      mergeInputsRan?: boolean;
6    };
7  }
8}
9