Query and build knowledge stores made with knowledge-store-builder. knowledge-store: ask a store questions (architecture, duplication, journeys, business…

Ask questions about a software estate and get answers that cite the code, the commits and the tickets behind them.
Which applications implement their own address formatting, and which tickets changed them?
<img src="docs/images/explorer-answering-a-question.png" alt="The explorer page answering 'how are addresses validated?'. A headline verdict reads: each application formats addresses with its own copy of AddressPipe; there is no shared implementation. Below it, sections headed How it works, Where it lives, and What this is NOT, then the business features in the area and the commits that changed them, each citing a repository, a file path or a ticket id." width="580">
That is explorer.html, one of the artefacts a build produces. It is a single file, it runs from file://, and it answers with no server, no network and no LLM. The screenshot is this repository's own test fixture, so you can produce it yourself: python3 tests/explorer/fixture.py.
Note the section headed What this is NOT. Two applications have a same-named AddressPipe and no edge connects them, so the store reports them as independent implementations rather than guessing they are shared. Absence of evidence is a finding here, not a silence.
Point the library at a GitHub organisation and choose the repositories. It drives graphify for the extraction itself, then enriches and indexes what comes back. The build writes static files you commit alongside the code:
| Artefact | What it holds |
|---|---|
graphify-out/graph.json | the estate graph, merged from graphify's per-repository extraction and enriched here with business features, package and deployment edges |
graphify-out/explorer.html | the self-contained page above |
knowledge/git-history/ | per-repository commit history as NDJSON |
knowledge/intent/ | which tickets changed which files |
knowledge/summaries/, docs/topics/, docs/deep-dives/ | prose an LLM wrote at build time from graph evidence, then reviewed |
Everything is a committed file. Consumers clone and read; nothing is computed at query time.
In a browser — open explorer.html. No install, no Claude licence, no network.
In Claude Code — install the plugin and ask in English.
/plugin marketplace add hmcts/knowledge-store-builder
/plugin install knowledge-store@knowledge-store-builder
/reload-plugins
/reload-plugins is not optional; the skills do not load without it. Then ask:
> how are addresses validated across these services?
Each application formats addresses with its own copy of AddressPipe.
There is no shared implementation.
Where it lives
demo-app-a src/pipes/address.pipe.ts (graph)
demo-app-b src/pipes/address.pipe.ts (graph)
What this is NOT
These two are not one component. They share a name and no edge
connects them, so the store reports them separately rather than
assuming they are the same.
Answered from: the graph, and the commit history for both files.
> export that as a finding I can send to the platform team
The repositories and the file are the ones in tests/explorer/fixture.py, which builds the page in the screenshot. Every line names the layer it came from, and an answer the store cannot support says so instead of filling the gap.
The plugin also carries skills for building and refreshing a store, exporting a finding, and assessing a backlog of tickets against what the platform now does.
From the terminal — graphify query against the committed graph.
The library ships the knowledgestore command, one stage per step. A build is that sequence run in order, and each stage writes files the next one reads:
knowledgestore discover # list the estate's repositories
knowledgestore sync # clone or update them
knowledgestore extract-ast # the code layer, one repository at a time
knowledgestore export-history # per-repository commit history
knowledgestore intent # join files to the tickets that changed them
knowledgestore explorer # build the page
knowledgestore status # what is present, what is stale
knowledgestore with no arguments lists every stage with a line each. knowledgestore <stage> --help explains one. The full sequence, the extraction extras and the authoring steps are in Creating a knowledge store, which also carries the install command — the guides own install detail so there is one copy to keep correct.
Building needs Python 3.10+, Git, the GitHub CLI and graphify, which does the extraction. This library prepares its inputs and enriches its output; it does not re-implement it.
| You want to | Go to |
|---|---|
| ask questions about a store someone built | Asking questions |
| build a store for your estate | Creating a knowledge store |
| refresh a store you maintain | Refreshing a store |
| see every command with nothing around it | CHEATSHEET.md |
Asking needs the plugin and nothing else — no Python, no pip. Without a Claude licence, explorer.html answers in a browser.
explorer.html is committed for them. Everything an LLM writes during the build is committed as reviewed static text. Claude Code reads the same evidence when a question needs a new prose answer.| Document | For |
|---|---|
CHEATSHEET.md | the commands, per surface, with nothing else around them |
docs/asking-questions.md | asking questions with Claude Code, explorer.html or graphify query |
docs/creating-a-store.md | creating, building and publishing a new store |
docs/refreshing-a-store.md | refreshing an existing store and changing its pinned library version |
docs/configuring-a-store.md | pipeline settings, BDD support and stage outputs |
docs/building-a-knowledge-store.md | the operator's judgement: defining an estate, what extraction yields, refresh economics, the traps |
docs/grounding-and-verification.md | whether a store's answers are fact-based, and how to verify subagent-authored content |
docs/retrieval-architecture.md | how this differs from vector RAG, and where each answer layer lives |
docs/how-it-works.md | the science: each mechanism, its constants, and where its behaviour is proven |
CLAUDE.md | working on this repository: the dev install, the checks, and what has bitten us |
MIT. See LICENSE.
hooks/register.ts 29 lines1import type { Register } from "claude-code";
2// @ts-expect-error - a plain ES module beside this one, checked by its own harness
3import { decide } from "./guards.mjs";
4
5// A literal ref, because `claude plugin validate` holds every key to the contract.
6const MERGE_INPUTS_RAN = { plugin: "knowledge-store", key: "mergeInputsRan" } as const;
7
8export const register: Register = (on) => {
9 on("tool.call", { tool: "Bash" }, async ($, e, next) => {
10 try {
11 const command = String(e.command ?? "");
12 const held = await $.state.get(MERGE_INPUTS_RAN);
13 const verdict = decide({
14 command,
15 state: { mergeInputsRan: held.value === true },
16 });
17 if (!verdict.allow) return { deny: verdict.deny };
18 if (/\bknowledgestore\s+merge-inputs\b/.test(command)) {
19 await $.state.set(MERGE_INPUTS_RAN, true);
20 }
21 } catch {
22 // Fail open. These guards are an early warning, not a gate: the real
23 // checks are in CI and the library, so a fault here must never stop a
24 // command the operator is entitled to run.
25 }
26 return next(e);
27 });
28};
29hooks/guards.mjs 233 lines1// Refuses command shapes this library has documented as destructive. Pure:
2// it reads nothing, writes nothing and calls nothing outside its arguments.
3
4/** @param {unknown} value @returns {string} */
5function asText(value) {
6 return typeof value === "string" ? value : "";
7}
8
9/** Blank the contents of quoted spans, keeping the quotes and the length, so
10 * a separator or a flag inside a string cannot manufacture a segment or a
11 * token. Masking hides text, so on its own it can only stop a guard firing.
12 * The one place that hiding would instead cause a refusal is the graphify-out
13 * exclusion, whose value may be quoted: that is read from the unmasked text
14 * (see chainLinks and unexcludedClean), so the invariant holds there too. */
15/** @param {unknown} command @returns {string} */
16export function maskQuoted(command) {
17 let out = "";
18 let quote = null;
19 const text = asText(command);
20 // Per UTF-16 unit, not per code point: an astral character is two units, and
21 // chainLinks slices the original at offsets found in this string.
22 for (const ch of text.split("")) {
23 if (quote) {
24 out += ch === quote ? ch : "X";
25 if (ch === quote) quote = null;
26 } else if (ch === "'" || ch === '"') {
27 quote = ch;
28 out += ch;
29 } else {
30 out += ch;
31 }
32 }
33 return out;
34}
35
36/** Segments with the separator that precedes each one. `text` is masked, `raw`
37 * is the same slice of the original command, and `before` is "&&", "||", ";"
38 * or "\n" (a bare line break), null for the first. */
39/** @param {unknown} command @returns {{ text: string, raw: string, before: string|null }[]} */
40export function chainLinks(command) {
41 const original = asText(command);
42 const masked = maskQuoted(original);
43 const links = [];
44 // An operator wins over a line break, so `a &&\nb` keeps its `&&`.
45 const separator = /(&&|\|\||;)[ \t\r\n]*|(\r?\n)/g;
46 let last = 0;
47 let before = null;
48 let match;
49 while ((match = separator.exec(masked)) !== null) {
50 links.push({
51 text: masked.slice(last, match.index).trim(),
52 raw: original.slice(last, match.index).trim(),
53 before,
54 });
55 before = match[1] ?? "\n";
56 last = match.index + match[0].length;
57 }
58 links.push({ text: masked.slice(last).trim(), raw: original.slice(last).trim(), before });
59 return links.filter((link) => link.text);
60}
61
62/** Split a command into its &&, ||, ; and line-break separated segments. */
63/** @param {unknown} command @returns {string[]} */
64export function chainSegments(command) {
65 return chainLinks(command).map((link) => link.text);
66}
67
68/** The first command of a segment's pipeline. */
69/** @param {unknown} segment @returns {string} */
70export function pipelineHead(segment) {
71 return asText(segment).split("|")[0].trim();
72}
73
74/** A segment's whitespace-separated tokens. Quoting is not honoured, which
75 * can only cause a guard to stay silent, never to fire wrongly. */
76/** @param {unknown} segment @returns {string[]} */
77export function argsOf(segment) {
78 const t = asText(segment).trim();
79 return t ? t.split(/\s+/) : [];
80}
81
82const EXCLUDES_GRAPHIFY_OUT = /(?:--exclude|-e)[=\s]*["']?graphify-out["']?/;
83
84/** @param {string} argument @returns {boolean} */
85function isFlag(argument) {
86 return argument.startsWith("-");
87}
88
89/** True when the clean names a path to clean, which cannot reach
90 * graphify-out at the repository root. The value after -e or --exclude is
91 * that flag's, not a pathspec. */
92/** @param {string[]} args @returns {boolean} */
93function hasPathspec(args) {
94 const after = args.slice(args.indexOf("clean") + 1);
95 for (let i = 0; i < after.length; i++) {
96 if (after[i] === "-e" || after[i] === "--exclude") {
97 i++;
98 continue;
99 }
100 if (!isFlag(after[i])) return true;
101 }
102 return false;
103}
104
105/** @param {string} segment @param {{ mergeInputsRan?: boolean }} _state @param {string|undefined} command @param {string} [raw] the segment's unmasked text @returns {string|null} */
106function unexcludedClean(segment, _state, command, raw) {
107 const args = argsOf(segment);
108 if (args[0] !== "git" || !args.includes("clean")) return null;
109 const flags = args.filter((a) => /^-[^-]/.test(a)).join("");
110 if (!(flags.includes("f") && flags.includes("d"))) return null;
111 if (flags.includes("n") || args.includes("--dry-run")) return null;
112 if (EXCLUDES_GRAPHIFY_OUT.test(raw ?? segment)) return null;
113 if (hasPathspec(args)) return null;
114 // The context is read from the whole command, not this segment: the form an
115 // agent writes is `cd repositories/<name> && git clean -fd`, where the cd is
116 // a different segment. A clone and a store root that merely sits under a
117 // directory named repositories share a path shape, so cwd cannot tell them
118 // apart and is not consulted.
119 if (!/(^|\s)repositories\//.test(maskQuoted(command))) return null;
120 return (
121 "This clean would delete the per-repo graphs. They live untracked at " +
122 "repositories/<name>/graphify-out/, and a clean without the exclusion " +
123 "forces a full re-extraction; it has destroyed 61 of 81 graphs here once. " +
124 "Run: git clean -fd -e graphify-out"
125 );
126}
127
128/** @param {string} segment @param {{ mergeInputsRan?: boolean }} _state @param {string|undefined} _command @param {string} [_raw] @returns {string|null} */
129function indiscriminateStage(segment, _state, _command, _raw) {
130 const args = argsOf(segment);
131 if (args[0] !== "git" || args[1] !== "add") return null;
132 const rest = args.slice(2);
133 if (!rest.some((a) => a === "-A" || a === "--all" || a === ".")) return null;
134 return (
135 "Stage explicit paths. The working tree here routinely holds generated " +
136 "pipeline output, another session's edits and scratch files at once, and " +
137 "-A cannot tell them apart. Read git status --short, then name the paths."
138 );
139}
140
141const EXTRACTION_VERBS = new Set(["update", "extract"]);
142
143/** @param {string[]} args @returns {boolean} */
144function asksForHelp(args) {
145 return args.includes("-h") || args.includes("--help");
146}
147
148/** @param {string} segment @param {{ mergeInputsRan?: boolean }} _state @param {string|undefined} _command @param {string} [_raw] @returns {string|null} */
149function outsideExtraction(segment, _state, _command, _raw) {
150 const args = argsOf(segment);
151 if (asksForHelp(args)) return null;
152 // Only the extraction verbs. merge-graphs is given repositories/*/... by the
153 // build skill itself, so matching every graphify subcommand would refuse the
154 // documented merge.
155 if (args[0] !== "graphify" || !EXTRACTION_VERBS.has(args[1])) return null;
156 // A flag's value is that flag's, not a path to extract.
157 const rest = args.slice(2);
158 if (!rest.some((a, i) => a.startsWith("repositories/") && !(i > 0 && isFlag(rest[i - 1])))) return null;
159 return (
160 "Extract from inside the repository. A path like repositories/<name> " +
161 "prefixes every source_file with it, which breaks the file-to-ticket join " +
162 "silently - the only symptom is that nodes lose their tickets. " +
163 "Run: ( cd repositories/<name> && graphify update . )"
164 );
165}
166
167// Real checkers only. A general-purpose interpreter in a pipeline is ordinary
168// work, not a gate, so node, npm, npx and bare python are deliberately absent.
169const CHECKERS = /^(?:pytest|ruff|pyright|tsc|eslint|mypy)\b|^python3?\s+-m\s+(?:unittest|pytest)\b/;
170const PUBLISHES = /^git\s+(?:push|commit)\b/;
171
172/** Reads the whole command rather than one segment: the hazard is the
173 * relationship between a piped checker and the commit or push straight after
174 * it. Only an immediate `&&` gates it - `;` and a line break run the next
175 * command whatever happened, and `||` runs it only on failure, so none of
176 * them can be masked by the pipe. */
177/** @param {string|undefined} command @returns {string|null} */
178function pipedGate(command) {
179 const links = chainLinks(command);
180 for (let i = 0; i < links.length; i++) {
181 if (!links[i].text.includes("|")) continue;
182 if (!CHECKERS.test(pipelineHead(links[i].text))) continue;
183 // Only a publish immediately `&&`-chained is ungated. Any link between the
184 // checker and the publish is reading the saved output, which is the remedy.
185 const next = links[i + 1];
186 if (next?.before === "&&" && PUBLISHES.test(next.text)) {
187 return (
188 "A pipeline's exit status is its last command's, not the checker's, " +
189 "so this commits or pushes whatever the checker did. Redirect the " +
190 "checker to a file and read it, or test ${PIPESTATUS[0]}."
191 );
192 }
193 }
194 return null;
195}
196
197/** @param {string} segment @param {{ mergeInputsRan?: boolean }} state @param {string|undefined} _command @param {string} [_raw] @returns {string|null} */
198function unreconciledMerge(segment, state, _command, _raw) {
199 const args = argsOf(segment);
200 if (asksForHelp(args)) return null;
201 if (args[0] !== "graphify" || args[1] !== "merge-graphs") return null;
202 if (state?.mergeInputsRan) return null;
203 return (
204 "Run knowledgestore merge-inputs first and read its output. The merge is " +
205 "driven by a shell glob, and a glob has picked up a previous run's outputs " +
206 "here before - the merge reported a healthy count over the wrong inputs. " +
207 "This guard does not inspect any graph; merge-inputs makes that judgement."
208 );
209}
210
211const GUARDS = [unexcludedClean, indiscriminateStage, outsideExtraction, unreconciledMerge];
212
213/** @typedef {{ allow: true } | { allow: false, deny: string }} Verdict */
214
215const RAN_MERGE_INPUTS = /\bknowledgestore\s+merge-inputs\b/;
216
217/** @param {{ command?: string, state?: { mergeInputsRan?: boolean } }} [input] @returns {Verdict} */
218export function decide({ command, state = {} } = {}) {
219 const whole = pipedGate(command);
220 if (whole) return { allow: false, deny: whole };
221 const masked = maskQuoted(command);
222 let mergeInputsRan = state.mergeInputsRan === true;
223 for (const link of chainLinks(command)) {
224 for (const guard of GUARDS) {
225 const deny = guard(link.text, { mergeInputsRan }, masked, link.raw);
226 if (deny) return { allow: false, deny };
227 }
228 // Checked before marking, so merge-inputs after the merge does not count.
229 if (RAN_MERGE_INPUTS.test(link.text)) mergeInputsRan = true;
230 }
231 return { allow: true };
232}
233types/index.d.ts 9 lines1declare module "claude-code" {
2 interface PluginState {
3 "knowledge-store": {
4 /** Set once `knowledgestore merge-inputs` has run in this session. */
5 mergeInputsRan?: boolean;
6 };
7 }
8}
9