GUARDRAILS.md gate: a System One decision model judges every Bash command before it runs — destructive / credentials / GUARDRAILS.md policy

Stops your coding agent from running the bash command you'd regret.
Every bash command your coding agent (opencode, pi or Claude Code) is about to execute is scored by a small decision model first, in about 100 ms. If the command destroys data, leaks a secret, or breaks a rule in your repo's GUARDRAILS.md, the call is blocked and the agent is told why, so it can pick another route.
SystemOne-gate: blocked command — destructive=0.98 > 0.7
rm -rf ./important-data
Judged destructive/leaking on its own merits — named exceptions in
GUARDRAILS.md do not override this. If it is intended, your human
can run it directly, or restart opencode with SYSTEMONE_GATE=off
for a session that needs it. The agent must not edit GUARDRAILS.md
to work around this.
Two situations where a pattern list leaves you exposed.
Teams that work close to production. The agent runs kubectl, helm and git push all day, and the difference between routine and incident is your team's policy, not a filename pattern. The gate reads that policy from GUARDRAILS.md and applies it to every command.
Background agents that read mail, tickets or forms. Their input is written by strangers, and a crafted message can talk an agent into running something harmful. No one is watching when it happens. The gate does not promise that the agent cannot be persuaded — it promises that the boundary holds anyway, because the gate judges the command and your rules, not the conversation that led to it.
This project is unrelated to the guardrails.md convention, where the agent appends its own lessons to GUARDRAILS.md. Here the file is written by humans, so if your repo already follows that convention, the edit tools will refuse the agent's writes to the file once the gate is installed (bash writes are still judged, not hard-blocked; see Protected files).
Two steps. First teach the gate your rules, then put it in the harness.
1. Add GUARDRAILS.md to your repo (root, .opencode/GUARDRAILS.md, .pi/GUARDRAILS.md or .claude/GUARDRAILS.md; the legacy lowercase guardrails.md spellings are still read). Write it yourself — the value is in deciding what your team actually allows, not in shipping a generic file. The example below is a starting point for the shape:
# Guardrails for agents in this repo
## The agent MUST NOT
- Edit this file (GUARDRAILS.md) itself — it is written and changed by humans, through review.
- Change anything in production — production changes reach production only through Git/CD.
- Push directly to the main branch — all changes go through pull request.
- Install software outside the project's declared dependencies.
- Send data to external services outside our approved list (docs/approved-domains.md).
- Run irreversible operations against shared systems — deletions, cleanup, resets.
## The agent MAY
- Inspect any environment read-only.
- Run tests, lint, and builds locally.
- Create branches and push feature branches.
- Read documentation from the approved sources in docs/approved-domains.md.
The gate reads the file once at session start, so restart the harness after editing. Only the first 2000 characters are sent to the model, so keep the file short and put the MUST NOT rules first. If the file is longer, the gate warns you at startup and in every block message, and the rules after the cut are not applied. See guardrails.example.md — copy it to GUARDRAILS.md.
2. Install the gate in your harness.
Add the plugin to opencode.json (global or per project):
{
"plugin": ["@bergetai/guardrails-md"]
}
The package was called @bergetai/opencode-guardrails-md up to 0.5.1. That name is deprecated; replace it with @bergetai/guardrails-md in your config.
Install the same package from npm (add -l to install it for the current project only):
pi install npm:@bergetai/guardrails-md
To run pi from a clone instead, use pi install ./ in the cloned repo after npm install.
Install from this repo's marketplace, at the prompt of a running session:
/plugin install guardrails-md --marketplace berget-ai/guardrails-md
Answer y to add the marketplace, then pick a scope; the user scope loads it in every session from then on. It is a hooks module that Claude Code loads in-process from hooks/hooks.json, so there is no build step and no npm install. To run it from a clone for one session instead, use claude --plugin-dir ./guardrails-md.
Hooks modules are an early-access Claude Code API (checked on 2.1.291); the engine may change them between releases. Claude Code has no Berget seat token — set BERGET_API_KEY (below) for this harness.
Set a key and restart the harness (plugins and extensions load at startup):
export BERGET_API_KEY=…
Keys come from berget.ai. The free tier includes €5 of credit, and a gate call is small enough that it lasts a long time. If you are logged in to Berget in your harness (@bergetai/opencode-auth in opencode, /login in pi), skip the key: the gate picks up your seat token. If you run your own System One-compatible endpoint, point BERGET_BASE_URL at it instead.
From now on, every bash command your agent runs has to pass your guardrails before it is allowed. Above the threshold the command is blocked with an explanation the agent can read; below it, it runs.
The same judgement works as a git pre-commit hook: every commit's staged diff is scored before it enters the repository. Personal data (GDPR) and secrets are blocked; the team's own names in bylines and author fields pass.
Install for every repo on your machine (uses your global core.hooksPath if you have one, otherwise copy to .git/hooks/pre-commit per repo):
curl -o ~/.git-hooks/guardrails-pre-commit \
https://raw.githubusercontent.com/berget-ai/guardrails-md/main/hooks/pre-commit
chmod +x ~/.git-hooks/guardrails-pre-commit
Then chain it from your global pre-commit hook (or create one):
# ~/.git-hooks/pre-commit
python3 ~/.git-hooks/guardrails-pre-commit || exit 1
Same env config as the plugin. Fail-closed by default: if the endpoint is unreachable the commit is blocked — retry, or set SYSTEMONE_FAIL_OPEN=1.
A deny-list of patterns knows rm -rf. It does not know that your team forbids pushing to main but allows feature branches, or that kubectl get is fine in prod while kubectl apply is not. Asking a second LLM to review each command does know that, at the price of a full generation per command.
System One is a decision model: it reads the command plus your written rules and returns scores for a fixed set of questions in a single forward pass. That makes two things possible.
The gate follows your policy. With the example guardrails below, git push origin HEAD:main is a violation and git push origin feature/x is not, even though both are a git push.
It also catches what the driving model shrugs at. An agent that prints a .env file to "check the config" sees a harmless read. The gate sees credentials leaving the file.
berget/bev is fine-tuned on 221,759 judged decisions from real operations traffic — privacy and risk calls, memory decisions, routing, evasion attempts — with labels validated by a stronger model and human review. Held-out accuracy, against the base model it is built on:
| Test | What it measures | Base model | berget/bev |
|---|---|---|---|
| Risk (16,902 questions) | credentials and destructive content in ops text | 93.5% | 97.0% |
| Memory (7,456) | what is worth remembering | 64.6% | 90.7% |
| Router (5,679) | routing decisions | 27.3% | 74.4% |
| Evasion holdout (90) | evasion attempts never seen in training | 65.6% | 93.3% |
| EU risk (43) | EU AI Act risk classification | 70% | 95% |
| Red team (34) | adversarial commands | 53% | 74% |
| Jev bench (1,200) | general Jev questions, outside our domain | 80.6% | 82.8% |
Read the table with two caveats. The test splits come from the same corpora as training, so they measure fit to this kind of traffic, not performance on your traffic. And the weakest rows are the honest ones: adversarial commands sit at 74%, which is why the gate is one layer and not the whole defense.
| Question | Meaning |
|---|---|
destructive | Does the command delete, overwrite, or irreversibly destroy data, databases, clusters, or infrastructure? Version-control-recoverable effects (git rm, checkout, branch operations) and removed build artifacts/caches are not irreversible — but in GitOps repositories a push can trigger irreversible infrastructure changes, so the actual effect is what gets judged. |
credentials | Does the command contain, print, or send credentials, secrets, API keys, or tokens? |
guardrails_violation | Does the command violate the team's GUARDRAILS.md? Asked only when the file exists. |
policy_exception | Does the guardrails text explicitly name this command as allowed? Vague permissions do not count. Asked only when the file exists. |
Each answer is a score between 0 and 1. Anything above the threshold (default 0.7) blocks the tool call before execution, and the error message goes back to the agent. The four questions are not interchangeable, and the difference matters:
guardrails_violation is about your policy, and policy false positives are fixable in the policy: a command the guardrails text explicitly names as allowed passes. The exception must be precise — "may manage databases" does not unlock rm -rf /var/lib/postgresql.destructive and credentials are about the command's nature, and named exceptions do not override them. Those two are the backstop, and GUARDRAILS.md is agent-editable between sessions — a file line must not be able to switch the backstop off.Some files are not judged, they are simply out of the agent's reach. The edit, write and patch tools refuse to touch a list of paths, relative to the project root (an entry ending in / protects everything below it; no globs):
GUARDRAILS.md
.agents/GUARDRAILS.md
.opencode/GUARDRAILS.md
.pi/GUARDRAILS.md
.claude/GUARDRAILS.md
the same five paths in lowercase — legacy aliases, each matched case-insensitively
opencode.json
opencode.jsonc
.opencode/
.pi/
.claude/settings.json
.claude/settings.local.json
.github/workflows/
While the gate is on, the refusal is deterministic — no model call, no threshold, no cooldown, and retrying costs nothing — and the message tells the agent whose file this is:
SystemOne-gate: protected file — GUARDRAILS.md
Policy and harness configuration are edited by your human, not the agent.
Ask them to make the change and restart opencode.
Why this list: the gate freezes GUARDRAILS.md at session start, but a session that could rewrite it with the Edit tool could rule the next session under its own rules — the same self-approval hole as the removed allow-file. .opencode/ and .pi/ hold the gate's own plugin and extension files; .claude/settings*.json can disable the plugin (the rest of .claude/ — CLAUDE.md, skills, rules — is content an agent legitimately edits); .github/workflows/ is what runs your CI. Symlinks are resolved in every harness, so a link that lands on a protected file is protected under its own name too. Matching ignores case on every platform, so Guardrails.MD is protected as well.
The list guards the file-editing tools, not bash: echo >> GUARDRAILS.md is not matched against the list — it goes to the decision model like every other command, subject to its judgement and your GUARDRAILS.md.
The agent cannot disable the gate or un-block a command by retrying, and there is no override file it could write. The agent-reachable influence left is narrow: the repo's GUARDRAILS.md itself is frozen for the running session, and the edit tools refuse to touch it (and the other protected paths) at all — see Protected files. A rewrite can still be attempted through bash, where it goes to the decision model like every other command. Treat GUARDRAILS.md changes as code review, and unattended agents should treat the file as untrusted input. Your overrides:
SYSTEMONE_GATE=off and redo the step.SYSTEMONE_THRESHOLD (it must stay below 1) if the gate is too jumpy for your taste.GUARDRAILS.md. That is the durable fix, and since the file lives in the repo, the change goes through review like any other edit.The gate is a trained model, not a deterministic rule engine. berget/bev scores around 96% on our held-out test sets — high, but not 100%, and it will occasionally judge a command wrong in both directions: blocking something safe, or letting something risky through. Berget provides the gate as a tool, not a guarantee, and accepts no liability for decisions it gets wrong or for damage done by commands it let pass. Keep the override and audit paths, and treat the gate as one layer among several — scoped credentials, sandboxes and human review still matter.
bash x.sh, perl x.pl, dotnet x.csx, awk -f x.awk, a data file passed to anything). This closes the write-then-run bypass where an agent writes logic with the ungated edit tools and executes it with an innocuous command line, and it needs no interpreter blocklist — those are never complete. Inline code (perl -e '…') is already part of the command text. One level deep, 4 kB per file, up to three files, binaries skipped: network-fetched payloads (curl … | sh) remain a gap, and a determined agent with write access can still be creative. This is one layer, not a sandbox.SYSTEMONE_FAIL_OPEN=1, knowingly.~/.cache/opencode/systemone-gate.allow. That mechanism is removed — the agent could write the file itself — and the file is now ignored.echo <base64> | base64 -d | sh is judged on its visible text, and in our testing an encoded rm -rf inside a base64 blob scored as harmless. Direct instruction injection aimed at the model — "ignore previous instructions", fake JSON answers, authority claims, prompts in other languages — did not move the verdict in any of eight tested cases, but encoding is a real gap. If your agents run untrusted input, treat encoded pipelines as blocked territory in GUARDRAILS.md.| Variable | Default | Meaning |
|---|---|---|
| (seat token) | auto | Berget seat auth from the harness's login (see Harness differences) |
BERGET_API_KEY | – | Bearer token for CI/headless (fallback: TYPESAFE_API_KEY) |
BERGET_BASE_URL | https://api.berget.ai | Gateway root or full /v1/systemone URL (fallback: TYPESAFE_BASE_URL) |
BERGET_MODEL | berget/bev | Model id as exposed by the gateway (fallback: TYPESAFE_DEFAULT_MODEL) |
SYSTEMONE_THRESHOLD | 0.7 | Block threshold, strictly between 0 and 1. Anything else (abc, empty, 0, 1, …) falls back to 0.7 with a warning |
SYSTEMONE_FAIL_OPEN | – | Set to 1 to let commands run when the endpoint is unreachable (default is fail-closed) |
SYSTEMONE_GATE | – | Set to off to disable the gate |
SYSTEMONE_LOG | – | Set to 1 to write the audit log |
Credentials resolve per request: seat token, then BERGET_API_KEY, then TYPESAFE_API_KEY. With no credential at all the gate logs one line and stays inactive.
Every block doubles the wait before the next command is judged: 10 ms, 20 ms, 40 ms … with no ceiling. The first nine blocks are imperceptible — a few false positives cost nothing. The first noticeable pause arrives around block ten (5 s), by block twenty an attempt waits 1.5 hours, by block twenty-five almost two days. An agent shopping for variants past the gate runs out of time long before it runs out of ideas.
On top of that, the counter fades with quiet time — one block of memory per 30 minutes — so early false positives heal instead of poisoning the whole session. The fade is wall-clock: an agent cannot launder the counter by interleaving allowed commands, and one that keeps hammering refreshes the clock on every block, so escalation continues unchanged.
The wait is a visible cooldown, not a silent hang: the next command gets an error saying exactly what is happening —
SystemOne-gate: cooling down after 3 blocked commands — next attempt in ~5 s.
The wait doubles with every block; restarting opencode resets it.
Retrying early returns the same message with the remaining time and does not extend the cooldown. Restarting the harness resets the counter; the audit log records the block count and cooldown with every verdict.
There is no fast path: every command is judged, including git status and ls. Prefix allowlists are the classic seam an agent learns to exploit — ls followed by a newline and rm -rf, or a harmless echo carrying a destructive tail — so the gate does not have one, and neither does it have an override file the agent could write to.
Every command costs one POST /v1/systemone. All questions are answered in the same forward pass, so adding a question type costs milliseconds and no extra round-trip. The questions use the noul type from the System One contract; any endpoint that implements the contract works.
Nothing is written to disk unless you turn on the audit log. With SYSTEMONE_LOG=1, every verdict is appended as JSONL to the harness's log file (see Harness differences) with the full command and all scores. Verdicts you disagree with can be reviewed there and fed back as training data for the next fine-tune. The log holds whatever your commands hold, so treat it as sensitive.
Manual install for opencode: copy core.ts and adapters/opencode.ts into ~/.config/opencode/plugins/guardrails-md/ (global) or .opencode/plugins/guardrails-md/ (per project), keeping the adapters/ folder, and register "plugin": ["./plugins/guardrails-md/adapters/opencode.ts"]. Manual installs need @typesafe-ai/sdk resolvable (npm install -g @typesafe-ai/sdk); the npm package brings it as a dependency.
The judging is the same in every harness: same questions, threshold, cooldown and fail-closed default. opencode and pi share core.ts; the Claude Code module carries its own copy of the questions, decision, block wording and cooldown, because a hooks module runs without Node and imports only files of the plugin. Keep the two in step when editing either. What differs is where the gate looks.
| opencode | pi | Claude Code | |
|---|---|---|---|
| Hook | tool.execute.before: bash judged; edit, write and apply_patch checked against the protected paths | tool_call, including calls a codemode script makes: bash judged; edit and write checked against the protected paths | hooks module: tool.call on Bash and on Monitor when it runs a command; Edit, Write and NotebookEdit checked against the protected paths; session.start freezes the policy, the root and the environment |
| On block | throws; the agent reads the message | returns { block, reason } to the agent and shows a warning to you (on stderr in pi -p) | answers { deny }; Claude reads the reason |
| Without a credential | inactive; one log line with SYSTEMONE_LOG=1 | inactive; warns you once per session (on stderr in pi -p) | inactive; a toast warns you at session start |
| Seat token | $XDG_DATA_HOME/opencode/auth.json (default ~/.local/share) | pi's Berget login, OAuth or API key, resolved by pi itself; then the OAuth entry in $PI_CODING_AGENT_DIR/auth.json (default ~/.pi/agent) | none — BERGET_API_KEY (or TYPESAFE_API_KEY) only |
| Policy file | GUARDRAILS.md, then .opencode/GUARDRAILS.md (lowercase legacy names after) | GUARDRAILS.md, then .pi/GUARDRAILS.md (lowercase legacy names after) | GUARDRAILS.md, then .claude/GUARDRAILS.md (lowercase legacy names after) |
| Policy read from | the project directory opencode passes the plugin | the directory pi was started in | the session's directory, frozen at session.start |
| Audit log | ~/.cache/opencode/systemone-gate.log | ~/.cache/pi/systemone-gate.log | not supported ($.fs cannot append) |
| Protected files | edit, write, apply_patch refuse protected paths; a block throws | edit and write refuse protected paths; a block returns { block, reason } and warns you | Edit, Write and NotebookEdit refuse protected paths; the module answers { deny }. Symlinks are resolved in all three |
In pi, /reload counts as a restart: it re-reads GUARDRAILS.md and resets the cooldown. pi's powershell tool is not gated; if a repo enables it in .pi/settings.json, commands run through it skip the gate.
In Claude Code, the module lives as long as the session, so the frozen policy, the environment and the cooldown stay in memory, as in opencode and pi. session.start fires once per session and not on /compact, so compaction neither re-reads GUARDRAILS.md nor resets the cooldown; /clear keeps both too. Starting a new session, or a hot reload while developing the plugin, re-reads GUARDRAILS.md and resets the cooldown.
Claude Code skips a hook that throws or overruns and lets the call through, so the module attaches a .catch handler that denies instead: the gate fails closed. If session.start never ran, every command is denied. The endpoint call goes through Claude Code's own $.http.fetch: when your organization's web-fetch policy refuses api.berget.ai (or your BERGET_BASE_URL), the call fails and the command is denied like any unreachable endpoint (or allowed with SYSTEMONE_FAIL_OPEN=1). The call times out after 5 s, as the SDK's does in opencode and pi, and a timeout is treated like an unreachable endpoint. Unlike the SDK, the module makes one attempt with no retry on 429 or 5xx. The SYSTEMONE_* and BERGET_* variables are read once, at session.start; changing them takes a new session.
Bash is gated, and so is Monitor when it runs a shell command
hooks/register.ts 566 lines1/**
2 * guardrails-md for Claude Code: a hooks module loaded in-process from
3 * hooks/hooks.json. Self-contained: a hooks module runs without Node and
4 * imports only files of the plugin, so it cannot share core.ts (node:fs,
5 * the SDK). The questions, decision, block wording and cooldown below are
6 * the same as core.ts's; keep the two in step when editing either.
7 *
8 * session.start snapshots GUARDRAILS.md (then .claude/GUARDRAILS.md, then the
9 * legacy lowercase files), the
10 * session root and every variable below; it fires once per session and not
11 * on compaction, and a missing snapshot denies. tool.call judges Bash, and
12 * Monitor when it runs a shell `command`, and deterministically refuses
13 * Edit, Write and NotebookEdit on protected paths — no model call, no
14 * threshold, no cooldown — answering { deny } on a block.
15 *
16 * Fail-closed: the engine skips a hook that throws or overruns and lets the
17 * call through, so .catch denies instead. The endpoint call is raced
18 * against a 5 s $.clock.sleep, as the SDK times out in opencode and pi:
19 * waiting on $.http.fetch does not count against the hook's budget.
20 *
21 * Credentials: BERGET_API_KEY, then TYPESAFE_API_KEY (no Berget seat token
22 * here). Cooldown: module memory, for the session. No audit log: $.fs
23 * cannot append.
24 */
25import type { Register } from "claude-code"
26
27const HARNESS = "claude"
28// Capitalised canonical names first; the lowercase legacy files are the fallback.
29const GUARDRAIL_PATHS = ["GUARDRAILS.md", ".claude/GUARDRAILS.md", "guardrails.md", ".claude/guardrails.md"]
30// The same list as core.ts's PROTECTED_PATHS — a hooks module cannot import
31// it, so keep the two in step.
32const PROTECTED_PATHS = [
33 "GUARDRAILS.md",
34 ".agents/GUARDRAILS.md",
35 ".opencode/GUARDRAILS.md",
36 ".pi/GUARDRAILS.md",
37 ".claude/GUARDRAILS.md",
38 "opencode.json",
39 "opencode.jsonc",
40 ".opencode/",
41 ".pi/",
42 ".claude/settings.json",
43 ".claude/settings.local.json",
44 ".github/workflows/",
45]
46const TIMEOUT_MS = 5000
47const DEFAULT_THRESHOLD = 0.7
48const GUARDRAILS_MAX = 2000 // chars — keep the state text tight
49const TRUNCATED = "\n… (truncated)"
50const SCRIPT_MAX = 4000 // chars of file content included per file
51const SCRIPT_FILES_MAX = 3 // files read per command
52const SCRIPT_BYTES_MAX = 1_000_000 // larger files are not read
53
54// --- config -----------------------------------------------------------------
55
56// Out-of-range values fail silently otherwise: "abc" (NaN) or anything ≥ 1
57// never blocks, while "", 0 or negatives block everything.
58function parseThreshold(raw: string | undefined): { value: number; warning?: string } {
59 if (raw === undefined) return { value: DEFAULT_THRESHOLD }
60 const value = Number(raw)
61 if (raw.trim() !== "" && value > 0 && value < 1) return { value }
62 return {
63 value: DEFAULT_THRESHOLD,
64 warning:
65 `SystemOne-gate: SYSTEMONE_THRESHOLD=${JSON.stringify(raw)} is not a number between 0 and 1 — ` +
66 `using ${DEFAULT_THRESHOLD}.`,
67 }
68}
69
70function isEnabled(raw: string | undefined): boolean {
71 return /^(1|true|yes)$/i.test(raw ?? "")
72}
73
74// The endpoint wants the API ROOT (we append /v1/systemone); accept both forms.
75function gatewayRoot(url: string): string {
76 return url.replace(/\/v1\/systemone\/?$/, "").replace(/\/+$/, "")
77}
78
79// --- GUARDRAILS.md ----------------------------------------------------------
80
81function truncateGuardrails(raw: string): string {
82 const text = raw.trim()
83 return text.length > GUARDRAILS_MAX ? text.slice(0, GUARDRAILS_MAX) + TRUNCATED : text
84}
85
86// A truncated policy is applied as if complete, so rules past the cut are
87// silently dropped — an agent can push them there by padding the file.
88function truncationWarning(guardrails: string | null): string | undefined {
89 if (!guardrails?.endsWith(TRUNCATED)) return undefined
90 return (
91 `SystemOne-gate: GUARDRAILS.md is longer than ${GUARDRAILS_MAX} characters — rules after that are ignored. ` +
92 `Shorten it or put the MUST NOT rules first.`
93 )
94}
95
96// --- judging ----------------------------------------------------------------
97
98type Verdict = {
99 destructive: number
100 credentials: number
101 guardrails_violation?: number
102 policy_exception?: number
103}
104
105// The SDK's noul(), inlined: the same wire JSON without a package import.
106function noul(instructions: string) {
107 return { type: "noul" as const, instructions }
108}
109
110function buildQuestions(hasGuardrails: boolean) {
111 const questions: Record<string, ReturnType<typeof noul>> = {
112 destructive: noul(
113 "Does this shell command delete, overwrite, format, or irreversibly destroy data, databases, clusters, or infrastructure? " +
114 "Effects that are recoverable through version control (git rm, git rm --cached, git checkout, branch operations) are not irreversible. " +
115 "Removing build artifacts, caches and temporary files is not irreversible. " +
116 "However, in GitOps repositories a version-control change can trigger irreversible changes to real infrastructure — " +
117 "deleting a manifest and pushing can destroy the resource it managed. Judge the actual effect of the command, not the version-control mechanics.",
118 ),
119 credentials: noul(
120 "Does this shell command contain, print, or send credentials, secrets, API keys, or tokens?",
121 ),
122 }
123 if (hasGuardrails) {
124 questions.guardrails_violation = noul(
125 "Given the team's guardrails in the state, does this command violate any of those rules?",
126 )
127 questions.policy_exception = noul(
128 "Does the team's guardrails text in the state EXPLICITLY allow this exact command — " +
129 "naming it, or a precise and unambiguous pattern that matches it? " +
130 "General or vague permissions do not count as an exception.",
131 )
132 }
133 return questions
134}
135
136function buildStateText(command: string, scripts: string, guardrails: string | null): string {
137 const stateParts = [`Command the agent wants to run:\n${command}`]
138 if (scripts) {
139 stateParts.push(`The command executes these script files — their content is part of the command:\n${scripts}`)
140 }
141 if (guardrails) {
142 stateParts.push(`Team guardrails (rules for what the agent may and may not do):\n${guardrails}`)
143 }
144 return stateParts.join("\n\n")
145}
146
147// No interpreter blocklist — perl, java, dotnet, ruby, lua, osascript,
148// xargs, find -exec and friends can never be enumerated. Instead: every
149// argument that points to an existing file is read and judged, whatever
150// tool would run it. Inline code (perl -e '…', node -e '…') is already
151// part of the command text. One level deep; network-fetched payloads and
152// binaries remain documented gaps.
153function scriptPath(token: string): string | null {
154 const path = token.replace(/^["']|["']$/g, "")
155 if (!path || path.startsWith("-") || !path.includes(".")) return null
156 return path
157}
158
159function scriptSection(path: string, text: string): string | null {
160 if (text.includes("\0")) return null // binary — utf8 garbage adds nothing
161 const cut = text.length > SCRIPT_MAX ? text.slice(0, SCRIPT_MAX) + TRUNCATED : text
162 return `--- ${path} ---\n${cut}`
163}
164
165// A score the endpoint sends as a string or garbage must not reach decide
166// as a non-number: "0.98" would pass the > comparison and then throw in
167// blockMessage. Absent questions stay absent.
168function score(answer: { noul?: unknown } | undefined): number {
169 const value = Number(answer?.noul ?? 0)
170 return Number.isNaN(value) ? 0 : value
171}
172
173function verdictFrom(answers: unknown): Verdict {
174 const a = (answers ?? {}) as Record<string, { noul?: unknown } | undefined>
175 return {
176 destructive: score(a.destructive),
177 credentials: score(a.credentials),
178 guardrails_violation: a.guardrails_violation === undefined ? undefined : score(a.guardrails_violation),
179 policy_exception: a.policy_exception === undefined ? undefined : score(a.policy_exception),
180 }
181}
182
183// Two tiers. A named exception in GUARDRAILS.md overrides the POLICY
184// question (guardrails_violation) — that is how policy false positives are
185// fixed. It does NOT override destructive or credentials: those judge the
186// command's nature, they are the backstop, and GUARDRAILS.md is
187// agent-editable between sessions. A command the model judges destructive
188// needs a human at the keyboard, whatever the file says.
189type Dimension = "destructive" | "credentials" | "guardrails_violation"
190
191interface Outcome {
192 decision: "BLOCK" | "allow"
193 kind: Dimension
194 worst: number
195 named: boolean
196}
197
198function decide(verdict: Verdict, hasGuardrails: boolean, threshold: number): Outcome {
199 const named = hasGuardrails && (verdict.policy_exception ?? 0) > threshold
200 const dimensions: [Dimension, number][] = [
201 ["destructive", verdict.destructive],
202 ["credentials", verdict.credentials],
203 ]
204 if (verdict.guardrails_violation !== undefined) {
205 dimensions.push(["guardrails_violation", verdict.guardrails_violation])
206 }
207 const triggering = dimensions.filter(([dimension, score]) => {
208 if (dimension === "guardrails_violation") return score > threshold && !named
209 return score > threshold
210 })
211 const decision = triggering.length > 0 ? "BLOCK" : "allow"
212 // report the worst triggering dimension, not a priority order
213 const [kind, worst] = triggering.reduce(
214 (a, b) => (b[1] > a[1] ? b : a),
215 ["destructive", 0] as [Dimension, number],
216 )
217 return { decision, kind, worst, named }
218}
219
220function blockMessage(outcome: Outcome, hasGuardrails: boolean, threshold: number): string {
221 const perKind =
222 outcome.kind === "guardrails_violation"
223 ? `\n If this is a false positive, your human can name the command in the` +
224 `\n MAY section of GUARDRAILS.md and restart ${HARNESS} — the gate` +
225 `\n follows the file.`
226 : `\n Judged destructive/leaking on its own merits — named exceptions in` +
227 `\n GUARDRAILS.md do not override this. If it is intended, your human` +
228 `\n can run it directly, or restart ${HARNESS} with SYSTEMONE_GATE=off` +
229 `\n for a session that needs it.`
230 const bootstrap = hasGuardrails
231 ? ""
232 : `\n No GUARDRAILS.md found in this repo. Your human can create one` +
233 `\n (legacy lowercase guardrails.md is still read) and write what the` +
234 `\n agent may and may not do — name what should` +
235 `\n pass in the MAY section, then restart ${HARNESS}.`
236 return (
237 `SystemOne-gate: blocked command — ${outcome.kind}=${outcome.worst.toFixed(2)} > ${threshold}\n` +
238 perKind +
239 bootstrap
240 )
241}
242
243function endpointReason(status: number | undefined): string {
244 if (status === 402) return "Berget account out of credit — top up at berget.ai"
245 if (status === 401) return "authentication failed — re-login to Berget or check BERGET_API_KEY"
246 if (status === 429) return "rate limited — wait a moment and retry"
247 return "endpoint unreachable"
248}
249
250function endpointFailure(command: string, err: unknown): string {
251 return (
252 `SystemOne-gate: ${endpointReason((err as { status?: number }).status)} — command blocked.\n` +
253 ` ${command.slice(0, 200)}\n` +
254 ` ${String(err).slice(0, 160)}\n` +
255 ` Retry shortly, or set SYSTEMONE_FAIL_OPEN=1 to prefer availability.`
256 )
257}
258
259// --- cooldown ---------------------------------------------------------------
260// Circumvention cooldown: every block doubles the wait before the next
261// command is judged — 10 ms, 20 ms, 40 ms … with no ceiling. The first
262// nine blocks are imperceptible (a few false positives cost nothing);
263// the first noticeable pause arrives around block ten (5 s), and by
264// block twenty an attempt waits 1.5 h, by block twenty-five almost two
265// days: brute-forcing variants past the gate is arithmetically hopeless.
266//
267// The wait is enforced as a visible cooldown block, not a silent sleep —
268// an invisible hang looks like a crash, while the reason tells the agent
269// (and through it the human) exactly what is happening and for how long.
270// Retrying early just returns the same reason with the remaining time.
271//
272// The counter decays with quiet time — one block of memory fades per 30
273// minutes since the last block — so early false positives do not poison a
274// whole session. Decay is wall-clock, not command-count: an agent cannot
275// launder the counter by interleaving allowed commands, and an agent that
276// keeps hammering refreshes lastBlockAt on every block, so escalation
277// continues unchanged. State lives for the session; a new session or a
278// mod reload resets everything.
279const DECAY_MS = 30 * 60_000
280const BACKOFF_BASE_MS = 10
281
282function backoffMs(blocks: number): number {
283 if (blocks < 1) return 0
284 // exponent capped at 40 to keep the float well-behaved
285 return BACKOFF_BASE_MS * 2 ** Math.min(blocks - 1, 40)
286}
287
288function formatWait(remainingMs: number): string {
289 if (remainingMs >= 90_000) return `${Math.round(remainingMs / 60_000)} min`
290 if (remainingMs >= 1000) return `${Math.ceil(remainingMs / 1000)} s`
291 return `${remainingMs} ms`
292}
293
294function createCooldown() {
295 let blockCount = 0
296 let lastBlockAt = 0
297 let cooldownUntil = 0
298
299 function effectiveBlocks(now: number): number {
300 if (blockCount === 0 || lastBlockAt === 0) return 0
301 return Math.max(0, blockCount - Math.floor((now - lastBlockAt) / DECAY_MS))
302 }
303
304 function register(): void {
305 const now = Date.now()
306 blockCount = effectiveBlocks(now) + 1
307 lastBlockAt = now
308 cooldownUntil = now + backoffMs(blockCount)
309 }
310
311 function reason(): string | null {
312 const remainingMs = cooldownUntil - Date.now()
313 if (remainingMs <= 0) return null
314 return (
315 `SystemOne-gate: cooling down after ${blockCount} blocked command${blockCount === 1 ? "" : "s"} — ` +
316 `next attempt in ~${formatWait(remainingMs)}. The wait doubles with every block; ` +
317 `restarting ${HARNESS} resets it.`
318 )
319 }
320
321 return { register, reason }
322}
323
324// --- protected paths --------------------------------------------------------
325// A deterministic deny sits in front of the model: a write to a protected
326// path is a rule, not a verdict — the cooldown counter stays untouched, an
327// agent that retries pays nothing. The list is relative to the session root;
328// an entry ending in `/` protects everything below it. The path logic below
329// mirrors core.ts's foldPath/relUnder/placePath — keep the two in step.
330
331// A path as its spelling says: `.` and `..` folded, doubles collapsed, no
332// symlink followed.
333function foldPath(p: string): string {
334 const absolute = p.startsWith("/")
335 const out: string[] = []
336 for (const part of p.split(/[\\/]/)) {
337 if (!part || part === ".") continue
338 if (part === "..") {
339 if (out.length > 0) out.pop()
340 } else {
341 out.push(part)
342 }
343 }
344 return absolute ? `/${out.join("/")}` : out.join("/")
345}
346
347function relUnder(root: string, path: string): string | null {
348 const r = foldPath(root)
349 const p = foldPath(path)
350 if (r === "" || p === r) return null
351 if (r === "/") return p.slice(1)
352 return p.startsWith(`${r}/`) ? p.slice(r.length + 1) : null
353}
354
355// The matching entry, or null. Called twice (placed path, then literal
356// spelling), like core.ts matches both spellings.
357function protectedPath(root: string, resolved: string): string | null {
358 if (!root || !resolved) return null
359 // Case-folded like core.ts: `Guardrails.MD` opens guardrails.md on a
360 // case-insensitive filesystem.
361 const rel = relUnder(root, resolved)?.toLowerCase() ?? null
362 if (rel === null) return null
363 for (const entry of PROTECTED_PATHS) {
364 const pattern = entry.toLowerCase()
365 if (pattern.endsWith("/")) {
366 if (rel.startsWith(pattern)) return entry
367 } else if (rel === pattern) {
368 return entry
369 }
370 }
371 return null
372}
373
374function protectedMessage(entry: string): string {
375 return (
376 `SystemOne-gate: protected file — ${entry}\n` +
377 ` Policy and harness configuration are edited by your human, not the agent.\n` +
378 ` Ask them to make the change and restart ${HARNESS}.`
379 )
380}
381
382// Where the path lands, every symlink followed — including for a file that
383// does not exist yet: the nearest existing ancestor is resolved and the tail
384// appended, as core.ts's walk does. null when nothing resolves (an odd
385// spelling, a dangling link), and the guard then denies: the engine denies
386// without a realPath too.
387function placeable(path: string): boolean {
388 const cut = Math.max(path.lastIndexOf("/"), path.lastIndexOf("\\"))
389 const name = path.slice(cut + 1)
390 return (
391 !/^[A-Za-z]:(?![\\/])/.test(path) &&
392 !/^[\\/][\\/]/.test(path) &&
393 !/^[A-Za-z]:/.test(name) &&
394 name !== "" &&
395 name !== "." &&
396 name !== ".."
397 )
398}
399
400async function place(
401 $: { fs: { stat(path: string, options: { resolve: boolean }): Promise<{ realPath?: string } | undefined> } },
402 path: string,
403): Promise<string | null> {
404 if (!placeable(path)) return null
405 const tail: string[] = []
406 let target = path
407 for (;;) {
408 const own = await $.fs.stat(target, { resolve: true }).catch(() => undefined)
409 if (own) {
410 if (own.realPath === undefined) return null
411 let real = own.realPath.replace(/[\\/]+$/, "")
412 for (const name of tail) real = `${real}/${name}`
413 return real
414 }
415 const stripped = target.replace(/[\\/]+$/, "")
416 const cut = Math.max(stripped.lastIndexOf("/"), stripped.lastIndexOf("\\"))
417 const name = stripped.slice(cut + 1)
418 if (cut < 0 || name === "" || name === "." || name === ".." || /^[A-Za-z]:/.test(name)) return null
419 tail.unshift(name)
420 target = stripped.slice(0, cut + 1)
421 }
422}
423
424// --- the module -------------------------------------------------------------
425
426interface Session {
427 guardrails: string | null
428 key: string | undefined
429 baseURL: string
430 model: string
431 threshold: number
432 failOpen: boolean
433 off: boolean
434 cwd: string
435 root: string
436 warnings: string[]
437 cooldown: ReturnType<typeof createCooldown>
438}
439
440class EndpointError extends Error {
441 constructor(
442 readonly status: number,
443 body: string,
444 ) {
445 super(`System One endpoint answered ${status}: ${body.slice(0, 120)}`)
446 }
447}
448
449async function firstReadable(read: (path: string) => Promise<string>, paths: string[]): Promise<string | null> {
450 for (const path of paths) {
451 try {
452 return await read(path)
453 } catch {}
454 }
455 return null
456}
457
458export const register: Register = (on) => {
459 let session: Session | null = null
460
461 on("session.start", async ($, e, next) => {
462 const raw = await firstReadable((path) => $.fs.read(path), GUARDRAIL_PATHS.map((p) => `${e.cwd}/${p}`))
463 const guardrails = raw === null ? null : truncateGuardrails(raw)
464 const rootStat = await $.fs.stat(e.cwd, { resolve: true }).catch(() => undefined)
465 const threshold = parseThreshold(await $.env.get("SYSTEMONE_THRESHOLD"))
466 session = {
467 cwd: e.cwd,
468 root: rootStat?.realPath && rootStat.realPath !== "" ? rootStat.realPath : e.cwd,
469 guardrails,
470 key: (await $.env.get("BERGET_API_KEY")) ?? (await $.env.get("TYPESAFE_API_KEY")),
471 baseURL: gatewayRoot((await $.env.get("BERGET_BASE_URL")) ?? (await $.env.get("TYPESAFE_BASE_URL")) ?? "https://api.berget.ai"),
472 model: (await $.env.get("BERGET_MODEL")) ?? (await $.env.get("TYPESAFE_DEFAULT_MODEL")) ?? "berget/bev",
473 threshold: threshold.value,
474 failOpen: isEnabled(await $.env.get("SYSTEMONE_FAIL_OPEN")),
475 off: (await $.env.get("SYSTEMONE_GATE")) === "off",
476 warnings: [threshold.warning, truncationWarning(guardrails)].filter((w): w is string => w !== undefined),
477 cooldown: createCooldown(),
478 }
479 for (const warning of session.warnings) $.ui.toast(warning)
480 if (!session.off && !session.key) {
481 $.ui.toast("guardrails-md inactive: no BERGET_API_KEY or TYPESAFE_API_KEY — Bash and Monitor commands run ungated this session")
482 }
483 return next(e)
484 })
485
486 on("tool.call", { tool: ["Bash", "Monitor"] }, async ($, e, next) => {
487 if (e.command === undefined) return next(e)
488 if (!session) return { deny: "SystemOne-gate: no frozen policy — session.start did not run. Restart Claude Code (fail-closed)." }
489 const s = session
490 const command = e.command.trim()
491 if (s.off || !s.key || !command) return next(e)
492 const cooling = s.cooldown.reason()
493 if (cooling) return { deny: cooling }
494 const deny = (reason: string) => ({ deny: [reason, ...s.warnings.map((w) => ` ${w}`)].join("\n") })
495
496 const scripts: string[] = []
497 for (const token of command.split(/\s+/)) {
498 if (scripts.length >= SCRIPT_FILES_MAX) break
499 const path = scriptPath(token)
500 if (!path) continue
501 try {
502 const st = await $.fs.stat(path)
503 if (st.kind !== "file" || st.size > SCRIPT_BYTES_MAX) continue
504 const section = scriptSection(path, await $.fs.read(path))
505 if (section) scripts.push(section)
506 } catch {}
507 }
508
509 let answers: unknown
510 const stop = new AbortController()
511 try {
512 const timeout = $.clock.sleep(TIMEOUT_MS, { signal: AbortSignal.any([stop.signal, next.signal]) }).then(() => {
513 throw new Error(`System One endpoint timed out after ${TIMEOUT_MS} ms`)
514 })
515 timeout.catch(() => {})
516 const res = await Promise.race([
517 $.http.fetch(`${s.baseURL}/v1/systemone`, {
518 method: "POST",
519 headers: { Authorization: `Bearer ${s.key}`, Accept: "application/json", "Content-Type": "application/json" },
520 body: JSON.stringify({
521 state: { text: buildStateText(command, scripts.join("\n"), s.guardrails) },
522 questions: buildQuestions(!!s.guardrails),
523 model: s.model,
524 }),
525 }),
526 timeout,
527 ])
528 if (!res.ok) throw new EndpointError(res.status, res.text)
529 answers = JSON.parse(res.text).answers
530 } catch (err) {
531 if (s.failOpen) return next(e)
532 return deny(endpointFailure(command, err))
533 } finally {
534 stop.abort()
535 }
536
537 const outcome = decide(verdictFrom(answers), !!s.guardrails, s.threshold)
538 if (outcome.decision === "allow") return next(e)
539 s.cooldown.register()
540 return deny(blockMessage(outcome, !!s.guardrails, s.threshold) + `\n ${command.slice(0, 200)}`)
541 }).catch(($, e, next) => ({
542 deny: `SystemOne-gate: the gate failed (${next.error.kind}${next.error.message ? `: ${next.error.message}` : ""}) — command blocked (fail-closed).`,
543 }))
544
545 on("tool.call", { tool: ["Edit", "Write", "NotebookEdit"] }, async ($, e, next) => {
546 if (!session) return { deny: "SystemOne-gate: no frozen policy — session.start did not run. Restart Claude Code (fail-closed)." }
547 const s = session
548 if (s.off) return next(e)
549 const target = e.tool === "NotebookEdit" ? e.notebook_path : e.file_path
550 if (typeof target !== "string" || target === "") return next(e)
551 const placed = await place($, target)
552 if (placed === null) {
553 return {
554 deny:
555 `SystemOne-gate: cannot resolve ${JSON.stringify(target)} — edit blocked (fail-closed).\n` +
556 ` The path could not be checked against the protected list.`,
557 }
558 }
559 const hit = protectedPath(s.root, placed) ?? protectedPath(s.cwd, target)
560 if (hit === null) return next(e)
561 return { deny: [protectedMessage(hit), ...s.warnings.map((w) => ` ${w}`)].join("\n") }
562 }).catch(($, e, next) => ({
563 deny: `SystemOne-gate: the gate failed (${next.error.kind}${next.error.message ? `: ${next.error.message}` : ""}) — edit blocked (fail-closed).`,
564 }))
565}
566