SLOPSHOPPER

guardrails-md

GUARDRAILS.md gate: a System One decision model judges every Bash command before it runs — destructive / credentials / GUARDRAILS.md policy

newguardtoastnetwork
v0.7.0MITupdated 2026-10-09berget-ai/guardrails-md
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · guardrails-md
› fix the failing auth test and add an audit log call ╭────────────────────────────────────────────╮ │ guardrails-md │ ⏺ Read(src/auth.ts) │ guardrails-md inactive: no BERGET_API_KEY │ ⎿ Read 6 lines │ or TYPESAFE_API_KEY — Bash and Monitor │ ⏺ Update(src/auth.ts) │ commands run ungated this session │ ⎿ Added 2 lines, removed 1 line ╰────────────────────────────────────────────╯ ⏺ Write(/work/app/src/audit.ts) ⎿ Denied by guardrails-md: SystemOne-gate: cannot resolve "/work/app/src/audit.ts" — edit blocked (fail-clos ● Done. refresh now rejects expired claims and logs an audit event. ✻ Worked for 42s · done 4:20 PM ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts
README

guardrails-md

Stops your coding agent from running the bash command you'd regret.

Every bash command your coding agent (opencode, pi or Claude Code) is about to execute is scored by a small decision model first, in about 100 ms. If the command destroys data, leaks a secret, or breaks a rule in your repo's GUARDRAILS.md, the call is blocked and the agent is told why, so it can pick another route.

SystemOne-gate: blocked command — destructive=0.98 > 0.7
  rm -rf ./important-data
Judged destructive/leaking on its own merits — named exceptions in
GUARDRAILS.md do not override this. If it is intended, your human
can run it directly, or restart opencode with SYSTEMONE_GATE=off
for a session that needs it. The agent must not edit GUARDRAILS.md
to work around this.

When it helps

Two situations where a pattern list leaves you exposed.

Teams that work close to production. The agent runs kubectl, helm and git push all day, and the difference between routine and incident is your team's policy, not a filename pattern. The gate reads that policy from GUARDRAILS.md and applies it to every command.

Background agents that read mail, tickets or forms. Their input is written by strangers, and a crafted message can talk an agent into running something harmful. No one is watching when it happens. The gate does not promise that the agent cannot be persuaded — it promises that the boundary holds anyway, because the gate judges the command and your rules, not the conversation that led to it.

Not the guardrails.md convention

This project is unrelated to the guardrails.md convention, where the agent appends its own lessons to GUARDRAILS.md. Here the file is written by humans, so if your repo already follows that convention, the edit tools will refuse the agent's writes to the file once the gate is installed (bash writes are still judged, not hard-blocked; see Protected files).

Quickstart

Two steps. First teach the gate your rules, then put it in the harness.

1. Add GUARDRAILS.md to your repo (root, .opencode/GUARDRAILS.md, .pi/GUARDRAILS.md or .claude/GUARDRAILS.md; the legacy lowercase guardrails.md spellings are still read). Write it yourself — the value is in deciding what your team actually allows, not in shipping a generic file. The example below is a starting point for the shape:

# Guardrails for agents in this repo

## The agent MUST NOT

- Edit this file (GUARDRAILS.md) itself — it is written and changed by humans, through review.
- Change anything in production — production changes reach production only through Git/CD.
- Push directly to the main branch — all changes go through pull request.
- Install software outside the project's declared dependencies.
- Send data to external services outside our approved list (docs/approved-domains.md).
- Run irreversible operations against shared systems — deletions, cleanup, resets.

## The agent MAY

- Inspect any environment read-only.
- Run tests, lint, and builds locally.
- Create branches and push feature branches.
- Read documentation from the approved sources in docs/approved-domains.md.

The gate reads the file once at session start, so restart the harness after editing. Only the first 2000 characters are sent to the model, so keep the file short and put the MUST NOT rules first. If the file is longer, the gate warns you at startup and in every block message, and the rules after the cut are not applied. See guardrails.example.md — copy it to GUARDRAILS.md.

2. Install the gate in your harness.

opencode

Add the plugin to opencode.json (global or per project):

{
  "plugin": ["@bergetai/guardrails-md"]
}

The package was called @bergetai/opencode-guardrails-md up to 0.5.1. That name is deprecated; replace it with @bergetai/guardrails-md in your config.

pi

Install the same package from npm (add -l to install it for the current project only):

pi install npm:@bergetai/guardrails-md

To run pi from a clone instead, use pi install ./ in the cloned repo after npm install.

Claude Code

Install from this repo's marketplace, at the prompt of a running session:

/plugin install guardrails-md --marketplace berget-ai/guardrails-md

Answer y to add the marketplace, then pick a scope; the user scope loads it in every session from then on. It is a hooks module that Claude Code loads in-process from hooks/hooks.json, so there is no build step and no npm install. To run it from a clone for one session instead, use claude --plugin-dir ./guardrails-md.

Hooks modules are an early-access Claude Code API (checked on 2.1.291); the engine may change them between releases. Claude Code has no Berget seat token — set BERGET_API_KEY (below) for this harness.

API key

Set a key and restart the harness (plugins and extensions load at startup):

export BERGET_API_KEY=…

Keys come from berget.ai. The free tier includes €5 of credit, and a gate call is small enough that it lasts a long time. If you are logged in to Berget in your harness (@bergetai/opencode-auth in opencode, /login in pi), skip the key: the gate picks up your seat token. If you run your own System One-compatible endpoint, point BERGET_BASE_URL at it instead.

From now on, every bash command your agent runs has to pass your guardrails before it is allowed. Above the threshold the command is blocked with an explanation the agent can read; below it, it runs.

Git pre-commit hook

The same judgement works as a git pre-commit hook: every commit's staged diff is scored before it enters the repository. Personal data (GDPR) and secrets are blocked; the team's own names in bylines and author fields pass.

Install for every repo on your machine (uses your global core.hooksPath if you have one, otherwise copy to .git/hooks/pre-commit per repo):

curl -o ~/.git-hooks/guardrails-pre-commit \
  https://raw.githubusercontent.com/berget-ai/guardrails-md/main/hooks/pre-commit
chmod +x ~/.git-hooks/guardrails-pre-commit

Then chain it from your global pre-commit hook (or create one):

# ~/.git-hooks/pre-commit
python3 ~/.git-hooks/guardrails-pre-commit || exit 1

Same env config as the plugin. Fail-closed by default: if the endpoint is unreachable the commit is blocked — retry, or set SYSTEMONE_FAIL_OPEN=1.

Why a model and not a regex

A deny-list of patterns knows rm -rf. It does not know that your team forbids pushing to main but allows feature branches, or that kubectl get is fine in prod while kubectl apply is not. Asking a second LLM to review each command does know that, at the price of a full generation per command.

System One is a decision model: it reads the command plus your written rules and returns scores for a fixed set of questions in a single forward pass. That makes two things possible.

The gate follows your policy. With the example guardrails below, git push origin HEAD:main is a violation and git push origin feature/x is not, even though both are a git push.

It also catches what the driving model shrugs at. An agent that prints a .env file to "check the config" sees a harmless read. The gate sees credentials leaving the file.

How well it judges

berget/bev is fine-tuned on 221,759 judged decisions from real operations traffic — privacy and risk calls, memory decisions, routing, evasion attempts — with labels validated by a stronger model and human review. Held-out accuracy, against the base model it is built on:

TestWhat it measuresBase modelberget/bev
Risk (16,902 questions)credentials and destructive content in ops text93.5%97.0%
Memory (7,456)what is worth remembering64.6%90.7%
Router (5,679)routing decisions27.3%74.4%
Evasion holdout (90)evasion attempts never seen in training65.6%93.3%
EU risk (43)EU AI Act risk classification70%95%
Red team (34)adversarial commands53%74%
Jev bench (1,200)general Jev questions, outside our domain80.6%82.8%

Read the table with two caveats. The test splits come from the same corpora as training, so they measure fit to this kind of traffic, not performance on your traffic. And the weakest rows are the honest ones: adversarial commands sit at 74%, which is why the gate is one layer and not the whole defense.

What it asks

QuestionMeaning
destructiveDoes the command delete, overwrite, or irreversibly destroy data, databases, clusters, or infrastructure? Version-control-recoverable effects (git rm, checkout, branch operations) and removed build artifacts/caches are not irreversible — but in GitOps repositories a push can trigger irreversible infrastructure changes, so the actual effect is what gets judged.
credentialsDoes the command contain, print, or send credentials, secrets, API keys, or tokens?
guardrails_violationDoes the command violate the team's GUARDRAILS.md? Asked only when the file exists.
policy_exceptionDoes the guardrails text explicitly name this command as allowed? Vague permissions do not count. Asked only when the file exists.

Each answer is a score between 0 and 1. Anything above the threshold (default 0.7) blocks the tool call before execution, and the error message goes back to the agent. The four questions are not interchangeable, and the difference matters:

  • guardrails_violation is about your policy, and policy false positives are fixable in the policy: a command the guardrails text explicitly names as allowed passes. The exception must be precise — "may manage databases" does not unlock rm -rf /var/lib/postgresql.
  • destructive and credentials are about the command's nature, and named exceptions do not override them. Those two are the backstop, and GUARDRAILS.md is agent-editable between sessions — a file line must not be able to switch the backstop off.

Protected files

Some files are not judged, they are simply out of the agent's reach. The edit, write and patch tools refuse to touch a list of paths, relative to the project root (an entry ending in / protects everything below it; no globs):

GUARDRAILS.md
.agents/GUARDRAILS.md
.opencode/GUARDRAILS.md
.pi/GUARDRAILS.md
.claude/GUARDRAILS.md
the same five paths in lowercase — legacy aliases, each matched case-insensitively
opencode.json
opencode.jsonc
.opencode/
.pi/
.claude/settings.json
.claude/settings.local.json
.github/workflows/

While the gate is on, the refusal is deterministic — no model call, no threshold, no cooldown, and retrying costs nothing — and the message tells the agent whose file this is:

SystemOne-gate: protected file — GUARDRAILS.md
  Policy and harness configuration are edited by your human, not the agent.
  Ask them to make the change and restart opencode.

Why this list: the gate freezes GUARDRAILS.md at session start, but a session that could rewrite it with the Edit tool could rule the next session under its own rules — the same self-approval hole as the removed allow-file. .opencode/ and .pi/ hold the gate's own plugin and extension files; .claude/settings*.json can disable the plugin (the rest of .claude/ — CLAUDE.md, skills, rules — is content an agent legitimately edits); .github/workflows/ is what runs your CI. Symlinks are resolved in every harness, so a link that lands on a protected file is protected under its own name too. Matching ignores case on every platform, so Guardrails.MD is protected as well.

The list guards the file-editing tools, not bash: echo >> GUARDRAILS.md is not matched against the list — it goes to the decision model like every other command, subject to its judgement and your GUARDRAILS.md.

Overriding a block

The agent cannot disable the gate or un-block a command by retrying, and there is no override file it could write. The agent-reachable influence left is narrow: the repo's GUARDRAILS.md itself is frozen for the running session, and the edit tools refuse to touch it (and the other protected paths) at all — see Protected files. A rewrite can still be attempted through bash, where it goes to the decision model like every other command. Treat GUARDRAILS.md changes as code review, and unattended agents should treat the file as untrusted input. Your overrides:

  • Once: restart the harness with SYSTEMONE_GATE=off and redo the step.
  • Tune: raise SYSTEMONE_THRESHOLD (it must stay below 1) if the gate is too jumpy for your taste.
  • Fix the policy: if the block is a false positive against your rules, change GUARDRAILS.md. That is the durable fix, and since the file lives in the repo, the change goes through review like any other edit.

Disclaimer

The gate is a trained model, not a deterministic rule engine. berget/bev scores around 96% on our held-out test sets — high, but not 100%, and it will occasionally judge a command wrong in both directions: blocking something safe, or letting something risky through. Berget provides the gate as a tool, not a guarantee, and accepts no liability for decisions it gets wrong or for damage done by commands it let pass. Keep the override and audit paths, and treat the gate as one layer among several — scoped credentials, sandboxes and human review still matter.

What it doesn't do

  • It gates bash commands only. Other tool calls pass through, except that the file-editing tools refuse to touch protected paths. Every argument that points to an existing file is read and judged together with the command — whatever tool would run it (bash x.sh, perl x.pl, dotnet x.csx, awk -f x.awk, a data file passed to anything). This closes the write-then-run bypass where an agent writes logic with the ungated edit tools and executes it with an innocuous command line, and it needs no interpreter blocklist — those are never complete. Inline code (perl -e '…') is already part of the command text. One level deep, 4 kB per file, up to three files, binaries skipped: network-fetched payloads (curl … | sh) remain a gap, and a determined agent with write access can still be creative. This is one layer, not a sandbox.
  • It fails closed. If the endpoint is unreachable the command is blocked until the endpoint responds again. That is deliberate: with fail-open, an agent could induce the failure itself — hammer the endpoint until it is rate-limited, or drain the account quota — and then run anything unchecked. Transient hiccups are absorbed by the SDK's retries with backoff; a persistent failure blocks, and the agent relays the error. Interactive users who prefer availability can set SYSTEMONE_FAIL_OPEN=1, knowingly.
  • It sends every command to the endpoint to be judged. Berget has a zero data retention policy and operates under EU data protection law, so a command that contains something sensitive is scored and not stored. If you would still rather keep it in-house, run your own endpoint.
  • It is one layer. The driving model's own refusals are another, and neither replaces scoped credentials or a sandbox.
  • Earlier versions read an override file at ~/.cache/opencode/systemone-gate.allow. That mechanism is removed — the agent could write the file itself — and the file is now ignored.
  • It does not decode obfuscated payloads. A command like echo <base64> | base64 -d | sh is judged on its visible text, and in our testing an encoded rm -rf inside a base64 blob scored as harmless. Direct instruction injection aimed at the model — "ignore previous instructions", fake JSON answers, authority claims, prompts in other languages — did not move the verdict in any of eight tested cases, but encoding is a real gap. If your agents run untrusted input, treat encoded pipelines as blocked territory in GUARDRAILS.md.

Configuration

VariableDefaultMeaning
(seat token)autoBerget seat auth from the harness's login (see Harness differences)
BERGET_API_KEY–Bearer token for CI/headless (fallback: TYPESAFE_API_KEY)
BERGET_BASE_URLhttps://api.berget.aiGateway root or full /v1/systemone URL (fallback: TYPESAFE_BASE_URL)
BERGET_MODELberget/bevModel id as exposed by the gateway (fallback: TYPESAFE_DEFAULT_MODEL)
SYSTEMONE_THRESHOLD0.7Block threshold, strictly between 0 and 1. Anything else (abc, empty, 0, 1, …) falls back to 0.7 with a warning
SYSTEMONE_FAIL_OPEN–Set to 1 to let commands run when the endpoint is unreachable (default is fail-closed)
SYSTEMONE_GATE–Set to off to disable the gate
SYSTEMONE_LOG–Set to 1 to write the audit log

Credentials resolve per request: seat token, then BERGET_API_KEY, then TYPESAFE_API_KEY. With no credential at all the gate logs one line and stays inactive.

Circumvention slows itself down

Every block doubles the wait before the next command is judged: 10 ms, 20 ms, 40 ms … with no ceiling. The first nine blocks are imperceptible — a few false positives cost nothing. The first noticeable pause arrives around block ten (5 s), by block twenty an attempt waits 1.5 hours, by block twenty-five almost two days. An agent shopping for variants past the gate runs out of time long before it runs out of ideas.

On top of that, the counter fades with quiet time — one block of memory per 30 minutes — so early false positives heal instead of poisoning the whole session. The fade is wall-clock: an agent cannot launder the counter by interleaving allowed commands, and one that keeps hammering refreshes the clock on every block, so escalation continues unchanged.

The wait is a visible cooldown, not a silent hang: the next command gets an error saying exactly what is happening —

SystemOne-gate: cooling down after 3 blocked commands — next attempt in ~5 s.
The wait doubles with every block; restarting opencode resets it.

Retrying early returns the same message with the remaining time and does not extend the cooldown. Restarting the harness resets the counter; the audit log records the block count and cooldown with every verdict.

Details

There is no fast path: every command is judged, including git status and ls. Prefix allowlists are the classic seam an agent learns to exploit — ls followed by a newline and rm -rf, or a harmless echo carrying a destructive tail — so the gate does not have one, and neither does it have an override file the agent could write to.

Every command costs one POST /v1/systemone. All questions are answered in the same forward pass, so adding a question type costs milliseconds and no extra round-trip. The questions use the noul type from the System One contract; any endpoint that implements the contract works.

Nothing is written to disk unless you turn on the audit log. With SYSTEMONE_LOG=1, every verdict is appended as JSONL to the harness's log file (see Harness differences) with the full command and all scores. Verdicts you disagree with can be reviewed there and fed back as training data for the next fine-tune. The log holds whatever your commands hold, so treat it as sensitive.

Manual install for opencode: copy core.ts and adapters/opencode.ts into ~/.config/opencode/plugins/guardrails-md/ (global) or .opencode/plugins/guardrails-md/ (per project), keeping the adapters/ folder, and register "plugin": ["./plugins/guardrails-md/adapters/opencode.ts"]. Manual installs need @typesafe-ai/sdk resolvable (npm install -g @typesafe-ai/sdk); the npm package brings it as a dependency.

Harness differences

The judging is the same in every harness: same questions, threshold, cooldown and fail-closed default. opencode and pi share core.ts; the Claude Code module carries its own copy of the questions, decision, block wording and cooldown, because a hooks module runs without Node and imports only files of the plugin. Keep the two in step when editing either. What differs is where the gate looks.

opencodepiClaude Code
Hooktool.execute.before: bash judged; edit, write and apply_patch checked against the protected pathstool_call, including calls a codemode script makes: bash judged; edit and write checked against the protected pathshooks module: tool.call on Bash and on Monitor when it runs a command; Edit, Write and NotebookEdit checked against the protected paths; session.start freezes the policy, the root and the environment
On blockthrows; the agent reads the messagereturns { block, reason } to the agent and shows a warning to you (on stderr in pi -p)answers { deny }; Claude reads the reason
Without a credentialinactive; one log line with SYSTEMONE_LOG=1inactive; warns you once per session (on stderr in pi -p)inactive; a toast warns you at session start
Seat token$XDG_DATA_HOME/opencode/auth.json (default ~/.local/share)pi's Berget login, OAuth or API key, resolved by pi itself; then the OAuth entry in $PI_CODING_AGENT_DIR/auth.json (default ~/.pi/agent)none — BERGET_API_KEY (or TYPESAFE_API_KEY) only
Policy fileGUARDRAILS.md, then .opencode/GUARDRAILS.md (lowercase legacy names after)GUARDRAILS.md, then .pi/GUARDRAILS.md (lowercase legacy names after)GUARDRAILS.md, then .claude/GUARDRAILS.md (lowercase legacy names after)
Policy read fromthe project directory opencode passes the pluginthe directory pi was started inthe session's directory, frozen at session.start
Audit log~/.cache/opencode/systemone-gate.log~/.cache/pi/systemone-gate.lognot supported ($.fs cannot append)
Protected filesedit, write, apply_patch refuse protected paths; a block throwsedit and write refuse protected paths; a block returns { block, reason } and warns youEdit, Write and NotebookEdit refuse protected paths; the module answers { deny }. Symlinks are resolved in all three

In pi, /reload counts as a restart: it re-reads GUARDRAILS.md and resets the cooldown. pi's powershell tool is not gated; if a repo enables it in .pi/settings.json, commands run through it skip the gate.

In Claude Code, the module lives as long as the session, so the frozen policy, the environment and the cooldown stay in memory, as in opencode and pi. session.start fires once per session and not on /compact, so compaction neither re-reads GUARDRAILS.md nor resets the cooldown; /clear keeps both too. Starting a new session, or a hot reload while developing the plugin, re-reads GUARDRAILS.md and resets the cooldown.

Claude Code skips a hook that throws or overruns and lets the call through, so the module attaches a .catch handler that denies instead: the gate fails closed. If session.start never ran, every command is denied. The endpoint call goes through Claude Code's own $.http.fetch: when your organization's web-fetch policy refuses api.berget.ai (or your BERGET_BASE_URL), the call fails and the command is denied like any unreachable endpoint (or allowed with SYSTEMONE_FAIL_OPEN=1). The call times out after 5 s, as the SDK's does in opencode and pi, and a timeout is treated like an unreachable endpoint. Unlike the SDK, the module makes one attempt with no retry on 429 or 5xx. The SYSTEMONE_* and BERGET_* variables are read once, at session.start; changing them takes a new session.

Bash is gated, and so is Monitor when it runs a shell command

Source 1 files
hooks/register.ts 566 lines
1/**
2 * guardrails-md for Claude Code: a hooks module loaded in-process from
3 * hooks/hooks.json. Self-contained: a hooks module runs without Node and
4 * imports only files of the plugin, so it cannot share core.ts (node:fs,
5 * the SDK). The questions, decision, block wording and cooldown below are
6 * the same as core.ts's; keep the two in step when editing either.
7 *
8 * session.start snapshots GUARDRAILS.md (then .claude/GUARDRAILS.md, then the
9 * legacy lowercase files), the
10 * session root and every variable below; it fires once per session and not
11 * on compaction, and a missing snapshot denies. tool.call judges Bash, and
12 * Monitor when it runs a shell `command`, and deterministically refuses
13 * Edit, Write and NotebookEdit on protected paths — no model call, no
14 * threshold, no cooldown — answering { deny } on a block.
15 *
16 * Fail-closed: the engine skips a hook that throws or overruns and lets the
17 * call through, so .catch denies instead. The endpoint call is raced
18 * against a 5 s $.clock.sleep, as the SDK times out in opencode and pi:
19 * waiting on $.http.fetch does not count against the hook's budget.
20 *
21 * Credentials: BERGET_API_KEY, then TYPESAFE_API_KEY (no Berget seat token
22 * here). Cooldown: module memory, for the session. No audit log: $.fs
23 * cannot append.
24 */
25import type { Register } from "claude-code"
26
27const HARNESS = "claude"
28// Capitalised canonical names first; the lowercase legacy files are the fallback.
29const GUARDRAIL_PATHS = ["GUARDRAILS.md", ".claude/GUARDRAILS.md", "guardrails.md", ".claude/guardrails.md"]
30// The same list as core.ts's PROTECTED_PATHS — a hooks module cannot import
31// it, so keep the two in step.
32const PROTECTED_PATHS = [
33  "GUARDRAILS.md",
34  ".agents/GUARDRAILS.md",
35  ".opencode/GUARDRAILS.md",
36  ".pi/GUARDRAILS.md",
37  ".claude/GUARDRAILS.md",
38  "opencode.json",
39  "opencode.jsonc",
40  ".opencode/",
41  ".pi/",
42  ".claude/settings.json",
43  ".claude/settings.local.json",
44  ".github/workflows/",
45]
46const TIMEOUT_MS = 5000
47const DEFAULT_THRESHOLD = 0.7
48const GUARDRAILS_MAX = 2000 // chars — keep the state text tight
49const TRUNCATED = "\n… (truncated)"
50const SCRIPT_MAX = 4000 // chars of file content included per file
51const SCRIPT_FILES_MAX = 3 // files read per command
52const SCRIPT_BYTES_MAX = 1_000_000 // larger files are not read
53
54// --- config -----------------------------------------------------------------
55
56// Out-of-range values fail silently otherwise: "abc" (NaN) or anything ≥ 1
57// never blocks, while "", 0 or negatives block everything.
58function parseThreshold(raw: string | undefined): { value: number; warning?: string } {
59  if (raw === undefined) return { value: DEFAULT_THRESHOLD }
60  const value = Number(raw)
61  if (raw.trim() !== "" && value > 0 && value < 1) return { value }
62  return {
63    value: DEFAULT_THRESHOLD,
64    warning:
65      `SystemOne-gate: SYSTEMONE_THRESHOLD=${JSON.stringify(raw)} is not a number between 0 and 1 — ` +
66      `using ${DEFAULT_THRESHOLD}.`,
67  }
68}
69
70function isEnabled(raw: string | undefined): boolean {
71  return /^(1|true|yes)$/i.test(raw ?? "")
72}
73
74// The endpoint wants the API ROOT (we append /v1/systemone); accept both forms.
75function gatewayRoot(url: string): string {
76  return url.replace(/\/v1\/systemone\/?$/, "").replace(/\/+$/, "")
77}
78
79// --- GUARDRAILS.md ----------------------------------------------------------
80
81function truncateGuardrails(raw: string): string {
82  const text = raw.trim()
83  return text.length > GUARDRAILS_MAX ? text.slice(0, GUARDRAILS_MAX) + TRUNCATED : text
84}
85
86// A truncated policy is applied as if complete, so rules past the cut are
87// silently dropped — an agent can push them there by padding the file.
88function truncationWarning(guardrails: string | null): string | undefined {
89  if (!guardrails?.endsWith(TRUNCATED)) return undefined
90  return (
91    `SystemOne-gate: GUARDRAILS.md is longer than ${GUARDRAILS_MAX} characters — rules after that are ignored. ` +
92    `Shorten it or put the MUST NOT rules first.`
93  )
94}
95
96// --- judging ----------------------------------------------------------------
97
98type Verdict = {
99  destructive: number
100  credentials: number
101  guardrails_violation?: number
102  policy_exception?: number
103}
104
105// The SDK's noul(), inlined: the same wire JSON without a package import.
106function noul(instructions: string) {
107  return { type: "noul" as const, instructions }
108}
109
110function buildQuestions(hasGuardrails: boolean) {
111  const questions: Record<string, ReturnType<typeof noul>> = {
112    destructive: noul(
113      "Does this shell command delete, overwrite, format, or irreversibly destroy data, databases, clusters, or infrastructure? " +
114        "Effects that are recoverable through version control (git rm, git rm --cached, git checkout, branch operations) are not irreversible. " +
115        "Removing build artifacts, caches and temporary files is not irreversible. " +
116        "However, in GitOps repositories a version-control change can trigger irreversible changes to real infrastructure — " +
117        "deleting a manifest and pushing can destroy the resource it managed. Judge the actual effect of the command, not the version-control mechanics.",
118    ),
119    credentials: noul(
120      "Does this shell command contain, print, or send credentials, secrets, API keys, or tokens?",
121    ),
122  }
123  if (hasGuardrails) {
124    questions.guardrails_violation = noul(
125      "Given the team's guardrails in the state, does this command violate any of those rules?",
126    )
127    questions.policy_exception = noul(
128      "Does the team's guardrails text in the state EXPLICITLY allow this exact command — " +
129        "naming it, or a precise and unambiguous pattern that matches it? " +
130        "General or vague permissions do not count as an exception.",
131    )
132  }
133  return questions
134}
135
136function buildStateText(command: string, scripts: string, guardrails: string | null): string {
137  const stateParts = [`Command the agent wants to run:\n${command}`]
138  if (scripts) {
139    stateParts.push(`The command executes these script files — their content is part of the command:\n${scripts}`)
140  }
141  if (guardrails) {
142    stateParts.push(`Team guardrails (rules for what the agent may and may not do):\n${guardrails}`)
143  }
144  return stateParts.join("\n\n")
145}
146
147// No interpreter blocklist — perl, java, dotnet, ruby, lua, osascript,
148// xargs, find -exec and friends can never be enumerated. Instead: every
149// argument that points to an existing file is read and judged, whatever
150// tool would run it. Inline code (perl -e '…', node -e '…') is already
151// part of the command text. One level deep; network-fetched payloads and
152// binaries remain documented gaps.
153function scriptPath(token: string): string | null {
154  const path = token.replace(/^["']|["']$/g, "")
155  if (!path || path.startsWith("-") || !path.includes(".")) return null
156  return path
157}
158
159function scriptSection(path: string, text: string): string | null {
160  if (text.includes("\0")) return null // binary — utf8 garbage adds nothing
161  const cut = text.length > SCRIPT_MAX ? text.slice(0, SCRIPT_MAX) + TRUNCATED : text
162  return `--- ${path} ---\n${cut}`
163}
164
165// A score the endpoint sends as a string or garbage must not reach decide
166// as a non-number: "0.98" would pass the > comparison and then throw in
167// blockMessage. Absent questions stay absent.
168function score(answer: { noul?: unknown } | undefined): number {
169  const value = Number(answer?.noul ?? 0)
170  return Number.isNaN(value) ? 0 : value
171}
172
173function verdictFrom(answers: unknown): Verdict {
174  const a = (answers ?? {}) as Record<string, { noul?: unknown } | undefined>
175  return {
176    destructive: score(a.destructive),
177    credentials: score(a.credentials),
178    guardrails_violation: a.guardrails_violation === undefined ? undefined : score(a.guardrails_violation),
179    policy_exception: a.policy_exception === undefined ? undefined : score(a.policy_exception),
180  }
181}
182
183// Two tiers. A named exception in GUARDRAILS.md overrides the POLICY
184// question (guardrails_violation) — that is how policy false positives are
185// fixed. It does NOT override destructive or credentials: those judge the
186// command's nature, they are the backstop, and GUARDRAILS.md is
187// agent-editable between sessions. A command the model judges destructive
188// needs a human at the keyboard, whatever the file says.
189type Dimension = "destructive" | "credentials" | "guardrails_violation"
190
191interface Outcome {
192  decision: "BLOCK" | "allow"
193  kind: Dimension
194  worst: number
195  named: boolean
196}
197
198function decide(verdict: Verdict, hasGuardrails: boolean, threshold: number): Outcome {
199  const named = hasGuardrails && (verdict.policy_exception ?? 0) > threshold
200  const dimensions: [Dimension, number][] = [
201    ["destructive", verdict.destructive],
202    ["credentials", verdict.credentials],
203  ]
204  if (verdict.guardrails_violation !== undefined) {
205    dimensions.push(["guardrails_violation", verdict.guardrails_violation])
206  }
207  const triggering = dimensions.filter(([dimension, score]) => {
208    if (dimension === "guardrails_violation") return score > threshold && !named
209    return score > threshold
210  })
211  const decision = triggering.length > 0 ? "BLOCK" : "allow"
212  // report the worst triggering dimension, not a priority order
213  const [kind, worst] = triggering.reduce(
214    (a, b) => (b[1] > a[1] ? b : a),
215    ["destructive", 0] as [Dimension, number],
216  )
217  return { decision, kind, worst, named }
218}
219
220function blockMessage(outcome: Outcome, hasGuardrails: boolean, threshold: number): string {
221  const perKind =
222    outcome.kind === "guardrails_violation"
223      ? `\n  If this is a false positive, your human can name the command in the` +
224        `\n  MAY section of GUARDRAILS.md and restart ${HARNESS} — the gate` +
225        `\n  follows the file.`
226      : `\n  Judged destructive/leaking on its own merits — named exceptions in` +
227        `\n  GUARDRAILS.md do not override this. If it is intended, your human` +
228        `\n  can run it directly, or restart ${HARNESS} with SYSTEMONE_GATE=off` +
229        `\n  for a session that needs it.`
230  const bootstrap = hasGuardrails
231    ? ""
232    : `\n  No GUARDRAILS.md found in this repo. Your human can create one` +
233      `\n  (legacy lowercase guardrails.md is still read) and write what the` +
234      `\n  agent may and may not do — name what should` +
235      `\n  pass in the MAY section, then restart ${HARNESS}.`
236  return (
237    `SystemOne-gate: blocked command — ${outcome.kind}=${outcome.worst.toFixed(2)} > ${threshold}\n` +
238    perKind +
239    bootstrap
240  )
241}
242
243function endpointReason(status: number | undefined): string {
244  if (status === 402) return "Berget account out of credit — top up at berget.ai"
245  if (status === 401) return "authentication failed — re-login to Berget or check BERGET_API_KEY"
246  if (status === 429) return "rate limited — wait a moment and retry"
247  return "endpoint unreachable"
248}
249
250function endpointFailure(command: string, err: unknown): string {
251  return (
252    `SystemOne-gate: ${endpointReason((err as { status?: number }).status)} — command blocked.\n` +
253    `  ${command.slice(0, 200)}\n` +
254    `  ${String(err).slice(0, 160)}\n` +
255    `  Retry shortly, or set SYSTEMONE_FAIL_OPEN=1 to prefer availability.`
256  )
257}
258
259// --- cooldown ---------------------------------------------------------------
260// Circumvention cooldown: every block doubles the wait before the next
261// command is judged — 10 ms, 20 ms, 40 ms … with no ceiling. The first
262// nine blocks are imperceptible (a few false positives cost nothing);
263// the first noticeable pause arrives around block ten (5 s), and by
264// block twenty an attempt waits 1.5 h, by block twenty-five almost two
265// days: brute-forcing variants past the gate is arithmetically hopeless.
266//
267// The wait is enforced as a visible cooldown block, not a silent sleep —
268// an invisible hang looks like a crash, while the reason tells the agent
269// (and through it the human) exactly what is happening and for how long.
270// Retrying early just returns the same reason with the remaining time.
271//
272// The counter decays with quiet time — one block of memory fades per 30
273// minutes since the last block — so early false positives do not poison a
274// whole session. Decay is wall-clock, not command-count: an agent cannot
275// launder the counter by interleaving allowed commands, and an agent that
276// keeps hammering refreshes lastBlockAt on every block, so escalation
277// continues unchanged. State lives for the session; a new session or a
278// mod reload resets everything.
279const DECAY_MS = 30 * 60_000
280const BACKOFF_BASE_MS = 10
281
282function backoffMs(blocks: number): number {
283  if (blocks < 1) return 0
284  // exponent capped at 40 to keep the float well-behaved
285  return BACKOFF_BASE_MS * 2 ** Math.min(blocks - 1, 40)
286}
287
288function formatWait(remainingMs: number): string {
289  if (remainingMs >= 90_000) return `${Math.round(remainingMs / 60_000)} min`
290  if (remainingMs >= 1000) return `${Math.ceil(remainingMs / 1000)} s`
291  return `${remainingMs} ms`
292}
293
294function createCooldown() {
295  let blockCount = 0
296  let lastBlockAt = 0
297  let cooldownUntil = 0
298
299  function effectiveBlocks(now: number): number {
300    if (blockCount === 0 || lastBlockAt === 0) return 0
301    return Math.max(0, blockCount - Math.floor((now - lastBlockAt) / DECAY_MS))
302  }
303
304  function register(): void {
305    const now = Date.now()
306    blockCount = effectiveBlocks(now) + 1
307    lastBlockAt = now
308    cooldownUntil = now + backoffMs(blockCount)
309  }
310
311  function reason(): string | null {
312    const remainingMs = cooldownUntil - Date.now()
313    if (remainingMs <= 0) return null
314    return (
315      `SystemOne-gate: cooling down after ${blockCount} blocked command${blockCount === 1 ? "" : "s"} — ` +
316      `next attempt in ~${formatWait(remainingMs)}. The wait doubles with every block; ` +
317      `restarting ${HARNESS} resets it.`
318    )
319  }
320
321  return { register, reason }
322}
323
324// --- protected paths --------------------------------------------------------
325// A deterministic deny sits in front of the model: a write to a protected
326// path is a rule, not a verdict — the cooldown counter stays untouched, an
327// agent that retries pays nothing. The list is relative to the session root;
328// an entry ending in `/` protects everything below it. The path logic below
329// mirrors core.ts's foldPath/relUnder/placePath — keep the two in step.
330
331// A path as its spelling says: `.` and `..` folded, doubles collapsed, no
332// symlink followed.
333function foldPath(p: string): string {
334  const absolute = p.startsWith("/")
335  const out: string[] = []
336  for (const part of p.split(/[\\/]/)) {
337    if (!part || part === ".") continue
338    if (part === "..") {
339      if (out.length > 0) out.pop()
340    } else {
341      out.push(part)
342    }
343  }
344  return absolute ? `/${out.join("/")}` : out.join("/")
345}
346
347function relUnder(root: string, path: string): string | null {
348  const r = foldPath(root)
349  const p = foldPath(path)
350  if (r === "" || p === r) return null
351  if (r === "/") return p.slice(1)
352  return p.startsWith(`${r}/`) ? p.slice(r.length + 1) : null
353}
354
355// The matching entry, or null. Called twice (placed path, then literal
356// spelling), like core.ts matches both spellings.
357function protectedPath(root: string, resolved: string): string | null {
358  if (!root || !resolved) return null
359  // Case-folded like core.ts: `Guardrails.MD` opens guardrails.md on a
360  // case-insensitive filesystem.
361  const rel = relUnder(root, resolved)?.toLowerCase() ?? null
362  if (rel === null) return null
363  for (const entry of PROTECTED_PATHS) {
364    const pattern = entry.toLowerCase()
365    if (pattern.endsWith("/")) {
366      if (rel.startsWith(pattern)) return entry
367    } else if (rel === pattern) {
368      return entry
369    }
370  }
371  return null
372}
373
374function protectedMessage(entry: string): string {
375  return (
376    `SystemOne-gate: protected file — ${entry}\n` +
377    `  Policy and harness configuration are edited by your human, not the agent.\n` +
378    `  Ask them to make the change and restart ${HARNESS}.`
379  )
380}
381
382// Where the path lands, every symlink followed — including for a file that
383// does not exist yet: the nearest existing ancestor is resolved and the tail
384// appended, as core.ts's walk does. null when nothing resolves (an odd
385// spelling, a dangling link), and the guard then denies: the engine denies
386// without a realPath too.
387function placeable(path: string): boolean {
388  const cut = Math.max(path.lastIndexOf("/"), path.lastIndexOf("\\"))
389  const name = path.slice(cut + 1)
390  return (
391    !/^[A-Za-z]:(?![\\/])/.test(path) &&
392    !/^[\\/][\\/]/.test(path) &&
393    !/^[A-Za-z]:/.test(name) &&
394    name !== "" &&
395    name !== "." &&
396    name !== ".."
397  )
398}
399
400async function place(
401  $: { fs: { stat(path: string, options: { resolve: boolean }): Promise<{ realPath?: string } | undefined> } },
402  path: string,
403): Promise<string | null> {
404  if (!placeable(path)) return null
405  const tail: string[] = []
406  let target = path
407  for (;;) {
408    const own = await $.fs.stat(target, { resolve: true }).catch(() => undefined)
409    if (own) {
410      if (own.realPath === undefined) return null
411      let real = own.realPath.replace(/[\\/]+$/, "")
412      for (const name of tail) real = `${real}/${name}`
413      return real
414    }
415    const stripped = target.replace(/[\\/]+$/, "")
416    const cut = Math.max(stripped.lastIndexOf("/"), stripped.lastIndexOf("\\"))
417    const name = stripped.slice(cut + 1)
418    if (cut < 0 || name === "" || name === "." || name === ".." || /^[A-Za-z]:/.test(name)) return null
419    tail.unshift(name)
420    target = stripped.slice(0, cut + 1)
421  }
422}
423
424// --- the module -------------------------------------------------------------
425
426interface Session {
427  guardrails: string | null
428  key: string | undefined
429  baseURL: string
430  model: string
431  threshold: number
432  failOpen: boolean
433  off: boolean
434  cwd: string
435  root: string
436  warnings: string[]
437  cooldown: ReturnType<typeof createCooldown>
438}
439
440class EndpointError extends Error {
441  constructor(
442    readonly status: number,
443    body: string,
444  ) {
445    super(`System One endpoint answered ${status}: ${body.slice(0, 120)}`)
446  }
447}
448
449async function firstReadable(read: (path: string) => Promise<string>, paths: string[]): Promise<string | null> {
450  for (const path of paths) {
451    try {
452      return await read(path)
453    } catch {}
454  }
455  return null
456}
457
458export const register: Register = (on) => {
459  let session: Session | null = null
460
461  on("session.start", async ($, e, next) => {
462    const raw = await firstReadable((path) => $.fs.read(path), GUARDRAIL_PATHS.map((p) => `${e.cwd}/${p}`))
463    const guardrails = raw === null ? null : truncateGuardrails(raw)
464    const rootStat = await $.fs.stat(e.cwd, { resolve: true }).catch(() => undefined)
465    const threshold = parseThreshold(await $.env.get("SYSTEMONE_THRESHOLD"))
466    session = {
467      cwd: e.cwd,
468      root: rootStat?.realPath && rootStat.realPath !== "" ? rootStat.realPath : e.cwd,
469      guardrails,
470      key: (await $.env.get("BERGET_API_KEY")) ?? (await $.env.get("TYPESAFE_API_KEY")),
471      baseURL: gatewayRoot((await $.env.get("BERGET_BASE_URL")) ?? (await $.env.get("TYPESAFE_BASE_URL")) ?? "https://api.berget.ai"),
472      model: (await $.env.get("BERGET_MODEL")) ?? (await $.env.get("TYPESAFE_DEFAULT_MODEL")) ?? "berget/bev",
473      threshold: threshold.value,
474      failOpen: isEnabled(await $.env.get("SYSTEMONE_FAIL_OPEN")),
475      off: (await $.env.get("SYSTEMONE_GATE")) === "off",
476      warnings: [threshold.warning, truncationWarning(guardrails)].filter((w): w is string => w !== undefined),
477      cooldown: createCooldown(),
478    }
479    for (const warning of session.warnings) $.ui.toast(warning)
480    if (!session.off && !session.key) {
481      $.ui.toast("guardrails-md inactive: no BERGET_API_KEY or TYPESAFE_API_KEY — Bash and Monitor commands run ungated this session")
482    }
483    return next(e)
484  })
485
486  on("tool.call", { tool: ["Bash", "Monitor"] }, async ($, e, next) => {
487    if (e.command === undefined) return next(e)
488    if (!session) return { deny: "SystemOne-gate: no frozen policy — session.start did not run. Restart Claude Code (fail-closed)." }
489    const s = session
490    const command = e.command.trim()
491    if (s.off || !s.key || !command) return next(e)
492    const cooling = s.cooldown.reason()
493    if (cooling) return { deny: cooling }
494    const deny = (reason: string) => ({ deny: [reason, ...s.warnings.map((w) => `  ${w}`)].join("\n") })
495
496    const scripts: string[] = []
497    for (const token of command.split(/\s+/)) {
498      if (scripts.length >= SCRIPT_FILES_MAX) break
499      const path = scriptPath(token)
500      if (!path) continue
501      try {
502        const st = await $.fs.stat(path)
503        if (st.kind !== "file" || st.size > SCRIPT_BYTES_MAX) continue
504        const section = scriptSection(path, await $.fs.read(path))
505        if (section) scripts.push(section)
506      } catch {}
507    }
508
509    let answers: unknown
510    const stop = new AbortController()
511    try {
512      const timeout = $.clock.sleep(TIMEOUT_MS, { signal: AbortSignal.any([stop.signal, next.signal]) }).then(() => {
513        throw new Error(`System One endpoint timed out after ${TIMEOUT_MS} ms`)
514      })
515      timeout.catch(() => {})
516      const res = await Promise.race([
517        $.http.fetch(`${s.baseURL}/v1/systemone`, {
518          method: "POST",
519          headers: { Authorization: `Bearer ${s.key}`, Accept: "application/json", "Content-Type": "application/json" },
520          body: JSON.stringify({
521            state: { text: buildStateText(command, scripts.join("\n"), s.guardrails) },
522            questions: buildQuestions(!!s.guardrails),
523            model: s.model,
524          }),
525        }),
526        timeout,
527      ])
528      if (!res.ok) throw new EndpointError(res.status, res.text)
529      answers = JSON.parse(res.text).answers
530    } catch (err) {
531      if (s.failOpen) return next(e)
532      return deny(endpointFailure(command, err))
533    } finally {
534      stop.abort()
535    }
536
537    const outcome = decide(verdictFrom(answers), !!s.guardrails, s.threshold)
538    if (outcome.decision === "allow") return next(e)
539    s.cooldown.register()
540    return deny(blockMessage(outcome, !!s.guardrails, s.threshold) + `\n  ${command.slice(0, 200)}`)
541  }).catch(($, e, next) => ({
542    deny: `SystemOne-gate: the gate failed (${next.error.kind}${next.error.message ? `: ${next.error.message}` : ""}) — command blocked (fail-closed).`,
543  }))
544
545  on("tool.call", { tool: ["Edit", "Write", "NotebookEdit"] }, async ($, e, next) => {
546    if (!session) return { deny: "SystemOne-gate: no frozen policy — session.start did not run. Restart Claude Code (fail-closed)." }
547    const s = session
548    if (s.off) return next(e)
549    const target = e.tool === "NotebookEdit" ? e.notebook_path : e.file_path
550    if (typeof target !== "string" || target === "") return next(e)
551    const placed = await place($, target)
552    if (placed === null) {
553      return {
554        deny:
555          `SystemOne-gate: cannot resolve ${JSON.stringify(target)} — edit blocked (fail-closed).\n` +
556          `  The path could not be checked against the protected list.`,
557      }
558    }
559    const hit = protectedPath(s.root, placed) ?? protectedPath(s.cwd, target)
560    if (hit === null) return next(e)
561    return { deny: [protectedMessage(hit), ...s.warnings.map((w) => `  ${w}`)].join("\n") }
562  }).catch(($, e, next) => ({
563    deny: `SystemOne-gate: the gate failed (${next.error.kind}${next.error.message ? `: ${next.error.message}` : ""}) — edit blocked (fail-closed).`,
564  }))
565}
566