SLOPSHOPPER

AntStreet

Fund an idea from Claude Code: you approve the checks first, a sandboxed gate outside the agent decides what passed, and a signed ledger records it.

newpanebandtoastprocess
v0.0.1Apache-2.0updated 2026-10-08kgorle1111/antstreet
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · antstreet
│ ┃ antstreet-approve ✕ › fix the failing auth test and add an audit log call │ ┃ No AntStreet run is awaiting approval. │ ┃ [ Close ] ⏺ Read(src/auth.ts) │ ⎿ Read 6 lines │ ⏺ Update(src/auth.ts) │ ⎿ Added 2 lines, removed 1 line │ ⏺ Bash(bun test) │ ⎿ 3 pass, 1 fail │ │ ● Done. refresh now rejects expired claims and logs an audit event. │ │ ✻ Worked for 42s · done 4:20 PM │ │ │ │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts

Draws

Pane · antstreet-approve
No AntStreet run is awaiting approval. [ Close ]
README

AntStreet: tests your coding agent never saw

<img src="docs/assets/hero.svg" alt="AntStreet: your AI agents get paid when the checks pass." width="100%">

license Apache-2.0 python 3.12+ tests 5700+ typed mypy strict status pre-release

The full technical tour: README-technical.md

AntStreet checks your AI coding agent's work against tests it never saw, so "all tests pass" actually means something.

Your agent writes the code and the tests, like a student who writes their own exam. AntStreet is the answer key the student never sees.

🚦 In three steps

  1. 🔏 Seal. Before the agent starts, antstreet audit plan drafts checks from your change request. You read and approve them, and they are sealed with a SHA-256, outside the repo.
  2. 🤖 Let any agent work. Claude Code, Codex, Cursor, a person: the audit never talks to it, and the agent never sees the checks.
  3. 🧱 Check the "done". antstreet audit check --claim done runs the sealed checks on the agent's commit in a sandbox and signs a verdict: refuted, unrefuted, inconclusive or no_claim.

The AI drafts. You and plain code decide.

🎬 See it run

$ antstreet audit plan --request ~/req.txt
Running the checks on the base...
Check c01 [t1] runs of symbols become one hyphen
Check c02 [t1] hyphens are trimmed at both ends
Check c03 [t1] a plain word is unchanged
Check c04 [t1] needs a library
  c01: fails on the base: counted
  c02: fails on the base: counted
  c03: passes on the base: shown, not counted
  c04: cannot run here: not counted (needs module 'nosuchlib_zq')
[a]pprove, [r]eject, or [e]dit files and re-check? a
Sealed audit run <run>: 2 of 4 checks fail on the base and will be counted.
Seal: <sha256>

# the agent works on its branch, commits, and says it is done

$ antstreet audit check --claim done
Verdict: REFUTED (claim: done, pre-registered)
Counted checks (failing on the base): 2; failing on the head: 2
  c01 failed: runs of symbols become one hyphen
  c02 failed: hyphens are trimmed at both ends

The agent said "done". Two checks it never saw say otherwise. This is the output of a test-suite run with a fake model, abridged (each check's code is left out), so the request is a toy one: make slugify collapse runs of symbols into one hyphen. Every output line is one the real command prints; a test re-runs it to keep it that way. Try it on your own repo in the quickstart.

🧨 The problem

Agents say "done" and mean "I stopped". We measured a narrower thing: how often a run can satisfy the checks the model drafted and still be wrong.

29 of 77 runs (38%, 95% interval 28-49%) passed every check the model had written and still failed a hidden check it never saw (written by Claude, separately from the agents measured). We read all 29 against the task text and judge each a real error; one class (tokenbucket's int-vs-float return) is debatable.

Caveats, kept on purpose: 17 small Python tasks, Haiku writing both the checks and the code, and runs of one task are not independent (resampling tasks widens the interval to 20-57%). 13 of the 29 failed only on non-ASCII input or a returned type; without those, 16 of 77 (21%). These were AntStreet's own build runs, not a general agent failure rate. Read the false-pass audit.

The lesson: tests written by the same mind that wrote the code share its blind spots. Checks sealed before the work starts, and never shown to the agent, do not.

🔭 How it flows

flowchart LR
  A["📝 Change request"] --> B["🔏 Checks drafted,<br/>run on the base,<br/>you approve, sealed"]
  B --> C["🤖 Any agent works<br/>(never sees the checks)"]
  C --> D["🧱 Sandboxed gate runs<br/>the sealed checks<br/>on its commit"]
  D --> E["🧾 Signed verdict:<br/>refuted / unrefuted /<br/>inconclusive"]

✨ Why it is different

🙈 The agent never sees the checksThey are drafted from your request and the base's public names only, and kept in a store outside the repo.
🎯 Only checks that bite are countedEach check runs on the base first. Only the ones that fail there, and so can tell a finished change from an unfinished one, decide the verdict.
🧱 A gate, not a model, decidespytest in an OS sandbox, with a signed proof that every test really ran. The agent's own words never reach the verdict.
🔎 It notices cheatingA change that quotes the sealed checks is flagged, and base tests the agent deleted or broke are listed beside the verdict.
🔏 A record you can checkEvery approval and verdict is on a signed, hash-chained ledger, and an automatic decision is never signed in your name.

🚀 Quickstart

You need macOS or Linux, Python 3.12+, uv and Claude Code, logged in (claude auth login). AntStreet is coming to PyPI; until then, run it from a clone:

git clone https://github.com/kgorle1111/antstreet.git ~/antstreet
cd ~/antstreet && uv sync

Then, in the Python repo your agent will change (a clean working tree):

alias antstreet="uv run --project ~/antstreet antstreet"
antstreet audit plan --request ~/req.txt   # the change request, in plain words, kept outside the repo
# ... the agent works and commits ...
antstreet audit check --claim done         # the latest sealed run, against HEAD
antstreet audit report                     # every verdict, and the false-pass rate

audit plan makes one model call (Haiku by default) to draft the checks. The checks run with your repo's own .venv packages, read-only inside the sandbox; nothing is installed. Every option: docs/CLI.md.

🧪 We measure, and we publish the "no"

Every experiment's rule is fixed in bench/PREREG.md before it runs. Four have run. None has shown its benefit, and every one is published.

The headline loss: a fair, blinded benchmark of AntStreet's build mode. 35 tasks, three runs each, $0.40 a task, hidden checks written separately from the agents being measured (by Claude, in a different session), validated against a reference solution and planted wrong solutions, and never shown to any agent.

passedcost per task
AntStreet (boss + ants)64 of 105 (61%)$0.2149
one agent62 of 105 (59%)$0.0883

Paired by task, the difference is not shown, and the firm cost about 2.4 times as much. So we do not claim a team of agents builds better than one agent. The earlier, rosier headline did not survive a fair re-run. The value is verification and control, which is why the audit leads. Details: blind35 notes, bench/METHOD.md.

The closest call so far, E4b: the build mode with a fresh-context critic delivered 74 of 105 (70%) against 61% for self-review, with the fewest false passes (24%), at about 1.56x the cost per delivered task. The pre-registered interval is +0.0952 [-0.0000, +0.1905]: its lower bound is not above zero, so the verdict is not shown, and we do not round it up. E4b write-up.

🛠️ How it's engineered

Built like something you would be happy to inherit.

  • 5,700+ tests, no model calls needed (recorded CLI output and fake claude executables).
  • Coverage floor 96% (line and branch) enforced in CI; about 98% measured when the floor was set.
  • mypy --strict over src/, plus ruff.
  • A threat model with 71 rows, each bound to the test for its control (THREAT_MODEL.md).
  • A signed, hash-chained ledger: edit a line and the chain breaks.
  • A sandboxed gate: macOS seatbelt; Linux bwrap, run and required in CI.
  • 7 pre-registered experiments, with every change to the plan dated (bench/PREREG.md).
  • 46 recorded design decisions, including what was rejected and why (DECISIONS.md).
  • Docs that cannot drift: tests fail when this README, the CLI docs or the ledger docs disagree with the code, and the audit output above is re-run and compared with what the command prints.

👋 Built by

Kannishk Naidu Gorle, an AI engineer working on evals, agent reliability and systems that hold up under a deep review. AntStreet was built with AI coding agents and held to the same discipline it sells: every change through a pull request, every claim bound to a test, every negative result published.

🧰 Advanced and experimental: antstreet fund

AntStreet started as a firm of agents: you give it an idea and a budget, the boss drafts checks you approve, and budget-capped worker agents build until the gate says pass, with every dollar on the ledger.

<img src="docs/assets/demo.svg" alt="Terminal replay: boss fund, the term sheet, approval, a worker slice, the gate verdict and the board report" width="100%">

It works end to end, and it is not shown to build better than a single agent, at about 2.4 times the cost (the table above). Try it if you are curious; do not pick it for results.

cd ~/antstreet
uv run antstreet doctor --live      # checks this machine; two paid calls of at most $0.05 each
uv run antstreet fund "A function is_palindrome(text) that ignores case, spaces and punctuation." --budget 0.40

The pig is you, the investor. The bull is the boss, the model that drafts the checks. The ants are the worker agents. Optional roles (the critic is the one we suggest) and model dispatch are off unless you ask for them. Every command: docs/CLI.md.

<img src="docs/assets/mascot-pig.svg" alt="The pig: the investor, you, who funds the idea and approves the checks" width="30%"> <img src="docs/assets/mascot-bull.svg" alt="The bull: the boss, which drafts the checks and runs the crew" width="30%"> <img src="docs/assets/mascot-ant.svg" alt="The ant: a worker agent, building in a capped slice" width="30%">

🧭 Status and honest limits

Early and pre-release. It works end to end, and:

  • unrefuted is not proof. The sealed checks catch only what they test, and a model drafts them: you reading them before you approve is the control.
  • A determined adversary can forge a pass (T12, accepted). Do not run ideas or code from sources you do not trust.
  • An agent running as your own OS user could read the audit store; keeping it out needs another user or a container (T51 in the threat model).
  • Python with pytest only; one machine, one user. Not on PyPI yet.

Everything else, with the threat rows: README-technical.md and THREAT_MODEL.md.

🗺️ Where it's going (planned, not built)

Recently shipped: audits that run in your repo's own environment (#64), the E4b result (#58), optional roles out of beta (#60), and a ledger that never signs an automatic decision in your name (#61).

Next: check strength, which shows which checks would catch a broken change; approval as a few yes/no questions instead of reading test code; a one-command, zero-install audit; a Claude Code Stop hook that audits the agent's "done" before the session ends; and CI export for the GitHub Action. Planned after that: the Agent Honesty Report, a pre-registered public measurement of how often coding agents pass their own tests while hidden checks fail.

The full plan, every experiment and what we will not do: ROADMAP.md.

🤝 Contributing

Bug reports, new benchmark tasks and sharper checks are welcome. Start with CONTRIBUTING.md. Security problems: SECURITY.md. If this made you think, a star helps others find it. ⭐

📄 License

Apache License 2.0. See LICENSE.

Source 2 files
hooks/approve-pane.tsx 181 lines
1// The approve pane: a Claude Code mod that only draws and calls the `antstreet` CLI. Every
2// decision (is a run waiting, what the sheet says, whether the value still matches) is the CLI's.
3// The one call that approves runs from the Approve Button's onPress and nowhere else: the mods API
4// has no method that presses a Button, so the model cannot reach it (tests/test_plugin.py holds
5// this source to that).
6import { atom, read, update } from 'claude-code'
7import type { EngineInterface, Register } from 'claude-code'
8
9import type { Pending, Result } from './approve-pane-state'
10
11const PANE = 'antstreet-approve'
12const TIMINGS = '.boss/approve-timings.jsonl'
13const RUN_ID = /^[A-Za-z0-9][A-Za-z0-9._-]*$/
14const SHEET = /--sheet ([0-9a-f]{16})\b/g
15
16const pending = atom({ plugin: 'antstreet', key: 'pending' } as const, null)
17const shownAt = atom({ plugin: 'antstreet', key: 'shownAt' } as const, null)
18const result = atom({ plugin: 'antstreet', key: 'result' } as const, null)
19const hidden = atom({ plugin: 'antstreet', key: 'hidden' } as const, null)
20const busy = atom({ plugin: 'antstreet', key: 'busy' } as const, false)
21
22async function antstreet($: EngineInterface, cli: readonly string[], args: readonly string[]) {
23  return $.process.run([...cli, ...args], { timeoutMs: 60_000 })
24}
25
26// What is waiting, read fresh from the CLI: `status --json`, then the sheet `approve RUN` shows.
27// Anything unexpected (no run, a refusal, output that does not parse) reads as nothing waiting.
28async function refresh($: EngineInterface, cli: readonly string[]): Promise<Pending | null> {
29  let found: Pending | null = null
30  try {
31    const status = await antstreet($, cli, ['status', '--json'])
32    const state = status.exitCode === 0 ? JSON.parse(status.stdout) : null
33    const run = state?.awaiting === true ? state.run : null
34    if (typeof run === 'string' && RUN_ID.test(run)) {
35      const shown = await antstreet($, cli, ['approve', run])
36      const values = new Set([...shown.stdout.matchAll(SHEET)].map(m => m[1]))
37      const [sheet] = values
38      if (shown.exitCode === 0 && values.size === 1 && sheet !== undefined) {
39        const cut = shown.stdout.indexOf(`\nRun ${run} is awaiting`)
40        const text = (cut < 0 ? shown.stdout : shown.stdout.slice(0, cut)).trimEnd()
41        found = { run, sheet, text }
42      }
43    }
44  } catch {
45    found = null
46  }
47  await update($, pending, () => found)
48  return found
49}
50
51async function review($: EngineInterface, cli: readonly string[]) {
52  await update($, result, () => null)
53  await refresh($, cli)
54  await $.ui.open({ id: PANE, title: 'AntStreet: approve', focus: true, closeOnEscape: true })
55  const now = await $.clock.now()
56  await update($, shownAt, () => now)
57}
58
59// kn: read-then-write append, not atomic; two sessions approving in the same instant could drop a
60// line. Fine for a friction metric; move the write into the CLI if it ever feeds a decision.
61async function recordTiming($: EngineInterface, line: object) {
62  const before = await $.fs.read(TIMINGS).catch(() => '')
63  const text = typeof before === 'string' ? before : ''
64  await $.fs.write(TIMINGS, `${text}${JSON.stringify(line)}\n`)
65}
66
67// The only place `approve --sheet` is run: from the Approve Button's onPress.
68async function approve($: EngineInterface, cli: readonly string[], surface: string) {
69  if (await read($, busy)) return
70  const shown = await read($, pending)
71  const openedAt = await read($, shownAt)
72  if (shown === null) return
73  await update($, busy, () => true)
74  try {
75    const pressedAt = await $.clock.now()
76    const done = await antstreet($, cli, ['approve', shown.run, '--sheet', shown.sheet])
77    const said = `${done.stdout}${done.stderr}`.trim()
78    await update($, result, (): Result => ({ exitCode: done.exitCode, text: said }))
79    const line = {
80      run: shown.run,
81      sheet: shown.sheet,
82      shown_at_ms: openedAt,
83      pressed_at_ms: pressedAt,
84      ms: openedAt === null ? null : pressedAt - openedAt,
85      exit_code: done.exitCode,
86      surface,
87    }
88    await recordTiming($, line).catch(() => $.ui.toast(`AntStreet: could not write ${TIMINGS}`))
89    await refresh($, cli)
90  } catch (error) {
91    const text = `Not approved: ${String(error)}`
92    await update($, result, (): Result => ({ exitCode: -1, text }))
93  } finally {
94    await update($, busy, () => false)
95  }
96}
97
98export const register: Register = (on, options) => {
99  const cli = String(options.command ?? 'uvx antstreet')
100    .split(/\s+/)
101    .filter(Boolean)
102
103  on('session.start', async ($, e, next) => {
104    const started = await next(e)
105    void refresh($, cli)
106    return started
107  })
108
109  on('turn.complete', async ($, e, next) => {
110    const done = await next(e)
111    void refresh($, cli) // the agent may have just run `antstreet fund` with no terminal (exit 4)
112    return done
113  })
114
115  on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
116    const waiting = await read($, pending)
117    if (e.props.hasSurvey || waiting === null || (await read($, hidden)) === waiting.run) {
118      return next(e)
119    }
120    const { Box, Button, Text } = $.ui.resolve(e)
121    return (
122      <Box>
123        <Text>AntStreet run {waiting.run} is awaiting your approval. </Text>
124        <Button key="review" label="Review" onPress={() => review($, cli)} />
125        <Text> </Text>
126        <Button key="hide" label="Hide" onPress={() => update($, hidden, () => waiting.run)} />
127      </Box>
128    )
129  })
130
131  on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
132    const { Box, Button, Text } = $.ui.resolve(e)
133    const shown = await read($, pending)
134    const done = await read($, result)
135    const isBusy = await read($, busy)
136    return (
137      <Box flexDirection="column">
138        {done !== null && (
139          <Text key="result" color={done.exitCode === 0 ? 'green' : 'red'} wrap="wrap">
140            {done.text}
141          </Text>
142        )}
143        {shown === null && done === null && (
144          <Text dimColor>No AntStreet run is awaiting approval.</Text>
145        )}
146        {shown !== null && (
147          <Box flexDirection="column">
148            <Text key="sheet" wrap="wrap">
149              {shown.text}
150            </Text>
151            <Text dimColor wrap="wrap">
152              Approve records your signed approval of exactly the text above (value {shown.sheet})
153              by running `antstreet approve {shown.run} --sheet {shown.sheet}`. It spends nothing;
154              build it with `antstreet resume {shown.run}`. Close leaves the run waiting.
155            </Text>
156          </Box>
157        )}
158        <Box>
159          <Button
160            key="close"
161            label="Close"
162            role="dismiss"
163            autoFocus
164            onPress={() => $.ui.close({ id: PANE })}
165          />
166          <Text> </Text>
167          {shown !== null && !isBusy && (
168            <Button
169              key="approve"
170              label="Approve"
171              variant="primary"
172              onPress={pe => approve($, cli, pe.surface)}
173            />
174          )}
175          {isBusy && <Text dimColor>Approving...</Text>}
176        </Box>
177      </Box>
178    )
179  })
180}
181
hooks/approve-pane-state.d.ts 15 lines
1export type Pending = { run: string; sheet: string; text: string }
2export type Result = { exitCode: number; text: string }
3
4declare module 'claude-code' {
5  interface PluginState {
6    antstreet: {
7      pending: Pending | null
8      shownAt: number | null
9      result: Result | null
10      hidden: string | null
11      busy: boolean
12    }
13  }
14}
15