Fund an idea from Claude Code: you approve the checks first, a sandboxed gate outside the agent decides what passed, and a signed ledger records it.

<img src="docs/assets/hero.svg" alt="AntStreet: your AI agents get paid when the checks pass." width="100%">
The full technical tour: README-technical.md
AntStreet checks your AI coding agent's work against tests it never saw, so "all tests pass" actually means something.
Your agent writes the code and the tests, like a student who writes their own exam. AntStreet is the answer key the student never sees.
antstreet audit plan drafts checks from your change request. You read and approve them, and they are sealed with a SHA-256, outside the repo.antstreet audit check --claim done runs the sealed checks on the agent's commit in a sandbox and signs a verdict: refuted, unrefuted, inconclusive or no_claim.The AI drafts. You and plain code decide.
$ antstreet audit plan --request ~/req.txt
Running the checks on the base...
Check c01 [t1] runs of symbols become one hyphen
Check c02 [t1] hyphens are trimmed at both ends
Check c03 [t1] a plain word is unchanged
Check c04 [t1] needs a library
c01: fails on the base: counted
c02: fails on the base: counted
c03: passes on the base: shown, not counted
c04: cannot run here: not counted (needs module 'nosuchlib_zq')
[a]pprove, [r]eject, or [e]dit files and re-check? a
Sealed audit run <run>: 2 of 4 checks fail on the base and will be counted.
Seal: <sha256>
# the agent works on its branch, commits, and says it is done
$ antstreet audit check --claim done
Verdict: REFUTED (claim: done, pre-registered)
Counted checks (failing on the base): 2; failing on the head: 2
c01 failed: runs of symbols become one hyphen
c02 failed: hyphens are trimmed at both ends
The agent said "done". Two checks it never saw say otherwise. This is the output of a test-suite run with a fake model, abridged (each check's code is left out), so the request is a toy one: make slugify collapse runs of symbols into one hyphen. Every output line is one the real command prints; a test re-runs it to keep it that way. Try it on your own repo in the quickstart.
Agents say "done" and mean "I stopped". We measured a narrower thing: how often a run can satisfy the checks the model drafted and still be wrong.
29 of 77 runs (38%, 95% interval 28-49%) passed every check the model had written and still failed a hidden check it never saw (written by Claude, separately from the agents measured). We read all 29 against the task text and judge each a real error; one class (tokenbucket's int-vs-float return) is debatable.
Caveats, kept on purpose: 17 small Python tasks, Haiku writing both the checks and the code, and runs of one task are not independent (resampling tasks widens the interval to 20-57%). 13 of the 29 failed only on non-ASCII input or a returned type; without those, 16 of 77 (21%). These were AntStreet's own build runs, not a general agent failure rate. Read the false-pass audit.
The lesson: tests written by the same mind that wrote the code share its blind spots. Checks sealed before the work starts, and never shown to the agent, do not.
flowchart LR
A["📝 Change request"] --> B["🔏 Checks drafted,<br/>run on the base,<br/>you approve, sealed"]
B --> C["🤖 Any agent works<br/>(never sees the checks)"]
C --> D["🧱 Sandboxed gate runs<br/>the sealed checks<br/>on its commit"]
D --> E["🧾 Signed verdict:<br/>refuted / unrefuted /<br/>inconclusive"]
| 🙈 The agent never sees the checks | They are drafted from your request and the base's public names only, and kept in a store outside the repo. |
| 🎯 Only checks that bite are counted | Each check runs on the base first. Only the ones that fail there, and so can tell a finished change from an unfinished one, decide the verdict. |
| 🧱 A gate, not a model, decides | pytest in an OS sandbox, with a signed proof that every test really ran. The agent's own words never reach the verdict. |
| 🔎 It notices cheating | A change that quotes the sealed checks is flagged, and base tests the agent deleted or broke are listed beside the verdict. |
| 🔏 A record you can check | Every approval and verdict is on a signed, hash-chained ledger, and an automatic decision is never signed in your name. |
You need macOS or Linux, Python 3.12+, uv and Claude Code, logged in (claude auth login). AntStreet is coming to PyPI; until then, run it from a clone:
git clone https://github.com/kgorle1111/antstreet.git ~/antstreet
cd ~/antstreet && uv sync
Then, in the Python repo your agent will change (a clean working tree):
alias antstreet="uv run --project ~/antstreet antstreet"
antstreet audit plan --request ~/req.txt # the change request, in plain words, kept outside the repo
# ... the agent works and commits ...
antstreet audit check --claim done # the latest sealed run, against HEAD
antstreet audit report # every verdict, and the false-pass rate
audit plan makes one model call (Haiku by default) to draft the checks. The checks run with your repo's own .venv packages, read-only inside the sandbox; nothing is installed. Every option: docs/CLI.md.
Every experiment's rule is fixed in bench/PREREG.md before it runs. Four have run. None has shown its benefit, and every one is published.
The headline loss: a fair, blinded benchmark of AntStreet's build mode. 35 tasks, three runs each, $0.40 a task, hidden checks written separately from the agents being measured (by Claude, in a different session), validated against a reference solution and planted wrong solutions, and never shown to any agent.
| passed | cost per task | |
|---|---|---|
| AntStreet (boss + ants) | 64 of 105 (61%) | $0.2149 |
| one agent | 62 of 105 (59%) | $0.0883 |
Paired by task, the difference is not shown, and the firm cost about 2.4 times as much. So we do not claim a team of agents builds better than one agent. The earlier, rosier headline did not survive a fair re-run. The value is verification and control, which is why the audit leads. Details: blind35 notes, bench/METHOD.md.
The closest call so far, E4b: the build mode with a fresh-context critic delivered 74 of 105 (70%) against 61% for self-review, with the fewest false passes (24%), at about 1.56x the cost per delivered task. The pre-registered interval is +0.0952 [-0.0000, +0.1905]: its lower bound is not above zero, so the verdict is not shown, and we do not round it up. E4b write-up.
Built like something you would be happy to inherit.
claude executables).mypy --strict over src/, plus ruff.bwrap, run and required in CI.Kannishk Naidu Gorle, an AI engineer working on evals, agent reliability and systems that hold up under a deep review. AntStreet was built with AI coding agents and held to the same discipline it sells: every change through a pull request, every claim bound to a test, every negative result published.
antstreet fundAntStreet started as a firm of agents: you give it an idea and a budget, the boss drafts checks you approve, and budget-capped worker agents build until the gate says pass, with every dollar on the ledger.
<img src="docs/assets/demo.svg" alt="Terminal replay: boss fund, the term sheet, approval, a worker slice, the gate verdict and the board report" width="100%">
It works end to end, and it is not shown to build better than a single agent, at about 2.4 times the cost (the table above). Try it if you are curious; do not pick it for results.
cd ~/antstreet
uv run antstreet doctor --live # checks this machine; two paid calls of at most $0.05 each
uv run antstreet fund "A function is_palindrome(text) that ignores case, spaces and punctuation." --budget 0.40
The pig is you, the investor. The bull is the boss, the model that drafts the checks. The ants are the worker agents. Optional roles (the critic is the one we suggest) and model dispatch are off unless you ask for them. Every command: docs/CLI.md.
<img src="docs/assets/mascot-pig.svg" alt="The pig: the investor, you, who funds the idea and approves the checks" width="30%"> <img src="docs/assets/mascot-bull.svg" alt="The bull: the boss, which drafts the checks and runs the crew" width="30%"> <img src="docs/assets/mascot-ant.svg" alt="The ant: a worker agent, building in a capped slice" width="30%">
Early and pre-release. It works end to end, and:
unrefuted is not proof. The sealed checks catch only what they test, and a model drafts them: you reading them before you approve is the control.Everything else, with the threat rows: README-technical.md and THREAT_MODEL.md.
Recently shipped: audits that run in your repo's own environment (#64), the E4b result (#58), optional roles out of beta (#60), and a ledger that never signs an automatic decision in your name (#61).
Next: check strength, which shows which checks would catch a broken change; approval as a few yes/no questions instead of reading test code; a one-command, zero-install audit; a Claude Code Stop hook that audits the agent's "done" before the session ends; and CI export for the GitHub Action. Planned after that: the Agent Honesty Report, a pre-registered public measurement of how often coding agents pass their own tests while hidden checks fail.
The full plan, every experiment and what we will not do: ROADMAP.md.
Bug reports, new benchmark tasks and sharper checks are welcome. Start with CONTRIBUTING.md. Security problems: SECURITY.md. If this made you think, a star helps others find it. ⭐
Apache License 2.0. See LICENSE.
hooks/approve-pane.tsx 181 lines1// The approve pane: a Claude Code mod that only draws and calls the `antstreet` CLI. Every
2// decision (is a run waiting, what the sheet says, whether the value still matches) is the CLI's.
3// The one call that approves runs from the Approve Button's onPress and nowhere else: the mods API
4// has no method that presses a Button, so the model cannot reach it (tests/test_plugin.py holds
5// this source to that).
6import { atom, read, update } from 'claude-code'
7import type { EngineInterface, Register } from 'claude-code'
8
9import type { Pending, Result } from './approve-pane-state'
10
11const PANE = 'antstreet-approve'
12const TIMINGS = '.boss/approve-timings.jsonl'
13const RUN_ID = /^[A-Za-z0-9][A-Za-z0-9._-]*$/
14const SHEET = /--sheet ([0-9a-f]{16})\b/g
15
16const pending = atom({ plugin: 'antstreet', key: 'pending' } as const, null)
17const shownAt = atom({ plugin: 'antstreet', key: 'shownAt' } as const, null)
18const result = atom({ plugin: 'antstreet', key: 'result' } as const, null)
19const hidden = atom({ plugin: 'antstreet', key: 'hidden' } as const, null)
20const busy = atom({ plugin: 'antstreet', key: 'busy' } as const, false)
21
22async function antstreet($: EngineInterface, cli: readonly string[], args: readonly string[]) {
23 return $.process.run([...cli, ...args], { timeoutMs: 60_000 })
24}
25
26// What is waiting, read fresh from the CLI: `status --json`, then the sheet `approve RUN` shows.
27// Anything unexpected (no run, a refusal, output that does not parse) reads as nothing waiting.
28async function refresh($: EngineInterface, cli: readonly string[]): Promise<Pending | null> {
29 let found: Pending | null = null
30 try {
31 const status = await antstreet($, cli, ['status', '--json'])
32 const state = status.exitCode === 0 ? JSON.parse(status.stdout) : null
33 const run = state?.awaiting === true ? state.run : null
34 if (typeof run === 'string' && RUN_ID.test(run)) {
35 const shown = await antstreet($, cli, ['approve', run])
36 const values = new Set([...shown.stdout.matchAll(SHEET)].map(m => m[1]))
37 const [sheet] = values
38 if (shown.exitCode === 0 && values.size === 1 && sheet !== undefined) {
39 const cut = shown.stdout.indexOf(`\nRun ${run} is awaiting`)
40 const text = (cut < 0 ? shown.stdout : shown.stdout.slice(0, cut)).trimEnd()
41 found = { run, sheet, text }
42 }
43 }
44 } catch {
45 found = null
46 }
47 await update($, pending, () => found)
48 return found
49}
50
51async function review($: EngineInterface, cli: readonly string[]) {
52 await update($, result, () => null)
53 await refresh($, cli)
54 await $.ui.open({ id: PANE, title: 'AntStreet: approve', focus: true, closeOnEscape: true })
55 const now = await $.clock.now()
56 await update($, shownAt, () => now)
57}
58
59// kn: read-then-write append, not atomic; two sessions approving in the same instant could drop a
60// line. Fine for a friction metric; move the write into the CLI if it ever feeds a decision.
61async function recordTiming($: EngineInterface, line: object) {
62 const before = await $.fs.read(TIMINGS).catch(() => '')
63 const text = typeof before === 'string' ? before : ''
64 await $.fs.write(TIMINGS, `${text}${JSON.stringify(line)}\n`)
65}
66
67// The only place `approve --sheet` is run: from the Approve Button's onPress.
68async function approve($: EngineInterface, cli: readonly string[], surface: string) {
69 if (await read($, busy)) return
70 const shown = await read($, pending)
71 const openedAt = await read($, shownAt)
72 if (shown === null) return
73 await update($, busy, () => true)
74 try {
75 const pressedAt = await $.clock.now()
76 const done = await antstreet($, cli, ['approve', shown.run, '--sheet', shown.sheet])
77 const said = `${done.stdout}${done.stderr}`.trim()
78 await update($, result, (): Result => ({ exitCode: done.exitCode, text: said }))
79 const line = {
80 run: shown.run,
81 sheet: shown.sheet,
82 shown_at_ms: openedAt,
83 pressed_at_ms: pressedAt,
84 ms: openedAt === null ? null : pressedAt - openedAt,
85 exit_code: done.exitCode,
86 surface,
87 }
88 await recordTiming($, line).catch(() => $.ui.toast(`AntStreet: could not write ${TIMINGS}`))
89 await refresh($, cli)
90 } catch (error) {
91 const text = `Not approved: ${String(error)}`
92 await update($, result, (): Result => ({ exitCode: -1, text }))
93 } finally {
94 await update($, busy, () => false)
95 }
96}
97
98export const register: Register = (on, options) => {
99 const cli = String(options.command ?? 'uvx antstreet')
100 .split(/\s+/)
101 .filter(Boolean)
102
103 on('session.start', async ($, e, next) => {
104 const started = await next(e)
105 void refresh($, cli)
106 return started
107 })
108
109 on('turn.complete', async ($, e, next) => {
110 const done = await next(e)
111 void refresh($, cli) // the agent may have just run `antstreet fund` with no terminal (exit 4)
112 return done
113 })
114
115 on('ui.render', { component: 'AbovePrompt' }, async ($, e, next) => {
116 const waiting = await read($, pending)
117 if (e.props.hasSurvey || waiting === null || (await read($, hidden)) === waiting.run) {
118 return next(e)
119 }
120 const { Box, Button, Text } = $.ui.resolve(e)
121 return (
122 <Box>
123 <Text>AntStreet run {waiting.run} is awaiting your approval. </Text>
124 <Button key="review" label="Review" onPress={() => review($, cli)} />
125 <Text> </Text>
126 <Button key="hide" label="Hide" onPress={() => update($, hidden, () => waiting.run)} />
127 </Box>
128 )
129 })
130
131 on('ui.render', { component: 'Pane', requestId: PANE }, async ($, e) => {
132 const { Box, Button, Text } = $.ui.resolve(e)
133 const shown = await read($, pending)
134 const done = await read($, result)
135 const isBusy = await read($, busy)
136 return (
137 <Box flexDirection="column">
138 {done !== null && (
139 <Text key="result" color={done.exitCode === 0 ? 'green' : 'red'} wrap="wrap">
140 {done.text}
141 </Text>
142 )}
143 {shown === null && done === null && (
144 <Text dimColor>No AntStreet run is awaiting approval.</Text>
145 )}
146 {shown !== null && (
147 <Box flexDirection="column">
148 <Text key="sheet" wrap="wrap">
149 {shown.text}
150 </Text>
151 <Text dimColor wrap="wrap">
152 Approve records your signed approval of exactly the text above (value {shown.sheet})
153 by running `antstreet approve {shown.run} --sheet {shown.sheet}`. It spends nothing;
154 build it with `antstreet resume {shown.run}`. Close leaves the run waiting.
155 </Text>
156 </Box>
157 )}
158 <Box>
159 <Button
160 key="close"
161 label="Close"
162 role="dismiss"
163 autoFocus
164 onPress={() => $.ui.close({ id: PANE })}
165 />
166 <Text> </Text>
167 {shown !== null && !isBusy && (
168 <Button
169 key="approve"
170 label="Approve"
171 variant="primary"
172 onPress={pe => approve($, cli, pe.surface)}
173 />
174 )}
175 {isBusy && <Text dimColor>Approving...</Text>}
176 </Box>
177 </Box>
178 )
179 })
180}
181hooks/approve-pane-state.d.ts 15 lines1export type Pending = { run: string; sheet: string; text: string }
2export type Result = { exitCode: number; text: string }
3
4declare module 'claude-code' {
5 interface PluginState {
6 antstreet: {
7 pending: Pending | null
8 shownAt: number | null
9 result: Result | null
10 hidden: string | null
11 busy: boolean
12 }
13 }
14}
15