SLOPSHOPPER

jev-vercel-sandbox

A Vercel Sandbox Claude sends work to on purpose: long test runs in the background, someone else's repository, an installer of unknown origin, a bulk change to…

newpaneguardcommandtoaststatus
A shopper browsing a rack in a slop shop
Preview · a replayed session in a sandbox
claude · ~/work/app · jev-vercel-sandbox
│ ┃ sandbox ✕ › fix the failing auth test and add an audit log call │ ┃ Vercel Sandbox │ ┃ – not configured ● jev-vercel-sandbox: [jev-vercel-sandbox] sandbox_run not registered │ ┃ ● jev-vercel-sandbox: [jev-vercel-sandbox] sandbox_jobs not registere │ ┃ Set vercelToken, vercelTeamId, ⏺ Read(src/auth.ts) │ ┃ vercelProjectId in the plugin's options ⎿ Read 6 lines │ ┃ (settings.json, ⏺ Update(src/auth.ts) │ ┃ pluginConfigs["jev-vercel-sandbox@skills-dir ⎿ Added 2 lines, removed 1 line │ ┃ "].options). ⏺ Bash(bun test) │ ┃ ⎿ 3 pass, 1 fail │ ┃ Jev │ ┃ built-in classifier · asks first ● Done. refresh now rejects expired claims and logs an audit event. │ ┃ 0 detected · 0 sandboxed · 0 local │ ┃ ✻ Worked for 42s · done 4:20 PM │ ┃ Jobs (0) │ ┃ none yet › /jev-vercel-sandbox │ ┃ ⎿ jev-vercel-sandbox: jev-vercel-sandbox: not configured, set verc │ ┃ [ close ] ● jev-vercel-sandbox: [jev-vercel-sandbox] sandbox_result not registe │ ● jev-vercel-sandbox: [jev-vercel-sandbox] sandbox_apply not register │ │ ────────────────────────────────────────────────────────────────────────────────────────────────────────────────────── › ? for shortcuts ⚠ jev-vercel-sandbox: sandbox · not configured · 0 jobs · asks first

Draws

Pane · sandbox
Vercel Sandbox – not configured Set vercelToken, vercelTeamId, vercelProjectId in the plugin's options (settings.json, pluginConfigs["jev-vercel-sandbox@skills-dir"].options). Jev built-in classifier · asks first 0 detected · 0 sandboxed · 0 local Jobs (0) none yet [ close ]
README

jev-vercel-sandbox

A Vercel Sandbox that Claude sends work to on purpose: a long test run that goes on in the background while Claude keeps working, someone else's repository, an installer of unknown origin, a bulk change you want to preview before it touches your files, a build from a clean checkout. The sandbox is an isolated Linux microVM (Firecracker); nothing a job does there reaches your disk unless you apply its patch yourself.

Whether a command may run on your machine at all is not this mod's call: auto mode and jev-guardrails / jev-auto-mode decide that. This one only moves the jobs a sandbox is good for.

How a job gets to the sandbox

Jev spots it. Every Bash command Claude is about to run goes to Jev, TypeSafe's System One decision model, with one yes/no question per use case (plain reads such as ls, cat, git status never do):

Use caseJev asks whether the command…Starts fromRunsHands back
testsis a full test, lint or typecheck run that takes a whilethe project as it is nowin the backgroundits output, as a new message when it ends
external_repoclones and runs another repositoryan empty folderwhile Claude waitsits output
untrusted_codedownloads and runs code of unknown origin (`curl … \sh`, a new package)an empty folderwhile Claude waitsits output, the files it wrote outside its folder, the processes it left running
repo_changechanges many project files at once (deleting folders, a codemod, a mass rename, a dependency upgrade)the project as it is nowwhile Claude waitsthe files it changed, and a patch that is not applied
clean_buildbuilds or installs the project from scratchthe project as it is nowin the backgroundits output

The likeliest use case at or above threshold (0.5) counts. useCases limits which ones are looked for.

You decide. With confirmSandbox on (the default), Claude Code's own question dialog asks:

Sandbox  Jev: `rm -rf src/legacy` looks like a bulk change to the project, to preview first.
         Run it in the Vercel Sandbox?
  1. Run in sandbox
  2. Run locally
  3. Always sandbox repo change

Run locally (or dismissing the dialog, or a claude -p run with no one to ask) lets the command go on as if the mod were not there, through Claude Code's usual permission prompt. Always sandbox … stops asking for that kind of job for the rest of the session.

With confirmSandbox: false every job Jev spots goes to the sandbox without a question, and the sandbox is started if it is not running (a start that failed is not retried until you run /jev-vercel-sandbox restart; those commands run locally meanwhile).

Claude can also ask for it with its own tool, mcp__jev-vercel-sandbox__sandbox_run:

sandbox_run { command: "npm ci && npm test", workspace: { git: "https://github.com/org/lib", ref: "v2.3.0" } }
sandbox_run { command: "npm test", background: true }
sandbox_run { command: "npx jscodeshift -t rename.js src", returns: "patch" }

workspace is "current" (default), "none", or { git, ref } (https only). Jev still reads the command: code of unknown origin never gets the project's copy, whatever Claude asked for. If the sandbox is not running, you are asked whether to start it (confirmSandbox off: it starts). sandbox_jobs lists the jobs, sandbox_result returns a job's whole output, and sandbox_apply applies a job's patch, always after asking you.

What every result says

The first line of every result names the job and what it started from, so neither you nor Claude mistakes it for something that happened on your disk:

[sandbox · j3 · tree 3f2a91c] rm -rf src/legacy: exit 0 · 1.2s
in the job's folder: 27 files changed: -src/legacy/a.ts, -src/legacy/b.ts, …
patch kept (0.1 MB); not applied to the user's project

Claude also reads a note: the command did not run on your machine, the copy has no .env files, keys or environment variables, and the patch can be applied with /jev-vercel-sandbox apply j3 or sandbox_apply (which asks you).

A background job answers at once (queued, runs in the background) and its result arrives later as a new message, with the same header.

Each job starts from the project as it is now

/jev-vercel-sandbox uploads the project once. Before each job, only what changed since goes up: the mod builds a git tree of the files that go through a throwaway index (GIT_INDEX_FILE=<tmp> git update-index --add then git write-tree; your own index and git status are untouched), sends git diff --binary <last tree> <this tree> (text even for binary files), and the sandbox applies it. The tree in the result's header is that one: if you edit a file while a background job runs, you can tell its result predates the edit.

Each job runs in a folder of its own (git worktree add of the synced copy, a clone, or an empty repository), so two jobs never share files and an experiment never dirties a test run.

Applying a patch is always yours: /jev-vercel-sandbox apply j3 (or the pane's apply button, or Claude's sandbox_apply) asks first, then runs git apply --check and git apply in your project. A patch that no longer applies (you changed those files since) is refused with git's reason, and one touching .git, .claude, .env*, keys or credentials is never applied.

What goes up. What git ls-files --cached lists: the files git tracks, as they are on disk (uncommitted edits included). Untracked files stay home, because an untracked file (secrets.yaml, service-account.json, .envrc) can hold credentials no name list recognizes. uploadUntracked: true adds the untracked files .gitignore does not exclude and, in a folder that is not a git repository, copies every file outside .git and node_modules (such a folder is copied once and not synced). Either way these are left out:

  • .env and .env.* (except .env.example, .sample, .template, .dist), .dev.vars, .npmrc, .pypirc, .netrc, .git-credentials, .pgpass, credentials(.json), Terraform state (*.tfstate)
  • *.pem, *.key, *.p12, *.pfx, *.jks, *.keystore, *.ppk, id_rsa/id_ed25519 and the like
  • anything inside .git, .ssh, .aws, .gnupg, .kube, .docker, .terraform, .vercel, .claude or node_modules

The first copy is packed with your tar, capped at workspaceMaxMB (10 MB compressed, also the cap on one sync), and uploaded as base64 through the command endpoint, 400 KB per request, because a mod's $.http.fetch sends text only. It needs git, tar, mktemp and split on your machine, and git in the sandbox image (the mod installs it with dnf or apt-get when missing). uploadWorkspace: false skips the copy: every job then runs in an empty folder.

The pane and the command

/jev-vercel-sandbox opens the side pane and starts the sandbox, the microVM booting while the project is packed:

Vercel Sandbox
● ready

✓ Starting the microVM
    sbx-quiet-owl · iad1 · 2 vCPU
✓ Packing the project
    412 files · 1.8 MB · 1 credential file left out
✓ Uploading it
✓ Unpacking in the sandbox

Jev checks each command Claude runs. Tests, someone else's repo, an unknown
installer, a bulk change or a clean build: it asks you first, and they run
here, not on your machine.

Jev
typesafe jev-latest · asks first
3 detected · 2 sandboxed · 1 local

Jobs (2)
◌ j2 npm test                      running
✓ j1 rm -rf src/legacy                   0
    tree 3f2a91c · 27 files changed · patch

[ restart ] [ stop ] [ apply j1 ] [ close ]
  • /jev-vercel-sandbox (or start): starts it and opens the pane; once ready, shows the state.
  • open, restart (a new sandbox with a fresh copy), stop, status
  • jobs, result <job>, apply <job>
  • run <command>: your own job on the current project, in the background; its result shows in the pane and the transcript (Claude is not interrupted)

Lifetime: sandboxTimeoutMinutes (45); Vercel stops it then. It is stopped when the session ends (stopOnExit). With persistent: false (default) nothing is kept. Billing is Vercel's Active CPU pricing plus provisioned memory while it runs: see pricing.

Options

Set them in user settings (~/.claude/settings.json, not the project's), --settings <file>, managed settings, or /config:

{
  "pluginConfigs": {
    "jev-vercel-sandbox@skills-dir": {
      "options": {
        "vercelToken": "<your Vercel access token>",
        "vercelTeamId": "team_…",
        "vercelProjectId": "prj_…",
        "typesafeApiKey": "<your TypeSafe key>",
        "confirmSandbox": true
      }
    }
  }
}

The key is the plugin's id: "jev-vercel-sandbox@skills-dir" when installed with --mod, "jev-vercel-sandbox" with --plugin-dir. Under the wrong key every option stays at its default and the pane says the credentials are missing.

Vercel authentication: an access token (vercel.com/account/tokens) plus the team id and project id it is scoped to, or an OIDC token (vercel env pull writes VERCEL_OIDC_TOKEN), which carries the team and project itself but expires after about 12 hours. Any of the three left empty falls back to VERCEL_TOKEN (then VERCEL_OIDC_TOKEN), VERCEL_TEAM_ID and VERCEL_PROJECT_ID in the environment Claude Code was started with. Without them, every command runs locally and nothing is asked.

  vercelToken:           string  access token or OIDC token (sensitive); env VERCEL_TOKEN / VERCEL_OIDC_TOKEN
  vercelTeamId:          string  team_…; env VERCEL_TEAM_ID; read from an OIDC token
  vercelProjectId:       string  prj_…; env VERCEL_PROJECT_ID; read from an OIDC token
  sandboxImage:          string  empty: vercel/sandbox/universal
  sandboxVcpus:          number  unset: Vercel's default
  sandboxTimeoutMinutes: number  lifetime before Vercel stops it (default 45)
  persistent:            boolean keep the filesystem across stops (default false)
  uploadWorkspace:       boolean copy the project into the sandbox on start (default true)
  uploadUntracked:       boolean copy untracked files too; needed outside git (default false)
  workspaceMaxMB:        number  cap on the compressed copy and on one sync (default 10)
  stopOnExit:            boolean stop when the session ends (default true)
  commandTimeoutMs:      number  a job's timeout when Claude sets none (default 120000, max 600000)
  confirmSandbox:        boolean ask before a detected job goes to the sandbox (default true)
  useCases:              string  comma-separated; empty: tests, external_repo, untrusted_code, repo_change, clean_build
  typesafeApiKey:        string  TypeSafe key (sensitive, preferred: calibrated probabilities)
  gatewayApiKey:         string  Vercel AI Gateway key (sensitive)
  provider:              string  "auto" | "typesafe" | "gateway" | "builtin"
  typesafeBaseUrl / typesafeModel / gatewayBaseUrl / gatewayModel
  threshold:             number  probability a use case needs (default 0.5)
  judgeTimeoutMs:        number  Jev latency budget; no answer runs the command locally (default 2000)
  columns:               number  pane width, 28-80 (default 44)
  vercelApiUrl:          string  empty: https://api.vercel.com
  logDecisions:          boolean log each detection and job (default true)

With no Jev key the engine's own $.model.classify picks the use case (or none) with the same questions as a rubric; that path has no probabilities, so threshold does not apply.

Privacy

With a Jev key, your latest message and each command that is not a plain read go to TypeSafe or the AI Gateway. /jev-vercel-sandbox uploads the project copy described above to Vercel, and each job sends what changed since and its command. No environment variable leaves the machine through this mod; untracked files (unless uploadUntracked is on) and the files listed as left out never do.

Install

npx claude-code-templates@latest --mod security/jev-vercel-sandbox
claude

--mod writes the plugin to .claude/skills/jev-vercel-sandbox/, which Claude Code auto-loads as jev-vercel-sandbox@skills-dir in a trusted project (accept the trust prompt on the first interactive claude there; -p never asks). In the fullscreen layout (/tui fullscreen) the pane docks beside the transcript.

For one session, or in a folder you do not want to trust:

claude --plugin-dir .claude/skills/jev-vercel-sandbox

claude plugin validate .claude/skills/jev-vercel-sandbox lists every event it hooks, every $ call, and the four environment variables it reads.

Tests

cd cli-tool/components/mods
claude plugin test security/jev-vercel-sandbox

They run the hooks against a fake Vercel API, a fake git/tar and a scripted question dialog: nothing starts with the session and the four tools are registered; a command that is no job, a plain read, "Run locally" and a dismissed question all run locally; the start uploads the project without its secrets; a confirmed job runs in its own worktree, is marked remote and keeps a patch it does not apply; an edit since the start goes up as a sync patch first, once; "Always" stops the questions; an installer runs in an empty folder and reports what it left; a test run goes to the background and reports back; a job confirmed before the start starts the sandbox; a hook that fails after the person chose the sandbox refuses the command; sandbox_run clones only https repositories and Jev keeps unknown code off the project copy; a patch is applied only after "Apply" and never when it touches .claude/; the size cap, an OIDC token, a stopped sandbox, the session end and the pane. tests/judge.test.ts and tests/workspace.test.ts cover the battery, the report reading and the URL checks. The sync, prepare and report scripts were run in bash against real git: the sandbox's tree after a sync matched the local tree hash, and the job's patch passed git apply --check in the original project.

Not yet checked against the live API: Vercel's limit on a request body (a failure shows as the upload or sync failing), and git in the default image.

Requirements. Mods are on by default in Claude Code 2.1.287+. Typed against Anthropic's declarations: https://github.com/anthropics/claude-code/tree/main/mods

A mod runs without node_modules, so @vercel/sandbox is not available: the Sandbox is driven over HTTP through $.http.fetch, with the endpoints of Vercel's REST API reference (POST /v3/sandboxes, GET /v2/sandboxes/sessions/{id}, POST …/cmd with wait and logs, answering application/x-ndjson, POST …/stop, each with ?teamId=), called the way @vercel/sandbox 3.5.0 calls them. The Jev wire shapes are jev-auto-mode's.

Source 4 files
hooks/jev-vercel-sandbox.tsx 1362 lines
1/**
2 * jev-vercel-sandbox — Claude Mod
3 *
4 * A Vercel Sandbox (an isolated Linux microVM) as a place Claude sends work to
5 * on purpose: a long test run, someone else's repository, an installer of
6 * unknown origin, a bulk change to preview, a clean build. Whether a command
7 * may run on the person's machine at all is not this mod's call (auto mode and
8 * jev-guardrails decide that); this one only moves the jobs a sandbox is for.
9 *
10 * TypeSafe's Jev reads every Bash command Claude is about to run and says
11 * whether it is one of those jobs. With confirmSandbox on (the default) the
12 * person is asked, in Claude Code's own question dialog, whether it runs in
13 * the sandbox or locally; with it off, a detected job always goes to the
14 * sandbox. Claude can also send work itself with the sandbox_run tool.
15 *
16 * Every job starts from the project as it is at that moment (the copy uploaded
17 * by /jev-vercel-sandbox, brought up to date with a patch) in its own git
18 * worktree, and every result says it is remote, which job it was and which
19 * tree it started from. Nothing a job does reaches the person's disk except a
20 * patch they apply themselves with /jev-vercel-sandbox apply, after a question.
21 *
22 *   session.start  registers /jev-vercel-sandbox and the sandbox_* tools, reads the credentials
23 *   command.run    /jev-vercel-sandbox: starts the sandbox and opens the pane (status|open|restart|stop|jobs|run|result|apply)
24 *   turn.start     records the user's request (the detector's context)
25 *   tool.call      Bash: Jev detects a sandbox job and asks where it runs;
26 *                  mcp__jev-vercel-sandbox__sandbox_*: Claude's own jobs
27 *   session.end    stops the sandbox
28 *   ui.render      the pane: the start's progress, then the jobs
29 *
30 * Vercel credentials and Jev keys come from the plugin's options
31 * (pluginConfigs["jev-vercel-sandbox@skills-dir"].options in user settings), with
32 * VERCEL_TOKEN / VERCEL_OIDC_TOKEN / VERCEL_TEAM_ID / VERCEL_PROJECT_ID as a
33 * fallback. Never from this code. Needs Claude Code >= 2.1.287.
34 *
35 * Privacy: with a Jev key set, the user's latest request and each command
36 * that is not a plain read go to that backend. /jev-vercel-sandbox uploads the
37 * project's files (what git tracks, minus .env*, keys and credentials;
38 * untracked files only with uploadUntracked) to Vercel, and each job sends
39 * what changed since; no environment variable goes.
40 */
41import type { ProcessRunInit, ProcessRunResult, Register } from 'claude-code'
42import {
43  DEFAULT_BASE_URL,
44  DEFAULT_MODEL,
45  PROFILES,
46  builtinLabels,
47  classifyText,
48  decide,
49  describeJudgement,
50  detectionReason,
51  endpoint,
52  isPlainRead,
53  parseUseCases,
54  readJudgement,
55  requestBody,
56  requestHeaders,
57  selectProvider,
58  shortCommand,
59  stateText,
60} from './judge.ts'
61import type { Detection, Judgement, Provider, Returns, UseCase } from './judge.ts'
62import { DEFAULT_API_URL, client, resolveCredentials } from './vercel.ts'
63import type { Client, CreateOptions, Credentials, Fetch, RunResult, SandboxInfo } from './vercel.ts'
64import {
65  APPEND_SCRIPT,
66  EXTRACT_SCRIPT,
67  PART_BYTES,
68  PATCH_SCRIPT,
69  PREPARE_SCRIPT,
70  REPORT_SCRIPT,
71  SYNC_FILE,
72  SYNC_SCRIPT,
73  TRUNCATE_SCRIPT,
74  UPLOAD_FILE,
75  describeChanges,
76  filesToUpload,
77  folderName,
78  isExcluded,
79  isGitRef,
80  isGitUrl,
81  megabytes,
82  patchTargets,
83  nulList,
84  readReport,
85  relativeCwd,
86  uploadBatches,
87} from './workspace.ts'
88import type { Report } from './workspace.ts'
89
90const PANE = 'jev-vercel-sandbox'
91const COMMAND = 'jev-vercel-sandbox'
92const TAG = '[jev-vercel-sandbox]'
93const TOOLS = { run: 'sandbox_run', jobs: 'sandbox_jobs', result: 'sandbox_result', apply: 'sandbox_apply' } as const
94const RECENT = 20
95// what the Bash tool itself keeps inline
96const MAX_OUTPUT = 30_000
97// what a job keeps for sandbox_result
98const KEEP_OUTPUT = 200_000
99// what a finished background job's message carries
100const SUBMIT_OUTPUT = 6_000
101const PATCH_MAX = 5 * 1_048_576
102const POLL_MS = 500
103const POLL_TRIES = 12
104const STEP_TIMEOUT_MS = 120_000
105
106type State = 'off' | 'idle' | 'starting' | 'ready' | 'failed' | 'stopped'
107type StepKey = 'vm' | 'pack' | 'upload' | 'prepare'
108type Step = { key: StepKey; label: string; state: 'wait' | 'run' | 'done' | 'fail' | 'skip'; detail: string }
109type Source = { kind: 'current' } | { kind: 'none' } | { kind: 'git'; url: string; ref: string }
110type Job = {
111  id: string
112  command: string
113  short: string
114  source: Source
115  useCase: UseCase | null
116  /** Who sent it: Jev's detection of a Bash command, Claude's sandbox_run, or the person's /jev-vercel-sandbox run. */
117  by: 'jev' | 'claude' | 'user'
118  background: boolean
119  returns: Returns
120  watch: boolean
121  timeoutMs: number
122  state: 'queued' | 'syncing' | 'running' | 'done' | 'error'
123  /** The tree it started from (7 characters), or what stands for one. */
124  tree: string
125  note: string
126  exitCode: number | null
127  ms: number
128  stdout: string
129  stderr: string
130  error?: string
131  report?: Report
132  dir?: string
133  /** The sandbox session it ran in: a patch lives only there. */
134  session?: string
135  applied?: boolean
136}
137type Workspace = {
138  /** The project's folder in the sandbox. */
139  dir: string
140  /** The session's working directory relative to the project root. */
141  rel: string
142  /** The project's root on this machine. */
143  root: string
144  files: number
145  excluded: number
146  bytes: number
147  /** Git works in the sandbox (the baseline, worktrees). */
148  git: boolean
149  /** The tree the sandbox's copy holds, when the project is a git repository here (else no sync). */
150  synced: string
151}
152type Choice = 'sandbox' | 'local'
153
154/** What a job needs from `$`, handed in as closures so it can run after the hook that started it returned. */
155type Host = {
156  fetch: Fetch
157  sleep: (ms: number) => Promise<void>
158  now: () => Promise<number>
159  run: (argv: readonly string[], init?: ProcessRunInit) => Promise<ProcessRunResult>
160  readBase64: (path: string) => Promise<string>
161  size: (path: string) => Promise<number>
162  list: (path: string) => Promise<string[]>
163  cwd: () => Promise<string>
164  log: (text: string) => void
165  toast: (text: string) => void
166  status: (text: string) => void
167  submit: (text: string) => void
168  redraw: () => void
169  classify: (text: string, labels: readonly string[]) => Promise<string | undefined>
170  ask: (question: string, options: readonly string[], header: string) => Promise<string>
171}
172
173let state: State = 'off'
174let info: SandboxInfo | undefined
175let workspace: Workspace | undefined
176let lastError: string | undefined
177let starting: Promise<void> | undefined
178let steps: Step[] = []
179let intent = ''
180let now = 0
181let isOpen = false
182let nextId = 1
183let syncChain: Promise<unknown> = Promise.resolve()
184const jobs: Job[] = []
185const detections = new Map<string, Detection>()
186/** Use cases the person said always go to the sandbox, this session. */
187const always = new Set<UseCase>()
188/** Bash calls the person (or confirmSandbox off) sent to the sandbox: they never fall through to the machine. */
189const chosen = new Set<string>()
190const tally = { detected: 0, sandboxed: 0, local: 0 }
191
192const plural = (n: number, word: string) => `${n} ${word}${n === 1 ? '' : 's'}`
193const messageOf = (err: unknown) => (err instanceof Error ? err.message : String(err))
194const shortTree = (sha: string) => sha.slice(0, 7)
195
196/** Both ends: how a run started and, where a test or build prints its failures and summary, how it ended. */
197function cap(text: string, max = MAX_OUTPUT): string {
198  if (text.length <= max) return text
199  const head = Math.floor(max / 2)
200  return `${text.slice(0, head)}\n… (${text.length - max} characters in the middle cut by jev-vercel-sandbox)\n${text.slice(-(max - head))}`
201}
202
203/** The tail, where a test run's summary is. */
204function tail(text: string, max: number): string {
205  return text.length > max ? `… (${text.length - max} earlier characters cut)\n${text.slice(-max)}` : text
206}
207
208function minutesLeft(): number | null {
209  if (!info || !info.timeout || !info.startedAt || !now) return null
210  return Math.max(0, Math.round((info.startedAt + info.timeout - now) / 60_000))
211}
212
213function freshSteps(upload: boolean): Step[] {
214  return [
215    { key: 'vm', label: 'Starting the microVM', state: 'wait', detail: '' },
216    { key: 'pack', label: 'Packing the project', state: upload ? 'wait' : 'skip', detail: upload ? '' : 'off (uploadWorkspace)' },
217    { key: 'upload', label: 'Uploading it', state: upload ? 'wait' : 'skip', detail: '' },
218    { key: 'prepare', label: 'Unpacking in the sandbox', state: upload ? 'wait' : 'skip', detail: '' },
219  ]
220}
221
222function step(key: StepKey, s: Step['state'], detail?: string): void {
223  // a failed start's other half may still be running: its progress no longer shows
224  if (state === 'failed') return
225  const found = steps.find(x => x.key === key)
226  if (!found) return
227  found.state = s
228  if (detail !== undefined) found.detail = detail
229}
230
231function findJob(id: string): Job | undefined {
232  const want = id.trim().toLowerCase().replace(/^#/, '')
233  return jobs.find(j => j.id === want || j.id === `j${want}`)
234}
235
236function sourceText(j: Job): string {
237  if (j.source.kind === 'git') return `clone of ${j.source.url}${j.source.ref ? `@${j.source.ref}` : ''}`
238  if (j.source.kind === 'none') return 'empty folder'
239  return j.tree ? `tree ${j.tree}` : 'the project'
240}
241
242/** The first line of every result: remote, which job, and what it started from. */
243function header(j: Job): string {
244  return `[sandbox · ${j.id} · ${sourceText(j)}]`
245}
246
247function exitText(j: Job): string {
248  if (j.state === 'error') return `failed: ${j.error ?? 'unknown error'}`
249  if (j.state !== 'done') return j.state
250  return j.exitCode === null ? 'no exit code' : `exit ${j.exitCode}`
251}
252
253function reportLines(j: Job): string[] {
254  const r = j.report
255  if (!r) return []
256  const lines: string[] = []
257  if (j.source.kind !== 'none' || r.changes.length) lines.push(`in the job's folder: ${describeChanges(r.changes)}`)
258  if (j.watch) {
259    lines.push(r.outside.length ? `written outside the job's folder (not /tmp): ${r.outside.slice(0, 30).join(', ')}${r.outside.length > 30 ? `, … ${r.outside.length - 30} more` : ''}` : 'nothing written outside the job\'s folder (besides /tmp)')
260    lines.push(r.processes.length ? `still running: ${r.processes.slice(0, 15).join(', ')}` : 'no process left running')
261  }
262  if (r.patchBytes) lines.push(`patch kept (${megabytes(r.patchBytes)}); not applied to the user's project`)
263  return lines
264}
265
266/** The whole result as text: header, exit, report, then the output. */
267function resultText(j: Job, max = MAX_OUTPUT): string {
268  const summary = [`${header(j)} ${j.short}: ${exitText(j)}${j.ms ? ` · ${Math.round(j.ms / 100) / 10}s` : ''}`, ...reportLines(j)]
269  const out = [j.stdout && `stdout:\n${cap(j.stdout, max)}`, j.stderr && `stderr:\n${cap(j.stderr, max)}`].filter(Boolean)
270  return [...summary, ...out].join('\n')
271}
272
273/** What Claude reads beside a finished job's result. */
274function jobContext(j: Job, sandboxName: string): string {
275  const from =
276    j.source.kind === 'git'
277      ? `in a fresh clone of ${j.source.url}${j.source.ref ? ` at ${j.source.ref}` : ''}`
278      : j.source.kind === 'none'
279        ? "in an empty folder with none of the project's files, environment variables or credentials"
280        : `on a copy of the project as it was when the job started (${j.tree ? `tree ${j.tree}` : 'the copy made at start'}; .env files, keys and credentials left out, and none of the user's environment variables)`
281  const patch = j.report?.patchBytes
282    ? ` The job's changes are kept as a patch and were NOT applied to the user's project. If they should be, tell the user they can apply it with /${COMMAND} apply ${j.id}, or call sandbox_apply (it asks the user first).`
283    : ''
284  return (
285    `This did not run on the user's machine. jev-vercel-sandbox ran it as job ${j.id} in Vercel Sandbox ${sandboxName} ${from}. ` +
286    `Nothing on the user's machine changed; files it wrote exist only in the sandbox, which is stopped when the session ends.${patch} ` +
287    `Do not run it again locally unless the user asks. sandbox_result ${j.id} returns its whole output.`
288  )
289}
290
291function remember(j: Job): Job {
292  jobs.push(j)
293  if (jobs.length > RECENT) {
294    // finished jobs go first; a running one is kept
295    const i = jobs.findIndex(x => x.state === 'done' || x.state === 'error')
296    if (i >= 0) jobs.splice(i, 1)
297  }
298  return j
299}
300
301export const register: Register = (on, options) => {
302  const text = (key: string, fallback = '') =>
303    typeof options[key] === 'string' && options[key] ? (options[key] as string).trim() : fallback
304  const number = (key: string, fallback: number) =>
305    typeof options[key] === 'number' && Number.isFinite(options[key]) ? (options[key] as number) : fallback
306  const flag = (key: string, fallback: boolean) => (typeof options[key] === 'boolean' ? (options[key] as boolean) : fallback)
307
308  // the detector
309  const typesafeKey = text('typesafeApiKey')
310  const gatewayKey = text('gatewayApiKey')
311  const active: Provider | null = selectProvider(text('provider', 'auto'), typesafeKey, gatewayKey)
312  const apiKey = active === 'typesafe' ? typesafeKey : active === 'gateway' ? gatewayKey : ''
313  const modelId = !active ? '' : active === 'typesafe' ? text('typesafeModel', DEFAULT_MODEL.typesafe) : text('gatewayModel', DEFAULT_MODEL.gateway)
314  const judgeUrl = !active
315    ? ''
316    : active === 'typesafe'
317      ? endpoint('typesafe', text('typesafeBaseUrl', DEFAULT_BASE_URL.typesafe))
318      : endpoint('gateway', text('gatewayBaseUrl', DEFAULT_BASE_URL.gateway))
319  const backend = active ? `${active} ${modelId}` : 'built-in classifier'
320  const useCases = parseUseCases(text('useCases'))
321  const threshold = number('threshold', 0.5)
322  const judgeTimeoutMs = number('judgeTimeoutMs', 2_000)
323  const confirmSandbox = flag('confirmSandbox', true)
324  const logDecisions = flag('logDecisions', true)
325
326  // the sandbox
327  const apiBase = text('vercelApiUrl', DEFAULT_API_URL)
328  const createOptions: CreateOptions = {
329    timeoutMs: Math.max(1, number('sandboxTimeoutMinutes', 45)) * 60_000,
330    image: text('sandboxImage') || undefined,
331    vcpus: number('sandboxVcpus', 0) || undefined,
332    persistent: flag('persistent', false),
333  }
334  const commandTimeoutMs = Math.min(600_000, Math.max(1_000, number('commandTimeoutMs', 120_000)))
335  const uploadWorkspace = flag('uploadWorkspace', true)
336  // untracked files are anything lying in the folder: never sent unless asked for
337  const uploadUntracked = flag('uploadUntracked', false)
338  const maxMB = Math.max(1, number('workspaceMaxMB', 10))
339  const columns = Math.min(80, Math.max(28, number('columns', 44)))
340
341  let credentials: Credentials | undefined
342  let credentialsKind = ''
343  let missing: string[] = []
344  const toolNames = new Map<string, keyof typeof TOOLS>()
345
346  const statusLine = () => {
347    const where =
348      state === 'ready' ? 'ready' : state === 'starting' ? 'starting…' : state === 'off' ? 'not configured' : state === 'idle' ? `/${COMMAND} to start` : state
349    const running = jobs.filter(j => j.state === 'queued' || j.state === 'syncing' || j.state === 'running').length
350    return `sandbox · ${where} · ${plural(jobs.length, 'job')}${running ? ` (${running} running)` : ''} · ${confirmSandbox ? 'asks first' : 'no questions'}`
351  }
352
353  const clampTimeout = (ms: unknown) => Math.min(600_000, typeof ms === 'number' && ms > 0 ? ms : commandTimeoutMs)
354
355  function newJob(fields: Pick<Job, 'command' | 'source' | 'useCase' | 'by' | 'background' | 'returns' | 'watch' | 'timeoutMs'> & { note?: string }): Job {
356    let source = fields.source
357    let note = fields.note ?? ''
358    if (source.kind === 'current' && !uploadWorkspace) {
359      source = { kind: 'none' }
360      note = [note, 'no copy of the project (uploadWorkspace is off)'].filter(Boolean).join('; ')
361    }
362    // a patch is of the project: another repo's or an empty folder's files never land in it
363    const returns = fields.returns === 'patch' && source.kind !== 'current' ? 'changes' : fields.returns
364    return remember({
365      ...fields,
366      returns,
367      source,
368      note,
369      id: `j${nextId++}`,
370      short: shortCommand(fields.command),
371      state: 'queued',
372      tree: '',
373      exitCode: null,
374      ms: 0,
375      stdout: '',
376      stderr: '',
377    })
378  }
379
380  // ── the project on this machine ───────────────────────────────────────────
381
382  /** The files that go: git's list minus deleted files and secrets (every file outside git, with uploadUntracked). */
383  async function listFiles(h: Host, root: string): Promise<{ files: string[]; excluded: string[]; git: boolean }> {
384    const listing = uploadUntracked ? ['git', 'ls-files', '-z', '--cached', '--others', '--exclude-standard'] : ['git', 'ls-files', '-z', '--cached']
385    const git = await h.run(listing, { cwd: root, timeoutMs: 60_000 }).catch(() => undefined)
386    if (git && git.exitCode === 0) {
387      // every entry ends in NUL: output cut at the limit ends mid-path
388      if (git.stdout && !git.stdout.endsWith('\0')) throw new Error('the file list was cut short; the project is too large to copy')
389      const gone = await h.run(['git', 'ls-files', '-z', '--deleted'], { cwd: root, timeoutMs: 60_000 }).catch(() => undefined)
390      return { ...filesToUpload(nulList(git.stdout), gone?.exitCode === 0 ? nulList(gone.stdout) : []), git: true }
391    }
392    // outside git every file is untracked
393    if (!uploadUntracked) throw new Error('the folder is not a git repository; set uploadUntracked to copy its files, or uploadWorkspace to false for an empty sandbox')
394    const found = await h.run(['find', '.', '-type', 'f', '-not', '-path', './.git/*', '-not', '-path', '*/node_modules/*', '-print0'], { cwd: root, timeoutMs: 60_000 })
395    if (found.exitCode !== 0) throw new Error(`listing the project failed: ${found.stderr.trim().slice(0, 200)}`)
396    if (found.stdout && !found.stdout.endsWith('\0')) throw new Error('the file list was cut short; the project is too large to copy')
397    return { ...filesToUpload(nulList(found.stdout)), git: false }
398  }
399
400  /** A folder of the machine's own temp dir, removed after `use`. */
401  async function withTemp<T>(h: Host, use: (tmp: string) => Promise<T>): Promise<T> {
402    const tmp = (await h.run(['mktemp', '-d'])).stdout.trim()
403    if (!tmp) throw new Error('mktemp -d gave no folder')
404    try {
405      return await use(tmp)
406    } finally {
407      await h.run(['rm', '-rf', tmp]).catch(() => undefined)
408    }
409  }
410
411  /**
412   * The tree of exactly these files as they are on disk, through a throwaway
413   * index: the person's own index and `git status` are untouched.
414   */
415  async function treeOf(h: Host, root: string, files: string[], tmp: string): Promise<string> {
416    const env = { GIT_INDEX_FILE: `${tmp}/index` }
417    const add = await h.run(['git', 'update-index', '--add', '-z', '--stdin'], { cwd: root, env, stdin: `${files.join('\0')}\0`, timeoutMs: STEP_TIMEOUT_MS })
418    if (add.exitCode !== 0) throw new Error(`git update-index failed: ${add.stderr.trim().slice(0, 200)}`)
419    const tree = await h.run(['git', 'write-tree'], { cwd: root, env, timeoutMs: 60_000 })
420    const sha = tree.stdout.trim()
421    if (tree.exitCode !== 0 || !/^[0-9a-f]{40,64}$/.test(sha)) throw new Error(`git write-tree failed: ${tree.stderr.trim().slice(0, 200)}`)
422    return sha
423  }
424
425  /** A file of the temp folder as base64, read in parts when it is over `$.fs.read`'s 4 MiB. */
426  async function readParts(h: Host, file: string, tmp: string): Promise<string> {
427    const bytes = await h.size(file)
428    let parts = [file]
429    if (bytes > PART_BYTES) {
430      const split = await h.run(['split', '-b', String(PART_BYTES), file, `${tmp}/part.`], { timeoutMs: 60_000 })
431      if (split.exitCode !== 0) throw new Error(`split failed: ${split.stderr.trim().slice(0, 200)}`)
432      parts = (await h.list(tmp)).filter(n => n.startsWith('part.')).sort().map(n => `${tmp}/${n}`)
433    }
434    let base64 = ''
435    for (const p of parts) base64 += await h.readBase64(p)
436    await h.run(['rm', '-f', ...parts.filter(p => p !== file)]).catch(() => undefined)
437    return base64
438  }
439
440  /** The project as one base64 gzip tar, and the tree it holds. */
441  async function pack(h: Host, root: string): Promise<{ base64: string; files: number; excluded: number; bytes: number; tree: string }> {
442    step('pack', 'run', 'listing files')
443    h.redraw()
444    const listed = await listFiles(h, root)
445    const { files, excluded } = listed
446    if (!files.length) throw new Error('the project has no files to copy')
447    step('pack', 'run', `${plural(files.length, 'file')}, compressing`)
448    h.redraw()
449    return withTemp(h, async tmp => {
450      const tree = listed.git ? await treeOf(h, root, files, tmp) : ''
451      const archive = `${tmp}/workspace.tgz`
452      const tar = await h.run(['tar', '-czf', archive, '--null', '-T', '-'], {
453        cwd: root,
454        stdin: `${files.join('\0')}\0`,
455        // macOS tar would add ._ AppleDouble files
456        env: { COPYFILE_DISABLE: '1' },
457        timeoutMs: STEP_TIMEOUT_MS,
458      })
459      if (tar.exitCode !== 0) throw new Error(`tar failed: ${tar.stderr.trim().slice(0, 200)}`)
460      const bytes = await h.size(archive)
461      if (bytes > maxMB * 1_048_576) throw new Error(`the project is ${megabytes(bytes)} compressed, over workspaceMaxMB (${maxMB} MB)`)
462      const base64 = await readParts(h, archive, tmp)
463      step('pack', 'done', `${plural(files.length, 'file')} · ${megabytes(bytes)}${excluded.length ? ` · ${excluded.length} credential file${excluded.length === 1 ? '' : 's'} left out` : ''}`)
464      h.redraw()
465      return { base64, files: files.length, excluded: excluded.length, bytes, tree }
466    })
467  }
468
469  // ── the sandbox ───────────────────────────────────────────────────────────
470
471  const must = (r: RunResult, what: string) => {
472    if (r.exitCode !== 0) throw new Error(`${what} failed in the sandbox: ${(r.error || r.stderr || `exit ${r.exitCode}`).trim().slice(0, 300)}`)
473    return r
474  }
475
476  /** base64 into a sandbox file, a few arguments per call. */
477  async function send(api: Client, s: SandboxInfo, base64: string, file: string, progress?: (done: number, total: number) => void): Promise<number> {
478    const batches = uploadBatches(base64)
479    must(await api.script(s, TRUNCATE_SCRIPT, [file], 30_000), 'preparing the upload')
480    for (let i = 0; i < batches.length; i++) {
481      must(await api.script(s, APPEND_SCRIPT, [file, ...batches[i]!], 60_000), 'uploading')
482      progress?.(i + 1, batches.length)
483    }
484    return batches.length
485  }
486
487  /** The microVM: create, then poll while it is pending. */
488  async function startVm(h: Host, api: Client): Promise<SandboxInfo> {
489    step('vm', 'run')
490    h.redraw()
491    let s = await api.create(createOptions)
492    // kept at once, so a sandbox whose polling fails can still be found and stopped
493    info = s
494    for (let i = 0; s.status === 'pending' && i < POLL_TRIES; i++) {
495      await h.sleep(POLL_MS)
496      s = await api.get(s)
497      info = s
498    }
499    if (s.status !== 'running') throw new Error(`the sandbox is ${s.status}`)
500    step('vm', 'done', `${s.name} · ${s.region} · ${s.vcpus} vCPU`)
501    h.redraw()
502    return s
503  }
504
505  /** The base64 into the sandbox, then unpacked with a git baseline. */
506  async function upload(h: Host, api: Client, s: SandboxInfo, base64: string, dir: string): Promise<boolean> {
507    step('upload', 'run', '0')
508    h.redraw()
509    const calls = await send(api, s, base64, UPLOAD_FILE, (done, total) => {
510      step('upload', 'run', `${done}/${total}`)
511      h.redraw()
512    })
513    step('upload', 'done', `${calls} request${calls === 1 ? '' : 's'}`)
514    step('prepare', 'run', dir)
515    h.redraw()
516    const r = must(await api.script(s, EXTRACT_SCRIPT, [UPLOAD_FILE, dir], STEP_TIMEOUT_MS), 'unpacking')
517    const git = r.stdout.trim().endsWith('git') && !r.stdout.trim().endsWith('no-git')
518    step('prepare', 'done', git ? dir : `${dir} · no git in the sandbox image: jobs cannot start from the project`)
519    h.redraw()
520    return git
521  }
522
523  /** The whole start, run after /jev-vercel-sandbox (or a confirmed job) asked for it; the pane follows every step. */
524  async function boot(h: Host, creds: Credentials): Promise<void> {
525    const api = client(h.fetch, creds, apiBase)
526    state = 'starting'
527    info = undefined
528    workspace = undefined
529    lastError = undefined
530    steps = freshSteps(uploadWorkspace)
531    h.status(statusLine())
532    h.redraw()
533    let vm: Promise<SandboxInfo> | undefined
534    try {
535      let root = ''
536      let cwd = ''
537      if (uploadWorkspace) {
538        cwd = await h.cwd()
539        const top = await h.run(['git', 'rev-parse', '--show-toplevel'], { cwd, timeoutMs: 10_000 }).catch(() => undefined)
540        root = top && top.exitCode === 0 && top.stdout.trim() ? top.stdout.trim() : cwd
541      }
542      // the microVM boots while the project is packed
543      const failing = (key: StepKey) => (err: unknown) => {
544        step(key, 'fail', messageOf(err))
545        throw err
546      }
547      vm = startVm(h, api).catch(failing('vm'))
548      const packing = uploadWorkspace ? pack(h, root).catch(failing('pack')) : undefined
549      // when one half fails the other runs on unobserved
550      vm.catch(() => undefined)
551      packing?.catch(() => undefined)
552      const [s, packed] = await Promise.all([vm, packing])
553      if (packed) {
554        const dir = `${(s.cwd || '/vercel/sandbox').replace(/\/+$/, '')}/${folderName(root)}`
555        const git = await upload(h, api, s, packed.base64, dir)
556        workspace = { dir, rel: relativeCwd(root, cwd) ?? '', root, files: packed.files, excluded: packed.excluded, bytes: packed.bytes, git, synced: packed.tree }
557      }
558      syncChain = Promise.resolve()
559      state = 'ready'
560      now = await h.now()
561      h.log(
562        `${TAG} sandbox ready: ${s.name} · ${s.region}${workspace ? ` · ${plural(workspace.files, 'file')} in ${workspace.dir}` : ''}. ` +
563          `Jev (${backend}) checks each command for sandbox jobs and ${confirmSandbox ? 'asks you before sending one' : 'sends them here without asking'}`,
564      )
565      h.toast('jev-vercel-sandbox: sandbox ready')
566    } catch (err) {
567      lastError = messageOf(err)
568      if (!steps.some(x => x.state === 'fail')) {
569        const running = steps.find(x => x.state === 'run')
570        step(running ? running.key : 'vm', 'fail', lastError)
571      }
572      state = 'failed'
573      h.log(`${TAG} the sandbox did not start: ${lastError}`)
574      h.toast(`jev-vercel-sandbox: the sandbox did not start (${lastError})`)
575      // never leave a half-started microVM running (and billed): when packing failed first,
576      // the create request may still be in flight, so wait for it before stopping what it made
577      await vm?.catch(() => undefined)
578      if (info) await api.stop(info).catch(() => undefined)
579    }
580    h.status(statusLine())
581    h.redraw()
582  }
583
584  function ensureStarted(h: Host, creds: Credentials): Promise<void> {
585    if (state === 'ready') return Promise.resolve()
586    if (state === 'starting' && starting) return starting
587    starting = boot(h, creds)
588    return starting
589  }
590
591  /** Brings the sandbox's copy to the project as it is now; the tree it then holds (empty when it cannot be synced). */
592  function sync(h: Host, api: Client, s: SandboxInfo): Promise<string> {
593    const run = syncChain.then(async () => {
594      const w = workspace
595      if (!w || !w.synced) return ''
596      const listed = await listFiles(h, w.root)
597      return withTemp(h, async tmp => {
598        const tree = await treeOf(h, w.root, listed.files, tmp)
599        if (tree === w.synced) return tree
600        const patch = `${tmp}/sync.patch`
601        const diff = await h.run(['git', 'diff', '--binary', '--full-index', `--output=${patch}`, w.synced, tree], { cwd: w.root, timeoutMs: STEP_TIMEOUT_MS })
602        if (diff.exitCode !== 0) throw new Error(`git diff failed: ${diff.stderr.trim().slice(0, 200)}`)
603        const bytes = await h.size(patch)
604        if (bytes > maxMB * 1_048_576) throw new Error(`the changes since the last sync are ${megabytes(bytes)}, over workspaceMaxMB; /${COMMAND} restart uploads the project again`)
605        if (bytes > 0) {
606          await send(api, s, await readParts(h, patch, tmp), SYNC_FILE)
607          must(await api.script(s, SYNC_SCRIPT, [SYNC_FILE, w.dir, `sync ${tree}`], STEP_TIMEOUT_MS), 'syncing the project')
608        }
609        w.synced = tree
610        return tree
611      })
612    })
613    syncChain = run.catch(() => undefined)
614    return run
615  }
616
617  /** One job, start to end: sync, its own folder, the command, the report. Never throws: a failure is the job's state. */
618  async function runJob(h: Host, api: Client, s: SandboxInfo, j: Job): Promise<void> {
619    const startedAt = await h.now()
620    try {
621      const jobsRoot = `${(s.cwd || '/vercel/sandbox').replace(/\/+$/, '')}/.jev-jobs`
622      const dir = `${jobsRoot}/${j.id}`
623      let rel = ''
624      if (j.source.kind === 'current') {
625        if (!workspace) throw new Error('the sandbox holds no copy of the project')
626        if (!workspace.git) throw new Error('git is not available in the sandbox image, so a job cannot start from the project')
627        j.state = 'syncing'
628        h.redraw()
629        const tree = await sync(h, api, s)
630        j.tree = tree ? shortTree(tree) : 'start copy'
631        if (!tree) j.note = [j.note, 'not a git repository here: the copy made at start, not synced'].filter(Boolean).join('; ')
632        rel = workspace.rel
633      }
634      const url = j.source.kind === 'git' ? j.source.url : ''
635      const ref = j.source.kind === 'git' ? j.source.ref : ''
636      must(await api.script(s, PREPARE_SCRIPT, [j.source.kind, workspace?.dir ?? '', dir, url, ref, rel], STEP_TIMEOUT_MS), 'preparing the job')
637      j.dir = dir
638      j.session = s.sessionId
639      j.state = 'running'
640      h.redraw()
641      const r = await api.run(s, j.command, j.timeoutMs, rel ? `${dir}/${rel}` : dir)
642      j.stdout = cap(r.stdout, KEEP_OUTPUT)
643      j.stderr = cap([r.stderr, r.error ? `jev-vercel-sandbox: ${r.error}` : ''].filter(Boolean).join('\n'), KEEP_OUTPUT)
644      j.exitCode = r.exitCode
645      const report = await api.script(s, REPORT_SCRIPT, [dir, j.watch ? '1' : '0', j.returns === 'patch' ? '1' : '0'], STEP_TIMEOUT_MS).catch(() => undefined)
646      if (report && report.exitCode === 0) j.report = readReport(report.stdout)
647      j.state = 'done'
648      j.ms = Math.round(r.durationMs ?? (await h.now()) - startedAt)
649    } catch (err) {
650      j.state = 'error'
651      j.error = messageOf(err)
652      j.ms = (await h.now()) - startedAt
653      // the session may have timed out or been stopped from the dashboard
654      const fresh = await api.get(s).catch(() => undefined)
655      if (fresh && fresh.status !== 'running') {
656        info = fresh
657        state = 'stopped'
658      }
659    }
660    now = await h.now()
661    tally.sandboxed += 1
662    if (logDecisions) h.log(`${TAG} ${header(j)} ${j.short}: ${exitText(j)}${j.report ? ` · ${describeChanges(j.report.changes, 3)}` : ''}`)
663    h.status(statusLine())
664    h.redraw()
665  }
666
667  /**
668   * A job that runs on after the hook answered (a background job, or any job
669   * while the sandbox starts): when it ends, Claude reads the result as a new
670   * message (the person's own jobs only log it).
671   */
672  function detach(h: Host, creds: Credentials, j: Job): void {
673    void (async () => {
674      await ensureStarted(h, creds).catch(() => undefined)
675      const s = info
676      if ((state as State) !== 'ready' || !s) {
677        j.state = 'error'
678        j.error = `the sandbox did not start (${lastError ?? state})`
679        h.redraw()
680      } else {
681        await runJob(h, client(h.fetch, creds, apiBase), s, j)
682      }
683      const done = `${header(j)} ${j.short}: ${exitText(j)}`
684      h.toast(`jev-vercel-sandbox: ${j.id} ${exitText(j)}`)
685      if (j.by === 'user') {
686        h.log(`${TAG} ${done}. /${COMMAND} result ${j.id} shows its output`)
687        return
688      }
689      h.submit(
690        [
691          `jev-vercel-sandbox: background job ${j.id} finished. It ran in the Vercel Sandbox, not on this machine; nothing on the user's machine changed.`,
692          tail(resultText(j, SUBMIT_OUTPUT), SUBMIT_OUTPUT * 2),
693          j.report?.patchBytes ? `Its changes are kept as a patch, not applied: the user applies it with /${COMMAND} apply ${j.id}, or call sandbox_apply (it asks the user).` : '',
694          `sandbox_result ${j.id} returns the whole output.`,
695        ]
696          .filter(Boolean)
697          .join('\n\n'),
698      )
699    })()
700  }
701
702  /** A patch a job kept, onto the person's project, after they say yes. */
703  async function applyJob(h: Host, j: Job): Promise<{ ok: boolean; text: string }> {
704    if (j.applied) return { ok: false, text: `${j.id}'s patch was already applied` }
705    if (j.state !== 'done' || !j.report?.patchBytes || !j.dir || j.source.kind !== 'current') return { ok: false, text: `${j.id} kept no patch to apply` }
706    if (!credentials || !info || state !== 'ready' || !workspace || info.sessionId !== j.session) return { ok: false, text: `the sandbox holding ${j.id} is no longer running` }
707    if (j.report.patchBytes > PATCH_MAX) return { ok: false, text: `${j.id}'s patch is ${megabytes(j.report.patchBytes)}, too large to bring back` }
708    // the patch itself, not the sandbox's word about it, says what it writes
709    const r = await client(h.fetch, credentials, apiBase).script(info, PATCH_SCRIPT, [`${j.dir}.patch`], 60_000)
710    if (r.exitCode !== 0 || r.error) return { ok: false, text: `the patch could not be read from the sandbox: ${(r.error || r.stderr).trim().slice(0, 200)}` }
711    if (r.stdout.length > PATCH_MAX) return { ok: false, text: `${j.id}'s patch is too large to bring back` }
712    const targets = patchTargets(r.stdout)
713    if (targets.symlink) return { ok: false, text: `${j.id}'s patch creates or changes a symlink; it was not applied` }
714    const refused = [...new Set([...targets.paths, ...j.report.changes.map(c => c.path)])].filter(isExcluded)
715    if (refused.length) return { ok: false, text: `${j.id}'s patch touches files this mod never writes (${refused.slice(0, 5).join(', ')}); it was not applied` }
716    const listed = targets.paths.length
717    let answer: string
718    try {
719      answer = await h.ask(
720        `Apply ${j.id}'s changes to your project? ${describeChanges(j.report.changes, 12)}${listed > j.report.changes.length ? ` (the patch names ${plural(listed, 'path')})` : ''} (from ${sourceText(j)}: ${j.short})`,
721        ['Apply', 'Cancel'],
722        'Apply',
723      )
724    } catch {
725      return { ok: false, text: 'not applied: nobody confirmed it' }
726    }
727    if (answer !== 'Apply') return { ok: false, text: 'not applied: the user said no' }
728    const check = await h.run(['git', 'apply', '--check', '-'], { cwd: workspace.root, stdin: r.stdout, timeoutMs: 60_000 })
729    if (check.exitCode !== 0) return { ok: false, text: `the patch no longer applies to your project (it changed since ${sourceText(j)}): ${check.stderr.trim().slice(0, 300)}` }
730    const applied = await h.run(['git', 'apply', '-'], { cwd: workspace.root, stdin: r.stdout, timeoutMs: 60_000 })
731    if (applied.exitCode !== 0) return { ok: false, text: `git apply failed: ${applied.stderr.trim().slice(0, 300)}` }
732    j.applied = true
733    h.redraw()
734    return { ok: true, text: `applied ${j.id}'s patch to ${workspace.root}: ${describeChanges(j.report.changes)}` }
735  }
736
737  function jobsText(): string {
738    if (!jobs.length) return 'no sandbox jobs yet'
739    return jobs
740      .map(j => `${j.id} · ${j.state === 'done' || j.state === 'error' ? exitText(j) : j.state} · ${sourceText(j)} · ${j.short}${j.report?.patchBytes ? ` · patch${j.applied ? ' applied' : ''}` : ''}`)
741      .join('\n')
742  }
743
744  // ── the detector ──────────────────────────────────────────────────────────
745
746  /** Jev's (or the built-in classifier's) answer for one command, cached per request. */
747  async function detect(h: Host, command: string, cwd: string, short: string): Promise<Detection> {
748    if (isPlainRead(command)) return { useCase: null, probability: null, by: 'read-only' }
749    const key = `${intent}\u0000${command}`
750    const cached = detections.get(key)
751    if (cached) return cached
752    const startedAt = await h.now()
753    const state_ = stateText(intent, command, cwd)
754    let found: Detection | undefined
755    let judgement: Judgement | null = null
756    try {
757      if (active) {
758        const response = await Promise.race([
759          h.fetch(judgeUrl, { method: 'POST', headers: requestHeaders(active, apiKey, modelId), body: requestBody(active, state_, modelId, useCases) }),
760          h.sleep(judgeTimeoutMs),
761        ])
762        if (response && response.ok) judgement = readJudgement(response.text, useCases)
763        else h.log(`${TAG} detector: ${response ? `${active} responded ${response.status}` : `no answer in ${judgeTimeoutMs}ms`}`)
764        if (judgement) found = decide(judgement, threshold)
765      } else {
766        const label = await h.classify(classifyText(state_, useCases), builtinLabels(useCases))
767        const useCase = useCases.find(c => c === label) ?? null
768        if (useCase || label === 'none') found = { useCase, probability: null, by: 'built-in' }
769      }
770    } catch (err) {
771      h.log(`${TAG} detector failed: ${messageOf(err)}`)
772    }
773    if (logDecisions && active) h.log(`${TAG} jev ${short}: ${describeJudgement(judgement, (await h.now()) - startedAt)}`)
774    if (!found) return { useCase: null, probability: null, by: 'no answer' }
775    detections.set(key, found)
776    if (detections.size > 300) detections.delete(detections.keys().next().value as string)
777    return found
778  }
779
780  const askLabels = (kind: UseCase) => ['Run in sandbox', 'Run locally', `Always sandbox ${kind.replace(/_/g, ' ')}`]
781
782  // ── hooks ─────────────────────────────────────────────────────────────────
783
784  on('session.start', async ($, e, next) => {
785    const r = await next(e)
786    await $.command
787      .register({
788        name: COMMAND,
789        description: 'Start the Vercel Sandbox and show its jobs (status|open|restart|stop|jobs|run <cmd>|result <job>|apply <job>)',
790        argumentHint: '[status|open|restart|stop|jobs|run <command>|result <job>|apply <job>]',
791        immediate: true,
792      })
793      .catch(err => $.ui.log(`${TAG} /${COMMAND} not registered: ${err}`))
794
795    const specs = [
796      {
797        key: 'run' as const,
798        description:
799          "Run a shell command in a Vercel Sandbox (an isolated Linux microVM), not on the user's machine. Use it for work a sandbox is for: a long test/lint/typecheck run in the background while you keep working, someone else's repository (workspace {git, ref}), an installer or script of unknown origin (workspace none; the result lists what it wrote and left running), a bulk change to preview (returns patch: the user decides whether it is applied), a clean build. " +
800          'workspace "current" starts from the project exactly as it is now (tracked files, secrets left out), in a folder of its own. Results are marked [sandbox · job · tree]. A background job reports back as a new message.',
801        inputSchema: {
802          type: 'object',
803          properties: {
804            command: { type: 'string', description: 'The bash command to run' },
805            workspace: {
806              description: '"current" (default): the project as it is now; "none": an empty folder; or { git, ref } to clone an https repository',
807              anyOf: [
808                { type: 'string', enum: ['current', 'none'] },
809                { type: 'object', properties: { git: { type: 'string' }, ref: { type: 'string' } }, required: ['git'] },
810              ],
811            },
812            background: { type: 'boolean', description: 'Return at once; the result arrives as a new message when it ends' },
813            returns: { type: 'string', enum: ['output', 'changes', 'patch'], description: '"patch" keeps a patch of the changes for the user to apply' },
814            timeoutMs: { type: 'number', description: `Longest it may run (default ${commandTimeoutMs}, max 600000)` },
815          },
816          required: ['command'],
817        },
818      },
819      { key: 'jobs' as const, description: 'List the Vercel Sandbox jobs of this session: id, state, what each started from.', inputSchema: { type: 'object', properties: {} } },
820      {
821        key: 'result' as const,
822        description: "A sandbox job's whole result: exit code, files it changed, and its output.",
823        inputSchema: { type: 'object', properties: { job: { type: 'string', description: 'The job id, e.g. j3' } }, required: ['job'] },
824      },
825      {
826        key: 'apply' as const,
827        description: "Apply a sandbox job's patch to the user's project. The user is always asked first; call it only when the task needs the change on their machine.",
828        inputSchema: { type: 'object', properties: { job: { type: 'string', description: 'The job id, e.g. j3' } }, required: ['job'] },
829      },
830    ]
831    toolNames.clear()
832    for (const spec of specs) {
833      await $.tool
834        .register({ name: TOOLS[spec.key], description: spec.description, inputSchema: spec.inputSchema })
835        .then(({ tool }) => void toolNames.set(tool, spec.key))
836        .catch(err => $.ui.log(`${TAG} ${TOOLS[spec.key]} not registered: ${err}`, { to: 'debug' }))
837    }
838
839    // a fresh load of the plugin starts from nothing
840    info = undefined
841    workspace = undefined
842    lastError = undefined
843    starting = undefined
844    steps = []
845    jobs.length = 0
846    nextId = 1
847    detections.clear()
848    always.clear()
849    chosen.clear()
850    tally.detected = tally.sandboxed = tally.local = 0
851
852    // env fallbacks, each read by its literal name (the engine lists what a module reads)
853    const orEmpty = (v: string | undefined) => v ?? ''
854    const token =
855      text('vercelToken') ||
856      orEmpty(await $.env.get('VERCEL_TOKEN').catch(() => undefined)) ||
857      orEmpty(await $.env.get('VERCEL_OIDC_TOKEN').catch(() => undefined))
858    const teamId = text('vercelTeamId') || orEmpty(await $.env.get('VERCEL_TEAM_ID').catch(() => undefined))
859    const projectId = text('vercelProjectId') || orEmpty(await $.env.get('VERCEL_PROJECT_ID').catch(() => undefined))
860    const resolved = resolveCredentials(token, teamId, projectId)
861    now = await $.clock.now()
862    if (resolved.ok) {
863      credentials = resolved.credentials
864      credentialsKind = resolved.kind
865      state = 'idle'
866    } else {
867      credentials = undefined
868      missing = resolved.missing
869      state = 'off'
870    }
871    $.ui.log(
872      credentials
873        ? `${TAG} configured (${credentialsKind}); Jev (${backend}) checks commands for sandbox jobs (${useCases.join(', ') || 'none turned on'}) and ${confirmSandbox ? 'asks you' : 'sends them without asking'}. /${COMMAND} starts the sandbox`
874        : `${TAG} not configured: set ${missing.join(', ')} in pluginConfigs["${$.plugin.name}@skills-dir"] (or "${$.plugin.name}" with --plugin-dir); every command runs locally until then`,
875    )
876    $.ui.status(statusLine())
877    return r
878  })
879
880  on('turn.start', async ($, e, next) => {
881    if (e.text.trim()) intent = e.text
882    now = await $.clock.now()
883    if (isOpen) $.ui.invalidate('ui.render')
884    return next(e)
885  })
886
887  on('tool.call', { tool: 'Bash' }, async ($, e, next) => {
888    const command = typeof e.command === 'string' ? e.command : ''
889    if (!command.trim() || !useCases.length) return next(e)
890    const short = shortCommand(command)
891    // what a job that outlives this hook needs, spelled here where `$` is
892    const host: Host = {
893      fetch: (url, init) => $.http.fetch(url, init),
894      sleep: ms => $.clock.sleep(ms),
895      now: () => $.clock.now(),
896      run: (argv, init) => $.process.run(argv, init),
897      readBase64: async path => (await $.fs.read(path, { as: 'bytes' })).base64,
898      size: async path => (await $.fs.stat(path)).size,
899      list: async path => (await $.fs.list(path)).map(x => x.name),
900      cwd: () => $.session.cwd(),
901      log: line => void $.ui.log(line),
902      toast: line => void $.ui.toast(line),
903      status: line => void $.ui.status(line),
904      submit: line => void $.prompt.submit({ text: line }).catch(err => $.ui.log(`${TAG} could not report back: ${err}`, { to: 'debug' })),
905      redraw: () => void $.ui.invalidate('ui.render'),
906      classify: (t, labels) => $.model.classify(t, labels),
907      ask: (question, labels, header) => $.ui.ask(question, { header, options: labels }),
908    }
909    const cwd = await $.session.cwd().catch(() => '')
910    const found = await detect(host, command, cwd, short)
911    if (!found.useCase) return next(e)
912
913    const kind = found.useCase
914    const profile = PROFILES[kind]
915    const reason = detectionReason(found)
916    tally.detected += 1
917    if (!credentials) {
918      if (logDecisions) $.ui.log(`${TAG} ${short} looks like ${profile.label} (${reason}), but no sandbox is configured: it runs locally`)
919      tally.local += 1
920      return next(e)
921    }
922
923    // a start that failed is not retried on every command without the person
924    if (state === 'failed' && !confirmSandbox) {
925      tally.local += 1
926      if (logDecisions) $.ui.log(`${TAG} local ${short} (${reason}): the sandbox did not start; /${COMMAND} restart tries again`)
927      return next(e)
928    }
929
930    // 1. where it runs: the person says, unless they turned the question off
931    let choice: Choice = 'sandbox'
932    if (confirmSandbox && !always.has(kind)) {
933      const labels = askLabels(kind)
934      const first = state === 'ready' ? '' : state === 'starting' ? ' (it is starting)' : ' (this starts it)'
935      try {
936        const answer = await host.ask(`Jev: \`${short}\` looks like ${profile.label}. Run it in the Vercel Sandbox${first}?`, labels, 'Sandbox')
937        if (answer === labels[2]) always.add(kind)
938        choice = answer === labels[0] || answer === labels[2] || /sandbox/i.test(answer) ? 'sandbox' : 'local'
939      } catch {
940        // dismissed, or a -p run with no one to ask: as if the mod were not here
941        choice = 'local'
942      }
943    }
944    if (choice === 'local') {
945      tally.local += 1
946      if (logDecisions) $.ui.log(`${TAG} local ${short} (${reason}; the user chose this machine)`)
947      $.ui.status(statusLine())
948      return next(e)
949    }
950    chosen.add(e.tool_use_id ?? command)
951
952    // 2. the job
953    const j = newJob({
954      command,
955      source: { kind: profile.workspace },
956      useCase: kind,
957      by: 'jev',
958      background: profile.background || e.run_in_background === true,
959      returns: profile.returns,
960      watch: profile.watch,
961      timeoutMs: clampTimeout(e.timeout),
962    })
963    if (logDecisions) $.ui.log(`${TAG} ${j.id} → sandbox: ${short} (${reason})`)
964    $.ui.invalidate('ui.render')
965    const s = info
966    if (state === 'ready' && s && !j.background) {
967      await runJob(host, client(host.fetch, credentials, apiBase), s, j)
968      const out = resultText(j)
969      return {
970        result: { stdout: out, stderr: '', interrupted: false },
971        context: [jobContext(j, s.name)],
972      }
973    }
974    // background, or the sandbox is not ready yet: it runs on and reports back
975    const why = state === 'ready' ? 'in the background' : state === 'starting' ? 'once the sandbox finishes starting' : 'once the sandbox has started (it is starting now)'
976    j.background = true
977    detach(host, credentials, j)
978    return {
979      result: { stdout: `${header(j)} ${short}: queued, runs ${why}. The result arrives as a new message.`, stderr: '', interrupted: false },
980      context: [
981        `This command did not run on the user's machine. It was sent to the Vercel Sandbox as job ${j.id} (${profile.label}); it runs there ${why} and its result arrives as a new message. Do not run it locally meanwhile; carry on with other work or wait for it.`,
982      ],
983    }
984  }).catch(async ($, e, next) => {
985    // over budget or thrown: a command sent to the sandbox never falls through to the machine
986    const key = typeof e.command === 'string' ? (e.tool_use_id ?? e.command) : ''
987    if (!chosen.has(key)) return next(e)
988    return { deny: `jev-vercel-sandbox could not run this command in the Vercel Sandbox (${next.error.kind}), and it was not run on the user's machine either. Tell the user; do not retry it another way.` }
989  })
990
991  on('tool.call', async ($, e, next) => {
992    const which = toolNames.get(e.tool)
993    if (!which) return next(e)
994    const input = e as unknown as Record<string, unknown>
995    // what a job that outlives this hook needs, spelled here where `$` is
996    const host: Host = {
997      fetch: (url, init) => $.http.fetch(url, init),
998      sleep: ms => $.clock.sleep(ms),
999      now: () => $.clock.now(),
1000      run: (argv, init) => $.process.run(argv, init),
1001      readBase64: async path => (await $.fs.read(path, { as: 'bytes' })).base64,
1002      size: async path => (await $.fs.stat(path)).size,
1003      list: async path => (await $.fs.list(path)).map(x => x.name),
1004      cwd: () => $.session.cwd(),
1005      log: line => void $.ui.log(line),
1006      toast: line => void $.ui.toast(line),
1007      status: line => void $.ui.status(line),
1008      submit: line => void $.prompt.submit({ text: line }).catch(err => $.ui.log(`${TAG} could not report back: ${err}`, { to: 'debug' })),
1009      redraw: () => void $.ui.invalidate('ui.render'),
1010      classify: (t, labels) => $.model.classify(t, labels),
1011      ask: (question, labels, header) => $.ui.ask(question, { header, options: labels }),
1012    }
1013
1014    if (which === 'jobs') return { result: `${statusLine()}\n${jobsText()}` }
1015    if (which === 'result') {
1016      const j = findJob(String(input.job ?? ''))
1017      if (!j) return { deny: `no sandbox job ${String(input.job ?? '')}; sandbox_jobs lists them` }
1018      return { result: resultText(j, KEEP_OUTPUT) }
1019    }
1020    if (which === 'apply') {
1021      const j = findJob(String(input.job ?? ''))
1022      if (!j) return { deny: `no sandbox job ${String(input.job ?? '')}; sandbox_jobs lists them` }
1023      const r = await applyJob(host, j)
1024      return r.ok ? { result: r.text } : { deny: `${r.text}.` }
1025    }
1026
1027    // sandbox_run
1028    const command = typeof input.command === 'string' ? input.command : ''
1029    if (!command.trim()) return { deny: 'sandbox_run needs a command' }
1030    if (!credentials) return { deny: `no Vercel Sandbox is configured (missing ${missing.join(', ')}); the user sets it in the plugin's options` }
1031    let source: Source = { kind: 'current' }
1032    const w = input.workspace
1033    if (w === 'none') source = { kind: 'none' }
1034    else if (w && typeof w === 'object') {
1035      const url = String((w as Record<string, unknown>).git ?? '')
1036      const ref = String((w as Record<string, unknown>).ref ?? '')
1037      if (!isGitUrl(url)) return { deny: `workspace.git must be an https URL, got ${JSON.stringify(url).slice(0, 100)}` }
1038      if (ref && !isGitRef(ref)) return { deny: `workspace.ref is not a branch, tag or commit: ${JSON.stringify(ref).slice(0, 100)}` }
1039      source = { kind: 'git', url, ref }
1040    }
1041    const returns: Returns = input.returns === 'patch' || input.returns === 'changes' ? input.returns : 'output'
1042    // Jev picks the profile: code of unknown origin never gets the project's copy
1043    const found = await detect(host, command, await $.session.cwd().catch(() => ''), shortCommand(command))
1044    let note = ''
1045    if (source.kind === 'current' && found.useCase && PROFILES[found.useCase].workspace === 'none') {
1046      source = { kind: 'none' }
1047      note = `Jev judged it ${PROFILES[found.useCase].label}, so it runs in an empty folder, not on the project's copy`
1048    }
1049    const j = newJob({
1050      command,
1051      source,
1052      useCase: found.useCase,
1053      by: 'claude',
1054      background: input.background === true,
1055      returns,
1056      watch: found.useCase ? PROFILES[found.useCase].watch : false,
1057      timeoutMs: clampTimeout(input.timeoutMs),
1058      note,
1059    })
1060    $.ui.invalidate('ui.render')
1061
1062    if (state !== 'ready' && state !== 'starting') {
1063      if (confirmSandbox) {
1064        let answer = ''
1065        try {
1066          answer = await host.ask(`Claude wants to run \`${j.short}\` in the Vercel Sandbox, which is not running. Start it?`, ['Start sandbox', 'Not now'], 'Sandbox')
1067        } catch {}
1068        if (answer !== 'Start sandbox') {
1069          j.state = 'error'
1070          j.error = 'the user did not start the sandbox'
1071          return { deny: `The Vercel Sandbox is not running and the user did not start it, so the command was not run. They can start it with /${COMMAND}.` }
1072        }
1073      }
1074      if (logDecisions) $.ui.log(`${TAG} starting the sandbox for ${j.id}`)
1075    }
1076    const s = info
1077    if (state === 'ready' && s && !j.background) {
1078      await runJob(host, client(host.fetch, credentials, apiBase), s, j)
1079      return { result: `${resultText(j)}${note ? `\n(${note})` : ''}`, context: [jobContext(j, s.name)] }
1080    }
1081    const queued = state === 'ready' ? ', running in the background' : ', runs once the sandbox has started'
1082    j.background = true
1083    detach(host, credentials, j)
1084    return {
1085      result: `${header(j)} ${j.short}: queued${queued}. The result arrives as a new message.${note ? ` (${note})` : ''}`,
1086    }
1087  })
1088
1089  on('session.end', async ($, e, next) => {
1090    if (credentials && info && (state === 'ready' || state === 'starting') && flag('stopOnExit', true)) {
1091      const api = client((url, init) => $.http.fetch(url, init), credentials, apiBase)
1092      const budget = Math.max(0, Math.min(2_000, next.budget.remainingMs - 500))
1093      await Promise.race([api.stop(info).catch(() => undefined), $.clock.sleep(budget)]).catch(() => undefined)
1094      state = 'stopped'
1095    }
1096    return next(e)
1097  })
1098
1099  on('command.run', { command: COMMAND }, async ($, e) => {
1100    const args = e.args.trim()
1101    const sub = (args.split(/\s+/)[0] ?? '').toLowerCase()
1102    const rest = args.slice(sub.length).trim()
1103    now = await $.clock.now()
1104    // what a job that outlives this hook needs, spelled here where `$` is
1105    const host: Host = {
1106      fetch: (url, init) => $.http.fetch(url, init),
1107      sleep: ms => $.clock.sleep(ms),
1108      now: () => $.clock.now(),
1109      run: (argv, init) => $.process.run(argv, init),
1110      readBase64: async path => (await $.fs.read(path, { as: 'bytes' })).base64,
1111      size: async path => (await $.fs.stat(path)).size,
1112      list: async path => (await $.fs.list(path)).map(x => x.name),
1113      cwd: () => $.session.cwd(),
1114      log: line => void $.ui.log(line),
1115      toast: line => void $.ui.toast(line),
1116      status: line => void $.ui.status(line),
1117      submit: line => void $.prompt.submit({ text: line }).catch(err => $.ui.log(`${TAG} could not report back: ${err}`, { to: 'debug' })),
1118      redraw: () => void $.ui.invalidate('ui.render'),
1119      classify: (t, labels) => $.model.classify(t, labels),
1120      ask: (question, labels, header) => $.ui.ask(question, { header, options: labels }),
1121    }
1122    const openPane = async () => {
1123      isOpen = true
1124      await $.ui.open({ id: PANE, title: 'sandbox', focus: true, columns }).catch(err => {
1125        isOpen = false
1126        $.ui.log(`${TAG} pane not opened: ${err}`, { to: 'debug' })
1127      })
1128      host.redraw()
1129    }
1130    const launch = (creds: Credentials) => {
1131      // runs on after this command answers: the pane shows the progress
1132      starting = boot(host, creds)
1133      void starting
1134    }
1135
1136    if (sub === 'jobs') return { text: `jev-vercel-sandbox: ${statusLine()}\n${jobsText()}` }
1137    if (sub === 'result') {
1138      const j = findJob(rest) ?? (rest ? undefined : jobs[jobs.length - 1])
1139      return { text: j ? resultText(j, KEEP_OUTPUT) : `jev-vercel-sandbox: no job ${rest}` }
1140    }
1141    if (sub === 'apply') {
1142      const j = findJob(rest) ?? (rest ? undefined : [...jobs].reverse().find(x => x.report?.patchBytes && !x.applied))
1143      if (!j) return { text: `jev-vercel-sandbox: no job ${rest || 'with a patch'}` }
1144      const r = await applyJob(host, j)
1145      return {
1146        text: `jev-vercel-sandbox: ${r.text}`,
1147        context: r.ok ? [`The user applied sandbox job ${j.id}'s patch to their project: ${describeChanges(j.report!.changes)}. These files changed on their machine.`] : undefined,
1148      }
1149    }
1150    if (sub === 'run') {
1151      if (!rest) return { text: `jev-vercel-sandbox: /${COMMAND} run <command>` }
1152      if (!credentials) return { text: `jev-vercel-sandbox: not configured (missing ${missing.join(', ')})` }
1153      const j = newJob({ command: rest, source: { kind: 'current' }, useCase: null, by: 'user', background: true, returns: 'patch', watch: false, timeoutMs: commandTimeoutMs })
1154      await openPane()
1155      detach(host, credentials, j)
1156      return { text: `jev-vercel-sandbox: ${j.id} ${state === 'ready' ? 'running' : 'queued until the sandbox has started'}; the pane follows it` }
1157    }
1158    if (sub === 'stop') {
1159      if (credentials && info && (state === 'ready' || state === 'starting')) {
1160        await client((url, init) => $.http.fetch(url, init), credentials, apiBase)
1161          .stop(info)
1162          .catch(err => $.ui.log(`${TAG} stop failed: ${err}`))
1163      }
1164      state = credentials ? 'stopped' : 'off'
1165      $.ui.status(statusLine())
1166      host.redraw()
1167      return { text: `jev-vercel-sandbox: stopped; detected jobs ${confirmSandbox ? 'ask to start a new one' : 'start a new one'}, or /${COMMAND}` }
1168    }
1169    if (sub === 'restart') {
1170      if (!credentials) {
1171        await openPane()
1172        return { text: `jev-vercel-sandbox: not configured (missing ${missing.join(', ')})` }
1173      }
1174      if (state === 'starting') {
1175        await openPane()
1176        return { text: 'jev-vercel-sandbox: already starting; the pane shows the progress' }
1177      }
1178      if (info && state === 'ready') await client((url, init) => $.http.fetch(url, init), credentials, apiBase).stop(info).catch(() => undefined)
1179      await openPane()
1180      launch(credentials)
1181      return { text: 'jev-vercel-sandbox: starting a new sandbox; the pane shows the progress' }
1182    }
1183    if (sub === '' || sub === 'start' || sub === 'open') {
1184      await openPane()
1185      if (!credentials) return { text: `jev-vercel-sandbox: not configured, set ${missing.join(', ')} in the plugin's options` }
1186      if (sub !== 'open' && state !== 'ready' && state !== 'starting') {
1187        launch(credentials)
1188        return { text: `jev-vercel-sandbox: starting the Vercel Sandbox${uploadWorkspace ? ' and copying the project into it' : ''}; the pane shows the progress` }
1189      }
1190      if (sub === 'open') return {}
1191    }
1192    const left = minutesLeft()
1193    return {
1194      text: [
1195        `jev-vercel-sandbox: ${statusLine()}`,
1196        credentials
1197          ? `sandbox: ${info ? `${info.name} · ${info.status} · ${info.region} · ${info.vcpus} vCPU · ${info.memory} MB${left !== null ? ` · stops in ${left}m` : ''}` : 'not started'} (${credentialsKind})`
1198          : `not configured: missing ${missing.join(', ')}`,
1199        workspace ? `project: ${plural(workspace.files, 'file')} · ${megabytes(workspace.bytes)} in ${workspace.dir}${workspace.synced ? ` · tree ${shortTree(workspace.synced)}` : ''}` : '',
1200        `jev: ${backend} · ${useCases.join(', ') || 'no use cases'} · threshold ${threshold.toFixed(2)} · ${confirmSandbox ? 'asks first' : 'no questions'}${always.size ? ` · always: ${[...always].join(', ')}` : ''}`,
hooks/judge.ts 257 lines
1/**
2 * jev-vercel-sandbox — the detector: TypeSafe's Jev asked whether one Bash
3 * command is one of the jobs a sandbox is for.
4 *
5 * Whether a command may run on the person's machine at all is not asked here:
6 * auto mode and jev-guardrails decide that. This asks only whether the command
7 * is better done in a Vercel Sandbox (a long test run, someone else's repo, an
8 * unknown installer, a bulk change to preview, a clean build), and each answer
9 * comes with the profile the job runs under.
10 *
11 * No `$` and no I/O here. The wire shapes are jev-auto-mode's (TypeSafe's
12 * System One API and the Vercel AI Gateway's evaluation-model endpoint, both
13 * with `ai-gateway-protocol-version`); the battery is this mod's own.
14 */
15
16export type Provider = 'typesafe' | 'gateway'
17
18export const DEFAULT_BASE_URL: Record<Provider, string> = {
19  typesafe: 'https://api.typesafe.ai',
20  gateway: 'https://ai-gateway.vercel.sh/v4/ai',
21}
22
23export const DEFAULT_MODEL: Record<Provider, string> = {
24  typesafe: 'jev-latest',
25  gateway: 'typesafe-ai/jev',
26}
27
28/** `@ai-sdk/gateway`'s AI_GATEWAY_PROTOCOL_VERSION. */
29const AI_GATEWAY_PROTOCOL_VERSION = '0.0.1'
30
31export function selectProvider(forced: string, typesafeKey: string, gatewayKey: string): Provider | null {
32  if (forced === 'builtin') return null
33  if (forced === 'typesafe') return typesafeKey ? 'typesafe' : null
34  if (forced === 'gateway') return gatewayKey ? 'gateway' : null
35  if (typesafeKey) return 'typesafe'
36  if (gatewayKey) return 'gateway'
37  return null
38}
39
40export function endpoint(provider: Provider, baseUrl: string): string {
41  const root = baseUrl.replace(/\/+$/, '')
42  return provider === 'typesafe' ? `${root}/v1/systemone` : `${root}/evaluation-model`
43}
44
45export type UseCase = 'tests' | 'external_repo' | 'untrusted_code' | 'repo_change' | 'clean_build'
46
47/** Which copy of the project a job starts from. */
48export type WorkspaceKind = 'current' | 'none'
49/** What a job hands back: its output, or also the files it changed and a patch of them. */
50export type Returns = 'output' | 'changes' | 'patch'
51
52export type Profile = {
53  /** How the person reads it in the question and the pane. */
54  label: string
55  workspace: WorkspaceKind
56  background: boolean
57  returns: Returns
58  /** Also list the files it wrote outside the job's folder and the processes it left. */
59  watch: boolean
60}
61
62export const PROFILES: Record<UseCase, Profile> = {
63  tests: { label: 'a long test, lint or typecheck run', workspace: 'current', background: true, returns: 'output', watch: false },
64  external_repo: { label: "someone else's repository to clone and run", workspace: 'none', background: false, returns: 'output', watch: false },
65  untrusted_code: { label: 'an installer or script of unknown origin', workspace: 'none', background: false, returns: 'output', watch: true },
66  repo_change: { label: 'a bulk change to the project, to preview first', workspace: 'current', background: false, returns: 'patch', watch: false },
67  clean_build: { label: 'a build from a clean checkout', workspace: 'current', background: true, returns: 'output', watch: false },
68}
69
70export const USE_CASES = Object.keys(PROFILES) as UseCase[]
71
72type Question = { instructions: string; yes: string; no: string }
73
74/** One yes/no question per use case: is this command that kind of job? */
75export const BATTERY: Record<UseCase, Question> = {
76  tests: {
77    instructions:
78      "Is this shell command a project's full test suite, lint or typecheck run that takes a while and whose result the agent can wait for (for example npm test, pytest, go test ./..., cargo test, tsc --noEmit, eslint .)? A single quick test file does not count.",
79    yes: 'It is a long check the agent could run elsewhere and keep working.',
80    no: 'It is not a test, lint or typecheck run, or it is a quick one.',
81  },
82  external_repo: {
83    instructions:
84      'Does this shell command clone, download or run a repository or project that is not the one the agent is working in (for example git clone of another repo followed by its install or tests)?',
85    yes: "It works on someone else's code.",
86    no: 'It works only on the current project.',
87  },
88  untrusted_code: {
89    instructions:
90      'Does this shell command download and run code whose origin is not already part of the project: a script piped to a shell, a package or tool installed or run for the first time, a downloaded binary (for example curl ... | sh, npx some-new-package, pip install from a URL)?',
91    yes: 'It runs code that has not been reviewed.',
92    no: 'It runs only code already in the project or on the machine.',
93  },
94  repo_change: {
95    instructions:
96      'Does this shell command change many of the project\'s files at once in a way worth previewing before it touches them: deleting folders, a codemod, a mass rename or search-and-replace, regenerating files, a dependency upgrade that rewrites lockfiles (for example rm -rf src/legacy, npx jscodeshift, sed -i across many files, npm update)?',
97    yes: 'It rewrites or removes many project files at once.',
98    no: 'It reads, or changes one or two files.',
99  },
100  clean_build: {
101    instructions:
102      'Is this shell command a full build or install of the project from scratch, where a clean machine would tell whether it builds without local caches (for example npm ci && npm run build, rm -rf node_modules && npm install, a docker-free release build)?',
103    yes: 'It is a from-scratch build or install.',
104    no: 'It is not a full build or install.',
105  },
106}
107
108const SAFE_PROGRAMS = new Set([
109  'ls', 'pwd', 'cat', 'head', 'tail', 'wc', 'echo', 'printf', 'which', 'type', 'file', 'stat', 'du', 'df',
110  'grep', 'egrep', 'rg', 'tree', 'date', 'whoami', 'uname', 'basename', 'dirname', 'realpath', 'sort', 'uniq', 'diff', 'true',
111])
112const SAFE_GIT = new Set(['status', 'log', 'diff', 'show', 'rev-parse', 'blame', 'ls-files', 'describe'])
113// `git branch -D x` and `git remote remove origin` write: these two are reads only with listing flags
114const LISTING_GIT = new Set(['branch', 'remote'])
115const LISTING_FLAGS = new Set(['-a', '-r', '-v', '-vv', '--all', '--remotes', '--list', '--verbose', '--show-current'])
116
117/**
118 * A command that plainly only reads: one program from a short list, no shell
119 * operators, redirects or substitutions. None of the use cases can be one of
120 * these, so they stay local without a call to the detector.
121 */
122export function isPlainRead(command: string): boolean {
123  const c = command.trim()
124  if (!c || /[;&|<>`$(){}\n\\]/.test(c)) return false
125  const words = c.split(/\s+/)
126  const program = words[0]!
127  if (program === 'git') {
128    const sub = words[1] ?? ''
129    if (LISTING_GIT.has(sub)) return words.slice(2).every(w => LISTING_FLAGS.has(w))
130    return SAFE_GIT.has(sub) && !words.some(w => /^--(output|exec)/.test(w))
131  }
132  if (program === 'find') return !words.some(w => /^-(delete|exec|execdir|ok|okdir|fprint|fls)/.test(w))
133  return SAFE_PROGRAMS.has(program)
134}
135
136/** What the detector reads: the user's request, then the command, whole. */
137export function stateText(intent: string, command: string, cwd: string): string {
138  return [
139    "The user's latest request to an AI coding agent:",
140    intent.trim() ? intent.trim().slice(0, 2000) : '(none recorded)',
141    '',
142    `The agent is about to run this shell command${cwd ? ` in ${cwd}` : ''}:`,
143    command,
144  ].join('\n')
145}
146
147function yesNo(provider: Provider, q: Question): Record<string, unknown> {
148  if (provider === 'typesafe') return { type: 'noul', instructions: q.instructions, criteria: { true: q.yes, false: q.no } }
149  return { type: 'boolean', instructions: `${q.instructions} Yes: ${q.yes} No: ${q.no}` }
150}
151
152/** The battery for the use cases turned on, one question each. */
153export function requestBody(provider: Provider, state: string, model: string, cases: readonly UseCase[] = USE_CASES): string {
154  const questions: Record<string, unknown> = {}
155  for (const c of cases) questions[c] = yesNo(provider, BATTERY[c])
156  return JSON.stringify(provider === 'typesafe' ? { model, state, questions } : { state, questions })
157}
158
159export function requestHeaders(provider: Provider, apiKey: string, model: string): Record<string, string> {
160  const common = { 'content-type': 'application/json', authorization: `Bearer ${apiKey}` }
161  if (provider === 'typesafe') return common
162  return {
163    ...common,
164    'ai-gateway-auth-method': 'api-key',
165    'ai-model-id': model,
166    'ai-gateway-protocol-version': AI_GATEWAY_PROTOCOL_VERSION,
167    'ai-evaluation-model-specification-version': '4',
168  }
169}
170
171export type Judgement = Partial<Record<UseCase, number>>
172
173/** Reads either backend's answer; one with any asked use case unanswered reads as none. */
174export function readJudgement(text: string, cases: readonly UseCase[] = USE_CASES): Judgement | null {
175  let parsed: unknown
176  try {
177    parsed = JSON.parse(text)
178  } catch {
179    return null
180  }
181  if (parsed === null || typeof parsed !== 'object') return null
182  const answers = (parsed as { answers?: Record<string, Record<string, unknown>> }).answers
183  if (!answers || typeof answers !== 'object') return null
184  const out: Judgement = {}
185  for (const c of cases) {
186    const a = answers[c]
187    const p = typeof a?.noul === 'number' ? a.noul : typeof a?.probability === 'number' ? a.probability : null
188    if (p === null) return null
189    out[c] = p
190  }
191  return out
192}
193
194export type Detection = {
195  /** The use case found, or null when the command is none of them. */
196  useCase: UseCase | null
197  probability: number | null
198  by: 'read-only' | 'jev' | 'built-in' | 'no answer'
199}
200
201/** The likeliest use case when it reaches `threshold`, else none. */
202export function decide(j: Judgement, threshold: number): Detection {
203  let top: UseCase | null = null
204  for (const c of Object.keys(j) as UseCase[]) if (top === null || j[c]! > j[top]!) top = c
205  const p = top === null ? null : j[top]!
206  if (top !== null && p !== null && p >= threshold) return { useCase: top, probability: p, by: 'jev' }
207  return { useCase: null, probability: p, by: 'jev' }
208}
209
210/** The labels the engine's small classifier picks from when no Jev key is set. */
211export function builtinLabels(cases: readonly UseCase[] = USE_CASES): string[] {
212  return ['none', ...cases]
213}
214
215/** The rubric the engine's small classifier reads when no Jev key is set. */
216export function classifyText(state: string, cases: readonly UseCase[] = USE_CASES): string {
217  return [
218    'You decide whether a shell command an AI coding agent is about to run is one of the jobs a remote sandbox is for. Answer with the job\'s label, or "none".',
219    ...cases.map(c => `- ${c}: ${BATTERY[c].instructions}`),
220    'Ordinary development work (reading, editing a file, a quick single test, git commits, starting the dev server) is "none".',
221    '',
222    state,
223  ].join('\n')
224}
225
226/** The use cases a comma-separated option turns on; empty or unknown names turn on none of their own. */
227export function parseUseCases(value: string): UseCase[] {
228  if (!value.trim()) return [...USE_CASES]
229  const wanted = new Set(value.split(',').map(s => s.trim().toLowerCase().replace(/-/g, '_')))
230  return USE_CASES.filter(c => wanted.has(c))
231}
232
233export function describeJudgement(j: Judgement | null, ms: number): string {
234  const took = ` · ${Math.round(ms)}ms`
235  if (!j) return `no answer${took}`
236  return (
237    (Object.entries(j) as [UseCase, number][])
238      .sort((a, b) => b[1] - a[1])
239      .map(([c, p]) => `${c} ${p.toFixed(2)}`)
240      .join(' · ') + took
241  )
242}
243
244export function detectionReason(d: Detection): string {
245  if (d.by === 'read-only') return 'plain read'
246  if (d.by === 'no answer') return 'the detector gave no answer'
247  if (!d.useCase) return d.by === 'built-in' ? 'built-in classifier: none' : `none${d.probability === null ? '' : ` (top ${d.probability.toFixed(2)})`}`
248  const p = d.probability === null ? '' : ` ${d.probability.toFixed(2)}`
249  return `${d.useCase.replace(/_/g, ' ')}${p}${d.by === 'built-in' ? ' (built-in classifier)' : ''}`
250}
251
252/** One line for the transcript, the pane and the status: the command, cut to fit. */
253export function shortCommand(command: string, max = 60): string {
254  const one = command.replace(/\s+/g, ' ').trim()
255  return one.length > max ? `${one.slice(0, max - 1)}…` : one
256}
257
hooks/vercel.ts 288 lines
1/**
2 * jev-vercel-sandbox — the Vercel Sandbox REST client.
3 *
4 * A mod runs without node_modules, so `@vercel/sandbox` is not available: this
5 * speaks its HTTP API directly, through a `fetch` the hooks module hands in
6 * (`(url, init) => $.http.fetch(url, init)`), so nothing here touches `$`.
7 *
8 * The endpoints are Vercel's REST API reference (vercel.com/docs/rest-api/sandboxes),
9 * called the way `@vercel/sandbox` 3.5.0 calls them (dist/api-client/api-client.js):
10 *
11 *   POST /v3/sandboxes?teamId=…                       create → { sandbox, session, routes }
12 *   GET  /v2/sandboxes/sessions/{id}?teamId=…         → { session, routes }
13 *   POST /v2/sandboxes/sessions/{id}/cmd?teamId=…     { command, args, cwd, env, sudo, wait: true, logs: true, timeout }
14 *                                                     → application/x-ndjson: { command } · { stream, data }… · { command: { exitCode } }
15 *   POST /v2/sandboxes/sessions/{id}/stop?teamId=…    → { session }
16 *
17 * Base URL https://api.vercel.com (the SDK uses the equivalent
18 * https://vercel.com/api), `Authorization: Bearer <token>`. An error body is
19 * `{ error: { message } }`. `persistent` defaults to true on the server, so it
20 * is always sent.
21 */
22import type { HttpInit, HttpResponse } from 'claude-code'
23
24export type Fetch = (url: string, init: HttpInit) => Promise<HttpResponse>
25
26export const DEFAULT_API_URL = 'https://api.vercel.com'
27
28export type Credentials = { token: string; teamId: string; projectId: string }
29
30export type SessionStatus = 'pending' | 'running' | 'stopping' | 'stopped' | 'failed' | 'aborted' | 'snapshotting'
31
32export type SandboxInfo = {
33  /** The sandbox's name (the SDK's `sandbox.name`). */
34  name: string
35  /** The current session: commands and stop go to it. */
36  sessionId: string
37  status: SessionStatus
38  region: string
39  vcpus: number
40  memory: number
41  /** The session's timeout, in ms. */
42  timeout: number
43  cwd: string
44  image?: string
45  /** Epoch ms the session started (or was requested, while pending). */
46  startedAt: number
47}
48
49export type CreateOptions = {
50  timeoutMs: number
51  image?: string
52  vcpus?: number
53  persistent: boolean
54}
55
56export type RunResult = {
57  exitCode: number | null
58  stdout: string
59  stderr: string
60  durationMs?: number
61  /** A `{ stream: "error" }` line, or a stream that ended before the command did. */
62  error?: string
63}
64
65/** base64url → text, without `atob` (the mod runtime's lib is es2023 only). */
66function decodeBase64Url(s: string): string {
67  const alphabet = 'ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789-_'
68  let bits = 0
69  let value = 0
70  const bytes: number[] = []
71  for (const ch of s.replace(/=+$/, '').replace(/\+/g, '-').replace(/\//g, '_')) {
72    const i = alphabet.indexOf(ch)
73    if (i < 0) throw new Error('not base64url')
74    value = (value << 6) | i
75    bits += 6
76    if (bits >= 8) {
77      bits -= 8
78      bytes.push((value >> bits) & 0xff)
79    }
80  }
81  // UTF-8 without TextDecoder (not in the es2023 lib either)
82  return decodeURIComponent(bytes.map(b => `%${b.toString(16).padStart(2, '0')}`).join(''))
83}
84
85/**
86 * The claims a Vercel OIDC token carries (`owner_id` is the team,
87 * `project_id` the project), as the SDK's `decodeUnverifiedToken` reads them.
88 * Null for an access token, which is not a JWT.
89 */
90export function oidcClaims(token: string): { teamId: string; projectId?: string } | null {
91  const parts = token.split('.')
92  if (parts.length !== 3) return null
93  try {
94    const payload = JSON.parse(decodeBase64Url(parts[1]!)) as { owner_id?: unknown; project_id?: unknown }
95    if (typeof payload.owner_id !== 'string') return null
96    return { teamId: payload.owner_id, projectId: typeof payload.project_id === 'string' ? payload.project_id : undefined }
97  } catch {
98    return null
99  }
100}
101
102/**
103 * Credentials from the options (or the env fallbacks): an access token needs
104 * the team and project ids beside it; an OIDC token carries both itself.
105 */
106export function resolveCredentials(
107  token: string,
108  teamId: string,
109  projectId: string,
110): { ok: true; credentials: Credentials; kind: 'access token' | 'oidc token' } | { ok: false; missing: string[] } {
111  const t = token.trim()
112  const claims = t ? oidcClaims(t) : null
113  const team = teamId.trim() || claims?.teamId || ''
114  const project = projectId.trim() || claims?.projectId || ''
115  const missing = [!t && 'vercelToken', !team && 'vercelTeamId', !project && 'vercelProjectId'].filter((m): m is string => !!m)
116  if (missing.length) return { ok: false, missing }
117  return { ok: true, credentials: { token: t, teamId: team, projectId: project }, kind: claims ? 'oidc token' : 'access token' }
118}
119
120export function apiUrl(base: string, path: string, teamId: string): string {
121  return `${base.replace(/\/+$/, '')}${path}?teamId=${encodeURIComponent(teamId)}`
122}
123
124export function headers(c: Credentials): Record<string, string> {
125  return { authorization: `Bearer ${c.token}`, 'content-type': 'application/json', 'user-agent': 'jev-vercel-sandbox (claude-code mod)' }
126}
127
128export function createBody(c: Credentials, o: CreateOptions): string {
129  const body: Record<string, unknown> = {
130    projectId: c.projectId,
131    ports: [],
132    timeout: o.timeoutMs,
133    persistent: o.persistent,
134  }
135  if (o.image) body.image = o.image
136  if (o.vcpus) body.resources = { vcpus: o.vcpus }
137  return JSON.stringify(body)
138}
139
140/** The Bash tool's command, run by bash in the sandbox with no local env forwarded. */
141export function commandBody(command: string, timeoutMs: number, cwd?: string): string {
142  return execBody('bash', ['-c', command], timeoutMs, cwd)
143}
144
145/** Any program by its argv (the mod's own upload and bookkeeping scripts). */
146export function execBody(command: string, args: readonly string[], timeoutMs: number, cwd?: string): string {
147  return JSON.stringify({
148    command,
149    args,
150    cwd: cwd || undefined,
151    env: {},
152    sudo: false,
153    wait: true,
154    logs: true,
155    timeout: timeoutMs,
156  })
157}
158
159/** The server's `{ error: { message } }`, else the status and the start of the body. */
160export function errorOf(r: HttpResponse): string {
161  try {
162    const m = (JSON.parse(r.text) as { error?: { message?: unknown } }).error?.message
163    if (typeof m === 'string' && m) return `${r.status}: ${m}`
164  } catch {}
165  return `${r.status}${r.text ? `: ${r.text.slice(0, 200)}` : ''}`
166}
167
168function num(v: unknown, fallback = 0): number {
169  return typeof v === 'number' && Number.isFinite(v) ? v : fallback
170}
171
172function readSessionObject(session: Record<string, unknown>, sandbox?: Record<string, unknown>): SandboxInfo | null {
173  if (typeof session.id !== 'string' || typeof session.status !== 'string') return null
174  return {
175    name: typeof sandbox?.name === 'string' ? sandbox.name : session.id,
176    sessionId: session.id,
177    status: session.status as SessionStatus,
178    region: typeof session.region === 'string' ? session.region : '',
179    vcpus: num(session.vcpus),
180    memory: num(session.memory),
181    timeout: num(session.timeout),
182    cwd: typeof session.cwd === 'string' ? session.cwd : '',
183    image: typeof sandbox?.image === 'string' ? sandbox.image : undefined,
184    startedAt: num(session.startedAt, num(session.requestedAt, num(session.createdAt))),
185  }
186}
187
188/** `{ sandbox, session, routes }` (create) or `{ session, routes }` (get). */
189export function readSandbox(text: string, previousName?: string): SandboxInfo | null {
190  let parsed: unknown
191  try {
192    parsed = JSON.parse(text)
193  } catch {
194    return null
195  }
196  if (!parsed || typeof parsed !== 'object') return null
197  const { session, sandbox } = parsed as { session?: unknown; sandbox?: unknown }
198  if (!session || typeof session !== 'object') return null
199  const info = readSessionObject(session as Record<string, unknown>, sandbox && typeof sandbox === 'object' ? (sandbox as Record<string, unknown>) : undefined)
200  if (info && !sandbox && previousName) info.name = previousName
201  return info
202}
203
204/**
205 * The command endpoint's ndjson, read whole: the first line names the
206 * command, log lines follow, and a line with `command.exitCode` ends it.
207 */
208export function readCommandStream(text: string): RunResult {
209  const out: RunResult = { exitCode: null, stdout: '', stderr: '' }
210  let finished = false
211  for (const line of text.split('\n')) {
212    if (!line.trim()) continue
213    let chunk: Record<string, unknown>
214    try {
215      chunk = JSON.parse(line) as Record<string, unknown>
216    } catch {
217      continue
218    }
219    const command = chunk.command as { exitCode?: unknown; durationMs?: unknown } | undefined
220    if (command && typeof command === 'object') {
221      if (typeof command.exitCode === 'number') {
222        out.exitCode = command.exitCode
223        if (typeof command.durationMs === 'number') out.durationMs = command.durationMs
224        finished = true
225      }
226      continue
227    }
228    if (chunk.stream === 'stdout' && typeof chunk.data === 'string') out.stdout += chunk.data
229    else if (chunk.stream === 'stderr' && typeof chunk.data === 'string') out.stderr += chunk.data
230    else if (chunk.stream === 'error') {
231      const d = chunk.data as { code?: unknown; message?: unknown } | undefined
232      out.error = `${typeof d?.code === 'string' ? d.code : 'error'}: ${typeof d?.message === 'string' ? d.message : 'the sandbox reported an error'}`
233    }
234  }
235  if (!finished && !out.error) out.error = 'the stream ended before the command finished'
236  return out
237}
238
239export type Client = {
240  create(o: CreateOptions): Promise<SandboxInfo>
241  get(s: SandboxInfo): Promise<SandboxInfo>
242  /** The Bash tool's command, by bash, in `cwd` (the sandbox's own when absent). */
243  run(s: SandboxInfo, command: string, timeoutMs: number, cwd?: string): Promise<RunResult>
244  /** A bash script of the mod's own with positional arguments (`bash -c script _ …args`). */
245  script(s: SandboxInfo, script: string, args: readonly string[], timeoutMs: number): Promise<RunResult>
246  stop(s: SandboxInfo): Promise<void>
247}
248
249/** The calls this mod makes, over the fetch it is handed. Each throws with the API's message on a non-2xx. */
250export function client(fetch: Fetch, c: Credentials, base = DEFAULT_API_URL): Client {
251  const call = async (path: string, init: HttpInit) => {
252    const r = await fetch(apiUrl(base, path, c.teamId), { ...init, headers: headers(c) })
253    if (!r.ok) throw new Error(`Vercel Sandbox ${errorOf(r)}`)
254    return r
255  }
256  return {
257    async create(o) {
258      const r = await call('/v3/sandboxes', { method: 'POST', body: createBody(c, o) })
259      const info = readSandbox(r.text)
260      if (!info) throw new Error('Vercel Sandbox: unreadable create response')
261      return info
262    },
263    async get(s) {
264      const r = await call(`/v2/sandboxes/sessions/${encodeURIComponent(s.sessionId)}`, { method: 'GET' })
265      const info = readSandbox(r.text, s.name)
266      if (!info) throw new Error('Vercel Sandbox: unreadable session response')
267      return { ...info, image: info.image ?? s.image }
268    },
269    async run(s, command, timeoutMs, cwd) {
270      const r = await call(`/v2/sandboxes/sessions/${encodeURIComponent(s.sessionId)}/cmd`, {
271        method: 'POST',
272        body: commandBody(command, timeoutMs, cwd || s.cwd),
273      })
274      return readCommandStream(r.text)
275    },
276    async script(s, script, args, timeoutMs) {
277      const r = await call(`/v2/sandboxes/sessions/${encodeURIComponent(s.sessionId)}/cmd`, {
278        method: 'POST',
279        body: execBody('bash', ['-c', script, 'jev', ...args], timeoutMs, s.cwd),
280      })
281      return readCommandStream(r.text)
282    },
283    async stop(s) {
284      await call(`/v2/sandboxes/sessions/${encodeURIComponent(s.sessionId)}/stop`, { method: 'POST' })
285    },
286  }
287}
288
hooks/workspace.ts 292 lines
1/**
2 * jev-vercel-sandbox — copying the project into the sandbox.
3 *
4 * `$.http.fetch` sends text only, so the SDK's `fs/write` (a gzip tar body) is
5 * out of reach. Instead the hooks module packs the project with the local
6 * `tar`, reads the archive back as base64, and this file's scripts rebuild it
7 * inside the sandbox through the command endpoint: the base64 is appended to a
8 * file a few arguments at a time, then decoded and unpacked, and a git
9 * baseline is committed. Before each job the project's current state is
10 * synced as a patch (`git diff --binary` between two trees the hooks module
11 * builds with a throwaway index), and the job runs in its own worktree, so a
12 * job always starts from the files as they are and two jobs never share a folder.
13 *
14 * Pure helpers only: nothing here touches `$`.
15 */
16
17/** Where the base64 is assembled inside the sandbox. */
18export const UPLOAD_FILE = '/tmp/jev-workspace.b64'
19
20/** One argument's size: well under Linux's 128 KiB per-argument limit. */
21export const ARG_BYTES = 100_000
22/** Arguments per command call, so one request body stays near 400 KB. */
23export const ARGS_PER_CALL = 4
24/**
25 * `$.fs.read` rejects files over 4 MiB, so a bigger archive is `split` into
26 * parts of this size. A multiple of 3, so every part but the last encodes to
27 * base64 without padding and the parts' base64 concatenates into one stream.
28 */
29export const PART_BYTES = 3_000_000
30
31/** Files that hold credentials: never copied, whatever git says about them. */
32const SECRET_NAMES = [
33  /^\.env$/,
34  /^\.env\.(?!example$|sample$|template$|dist$)[^/]+$/,
35  /^\.dev\.vars$/,
36  /^\.npmrc$/,
37  /^\.pypirc$/,
38  /^\.netrc$/,
39  /^\.git-credentials$/,
40  /^\.pgpass$/,
41  /\.tfstate(\.backup)?$/,
42  /\.(pem|key|p12|pfx|jks|keystore|ppk)$/i,
43  /^id_(rsa|dsa|ecdsa|ed25519)(\.pub)?$/,
44  /^credentials(\.json)?$/,
45]
46const SECRET_DIRS = ['.git', '.ssh', '.aws', '.gnupg', '.kube', '.docker', '.terraform', 'node_modules', '.vercel', '.claude']
47
48/** True for a path the upload leaves out: a credential file, or one inside a folder that never goes. */
49export function isExcluded(path: string): boolean {
50  const parts = path.split('/')
51  const name = parts[parts.length - 1] ?? ''
52  if (parts.slice(0, -1).some(dir => SECRET_DIRS.includes(dir))) return true
53  return SECRET_NAMES.some(re => re.test(name))
54}
55
56/** `-z` / `-print0` output → paths, `./` dropped. */
57export function nulList(text: string): string[] {
58  return text
59    .split('\0')
60    .map(p => p.replace(/^\.\//, ''))
61    .filter(Boolean)
62}
63
64/** What goes up: listed files, minus the deleted ones git still lists, minus the excluded. */
65export function filesToUpload(listed: string[], deleted: string[] = []): { files: string[]; excluded: string[] } {
66  const gone = new Set(deleted)
67  const files: string[] = []
68  const excluded: string[] = []
69  for (const p of new Set(listed)) {
70    if (gone.has(p)) continue
71    if (isExcluded(p)) excluded.push(p)
72    else files.push(p)
73  }
74  return { files: files.sort(), excluded: excluded.sort() }
75}
76
77/** The base64 cut into argument lists, one list per command call. */
78export function uploadBatches(base64: string, argBytes = ARG_BYTES, perCall = ARGS_PER_CALL): string[][] {
79  const args: string[] = []
80  for (let i = 0; i < base64.length; i += argBytes) args.push(base64.slice(i, i + argBytes))
81  const batches: string[][] = []
82  for (let i = 0; i < args.length; i += perCall) batches.push(args.slice(i, i + perCall))
83  return batches
84}
85
86/** bash -c script, `$1` the file: empties it. */
87export const TRUNCATE_SCRIPT = ': > "$1"'
88/** bash -c script, `$1` the file, the rest base64 pieces: appends them in order. */
89export const APPEND_SCRIPT = 'f="$1"; shift; printf %s "$@" >> "$f"'
90/** Git inside the sandbox: the baseline, every sync and every job's worktree need it. */
91const ENSURE_GIT = [
92  'if ! command -v git >/dev/null 2>&1; then',
93  '  (sudo -n dnf install -y -q git || sudo -n apt-get install -y -q git) >/dev/null 2>&1 || true',
94  'fi',
95]
96const IDENTITY = '-c user.name=jev-vercel-sandbox -c user.email=jev@sandbox.invalid'
97/**
98 * bash -c script, `$1` the base64 file, `$2` the folder: unpacks the project
99 * there and commits a baseline when git exists. Prints `git` or `no-git` last.
100 */
101export const EXTRACT_SCRIPT = [
102  'set -e',
103  'mkdir -p "$2"',
104  'base64 -d "$1" | tar -xzf - -C "$2"',
105  'rm -f "$1"',
106  'cd "$2"',
107  ...ENSURE_GIT,
108  'if command -v git >/dev/null 2>&1; then',
109  '  git init -q',
110  '  git add -A',
111  `  git ${IDENTITY} commit -q --allow-empty -m baseline`,
112  '  echo git',
113  'else',
114  '  echo no-git',
115  'fi',
116].join('\n')
117
118/** Where a sync's patch is assembled inside the sandbox. */
119export const SYNC_FILE = '/tmp/jev-sync.b64'
120/**
121 * bash -c script, `$1` the base64 patch, `$2` the project folder, `$3` the
122 * commit message: brings the sandbox's copy to the project's current state.
123 */
124export const SYNC_SCRIPT = [
125  'set -e',
126  'base64 -d "$1" > /tmp/jev-sync.patch',
127  'rm -f "$1"',
128  'cd "$2"',
129  'git apply --binary --whitespace=nowarn /tmp/jev-sync.patch',
130  'rm -f /tmp/jev-sync.patch',
131  'git add -A',
132  `git ${IDENTITY} commit -q --allow-empty -m "$3"`,
133].join('\n')
134
135/** The processes running now, without this script's own. */
136const PROCESSES = `ps -eo pid=,ppid=,comm= 2>/dev/null | awk -v me=$$ '$1!=me && $2!=me && $1!=2 && $2!=2 {print $1" "$3}' | sort`
137/**
138 * bash -c script, `$1` how the job starts (`current`, `none`, `git`), `$2` the
139 * project folder, `$3` the job's folder, `$4` a git URL, `$5` its ref, `$6`
140 * the working directory inside the job's folder: makes
141 * the job's folder (a worktree of the synced project, a clone, or an empty
142 * repository), then marks the time and the processes so the report can tell
143 * what the job did.
144 */
145export const PREPARE_SCRIPT = [
146  'set -e',
147  'rm -rf "$3"',
148  'mkdir -p "$(dirname "$3")"',
149  ...ENSURE_GIT,
150  'case "$1" in',
151  '  current) git -C "$2" worktree add -q --detach "$3" HEAD ;;',
152  '  git) if [ -n "$5" ]; then git clone -q -- "$4" "$3" && git -C "$3" checkout -q "$5" --; else git clone -q --depth 1 -- "$4" "$3"; fi ;;',
153  `  *) mkdir -p "$3" && git -C "$3" init -q && git -C "$3" ${IDENTITY} commit -q --allow-empty -m empty ;;`,
154  'esac',
155  'if [ -n "$6" ]; then mkdir -p "$3/$6"; fi',
156  'touch "$3.marker"',
157  `${PROCESSES} > "$3.ps" || true`,
158].join('\n')
159/**
160 * bash -c script, `$1` the job's folder, `$2` 1 to list what the job did
161 * outside it, `$3` 1 to keep a patch of its changes: prints `@@changes` and
162 * git's porcelain status, `@@patch <bytes>`, then `@@outside` (files written
163 * since the job began, outside the jobs, /tmp and the kernel's folders) and
164 * `@@processes` (processes started since then and still running).
165 */
166export const REPORT_SCRIPT = [
167  'cd "$1" || exit 0',
168  'if git rev-parse --git-dir >/dev/null 2>&1; then',
169  '  git add -A >/dev/null 2>&1',
170  '  echo "@@changes"',
171  '  git status --porcelain=v1 --no-renames',
172  '  if [ "$3" = 1 ]; then',
173  '    git diff --cached --binary --full-index HEAD > "$1.patch"',
174  `    echo "@@patch $(wc -c < "$1.patch" | tr -d ' ')"`,
175  '  fi',
176  'fi',
177  'if [ "$2" = 1 ]; then',
178  '  echo "@@outside"',
179  '  find / -xdev \\( -path /proc -o -path /sys -o -path /dev -o -path /tmp -o -path /run -o -path /var/tmp -o -path /var/cache -o -path /var/log -o -path "$(dirname "$1")" \\) -prune -o -newer "$1.marker" \\( -type f -o -type l \\) -print 2>/dev/null | head -n 200',
180  '  echo "@@processes"',
181  `  ${PROCESSES} > "$1.ps2" || true`,
182  '  comm -13 "$1.ps" "$1.ps2" 2>/dev/null | head -n 50',
183  'fi',
184].join('\n')
185/** bash -c script, `$1` a job's patch: prints it, refusing one that is not UTF-8 text (it would not survive the trip). */
186export const PATCH_SCRIPT = [
187  'test -f "$1" || { echo "the job kept no patch" >&2; exit 3; }',
188  'if command -v iconv >/dev/null 2>&1 && ! iconv -f UTF-8 -t UTF-8 "$1" >/dev/null 2>&1; then echo "the patch holds text that is not UTF-8; apply it by hand" >&2; exit 3; fi',
189  'cat "$1"',
190].join('\n')
191
192export type Report = { changes: Change[]; patchBytes: number | null; outside: string[]; processes: string[] }
193
194/** REPORT_SCRIPT's output → its parts. */
195export function readReport(text: string): Report {
196  const out: Report = { changes: [], patchBytes: null, outside: [], processes: [] }
197  let section = ''
198  const status: string[] = []
199  for (const line of text.split('\n')) {
200    const m = /^@@(changes|patch|outside|processes)(?: (\d+))?$/.exec(line.trim())
201    if (m) {
202      section = m[1]!
203      if (section === 'patch') out.patchBytes = Number(m[2] ?? 0)
204      continue
205    }
206    if (!line.trim()) continue
207    if (section === 'changes') status.push(line)
208    else if (section === 'outside') out.outside.push(line.trim())
209    else if (section === 'processes') out.processes.push(line.trim().replace(/^\d+\s+/, ''))
210  }
211  out.changes = readChanges(status.join('\n'))
212  return out
213}
214
215/**
216 * The paths a patch writes, read from the patch itself (the sandbox's own
217 * file list could lie), and whether it makes a symlink, which a later write
218 * could follow out of the project.
219 */
220export function patchTargets(patch: string): { paths: string[]; symlink: boolean } {
221  const paths = new Set<string>()
222  let symlink = false
223  for (const line of patch.split('\n')) {
224    const header = /^diff --git a\/(.+) b\/(.+)$/.exec(line)
225    if (header) {
226      paths.add(header[1]!)
227      paths.add(header[2]!)
228      continue
229    }
230    const moved = /^(?:rename|copy) (?:from|to) (.+)$/.exec(line)
231    if (moved) paths.add(moved[1]!)
232    const file = /^(?:\+\+\+|---) (?:[ab]\/)?(.+)$/.exec(line)
233    if (file && file[1] !== '/dev/null') paths.add(file[1]!)
234    if (/^(?:new file mode|new mode|old mode|deleted file mode) 120000$/.test(line) || /^index [0-9a-f]+\.\.[0-9a-f]+ 120000$/.test(line)) symlink = true
235  }
236  return { paths: [...paths].map(p => p.replace(/^"(.*)"$/, '$1')), symlink }
237}
238
239/** A URL the sandbox may clone: https only (no file://, ext:: or option-looking strings). */
240export function isGitUrl(url: string): boolean {
241  return /^https:\/\/[^\s'"`]+$/.test(url)
242}
243
244/** A branch, tag or commit to check out: no spaces, never an option. */
245export function isGitRef(ref: string): boolean {
246  return /^[A-Za-z0-9._/-]+$/.test(ref) && !ref.startsWith('-') && !ref.includes('..')
247}
248
249export type Change = { path: string; kind: 'added' | 'modified' | 'deleted' }
250
251/** `git status --porcelain=v1` after `add -A` → the changed files. */
252export function readChanges(text: string): Change[] {
253  const out: Change[] = []
254  for (const line of text.split('\n')) {
255    if (line.length < 4) continue
256    const code = line.slice(0, 2)
257    let path = line.slice(3)
258    if (path.startsWith('"') && path.endsWith('"')) path = path.slice(1, -1)
259    const kind = code.includes('D') ? 'deleted' : code.includes('A') || code.includes('?') ? 'added' : 'modified'
260    out.push({ path, kind })
261  }
262  return out
263}
264
265/** "3 files: +a.txt, ~b.ts, -c.md" (at most `max` named). */
266export function describeChanges(changes: Change[], max = 8): string {
267  if (!changes.length) return 'no files changed'
268  const mark = { added: '+', modified: '~', deleted: '-' } as const
269  const named = changes.slice(0, max).map(c => `${mark[c.kind]}${c.path}`)
270  const more = changes.length > max ? `, … ${changes.length - max} more` : ''
271  return `${changes.length} file${changes.length === 1 ? '' : 's'} changed: ${named.join(', ')}${more}`
272}
273
274/** The folder name the project gets inside the sandbox. */
275export function folderName(root: string): string {
276  const base = root.replace(/[\\/]+$/, '').split(/[\\/]/).pop() || 'workspace'
277  return base.replace(/[^A-Za-z0-9._-]/g, '_') || 'workspace'
278}
279
280/** The session's working directory relative to the project root ('' at the root, null outside it). */
281export function relativeCwd(root: string, cwd: string): string | null {
282  const r = root.replace(/[\\/]+$/, '')
283  const c = cwd.replace(/[\\/]+$/, '')
284  if (c === r) return ''
285  if (c.startsWith(`${r}/`) || c.startsWith(`${r}\\`)) return c.slice(r.length + 1).replace(/\\/g, '/')
286  return null
287}
288
289export function megabytes(bytes: number): string {
290  return `${(bytes / 1_048_576).toFixed(bytes < 10_485_760 ? 1 : 0)} MB`
291}
292