Autonomous feedback loop for Claude Code: plan -> implement -> review -> adversarial final verification, with model routing by complexity

Dai un task a Claude Code e lascialo lavorare finché non è davvero finito.
Perseveranza è un plugin per Claude Code che trasforma una richiesta in un ciclo autonomo a feedback: Claude esplora il codice, scrive un piano, implementa uno step alla volta, fa revisionare ogni step da un agente con contesto pulito, e può dirsi "finito" solo dopo una verifica finale avversariale che prova a smontare il lavoro. Alla fine trovi commit e push verificati, un archivio del run e una notifica sul desktop. Se serve un umano, il loop si ferma e ti lascia un passaggio di consegne scritto.
Zero dipendenze: gira su Node.js, lo stesso runtime di Claude Code. Dormiente finché non lo armi: nelle chat normali non esiste.
Requisiti: Claude Code 2.1.287 o successivo, nella CLI (claude, anche claude -p) e Node.js ≥ 20. Dalla 3.0 Perseveranza è una mod: l'app Desktop e l'estensione VS Code non guidano il loop (requisiti completi). Tre modi, mai due insieme: due copie della mod guiderebbero lo stesso loop.
/plugin marketplace add https://github.com/ilmondovero/perseveranza
/plugin install perseveranza@perseveranza
Aggiornamento: claude plugin update perseveranza@perseveranza. Disinstallazione: dal pannello /plugin.
claude --plugin-dir <checkout>. Claude Code scrive nella cartella .claude-plugin/types/ e, se manca, un tsconfig.json (nel repository sono in .gitignore).node install.mjs copia il plugin in ~/.claude/perseveranza/ e lo fa caricare con CLAUDE_CODE_PLUGIN_DIRS nell'env di ~/.claude/settings.json, il modo documentato per un plugin che non viene da un marketplace. Non scrive hook di impostazioni e toglie quelli lasciati da un'installazione manuale 1.x o 2.x: solo un hook che esegue esattamente uno dei loro script, e solo se quello script è ancora il loro (non c'è più, o è identico a una release). Un tuo script con un nome simile resta; un hook che esegue un loro script che hai modificato resta anche lui, e viene elencato.~/.claude/perseveranza/.perseveranza-install.json (ogni file con dimensione e sha256), se sono ancora come li ha scritti, e i file che Claude Code genera in una cartella di plugin che riconosce come sua. (b) I file lasciati da un'installazione 1.x o 2.x che sono identici byte per byte a quelli di una release passata: le impronte sono in src/shell/legacy-hashes.mjs, calcolate dalla storia git e verificate dai test. (c) Gli avanzi di una sua esecuzione interrotta, riconosciuti dal file che ci scrive per primo (.perseveranza-staging.json, con pid, ora e un valore casuale che compaiono anche nel nome della cartella), e solo se quel processo non c'è più. Tutto il resto lo segnala e basta: un tuo file con il nome di uno vecchio, una vecchia copia che hai modificato, una cartella chiamata come un avanzo restano dove sono, elencati con "Left alone ... remove it yourself". Se ~/.claude/hooks, agents o commands è un link o una junction (per esempio verso un repository di dotfiles), lì dentro non cancella niente e lo dice.--uninstall toglie solo i file del marcatore ancora intatti e lascia i tuoi (e il marcatore se ne va: una nuova installazione su quella cartella la rifiuta finché non sposti i tuoi file). Se l'installazione c'è già, uguale file per file, non ricopia niente e lo dice.settings.json si scrive solo dopo. Se un'installazione si interrompe, la successiva rimette a posto la copia vecchia più recente se è completa, o pulisce i suoi avanzi. Due install.mjs sulla stessa cartella non girano insieme: il secondo aspetta il lucchetto ~/.claude/.perseveranza-install.lock (fino a 10 s) e poi si ferma spiegando cosa fare. Un lucchetto il cui processo non c'è più viene ripreso subito.settings.json. Lo modifica come testo: scrive o toglie la sua voce e toglie gli hook vecchi; ogni altro byte resta com'era (numeri come li hai scritti, array in linea, spazi, escape, chiavi ripetute). Il risultato deve dare esattamente le impostazioni attese, altrimenti rifiuta senza toccare niente (succede solo se una chiave che deve cambiare è scritta due volte). Installare e poi disinstallare restituisce lo stesso file byte per byte, con tre eccezioni: la voce che aggiunge è scritta come i membri accanto; il valore di CLAUDE_CODE_PLUGIN_DIRS, quando cambia, è scritto in JSON standard (un escape non necessario diventa il carattere); un file vuoto diventa {}, e un "env": {} vuoto che c'era già sparisce. La voce si confronta sul percorso reale: la stessa cartella scritta in un altro modo (nome 8.3, junction, subst) non viene aggiunta due volte, e --uninstall le toglie tutte. La prima volta che cambia un file che esisteva già ne fa una copia, settings.json.bak-perseveranza-3.0, anche se c'è il settings.json.bak-perseveranza lasciato dalla 2.x (che non tocca). Non la sovrascrive mai e non scrive attraverso niente che abbia già quel nome (nemmeno un link rotto). Se il file l'ha creato lei e contiene solo la sua voce, --uninstall lo cancella. Tiene i permessi, un BOM iniziale e un link simbolico (scrive nel file a cui punta). Rifiuta senza cambiare niente un file che non è JSON valido. Rilanciato non cambia nulla.--claude-dir <cartella> sceglie un'altra cartella di configurazione (altrimenti CLAUDE_CONFIG_DIR, poi ~/.claude). Disinstallazione: node install.mjs --uninstall (prima hud off se attivo).
Per controllare: in una sessione nuova /pf status risponde subito, senza un turno del modello. Se /pf non esiste la mod non è caricata: vedi risoluzione dei problemi.
Nel progetto su cui vuoi lavorare:
/perseveranza aggiungi la paginazione all'endpoint /orders, con test
Da qui Claude arma il loop, scrive il piano in .perseveranza/plan.md e il ciclo va avanti da solo: a ogni fine risposta la mod inietta l'istruzione della fase successiva, in italiano, con una riga di avanzamento in testa:
[perseveranza v3.0.2 · ▸impl ▰▰▱▱▱ 2/5 · it7/23 · 84k tok] Task: aggiungi la paginazione…
Quando è finito ricevi la notifica «Progetto finito e verificato · commit+push confermati». Intanto /pf status ti dice a che punto è, anche mentre Claude lavora; /pf help elenca gli altri verbi (il comando /pf).
Un agente che lavora da solo tende a dichiararsi finito troppo presto: il caso comune funziona, i casi limite no, i test "passano" nella sua testa. Un principal ex-Meta che ha messo un validatore davanti al suo agente misura che il 68% delle modifiche conteneva bug da correggere prima della PR (Kun Chen, no-mistakes).
Perseveranza nasce da tre principi:
flowchart TD
START(["/perseveranza «task»"]) --> PLAN
PLAN["<b>plan</b><br/>esplora il codice → checklist<br/>critica del piano: modello esterno<br/>o advisor interno pf-advisor<br/>registra la complessità"] --> IMPL
IMPL["<b>implement</b><br/>uno step della checklist"] --> REV
REV["<b>review</b><br/>agente pf-reviewer, contesto pulito<br/>verdetto in review.json"] -- "blocking > 0" --> FIX
FIX["<b>fix</b> · stesso step, ri-revisionato<br/>dal 2º fallimento: diagnosi esterna<br/>e/o advisor interno"] --> REV
REV -- "blocking = 0" --> NEXT{"restano step?"}
NEXT -- "sì" --> IMPL
NEXT -- "no → test verde fresco<br/>+ claim-done" --> CLEAN
CLEAN["<b>cleanup</b> · una tantum"] --> VERIFY
VERIFY["<b>verifica finale avversariale</b><br/>agente pf-verifier prova a falsificare<br/>+ modello esterno + lente security<br/>verdetto in verify.json<br/>(o uno per lente)"] -- "pass" --> DONE
VERIFY -- "fail" --> POSTFIX["fix post-verifica"] --> IMPL
FIX -. "fix esauriti" .-> PAUSE
VERIFY -. "bocciature esaurite" .-> PAUSE
DONE(["✅ commit + push verificati<br/>run archiviato · notifica"])
PAUSE(["⏸️ pausa + ESCALATION.md<br/>serve intervento umano"])
style DONE fill:#1a7f37,color:#fff
style PAUSE fill:#9a6700,color:#fff
style VERIFY fill:#0969da,color:#fff
| fase | chi la fa | cosa produce |
|---|---|---|
| plan | Claude, dopo aver esplorato il codice; critica di un modello esterno o di pf-advisor | plan.md come checklist, complessità registrata |
| implement | Claude (o pf-executor con opus se la complessità è alta) | uno step, con i suoi casi limite |
| review | pf-reviewer, contesto pulito, modello per complessità | review.json con requestId, blocking e findings |
| fix | Claude, sullo stesso step; dal 2º tentativo con il parere di pf-advisor | il fix, che torna in review |
| cleanup | Claude, una volta sola | codice morto e duplicazioni rimossi, docs aggiornati |
| verifica finale | pf-verifier che assume che il lavoro sia sbagliato, uno per lente | verify.json (o un verify-<lente>.json per lente) con requestId, pass e findings |
| chiusura | la mod allo Stop, non Claude | commit, push, archivio del run, notifica |
Il modello dei revisori segue la complessità che Claude registra: haiku / sonnet / opus per la review, sonnet / opus / opus per la verifica finale. Con high la verifica finale si divide in tre lenti che lavorano in parallelo (vedi sotto).
test lancia la suite, registra l'exit code reale e un'impronta del working tree. Il claim-done è accettato solo con un run verde per l'albero attuale: nella stessa iterazione, o di un'iterazione precedente se il codice non è cambiato da allora (una modifica ai soli file di documentazione non conta).test --if-needed non rilancia una suite il cui verde è già registrato per lo stesso albero, e ogni fase riceve la "prova dei test" corrente, così Claude e i suoi subagent eseguono i test mirati invece di ripetere la suite intera per prudenza. Il verbo annota anche quali test sono falliti e segnala un rosso che sullo stesso albero non si ripete (un test instabile, non un bug).review.json e verify.json sono rinominati in review-<n>.json / verify-<n>.json quando il loop li legge: la fase di fix rilegge i findings da lì invece di richiederli al reviewer.review.json e verify.json sono validati; se il verdetto dichiarato e i findings non concordano vince la lettura più severa; un file malformato o mancante conta come bocciatura, in review come al gate finale.git-finish e ti dice cosa manca; resume ritenta. .perseveranza/ non finisce mai nel commit..perseveranza/ (journal, piano, note, pareri esterni) viene archiviata in ~/.perseveranza/runs/ con un summary.json: runs list, runs show. Se l'archiviazione fallisce, i file restano in .perseveranza e il loop viene disattivato; status indica il recupero. Sistemata la destinazione, disarm ritenta l'archiviazione. Fino al recupero, arm impedisce di sovrascrivere il run conservato.--max, token reali con --budget-tokens, fix per step con --max-retries. Kill switch da qualunque sessione: il file .perseveranza/STOP o PERSEVERANZA_KILL=1.resume --takeover). N git worktree = N loop paralleli.resume --takeover) o fermarlo (disarm). status, la HUD e il riepilogo di disarm mostrano l'età dell'ultimo fire (STALE oltre trenta minuti, PERSEVERANZA_STALE_MS); il journal registra il buco (gap)..perseveranza/activity.json: l'età del silenzio si misura dall'ultimo segno di vita, non dall'ultimo Stop, e un subagent delegato e mai tornato si legge come tale ("delegato a pf-reviewer alle 11:10, non ancora tornato") in status, nella HUD e nell'avviso all'avvio di una sessione.arm) parte un processo staccato che dorme fino alla soglia, si riallinea finché il loop dà segni di vita e, se il silenzio è vero, manda la notifica desktop e scrive watchdog nel journal e nel summary.json. Nessun cron, nessuna sessione da tenere aperta; PERSEVERANZA_NO_WATCHDOG=1 la spegne.PERSEVERANZA_RESTORE=1 la sentinella sblocca. Non esiste un'interfaccia per interrompere un tool in corso in Claude Code: l'unico interrupt è Esc. La sentinella fa quello che farebbe Esc: termina il processo Claude Code che guidava il loop (registrato allo Stop) e riapre la stessa sessione, stesso id, in una nuova console con un prompt di ripristino. Un turno appeso costa la soglia, non una notte. Tre segni di vita (Stop, attività dei tool, scrittura della trascrizione) e il battito del verbo test durante la suite fanno sì che trenta minuti di silenzio siano un turno morto, non lento. Due stadi: avviso a trenta minuti, kill e ripristino a sessanta (PERSEVERANZA_RESTORE_AFTER_MS), perché una sessione ferma su una domanda all'utente non scrive nulla e l'avviso è la sua occasione. Al massimo tre ripristini per run; nessun rilancio senza un processo registrato. Prima di terminare o lanciare qualcosa crea in esclusiva .perseveranza/restore-launched.json (chi lo crea vince il ripristino: due sentinelle che si credono entrambe proprietarie lanciano una volta sola) e scrive l'interruzione in state.json: se una delle due scritture fallisce, o se lo stato c'è solo nella sua copia in sospeso (una scrittura interrotta), avvisa e basta. Un secondo ripristino non parte finché la sessione riaperta non arriva a uno Stop. Quel file non blocca per sempre: datato nel futuro (un orologio tornato indietro) è scartato; uno rimasto da una sentinella morta prima del lancio è considerato abbandonato dopo il doppio della soglia di ripristino (almeno dieci minuti) e il tentativo conta fra i tre; una cartella al suo posto è spostata da parte. Uno di un lancio riuscito aspetta lo Stop della sessione riaperta, che può essere ferma su un prompt. Disattivo di default; verificato a mano su Windows prima di scriverlo. PERSEVERANZA_CLAUDE_BIN indica il binario claude se non è nel PATH; PERSEVERANZA_ACTIVITY_HEARTBEAT_MS regola il battito del verbo test..perseveranza/reconcile.json (complete, partial, uncertain, con i comandi ancora in esecuzione). Finché non lo fa, la mod rifiuta modifiche, deleghe e comandi non di sola lettura: le istruzioni nel prompt non bastano, un tool rifiutato sì. partial continua lo step senza rifare il fatto, complete va in review, uncertain o un comando ancora vivo mettono in pausa per un umano. I contatori di retry non si azzerano: l'interruzione conta, non regala budget. La nuova sentinella concede alla sessione ripristinata un intero intervallo di avvio prima di valutarla di nuovo.requestId, che l'agente copia nel file. Un review.json o verify.json con un id diverso risponde a una richiesta precedente (un subagent di un turno ucciso o di un giro precedente, un file rimasto attraverso un takeover), anche se è stato scritto dopo: viene messo da parte come review-stale-<n>.json e la fase richiede il verdetto, una volta. Un verdetto senza id (un pack di prompt più vecchio) vale se scritto dopo la richiesta; con l'id giusto vale anche se l'orologio del file è indietro.--verifiers correctness,security,tests (e di default con complessità high) Claude delega nello stesso messaggio, in primo piano, un verificatore per lente: correctness (logica, casi limite, input ostili, regressioni dei fix), security (segreti, input non fidati, injection, path traversal), tests (test mirati, casi non coperti, commenti e documentazione che dicono il vero). Ognuno scrive verify-<lente>.json con lo stesso requestId. Il giro passa solo se tutte le lenti hanno scritto e nessuna ha un finding critical; un pass: false senza critical è un pass con warning, annotato nel journal, salvo un finding high/major (quello blocca). Una lente mancante viene richiesta da sola, le altre restano valide; la seconda volta è una bocciatura. I findings di tutte le lenti finiscono, ciascuno con la sua lente, in un solo verify-<n>.json. Un verify.json del giro (un pack di prompt personalizzato che non conosce le lenti) copre ogni lente senza un file suo, e boccia il giro se boccia (lens-fallback nel journal). Con la sola lente general (il default fino a complessità medium) tutto resta com'era: un verificatore, verify.json.verify-<n>.json degli ultimi tre giri bocciati entrano nel prompt della verifica successiva: ogni verificatore deve controllare che quei difetti siano risolti e che i fix non abbiano introdotto regressioni.pf-advisor (contesto pulito, sola lettura sul sorgente, modello opus di default) interviene nei momenti in cui un secondo parere serve davvero: prima di consegnare il piano e quando lo stesso errore si ripresenta (dal 2º fix dopo una review bocciata, dalla 2ª bocciatura della verifica finale). Con modelli esterni rilevati è il ripiego se nessuno risponde; senza, è il secondo parere. Scrive un file di testo libero .perseveranza/advisor-<slot>-<n>.md (critica, tre rischi, cosa cambierebbe, cosa non ha potuto verificare): niente JSON, niente verdetto. Nel fix riceve tutti i tentativi già falliti sullo stesso step (review-<n>.json, verify-<n>.json), non deve riproporre un approccio già fallito e dice se il problema è lo step stesso: allora Claude riscrive lo step in plan.md prima di riprovare. È consultivo: non instrada mai il loop, e un parere mancante, vuoto o in errore non è un finding e non blocca (Claude lo annota in notes.md). Il journal registra ogni suggerimento (advisor-hint: slot e motivo no-external/fallback/off) per misurare quanto serve; status mostra Advisor: on (opus) o off. --advisor off lo spegne, --advisor-model o PERSEVERANZA_ADVISOR_MODEL scelgono il modello..gitignore, altrimenti i run del verificatore stesso lo cambiano. Lo stato di altri strumenti non conta finché non è tracciato: la cartella di stato di oh-my-claudecode, .claude/settings.local.json, .claude/scheduled_tasks.lock e ciò che aggiungono arm --ignore o PERSEVERANZA_FINGERPRINT_IGNORE (un file tracciato lì conta ancora, e il commit di chiusura ne prende le modifiche; quelli non tracciati li lascia stare). Dopo quattro pass che non coprono l'albero, senza bocciature in mezzo, il loop si mette in pausa per un umano. Fuori da git non c'è impronta da confrontare.Le opzioni di /perseveranza:
| opzione | effetto |
|---|---|
--max N | tetto di iterazioni (altrimenti adattivo: 8 + 3 × step, massimo 60); pulizia, verifica finale e chiusura git ne hanno 3 in più, e così lo Stop che porta un claim-done; una verifica che rimanda in implement torna sotto il tetto pieno |
--budget-tokens N | tetto di token, misurati dalla trascrizione della sessione e da quelle dei suoi subagent |
--max-retries N | fix concessi per step prima della pausa (default 3) |
--commit | commit atomico dopo ogni step validato |
--test "cmd" | la suite (se non la passi, Claude la individua) |
| --approve-plan | pausa dopo il piano: approvi tu con /pf resume (Claude
hooks/register.js 109 lines1// perseveranza as a Claude Code mod (2.1.287 or later): the entry point. Only the on(...)
2// calls, with literal event names, and the `io` object below; what each hook does lives in
3// ./lib (one adapter per event), and the loop itself in the Node shell, reached through the
4// bridge (src/shell/mod-bridge.mjs) with $.process.run, which is CLI only.
5//
6// io: the mods API calls the adapters use, as plain closures. A hooks module may pass `$` only
7// to a function declared at the top of this same file, never to an imported one, so the
8// adapters get these closures instead (`claude plugin validate` lists them "via io").
9//
10// Failure policy (PIANO-MOD, section 4): every hook is fail-open (a hook that throws is skipped
11// and Claude Code goes on) except classic.Stop and classic.SubagentStop, whose .catch answers
12// explicitly; their adapters say when that may block and when it may not.
13import { createMod } from './lib/core.js';
14import { onStop, onStopFailed } from './lib/stop.js';
15import { onSubagentStop, onSubagentStopFailed } from './lib/subagent.js';
16import { routeFor, rememberAgent } from './lib/spawn.js';
17import { onToolCall } from './lib/tool.js';
18import { addStep, scheduleUsage } from './lib/usage.js';
19import { onSessionStart, sessionNotice } from './lib/session.js';
20import { serveTool, serveCommand } from './lib/verbs.js';
21import { isOwnTool } from './lib/core.js';
22
23function io($) {
24 return {
25 run: (argv, init) => $.process.run(argv, init),
26 after: (ms, fn) => $.clock.after(ms, fn),
27 now: () => $.clock.now(),
28 cwd: () => $.session.cwd(),
29 sessionId: () => $.session.id(),
30 version: () => $.session.version(),
31 nodeEnv: () => $.env.get('PERSEVERANZA_NODE'),
32 root: () => $.plugin.root,
33 sleep: (ms) => $.clock.sleep(ms),
34 exists: (path) => $.fs.exists(path),
35 read: (path) => $.fs.read(path),
36 write: (path, text) => $.fs.write(path, text),
37 list: (path) => $.fs.list(path),
38 surfaces: () => $.session.surfaces(),
39 log: (text) => $.ui.log(text, { to: 'debug' }),
40 status: (text) => $.ui.status(text),
41 registerTool: (spec) => $.tool.register(spec),
42 registerCommand: (spec) => $.command.register(spec),
43 pluginName: () => $.plugin.name,
44 perseveranzaHome: () => $.env.get('PERSEVERANZA_HOME'),
45 userProfile: () => $.env.get('USERPROFILE'),
46 home: () => $.env.get('HOME'),
47 };
48}
49
50export function register(on) {
51 const mod = createMod();
52
53 on('session.start', async ($, e, next) => {
54 await onSessionStart(io($), mod, e);
55 return next(e);
56 });
57
58 on('classic.SessionStart', async ($, e, next) => {
59 const text = await sessionNotice(io($), mod, e);
60 const r = await next(e);
61 if (!text) return r;
62 const base = r && typeof r === 'object' ? r : {};
63 return { ...base, additionalContext: [...(Array.isArray(base.additionalContext) ? base.additionalContext : []), text] };
64 });
65
66 on('classic.Stop', async ($, e, next) => {
67 const decision = await onStop(io($), mod, e);
68 const r = await next(e);
69 return decision ? { ...(r && typeof r === 'object' ? r : {}), block: decision.block } : r;
70 }).catch(async ($, e, next) => {
71 const decision = await onStopFailed(io($), mod, e, next.error && next.error.message ? next.error.message : next.error && next.error.kind);
72 const r = await next(e);
73 return decision ? { ...(r && typeof r === 'object' ? r : {}), block: decision.block } : r;
74 });
75
76 on('classic.SubagentStop', async ($, e, next) => {
77 const decision = await onSubagentStop(io($), mod, e);
78 const r = await next(e);
79 return decision ? { ...(r && typeof r === 'object' ? r : {}), block: decision.block } : r;
80 }).catch(async ($, e, next) => {
81 const decision = await onSubagentStopFailed(io($), mod, e, next.error && next.error.message ? next.error.message : next.error && next.error.kind);
82 const r = await next(e);
83 return decision ? { ...(r && typeof r === 'object' ? r : {}), block: decision.block } : r;
84 });
85
86 on('agent.spawn', async ($, e, next) => {
87 const model = await routeFor(io($), mod, e);
88 const r = await next(model ? { ...e, model } : e);
89 rememberAgent(mod, e, r);
90 return r;
91 });
92
93 on('tool.call', async ($, e, next) => {
94 // the mod's own tool: answered here (validated, gated, run through the CLI), never passed on
95 if (isOwnTool(mod, e)) return serveTool(io($), mod, e);
96 const refused = await onToolCall(io($), mod, e);
97 return refused ? { deny: refused.deny } : next(e);
98 });
99
100 // /pf (verbs.js COMMAND_NAME): the matcher is a literal, like the event names
101 on('command.run', { command: 'pf' }, async ($, e) => serveCommand(io($), mod, e));
102
103 on('turn.step', async function* ($, e, next) {
104 const result = yield* next(e);
105 if (addStep(mod, e, result)) scheduleUsage(io($), mod);
106 return result;
107 });
108}
109hooks/lib/core.js 131 lines1// What the mod's adapters share: the timings, the mod's ephemeral facts, and the pure helpers.
2// No Node API and no timer globals here (a hooks module has neither): every call that reaches
3// outside goes through the `io` object register.js builds from `$` for each hook.
4//
5// The mod keeps NO loop state in memory. What it keeps is ephemeral and safe to lose on a hot
6// reload or a restart: tokens measured since the last flush, the last tool activity and the
7// delegations not back yet, how many times a judge was sent back, which subagent is which.
8import { loopAgentName, MAX_VERDICT_ASKS } from '../../src/core/subagents.mjs';
9import { LENSES } from '../../src/core/state.mjs';
10import { GATE, MAX_SKIP_NOTES } from './gate.js';
11
12export { loopAgentName, MAX_VERDICT_ASKS, GATE, MAX_SKIP_NOTES };
13
14// The first Claude Code with mods (the docs: "Mods require Claude Code v2.1.287 or later").
15export const MIN_CLAUDE_CODE = '2.1.287';
16
17// $.process.run timeouts. Time spent inside a $ call does not count against a hook's own
18// 10 s, so a stop may wait for the bridge as long as the settings hook could (stop-core keeps
19// its own deadline: PERSEVERANZA_HOOK_TIMEOUT_MS, 120 s, minus a margin).
20export const STOP_TIMEOUT_MS = 125000;
21export const FLUSH_TIMEOUT_MS = 8000;
22export const QUICK_TIMEOUT_MS = 8000;
23// in a .catch handler (1 s of grace for its own code; the $ call itself does not count)
24export const CATCH_TIMEOUT_MS = 5000;
25// A stop waits at most this long for a flush already on its way. This wait IS the hook's own
26// time ($.clock.sleep runs the budget on, unlike the other $ calls): half of the 10 s, so the
27// stop keeps the rest for its own code (its bridge call does not count).
28export const STOP_FLUSH_WAIT_MS = 5000;
29
30// Debounce of the flushes ($.clock.after): tokens every 15 s while a turn runs (and with every
31// stop); a delegation or a subagent's return within 2 s; a plain tool call at most every 30 s
32// (the activity hook's ACTIVITY_THROTTLE_MS).
33export const USAGE_FLUSH_MS = 15000;
34export const ACTIVITY_URGENT_MS = 2000;
35export const ACTIVITY_HEARTBEAT_MS = 30000;
36// a heartbeat the bridge could not write (state.json busy or corrupt) is sent again this many times
37export const MAX_ACTIVITY_RETRIES = 5;
38
39// the mod-start line is offered to a loop's journal this many times per process
40export const MAX_HELLO_TRIES = 3;
41
42// The tools a reconciliation refuses (activity-hook.mjs reconcileDecision): only these ask the
43// bridge before running.
44export const MUTATING_TOOLS = ['Edit', 'Write', 'MultiEdit', 'NotebookEdit', 'Agent', 'Task', 'Bash', 'PowerShell'];
45
46// The mod's own tool (verbs.js): its full name as the model calls it, learned from
47// $.tool.register (mcp__<plugin>__perseveranza), and its verbs the reconciliation lets run
48// without asking (they change nothing: activity-hook.mjs RECONCILE_TOOL_VERBS has these and pause).
49export const OWN_TOOL = 'mcp__perseveranza__perseveranza';
50export const READ_ONLY_VERBS = ['status', 'history', 'explain'];
51export const isOwnTool = (mod, e) => isObj(e) && typeof e.tool === 'string' && e.tool === (mod.toolName || OWN_TOOL);
52
53export function createMod() {
54 return {
55 node: null, // the node binary (PERSEVERANZA_NODE, else node on the PATH)
56 bridge: null, // <plugin root>/src/shell/mod-bridge.mjs
57 tail: Promise.resolve(), // the serial queue of the flushes
58 usage: {}, // { <agentId|'main'>: counts } measured since the last flush
59 usageTimer: null,
60 activity: null, // the last activity record not yet written
61 lines: [], // activity journal lines not yet written
62 inflight: new Set(), // the flushes on their way (a stop waits for them, see stop.js)
63 driving: false, // a stop of this process reached a live loop of this session
64 pending: [], // delegations not back yet [{ at, agent, id: the Agent call's tool_use_id }]
65 agentTool: {}, // agentId -> the tool_use_id of the Agent call that spawned it
66 settled: {}, // agent_id -> its delegation was closed (once per subagent)
67 activityTimer: null,
68 activityDue: 0,
69 activityRetries: 0,
70 lastActivityFlush: 0,
71 transcript: '',
72 asks: {}, // agent_id -> times this judge was sent back for its verdict
73 agents: {}, // agentId -> { type, lens } learned at agent.spawn
74 hello: null, // { claudeCode, ok, min, problem } from session.start
75 helloSent: false,
76 helloTries: 0,
77 skips: {}, // hook -> fail-open skips seen
78 toolName: null, // the full name of the mod's tool, once registered (session.start)
79 commandName: null, // the /pf command, once registered
80 pluginName: '', // $.plugin.name
81 alive: {}, // session id -> when its sign of life was last written or tried (verbs.js writeAlive, refreshAlive)
82 alivePruned: false, // the old signs of life were looked at (verbs.js pruneAliveIfNeeded)
83 };
84}
85
86const isObj = (v) => !!v && typeof v === 'object' && !Array.isArray(v);
87export { isObj };
88export const errText = (e) => String((e && e.message) || e || 'unknown error').slice(0, 300);
89
90// "2.1.289" or "2.1.289-dev..." -> [2, 1, 289] | null
91function parts(v) {
92 const m = typeof v === 'string' ? v.trim().match(/^(\d+)\.(\d+)\.(\d+)/) : null;
93 return m ? [Number(m[1]), Number(m[2]), Number(m[3])] : null;
94}
95
96// $.session.version() -> what the journal and `status` say about the Claude Code the mod runs on
97export function versionCheck(v) {
98 const raw = isObj(v) ? (typeof v.base === 'string' && v.base ? v.base : v.version) : null;
99 const claudeCode = typeof raw === 'string' && raw ? raw.slice(0, 60) : 'unknown';
100 const have = parts(raw);
101 const min = parts(MIN_CLAUDE_CODE);
102 if (!have) return { claudeCode, ok: false, min: MIN_CLAUDE_CODE, problem: 'the version could not be read' };
103 for (let i = 0; i < 3; i++) {
104 if (have[i] > min[i]) break;
105 if (have[i] < min[i]) return { claudeCode, ok: false, min: MIN_CLAUDE_CODE, problem: `older than ${MIN_CLAUDE_CODE}, the first version with mods: hooks may be missing` };
106 }
107 return { claudeCode, ok: true, min: MIN_CLAUDE_CODE, problem: '' };
108}
109
110// A model response's usage (turn.step's result.usage) in the loop's counts; null when absent.
111export function usageCounts(u) {
112 if (!isObj(u)) return null;
113 const n = (v) => (typeof v === 'number' && Number.isFinite(v) && v > 0 ? Math.floor(v) : 0);
114 const c = { inputTokens: n(u.input_tokens), outputTokens: n(u.output_tokens), cacheReadTokens: n(u.cache_read_input_tokens), cacheCreationTokens: n(u.cache_creation_input_tokens) };
115 return c.inputTokens || c.outputTokens || c.cacheReadTokens || c.cacheCreationTokens ? c : null;
116}
117
118// The lens a verifier was given, when its prompt names exactly one lens file (verify-<lens>.json).
119export function lensOf(prompt) {
120 if (typeof prompt !== 'string') return undefined;
121 const found = new Set();
122 for (const m of prompt.matchAll(/verify-([a-z]+)\.json/g)) if (LENSES.includes(m[1])) found.add(m[1]);
123 return found.size === 1 ? [...found][0] : undefined;
124}
125
126// The judges the loop waits a verdict from.
127export function isJudge(agentType) {
128 const n = loopAgentName(agentType);
129 return n === 'pf-reviewer' || n === 'pf-verifier';
130}
131hooks/lib/stop.js 103 lines1// classic.Stop: the loop's driver. The bridge runs the Stop logic of the settings hook
2// (stop-core.mjs) with the facts only the mod sees: the background tasks of the Stop input,
3// the tokens measured since the last flush, the loop mode. Its answer is a block (the next
4// instruction) or nothing (let Claude stop).
5//
6// Fail-closed, within limits: when the bridge gives no usable answer the hook throws, and its
7// .catch answers ONE recovery block, so a broken helper does not end a loop in silence:
8// - only where a loop is armed (state.json, see gate.js), never because a .perseveranza/
9// folder exists (~/.perseveranza is perseveranza's own home);
10// - only for the session that drives it (the owner read from state.json; if unreadable, a
11// session this mod has seen driving it), never for another session's loop;
12// - never with stop_hook_active already true (the last stop was blocked), so never an endless
13// block. The price, said plainly: after the first block of a user turn every stop arrives
14// with stop_hook_active true, so a bridge that fails in the middle of a loop lets Claude
15// stop at once and the watchdog takes over. That stop is never silent: the failure goes to
16// the journal (through the bridge), or, when the bridge cannot even journal (node
17// unreachable), to .perseveranza/mod-fault.json, which `status` shows and the next stop
18// journals; and to the line under the prompt.
19import { STOP_TIMEOUT_MS, CATCH_TIMEOUT_MS, STOP_FLUSH_WAIT_MS, isObj, errText } from './core.js';
20import { call, hasGate, readOwner, writeFault } from './gate.js';
21import { takeUsage, giveBack, cancelUsageTimer } from './usage.js';
22import { sendHello, loopModeOf } from './session.js';
23
24export class BridgeFailure extends Error {}
25
26// The bridge's answer -> { block } | null. Throws BridgeFailure when there is none (the
27// .catch decides). The delta it carried: sent again only when the bridge says it surely
28// did not queue it (usageDropped); never after no answer, never when unverified.
29export function settleStop(mod, delta, r) {
30 if (!r || !r.answer) throw new BridgeFailure(r && r.error ? r.error : 'no answer from the bridge');
31 const a = r.answer;
32 if (a.ok !== true) throw new BridgeFailure(`the bridge failed: ${String(a.error || 'unknown').slice(0, 200)}`);
33 if (a.usageDropped && delta) giveBack(mod, delta);
34 if (a.outcome !== 'dormant' && a.outcome !== 'foreign-session') mod.driving = true;
35 const block = isObj(a.decision) && typeof a.decision.block === 'string' && a.decision.block.trim() ? a.decision.block : null;
36 return block ? { block } : null;
37}
38
39// A token flush already on its way (its delta comes back to memory if the bridge drops it) is
40// waited for, up to STOP_FLUSH_WAIT_MS, before the stop takes what is in memory: the stop that
41// archives the run then counts it. Only when one is on its way: otherwise no wait at all.
42export async function waitFlushes(io, mod) {
43 if (!mod.inflight.size) return 'idle';
44 const all = Promise.allSettled([...mod.inflight]).then(() => 'flushed');
45 let sleep = null;
46 try { sleep = Promise.resolve(io.sleep(STOP_FLUSH_WAIT_MS)).then(() => 'timeout', () => 'timeout'); } catch { /* no clock: the flush's own timeout bounds it */ }
47 return Promise.race(sleep ? [all, sleep] : [all]);
48}
49
50export async function onStop(io, mod, e) {
51 const cwd = isObj(e) && typeof e.cwd === 'string' && e.cwd ? e.cwd : await io.cwd();
52 if (isObj(e) && typeof e.transcript_path === 'string' && e.transcript_path) mod.transcript = e.transcript_path;
53 // DORMANT: no loop armed, no node process (a failed check calls the bridge, which knows)
54 if (!(await hasGate(io, cwd, true))) return null;
55 await waitFlushes(io, mod);
56 cancelUsageTimer(mod);
57 const delta = takeUsage(mod);
58 const facts = {
59 backgroundTasks: isObj(e) && Array.isArray(e.background_tasks) ? e.background_tasks : [],
60 // the tool, when this process registered it (the bridge names it only if the run was armed for it)
61 loopMode: loopModeOf(mod),
62 // always a mod reading, even an empty one: the transcripts are never read under the mod
63 usage: delta ? { byAgent: delta } : {},
64 };
65 const r = await call(io, mod, 'stop', { cwd, event: e, facts, timeoutMs: STOP_TIMEOUT_MS });
66 const decision = settleStop(mod, delta, r);
67 // the Claude Code version, once, in the journal of a live loop (status shows it)
68 if (!mod.helloSent && r.answer.outcome !== 'dormant' && r.answer.outcome !== 'foreign-session') await sendHello(io, mod, cwd, isObj(e) ? e.session_id : '');
69 return decision;
70}
71
72export function recoveryText(error, root, tool = false) {
73 const cli = `node "${String(root || '<perseveranza root>').replace(/[\\/]+$/, '')}/src/cli/perseveranza.mjs"`;
74 // the helper is down, so the tool (which runs the same node) may be too: the CLI is named either way
75 const how = tool ? `Call the \`perseveranza\` tool with {"verb":"status"} (if it fails too, run \`${cli} status\`)` : `Run \`${cli} status\``;
76 return `perseveranza (mod): the Stop hook could not reach its helper (${error}). The loop is still armed. ${how} to see the current phase and its instruction, then carry on with that phase. If this happens again at the next stop, the loop lets you stop and the watchdog takes over.`;
77}
78
79// The .catch of classic.Stop -> { block } | null (see the head of this file).
80export async function onStopFailed(io, mod, e, error) {
81 const cwd = isObj(e) && typeof e.cwd === 'string' && e.cwd ? e.cwd : '';
82 // a wrong "yes" here would block a session that has no loop: unknown counts as no
83 if (!(await hasGate(io, cwd, false))) return null;
84 const session = isObj(e) && typeof e.session_id === 'string' ? e.session_id : '';
85 const owner = await readOwner(io, cwd);
86 const ours = owner === '' || (owner !== null && owner === session) || (owner === null && mod.driving === true);
87 const why = errText(error);
88 try { await io.log(`perseveranza: the Stop hook could not reach its helper (${why})`); } catch { /* best effort */ }
89 if (!ours) return null; // another session's loop: not this session's to resume, nor to report
90 const active = isObj(e) && e.stop_hook_active === true;
91 const recovery = !active;
92 let now = 0;
93 try { now = await io.now(); } catch { /* 0 */ }
94 const r = await call(io, mod, 'journal', { cwd, event: { session_id: session }, facts: { lines: [{ type: 'mod-hook-skipped', hook: 'classic.Stop', error: why, recovery }] }, timeoutMs: CATCH_TIMEOUT_MS });
95 const journaled = !!(r.answer && r.answer.ok === true && r.answer.journaled > 0);
96 if (!journaled) await writeFault(io, cwd, { at: now, hook: 'classic.Stop', error: why, session, stopHookActive: active, recovery });
97 try { await io.status(`perseveranza: the Stop hook could not reach its helper (${why})${recovery ? '' : '; the loop was left at this stop (the watchdog takes over)'}`); } catch { /* best effort */ }
98 if (!recovery) return null;
99 let root = '';
100 try { root = io.root(); } catch { /* the text says where without it */ }
101 return { block: recoveryText(why, root, loopModeOf(mod) === 'tool') };
102}
103hooks/lib/subagent.js 74 lines1// classic.SubagentStop: for a judge (pf-reviewer, pf-verifier), the check that the verdict it
2// was asked for is on disk, then the heartbeat (a delegation came back). When the verdict is
3// missing, stale or malformed the judge is sent back ({ block }) with the reason, at most
4// MAX_VERDICT_ASKS times per subagent; then it goes, and the machine's own `missing` outcome is
5// the net. Fail-closed like the stop: no answer from the bridge -> the .catch sends the judge
6// back once, unless stop_hook_active says it was already sent back (never an endless block).
7//
8// A judge sent back is still running: its delegation stays open. It closes once, when the
9// subagent is let go (by its agent_id, see activity.js noteSubagentStop).
10import { MAX_VERDICT_ASKS, QUICK_TIMEOUT_MS, CATCH_TIMEOUT_MS, isObj, isJudge, errText } from './core.js';
11import { call, hasGate } from './gate.js';
12import { noteSubagentStop, scheduleActivity } from './activity.js';
13
14const idOf = (e) => (isObj(e) && typeof e.agent_id === 'string' ? e.agent_id : '');
15
16// The subagent is let go: its delegation closes (once per agent_id) and the heartbeat goes.
17async function letGo(io, mod, e) {
18 const id = idOf(e);
19 if (id && mod.settled[id]) return;
20 if (id) mod.settled[id] = true;
21 let now = 0;
22 try { now = await io.now(); } catch { /* 0: the record is still written */ }
23 scheduleActivity(io, mod, noteSubagentStop(mod, e, now), now);
24}
25
26// -> { block } to send the judge back, or null. Throws when the bridge gives no answer.
27async function judge(io, mod, e) {
28 if (!isObj(e) || !isJudge(e.agent_type)) return null;
29 const id = idOf(e);
30 const asked = mod.asks[id] || 0;
31 if (asked >= MAX_VERDICT_ASKS) return null;
32 const cwd = typeof e.cwd === 'string' && e.cwd ? e.cwd : await io.cwd();
33 if (!(await hasGate(io, cwd, true))) return null;
34 const known = mod.agents[id];
35 const facts = { askedTimes: asked, ...(known && known.lens ? { lens: known.lens } : {}) };
36 const r = await call(io, mod, 'subagent-stop', { cwd, event: e, facts, timeoutMs: QUICK_TIMEOUT_MS });
37 if (!r.answer) throw new Error(r.error || 'no answer from the bridge');
38 if (r.answer.ok !== true) throw new Error(`the bridge failed: ${String(r.answer.error || 'unknown').slice(0, 200)}`);
39 const block = isObj(r.answer.decision) && typeof r.answer.decision.block === 'string' && r.answer.decision.block.trim() ? r.answer.decision.block : null;
40 if (!block) return null;
41 mod.asks[id] = asked + 1;
42 return { block };
43}
44
45export async function onSubagentStop(io, mod, e) {
46 const decision = await judge(io, mod, e);
47 if (!decision) await letGo(io, mod, e);
48 return decision;
49}
50
51export function verdictUncheckedText(error) {
52 return `perseveranza (mod): your verdict could not be checked (${error}). Before you finish, make sure the verdict file you were asked to write (.perseveranza/review.json, or .perseveranza/verify.json / verify-<lens>.json) exists, is valid JSON, and carries the requestId you were given. Then finish.`;
53}
54
55// The .catch of classic.SubagentStop -> { block } | null.
56async function failed(io, mod, e, error) {
57 if (!isObj(e) || !isJudge(e.agent_type) || e.stop_hook_active === true) return null;
58 const id = idOf(e);
59 const asked = mod.asks[id] || 0;
60 if (asked >= MAX_VERDICT_ASKS) return null;
61 const cwd = typeof e.cwd === 'string' ? e.cwd : '';
62 if (!(await hasGate(io, cwd, false))) return null;
63 const session = typeof e.session_id === 'string' ? e.session_id : '';
64 await call(io, mod, 'journal', { cwd, event: { session_id: session }, facts: { lines: [{ type: 'mod-hook-skipped', hook: 'classic.SubagentStop', error: errText(error), recovery: true }] }, timeoutMs: CATCH_TIMEOUT_MS });
65 mod.asks[id] = asked + 1;
66 return { block: verdictUncheckedText(errText(error)) };
67}
68
69export async function onSubagentStopFailed(io, mod, e, error) {
70 const decision = await failed(io, mod, e, error);
71 if (!decision) await letGo(io, mod, e);
72 return decision;
73}
74hooks/lib/spawn.js 40 lines1// agent.spawn: a pf-* subagent runs on the model MODEL_ROUTING gives the run's complexity
2// (routeModel, through the bridge's 'route-model', which journals the route). Fail-open: when
3// the route cannot be had (no loop, another session's loop, the bridge down) the model the
4// prompt asked for stands, and a failure is journaled as mod-hook-skipped.
5//
6// The subagent's id comes back from the spawn: the mod remembers its type and, for a
7// verifier whose prompt names one lens file, its lens (the SubagentStop input does not say it).
8import { QUICK_TIMEOUT_MS, isObj, loopAgentName, lensOf } from './core.js';
9import { call, hasGate, skipped } from './gate.js';
10
11// -> the model alias to spawn with, or null (keep e.model)
12export async function routeFor(io, mod, e) {
13 if (!isObj(e) || !loopAgentName(e.subagentType) || e.fork === true) return null;
14 let cwd = '';
15 let session = '';
16 try {
17 cwd = typeof e.cwd === 'string' && e.cwd ? e.cwd : await io.cwd();
18 session = await io.sessionId();
19 } catch (err) {
20 await skipped(io, mod, cwd, 'agent.spawn', String((err && err.message) || err));
21 return null;
22 }
23 if (!(await hasGate(io, cwd, true))) return null;
24 const r = await call(io, mod, 'route-model', { cwd, event: { subagentType: e.subagentType, model: typeof e.model === 'string' ? e.model : '', session_id: session }, timeoutMs: QUICK_TIMEOUT_MS });
25 if (!r.answer || r.answer.ok !== true) {
26 await skipped(io, mod, cwd, 'agent.spawn', r.error || String(r.answer && r.answer.error));
27 return null;
28 }
29 return typeof r.answer.model === 'string' && r.answer.model ? r.answer.model : null;
30}
31
32// After the spawn: which Agent call started which agent id (its SubagentStop closes that
33// delegation, activity.js) and, for a loop subagent, which one it is.
34export function rememberAgent(mod, e, result) {
35 if (!isObj(e) || !isObj(result) || typeof result.agentId !== 'string' || !result.agentId) return;
36 if (typeof e.tool_use_id === 'string' && e.tool_use_id) mod.agentTool[result.agentId] = e.tool_use_id;
37 if (!loopAgentName(e.subagentType)) return;
38 mod.agents[result.agentId] = { type: String(e.subagentType).slice(0, 80), lens: lensOf(e.prompt) };
39}
40hooks/lib/tool.js 68 lines1// tool.call: the reconciliation guard (while a restored loop is being reconciled, a tool that
2// would mutate is refused: activity-hook.mjs's reconcileDecision, asked of the bridge's
3// 'tool-check'), then the heartbeat. The guard asks only for the tools it could refuse, only
4// where a loop is armed (gate.js hasGate), and only when state.json does not already say the
5// loop is not being reconciled (needsGuard: read with $.fs.read, no process); it is fail-open
6// like the settings hook was (a failure is journaled as mod-hook-skipped and the tool runs).
7import { MUTATING_TOOLS, QUICK_TIMEOUT_MS, isObj } from './core.js';
8import { call, hasGate, skipped, gateFile, STATE_FILES } from './gate.js';
9import { noteTool, scheduleActivity } from './activity.js';
10import { refreshAlive } from './verbs.js';
11
12// tool.call's input is the tool's input with tool, tool_use_id (and agentId) beside it
13const RESERVED = new Set(['tool', 'tool_use_id', 'agentId', 'consent']);
14export function toolInput(e) {
15 return Object.fromEntries(Object.entries(isObj(e) ? e : {}).filter(([k]) => !RESERVED.has(k)));
16}
17
18// -> the reason to refuse, or null
19export async function guardTool(io, mod, e, session) {
20 if (!isObj(e) || !MUTATING_TOOLS.includes(e.tool)) return null;
21 let cwd = '';
22 try { cwd = await io.cwd(); } catch (err) {
23 await skipped(io, mod, '', 'tool.call', String((err && err.message) || err));
24 return null;
25 }
26 if (!(await hasGate(io, cwd, true))) return null;
27 return askGuard(io, mod, cwd, session, e.tool, toolInput(e));
28}
29
30// Does the guard need the bridge at all? Only a loop being reconciled after a restore
31// (signals.interrupted, written by the watchdog) refuses anything, and that is in state.json,
32// which $.fs.read reads without a process. A node process per Edit/Write/Bash/Agent call was
33// ~130 ms each on an idle machine and timed out (8 s) under load in a real run. -> true unless
34// state.json reads as a loop state that is not being reconciled: absent (a pending copy only),
35// unreadable, truncated, not a state (no phase): the bridge decides, with its retries, as before.
36export async function needsGuard(io, cwd) {
37 let raw;
38 try { raw = JSON.parse(await io.read(gateFile(cwd, STATE_FILES.state))); } catch { return true; }
39 if (!isObj(raw) || typeof raw.phase !== 'string') return true;
40 const i = isObj(raw.signals) ? raw.signals.interrupted : null;
41 // the bridge's normalizeState: any object (an array too) is an interruption, anything else none
42 return !!i && typeof i === 'object';
43}
44
45// The bridge's reconciliation guard for one call (op 'tool-check'). Fail-open. -> reason | null
46export async function askGuard(io, mod, cwd, session, tool, input) {
47 if (!(await needsGuard(io, cwd))) return null;
48 const r = await call(io, mod, 'tool-check', { cwd, event: { session_id: session, tool, input }, timeoutMs: QUICK_TIMEOUT_MS });
49 if (!r.answer || r.answer.ok !== true) {
50 await skipped(io, mod, cwd, 'tool.call', r.error || String(r.answer && r.answer.error));
51 return null;
52 }
53 return typeof r.answer.deny === 'string' && r.answer.deny ? r.answer.deny : null;
54}
55
56// -> { deny } to refuse the call, or null to let it run (and count it as life)
57export async function onToolCall(io, mod, e) {
58 let session = '';
59 let now = 0;
60 try { session = await io.sessionId(); now = await io.now(); } catch { /* the record goes without them */ }
61 // the sign of life for `arm`, kept fresh (verbs.js refreshAlive): before the tool runs
62 await refreshAlive(io, mod, session);
63 const deny = await guardTool(io, mod, e, session);
64 if (deny) return { deny };
65 scheduleActivity(io, mod, noteTool(mod, e, now, session), now);
66 return null;
67}
68hooks/lib/usage.js 88 lines1// Tokens per agent, exact, from turn.step (every model request, main's and every subagent's,
2// with its cache counts). They wait in memory and go to the bridge's inbox ('usage-flush')
3// every USAGE_FLUSH_MS while a turn runs, and with every stop (facts.usage). The next stop
4// adds them to state.json (usage-inbox.mjs): the mod never writes state.json.
5//
6// The bridge's contract for a delta (mod-bridge.mjs, PIANO-MOD "Stato fase 1"):
7// queued counted once by a later stop -> done
8// dropped { ok: false, error } (not 'no-loop'): no file under an inbox name can exist,
9// nobody ever counts it -> sent again (back in memory)
10// unverified unverified: true: a file may exist and may already be counted
11// -> NOT sent again
12// no answer the bridge died or timed out: it may have queued it first
13// -> NOT sent again
14// no-loop no run to count it for (never armed, archived) -> dropped for good
15// busy / corrupt state.json: the bridge queues it as armUnknown (usage-queued) -> done
16// foreign-session: another session's loop -> dropped for good
17import { usageCounts, isObj, USAGE_FLUSH_MS, FLUSH_TIMEOUT_MS } from './core.js';
18import { call, enqueue, hasGate, track } from './gate.js';
19
20const KEYS = ['inputTokens', 'outputTokens', 'cacheReadTokens', 'cacheCreationTokens'];
21
22function add(into, key, counts) {
23 const prev = into[key] || { inputTokens: 0, outputTokens: 0, cacheReadTokens: 0, cacheCreationTokens: 0 };
24 into[key] = Object.fromEntries(KEYS.map((k) => [k, prev[k] + (counts[k] || 0)]));
25}
26
27// One model request: e is turn.step's input (agentId for a subagent), result its response.
28export function addStep(mod, e, result) {
29 const c = usageCounts(isObj(result) ? result.usage : null);
30 if (!c) return false;
31 const key = isObj(e) && typeof e.agentId === 'string' && e.agentId ? e.agentId.slice(0, 80) : 'main';
32 add(mod.usage, key, c);
33 return true;
34}
35
36// The delta since the last flush, taken out of memory (null: nothing measured).
37export function takeUsage(mod) {
38 const d = mod.usage;
39 mod.usage = {};
40 return Object.keys(d).length ? d : null;
41}
42
43// A delta the bridge surely did not queue goes back, to leave with the next flush or stop.
44export function giveBack(mod, delta) {
45 for (const [k, c] of Object.entries(isObj(delta) ? delta : {})) if (isObj(c)) add(mod.usage, k, c);
46}
47
48// What happens to a flushed delta, by the bridge's answer (the contract above).
49export function settleUsage(mod, delta, r) {
50 if (!r || !r.answer) return 'no-answer';
51 const a = r.answer;
52 if (a.ok !== true) {
53 if (a.error === 'no-loop') return 'no-loop';
54 giveBack(mod, delta);
55 return 'dropped';
56 }
57 if (a.unverified === true) return 'unverified';
58 return typeof a.outcome === 'string' ? a.outcome : 'queued';
59}
60
61export async function flushUsage(io, mod) {
62 const delta = takeUsage(mod);
63 if (!delta) return 'empty';
64 let cwd = '';
65 let session = '';
66 try { cwd = await io.cwd(); session = await io.sessionId(); } catch { /* below */ }
67 // a project without an armed loop: these tokens are no run's (what 'no-loop' would say)
68 if (!(await hasGate(io, cwd, true))) return 'no-gate';
69 const r = await enqueue(mod, () => call(io, mod, 'usage-flush', { cwd, event: { session_id: session }, facts: { usage: { byAgent: delta } }, timeoutMs: FLUSH_TIMEOUT_MS }));
70 return settleUsage(mod, delta, r);
71}
72
73// After a request: a flush within USAGE_FLUSH_MS, one timer at a time.
74export function scheduleUsage(io, mod) {
75 if (mod.usageTimer) return;
76 mod.usageTimer = io.after(USAGE_FLUSH_MS, () => {
77 mod.usageTimer = null;
78 // tracked: a stop that comes meanwhile waits for it (its delta may come back to memory)
79 return track(mod, () => flushUsage(io, mod)).catch(() => 'error');
80 });
81}
82
83// A stop carries the delta itself: no flush of it is left pending.
84export function cancelUsageTimer(mod) {
85 if (mod.usageTimer) { try { mod.usageTimer.cancel(); } catch { /* gone */ } }
86 mod.usageTimer = null;
87}
88hooks/lib/session.js 102 lines1// The start of a session.
2// session.start (once per process, before the first prompt): the Claude Code version the mod
3// runs on. One older than MIN_CLAUDE_CODE (or unreadable) is said under the prompt
4// ($.ui.status) and, where a loop is armed, in the journal (mod-start), which `status` shows.
5// A loop armed later gets the line at its first stop.
6// classic.SessionStart (startup, resume, clear, compact): the notice of session-start.mjs for
7// a loop this session does not own (or for the owner after a compaction), as additional
8// context. Not prompt.context: that event has neither the session id nor the source, and the
9// classic event's additionalContext reaches the model (verified on 2.1.289).
10// Both also leave the mod's sign of life for `arm` (verbs.js writeAlive: a /clear or a resume
11// brings a new session id, and classic.SessionStart has it), and session.start registers the
12// `perseveranza` tool and the /pf command (registered there, they are listed from the first
13// turn); at most once per process it also has the old signs of life pruned (pruneAliveIfNeeded).
14// Both fail-open: nothing they do may hold a session back. A tool that could not be registered
15// leaves the loop's instructions in the shell's words (the stop says loopMode 'shell').
16import { QUICK_TIMEOUT_MS, MAX_HELLO_TRIES, OWN_TOOL, isObj, versionCheck, errText } from './core.js';
17import { call, hasGate, skipped } from './gate.js';
18import { TOOL_NAME, COMMAND_NAME, toolSchema, toolDescription, cliCommand, writeAlive, pruneAliveIfNeeded } from './verbs.js';
19
20export const COMMAND_SPEC = {
21 name: COMMAND_NAME,
22 description: 'perseveranza loop verbs, run now: status, arm, disarm, report, claim-done, pause, resume, test, ask... (/perseveranza <task> starts a task)',
23 argumentHint: '<verb> [args]: status | arm "<task>" [flags] | disarm | resume [--takeover] | help',
24 // status and disarm must answer while Claude works (a runaway loop is stopped mid-turn); the
25 // flag is the command's, not a verb's: every verb runs at once, as the user typed it
26 immediate: true,
27};
28
29// The tool for Claude and the command for the user. Each registration on its own: a failure of
30// one (a name taken, an older Claude Code) leaves the other, and is said in the debug log.
31export async function registerVerbs(io, mod) {
32 try { mod.pluginName = String((await io.pluginName()) || ''); } catch { /* '' */ }
33 let root = '';
34 try { root = io.root(); } catch { /* the description names the CLI without it */ }
35 try {
36 const r = await io.registerTool({ name: TOOL_NAME, description: toolDescription(cliCommand(root || '<perseveranza>')), inputSchema: toolSchema() });
37 mod.toolName = r && typeof r.tool === 'string' && r.tool ? r.tool : OWN_TOOL;
38 } catch (err) {
39 mod.toolName = null;
40 try { await io.log(`perseveranza: the perseveranza tool could not be registered (${errText(err)}): the loop's instructions name the CLI`); } catch { /* best effort */ }
41 }
42 try {
43 const c = await io.registerCommand(COMMAND_SPEC);
44 mod.commandName = c && typeof c.command === 'string' ? c.command : COMMAND_NAME;
45 } catch (err) {
46 mod.commandName = null;
47 try { await io.log(`perseveranza: /${COMMAND_NAME} could not be registered (${errText(err)}): the CLI runs the verbs`); } catch { /* best effort */ }
48 }
49}
50
51// how the instructions may name the verbs in this process: the tool, once registered
52export const loopModeOf = (mod) => (mod.toolName ? 'tool' : 'shell');
53
54export async function checkVersion(io, mod) {
55 let v = null;
56 try { v = await io.version(); } catch { /* unreadable: said as such */ }
57 mod.hello = versionCheck(v);
58 if (!mod.hello.ok) {
59 try { await io.status(`perseveranza: Claude Code ${mod.hello.claudeCode} - ${mod.hello.problem}`); } catch { /* best effort */ }
60 }
61 return mod.hello;
62}
63
64// The mod-start line, once per process, in the journal of the loop in cwd (none yet: the first
65// stop of a live loop sends it). At most MAX_HELLO_TRIES tries per process.
66// The session goes with it: the bridge journals only for the loop's owner (op 'journal').
67export async function sendHello(io, mod, cwd, session = '') {
68 if (mod.helloSent || mod.helloTries >= MAX_HELLO_TRIES) return mod.helloSent;
69 mod.helloTries += 1;
70 if (!mod.hello) await checkVersion(io, mod);
71 const r = await call(io, mod, 'journal', { cwd, event: { session_id: typeof session === 'string' ? session : '' }, facts: { lines: [{ type: 'mod-start', ...mod.hello }] }, timeoutMs: QUICK_TIMEOUT_MS });
72 if (r.answer && r.answer.ok === true && r.answer.journaled > 0) mod.helloSent = true;
73 return mod.helloSent;
74}
75
76export async function onSessionStart(io, mod, e) {
77 await checkVersion(io, mod);
78 await registerVerbs(io, mod);
79 const cwd = isObj(e) && typeof e.cwd === 'string' ? e.cwd : '';
80 let session = isObj(e) && typeof e.session_id === 'string' ? e.session_id : '';
81 if (!session) { try { session = (await io.sessionId()) || ''; } catch { /* '' */ } }
82 if (await writeAlive(io, mod, session, cwd)) await pruneAliveIfNeeded(io, mod, session, cwd);
83 if (!(await hasGate(io, cwd, false))) return;
84 await sendHello(io, mod, cwd, session);
85}
86
87// -> the notice, or null
88export async function sessionNotice(io, mod, e) {
89 if (!isObj(e)) return null;
90 if (typeof e.transcript_path === 'string' && e.transcript_path) mod.transcript = e.transcript_path;
91 const cwd = typeof e.cwd === 'string' ? e.cwd : '';
92 // the id of this session as Bash will see it (CLAUDE_CODE_SESSION_ID): after a /clear, a new one
93 if (typeof e.session_id === 'string' && e.session_id) await writeAlive(io, mod, e.session_id, cwd);
94 if (!(await hasGate(io, cwd, false))) return null;
95 const r = await call(io, mod, 'session-start', { cwd, event: e, facts: { loopMode: loopModeOf(mod) }, timeoutMs: QUICK_TIMEOUT_MS });
96 if (!r.answer || r.answer.ok !== true) {
97 await skipped(io, mod, cwd, 'classic.SessionStart', r.error || String(r.answer && r.answer.error));
98 return null;
99 }
100 return typeof r.answer.context === 'string' && r.answer.context ? r.answer.context : null;
101}
102hooks/lib/verbs.js 487 lines1// The loop's verbs from the mod: a typed tool for Claude (mcp__perseveranza__perseveranza) and
2// the /pf command for the user. Neither reimplements a verb: both run the CLI
3// (src/cli/perseveranza.mjs) with $.process.run, an argument list and no shell, in the
4// session's own folder, so a verb has the same code, effects and tests whichever way it
5// comes. The CLI journals the way it came (PERSEVERANZA_VIA: `via: 'tool'` or 'command').
6//
7// The tool, for Claude. A mod's tool runs without a permission prompt, so it runs NOTHING the
8// user's Bash permissions would govern: only the verbs that read or move the loop's own state
9// (TOOL_VERBS), on a loop of this session (or one no session claimed yet):
10// - `test` (it runs the suite, a shell command) and `ask` (it starts an external agent CLI)
11// are refused: they stay shell commands, run with Bash, where the user's permissions decide;
12// - arm, disarm and resume (a takeover, resume --takeover, included) are the user's: /pf, or
13// the CLI when the user asks; the tool refuses them, and so does the CLI for a run that says
14// via=tool. resume is the user's because a pause is where the loop waits for a human: the
15// approval of a plan (--approve-plan) and the escalation after maxRetries. resume lifts the
16// pause and resets the retry counters, so a model that could resume its own loop would
17// approve its own plan and go past the limit that asks for a human;
18// - a loop owned by another session: every verb that changes something is refused (taking
19// a loop over is the user's /pf resume --takeover);
20// - its arguments are checked HERE, before any process starts: a closed list of verbs and of
21// arguments, each with its type and its values (Claude Code does not enforce the schema: on
22// 2.1.289 a verb outside the enum and an extra argument reached the hook). The words of a
23// loop instruction go as they are ({"verb": "report", "args": "pass"}) or as typed fields
24// ({"verb": "report", "outcome": "pass"}): both read the same way, and disagreeing is an error;
25// - no loop armed in this folder (the bridge's rule, gate.js hasGate): every verb but
26// `status` is refused with the reason, and no node process starts;
27// - no argument is a path: the verb runs in the session's folder and nowhere else;
28// - what it answers is the verb's output, cut to MAX_OUTPUT, with the exit code said in words;
29// a refusal or a failure to start is an error result ({ deny }), never a throw: a broken
30// tool does not break the session, and the CLI stays the named fallback.
31// The command, for the user: `/pf <verb> [args]`, the tool's verbs plus arm, disarm, test, ask,
32// a takeover and runs; its words split like a shell would split them (quotes) but never given
33// to one; a verb that takes no words refuses extra ones (`/pf disarm the alarm` disarms
34// nothing). Claude cannot run it (Claude Code refuses a mod's command from the Skill tool:
35// "a built-in CLI command, not a skill"), and the verbs that start something or end a loop
36// run only for what the user typed (origin composer, bridge or sdk), never for another plugin.
37// (Named /pf, not /perseveranza: a registered command replaces the markdown command of the
38// same name, and `/perseveranza <task>` stays the markdown one that starts a task.)
39import { COMPLEXITIES } from '../../src/core/state.mjs';
40import { isObj, errText, READ_ONLY_VERBS } from './core.js';
41import { hasGate, setup, readOwner } from './gate.js';
42import { toolInput, askGuard } from './tool.js';
43import { noteTool, scheduleActivity } from './activity.js';
44
45export const TOOL_NAME = 'perseveranza';
46export const COMMAND_NAME = 'pf';
47// the verbs Claude may run through the tool: the loop's own state, nothing else. pause is one
48// (a model that stops itself to ask the user hands control to a human: nothing is bypassed);
49// report and claim-done are the loop's own signals, judged by the machine (see the README,
50// "Security"): neither lifts a pause, and on a paused loop the CLI refuses both (an accepted
51// pass or claim resets the retries: not an outcome to store while a human is awaited)
52export const TOOL_VERBS = ['status', 'history', 'explain', 'report', 'complexity', 'claim-done', 'pause'];
53// the verbs that run something the user's Bash permissions govern: never through the tool
54export const SHELL_VERBS = ['test', 'ask'];
55// the user's: arm, disarm, resume (a pause waits for a human; a takeover is an argument of it)
56export const USER_VERBS = ['arm', 'disarm', 'resume'];
57// the verbs /pf runs (the user's): the tool's, the shell's, the user's, the archive
58export const COMMAND_VERBS = [...TOOL_VERBS, ...SHELL_VERBS, ...USER_VERBS, 'runs'];
59// the verbs of the command that act on an armed loop (refused without one, like the tool's)
60export const GATED_COMMAND_VERBS = ['report', 'complexity', 'claim-done', 'pause', 'resume', 'test', 'ask'];
61// where a run of /pf comes from (command.run's e.origin.kind) when the user typed it: Enter at
62// the prompt, the Remote Control bridge, the host of `claude -p`. Another plugin's
63// $.command.run reads 'plugin', and the verbs below are not run for it.
64export const USER_ORIGINS = ['composer', 'bridge', 'sdk'];
65export const ORIGIN_VERBS = ['arm', 'disarm', 'test', 'ask'];
66export const MAX_ARGS = 200;
67export const MAX_TAIL = 500;
68export const MAX_OUTPUT = 20000;
69export const MAX_COMMAND_ARGS = 8000;
70export const MAX_TOKENS = 64;
71// $.process.run allows ten minutes at most: a suite or a provider longer than that is killed
72export const LONG_TIMEOUT_MS = 600000;
73export const VERB_TIMEOUT_MS = 60000;
74export const ARM_TIMEOUT_MS = 180000;
75
76// The tool's typed arguments: the verb each belongs to and its JSON schema. `args` (the words
77// after the verb, as a loop instruction writes them) is the other way to say the same.
78export const TOOL_ARGS = {
79 args: { verbs: TOOL_VERBS, schema: { type: 'string', maxLength: MAX_ARGS, description: 'the words after the verb in the loop instruction, as written: "pass" for `report pass`, "low" for `complexity low`, "--tail 20" for `history --tail 20`; leave it out when the verb has none' } },
80 outcome: { verbs: ['report'], schema: { type: 'string', enum: ['pass', 'fail'], description: 'report: the outcome (the same as args "pass" or "fail")' } },
81 level: { verbs: ['complexity'], schema: { type: 'string', enum: [...COMPLEXITIES], description: 'complexity: the task complexity (the same as args "low", "medium" or "high")' } },
82 tail: { verbs: ['history'], schema: { type: 'integer', minimum: 1, maximum: MAX_TAIL, description: 'history: only the last N entries (the same as args "--tail N")' } },
83};
84
85export function toolSchema() {
86 const properties = { verb: { type: 'string', enum: TOOL_VERBS, description: 'the verb: the first word after "the `perseveranza` tool" in a loop instruction' } };
87 for (const [k, a] of Object.entries(TOOL_ARGS)) properties[k] = a.schema;
88 return { type: 'object', properties, required: ['verb'], additionalProperties: false };
89}
90
91// What Claude reads about the tool: the exact forms, and what it does not do.
92export function toolDescription(cli) {
93 return [
94 'The verbs of the perseveranza loop armed in this project (.perseveranza/), for the instructions that say "the `perseveranza` tool": the first word is "verb", the words after it go whole in "args".',
95 'Exactly: `report pass` -> {"verb": "report", "args": "pass"}; `report fail` -> {"verb": "report", "args": "fail"}; `complexity low` -> {"verb": "complexity", "args": "low"}; `claim-done` -> {"verb": "claim-done"}; `pause`, `status`, `explain` -> {"verb": "<it>"}; `history --tail 20` -> {"verb": "history", "args": "--tail 20"}.',
96 'It answers with the verb\'s output and its exit code; a verb that refuses (claim-done without a green test, a verb without an armed loop) says why.',
97 `It does NOT run the suite (test) or an external model (ask): those are shell commands, run with Bash: ${cli} test --if-needed -- <the suite>, ${cli} ask <provider> <slot> -- "<prompt>".`,
98 'arm, disarm and resume are the user\'s, and so is taking over a loop of another session (resume --takeover): a paused loop (a plan to approve, an escalation) waits for the user, who types /pf resume; the user types /pf arm|disarm|resume --takeover too.',
99 `Only if this tool is missing or cannot start, run the same verb as a shell command: ${cli} <verb> <args>.`,
100 ].join(' ');
101}
102
103const CONTROL = /[\u0000-\u0008\u000b-\u001f\u007f]/;
104const NUMBER = /^\d{1,6}$/;
105
106function asInt(v) {
107 const n = typeof v === 'number' ? v : typeof v === 'string' && NUMBER.test(v.trim()) ? Number(v.trim()) : NaN;
108 return Number.isInteger(n) ? n : null;
109}
110
111// A run's id as `runs list` prints it, <project>/<stamp>, or the stamp alone: each part the
112// archive's own characters (archive.mjs safe(): letters, digits, '.', '_', '-'), not starting
113// with a dot (so never '.' or '..'), one '/' at most, never '\', ':' or a leading '/'. The CLI
114// only looks the id up among the listed runs; this keeps the words of /pf to what it can match.
115const RUN_PART = /^[A-Za-z0-9_-][A-Za-z0-9._-]{0,119}$/;
116export function isRunId(id) {
117 if (typeof id !== 'string') return false;
118 const parts = id.split('/');
119 return parts.length <= 2 && parts.every((p) => RUN_PART.test(p));
120}
121
122// The words of a verb that takes few -> { argv } | { error }. Pure. Used by the tool (every verb
123// it runs) and by /pf (every verb but arm, test and ask, whose words are the CLI's to check).
124// opts.user: the user's command (resume --takeover and disarm --no-archive are allowed)
125export function verbArgs(verb, tokens, opts = {}) {
126 const t = Array.isArray(tokens) ? tokens : [];
127 const none = (allowed = []) => (t.every((w) => allowed.includes(w)) && new Set(t).size === t.length ? { argv: [verb, ...t] } : { error: `${verb} takes ${allowed.length ? `only ${allowed.join(', ')}` : 'no words'} (given: ${t.join(' ').slice(0, 80)})` });
128 switch (verb) {
129 case 'status': return none(['--json']);
130 case 'explain': return none(['--markdown']);
131 case 'claim-done': case 'pause': return none();
132 case 'resume': return none(opts.user ? ['--takeover'] : []);
133 case 'disarm': return none(['--no-archive']);
134 case 'report':
135 return t.length === 1 && ['pass', 'fail'].includes(t[0]) ? { argv: ['report', t[0]] } : { error: 'report takes one word: pass or fail' };
136 case 'complexity':
137 return t.length === 1 && COMPLEXITIES.includes(t[0]) ? { argv: ['complexity', t[0]] } : { error: `complexity takes one word: ${COMPLEXITIES.join(', ')}` };
138 case 'history': {
139 const argv = ['history'];
140 let tail = false;
141 let json = false;
142 for (let i = 0; i < t.length; i++) {
143 const w = t[i];
144 const n = w === '--tail' ? asInt(t[++i]) : NUMBER.test(w) && !opts.user ? asInt(w) : null;
145 if ((w === '--tail' || NUMBER.test(w)) && n !== null && !tail) {
146 if (n < 1 || n > MAX_TAIL) return { error: `history: the tail must be a whole number from 1 to ${MAX_TAIL}` };
147 argv.push('--tail', String(n));
148 tail = true;
149 } else if (w === '--json' && !json) { argv.push('--json'); json = true; } else return { error: `history takes --tail N and --json (given: ${t.join(' ').slice(0, 80)})` };
150 }
151 return { argv };
152 }
153 case 'runs': {
154 if (t.length === 0 || (t.length === 1 && t[0] === 'list')) return { argv: ['runs', ...t] };
155 if (t[0] === 'show' && isRunId(t[1]) && (t.length === 2 || (t.length === 3 && t[2] === '--all'))) return { argv: ['runs', ...t] };
156 return { error: 'runs takes: list | show <id> [--all]' };
157 }
158 default: return { error: `unknown verb "${String(verb).slice(0, 40)}"` };
159 }
160}
161
162// The tool's input (tool.call's event) -> { verb, argv } | { error, shell?, user? }. Pure.
163export function validateToolInput(e) {
164 if (!isObj(e)) return { error: 'the input is not an object' };
165 const input = toolInput(e);
166 if (typeof input.verb !== 'string' || !input.verb.trim()) return { error: `"verb" is missing: one of ${TOOL_VERBS.join(', ')}` };
167 // "report pass" in the verb: the first word is the verb, the rest its words
168 const head = splitArgs(input.verb);
169 if (head.error) return { error: `"verb": ${head.error}` };
170 const [verb, ...inVerb] = head.tokens;
171 // refused by kind whatever the case ("Resume", "ARM"): the answer names the door that runs it
172 const kind = verb.toLowerCase();
173 if (SHELL_VERBS.includes(kind)) return { error: `"${kind}" does not run through the perseveranza tool`, shell: kind };
174 if (kind === 'resume') {
175 // a takeover in any form (the field, the words in args or in the verb) is named as one
176 const sp = typeof input.args === 'string' ? splitArgs(input.args) : { tokens: [] };
177 const takeover = Object.prototype.hasOwnProperty.call(input, 'takeover') || [...inVerb, ...(sp.tokens || [])].includes('--takeover');
178 return { error: takeover ? 'a takeover is the user\'s' : 'resume is the user\'s, not the tool\'s', user: takeover ? 'resume --takeover' : 'resume' };
179 }
180 if (USER_VERBS.includes(kind)) return { error: `"${kind}" is the user's, not the tool's`, user: kind };
181 if (!TOOL_VERBS.includes(verb)) return { error: `unknown verb "${verb.slice(0, 40)}": one of ${TOOL_VERBS.join(', ')}` };
182 let tokens = inVerb;
183 for (const k of Object.keys(input)) {
184 if (k === 'verb') continue;
185 // own keys only: "constructor" or "__proto__" is not an argument
186 const spec = Object.prototype.hasOwnProperty.call(TOOL_ARGS, k) ? TOOL_ARGS[k] : null;
187 if (k === 'takeover') return { error: 'a takeover is the user\'s', user: 'resume --takeover' };
188 if (!spec) return { error: `unknown argument "${k.slice(0, 40)}" (the tool takes: verb, ${Object.keys(TOOL_ARGS).join(', ')})` };
189 if (!spec.verbs.includes(verb)) return { error: `"${k}" is not an argument of ${verb}` };
190 }
191 if (input.args !== undefined && input.args !== null && input.args !== '') {
192 if (typeof input.args !== 'string') return { error: '"args" must be a string (the words after the verb)' };
193 if (input.args.length > MAX_ARGS) return { error: `"args" is longer than ${MAX_ARGS} characters` };
194 const sp = splitArgs(input.args);
195 if (sp.error) return { error: `"args": ${sp.error}` };
196 tokens = [...tokens, ...sp.tokens];
197 }
198 // a typed field says the same as the words, or is the only one to say it
199 const typed = (k, word) => {
200 if (input[k] === undefined) return null;
201 if (tokens.length === 0) { tokens = word; return null; }
202 return tokens.length === word.length && tokens.every((w, i) => w === word[i]) ? null : `"${k}" and the words disagree (${k}: ${String(input[k]).slice(0, 20)}, words: ${tokens.join(' ').slice(0, 60)})`;
203 };
204 if (input.outcome !== undefined && !['pass', 'fail'].includes(input.outcome)) return { error: '"outcome" must be one of pass, fail' };
205 if (input.level !== undefined && !COMPLEXITIES.includes(input.level)) return { error: `"level" must be one of ${COMPLEXITIES.join(', ')}` };
206 let conflict = null;
207 if (input.outcome !== undefined) conflict = typed('outcome', [input.outcome]);
208 if (input.level !== undefined) conflict = typed('level', [input.level]);
209 if (input.tail !== undefined) {
210 const n = asInt(input.tail);
211 if (n === null || n < 1 || n > MAX_TAIL) return { error: `"tail" must be a whole number from 1 to ${MAX_TAIL}` };
212 if (tokens.length) conflict = `"tail" and the words both say what to show (words: ${tokens.join(' ').slice(0, 60)})`;
213 else tokens = ['--tail', String(n)];
214 }
215 if (conflict) return { error: conflict };
216 const a = verbArgs(verb, tokens);
217 return a.error ? { error: a.error } : { verb, argv: a.argv };
218}
219
220// A valid tool call -> what the CLI runs: { argv (after the CLI path), stdin, timeoutMs }. Pure.
221export function toolRun({ argv }) {
222 return { argv, stdin: '', timeoutMs: VERB_TIMEOUT_MS };
223}
224
225// The words after /pf, split as a shell would split them (double quotes with \" and \\,
226// single quotes literal), never handed to one. -> { tokens } | { error }. Pure.
227export function splitArgs(text) {
228 const s = String(text ?? '');
229 if (s.length > MAX_COMMAND_ARGS) return { error: `the arguments are longer than ${MAX_COMMAND_ARGS} characters` };
230 if (CONTROL.test(s.replace(/[\r\n]/g, ' '))) return { error: 'the arguments hold control characters' };
231 const tokens = [];
232 let cur = '';
233 let has = false;
234 let quote = null;
235 for (let i = 0; i < s.length; i++) {
236 const c = s[i];
237 if (quote === "'") {
238 if (c === "'") quote = null; else cur += c;
239 } else if (quote === '"') {
240 if (c === '\\' && (s[i + 1] === '"' || s[i + 1] === '\\')) { cur += s[i + 1]; i++; } else if (c === '"') quote = null; else cur += c;
241 } else if (c === '"' || c === "'") { quote = c; has = true; } else if (/\s/.test(c)) {
242 if (has) { tokens.push(cur); cur = ''; has = false; }
243 } else { cur += c; has = true; }
244 }
245 if (quote) return { error: `a ${quote === '"' ? 'double' : 'single'} quote is not closed` };
246 if (has) tokens.push(cur);
247 if (tokens.length > MAX_TOKENS) return { error: `more than ${MAX_TOKENS} words` };
248 return { tokens };
249}
250
251// /pf's words -> { verb, argv, stdin, timeoutMs } | { help } | { error }. Pure.
252// arm, test and ask take the CLI's own words (the CLI checks them); every other verb is checked
253// here, word by word (verbArgs): `/pf disarm the legacy alarm` disarms nothing.
254export function commandRun(text) {
255 const sp = splitArgs(text);
256 if (sp.error) return { error: sp.error };
257 const [verb = 'status', ...rest] = sp.tokens;
258 if (verb === 'help' || verb === '--help' || verb === '-h') return { help: true };
259 if (!COMMAND_VERBS.includes(verb)) return { error: `unknown verb "${verb.slice(0, 40)}"`, unknown: true };
260 const timeoutMs = verb === 'test' || verb === 'ask' ? LONG_TIMEOUT_MS : verb === 'arm' ? ARM_TIMEOUT_MS : VERB_TIMEOUT_MS;
261 if (verb === 'arm' || verb === 'test' || verb === 'ask') return { verb, argv: [verb, ...rest], stdin: '', timeoutMs };
262 const a = verbArgs(verb, rest, { user: true });
263 if (a.error) return { error: a.error };
264 return { verb, argv: a.argv, stdin: '', timeoutMs };
265}
266
267export function commandHelp() {
268 return [
269 `/pf <verb> [args]: the perseveranza loop's verbs, run here and now (even while Claude is working).`,
270 ` status | history [--tail N] [--json] | explain | runs [list | show <id>]`,
271 ` arm "<task>" [--complexity low|medium|high] [--test "cmd"] [--max N] [--no-push] ... (the CLI's flags) disarm [--no-archive]`,
272 ` report pass|fail | complexity low|medium|high | claim-done | pause | resume [--takeover] | test [--if-needed] -- <cmd> | ask <provider> <slot> -- <prompt>`,
273 `To start a task with Claude (it arms the loop and writes the plan): /perseveranza <task>.`,
274 ].join('\n');
275}
276
277// The verb's run -> the text Claude (or the user) reads: the exit said in words, then the output,
278// cut to MAX_OUTPUT (the head and the tail: a verdict is at its end).
279export function formatRun(verb, r) {
280 const code = Number.isInteger(r.exitCode) ? r.exitCode : 1;
281 const out = [String(r.stdout || '').trimEnd(), r.stderr && String(r.stderr).trim() ? `[stderr]\n${String(r.stderr).trimEnd()}` : ''].filter(Boolean).join('\n');
282 const head = `perseveranza ${verb}: ${code === 0 ? 'done' : 'FAILED or REFUSED'} (exit ${code})${r.isStdoutTruncated ? ' [output truncated by Claude Code]' : ''}`;
283 return `${head}\n${cut(out, MAX_OUTPUT)}`.trimEnd();
284}
285
286export function cut(text, max) {
287 const t = String(text || '');
288 if (t.length <= max) return t;
289 const keepHead = Math.floor(max / 5);
290 const keepTail = max - keepHead;
291 return `${t.slice(0, keepHead)}\n[... ${t.length - max} characters cut ...]\n${t.slice(-keepTail)}`;
292}
293
294// an exit code Claude Code can take (0..255); anything else, a failure
295export const exitCodeOf = (r) => (Number.isInteger(r.exitCode) && r.exitCode >= 0 && r.exitCode <= 255 ? r.exitCode : 1);
296
297export const cliPath = (root) => `${String(root).replace(/[\\/]+$/, '')}/src/cli/perseveranza.mjs`;
298// the CLI as the fallback a loop instruction names
299export const cliCommand = (root) => `node "${cliPath(root)}"`;
300
301// One run of the CLI: [node, cli, ...argv], no shell. -> { exitCode, stdout, stderr } | { error }
302export async function runCli(io, mod, cwd, run, via, extraEnv = {}) {
303 try {
304 await setup(io, mod);
305 const r = await io.run([mod.node, cliPath(io.root()), ...run.argv], { cwd, env: { PERSEVERANZA_VIA: via, ...extraEnv }, stdin: run.stdin || '', timeoutMs: run.timeoutMs });
306 if (!r || typeof r !== 'object') return { error: 'no result from the process' };
307 return r;
308 } catch (err) { return { error: errText(err) }; }
309}
310
311const noLoop = (cwd, verb) => `perseveranza: no loop is armed in ${cwd} (.perseveranza/state.json is missing), so there is nothing to ${verb}. Only status runs without one; the user arms a loop with /pf arm "<task>" or /perseveranza <task>.`;
312const short = (id) => String(id || '').slice(0, 8) || 'unknown';
313
314// The tool's refusals that name another way: the shell for test and ask, the user for the rest.
315function toolRefusal(v, cli) {
316 if (v.shell) return `perseveranza tool: "${v.shell}" does not run through this tool, which runs nothing the user's Bash permissions would govern. Run it as a shell command with Bash: ${cli} ${v.shell === 'test' ? 'test --if-needed -- <the suite>' : 'ask <provider> <slot> -- "<prompt>"'}. Nothing was run.`;
317 if (v.user === 'resume') return 'perseveranza tool: resuming the loop is the user\'s decision, not the tool\'s: a paused loop waits for a human (a plan to approve, an escalation after the retries), and resume lifts the pause and resets the retry counters. Ask the user to type /pf resume, and do not resume it any other way. Nothing was run.';
318 if (v.user) return `perseveranza tool: ${v.user === 'resume --takeover' ? 'taking a loop over (resume --takeover)' : `"${v.user}"`} is the user's decision, not the tool's: ask the user to type /pf ${v.user}${v.user === 'resume --takeover' ? '' : ' ...'}. Nothing was run.`;
319 return `perseveranza tool: ${v.error}. Nothing was run.`;
320}
321
322// tool.call of the tool -> { result } | { deny }. Never throws.
323// 1. the arguments, checked here (nothing runs for a call that is not valid);
324// 2. the gate: no loop armed, no process (status excepted);
325// 3. for a verb that changes something: the loop must be this session's, or nobody's yet
326// (read with $.fs.read; unreadable: refused, the CLI with Bash stays);
327// 4. the reconciliation guard of a restored loop (op 'tool-check'): the same rule as a Bash
328// call of the CLI; fail-open like the guard;
329// 5. the verb, through the CLI; the call counts as the turn's activity (the heartbeat).
330export async function serveTool(io, mod, e) {
331 try {
332 let cli = 'node <perseveranza>/src/cli/perseveranza.mjs';
333 try { cli = cliCommand(io.root()); } catch { /* the placeholder */ }
334 const v = validateToolInput(e);
335 if (v.error) return { deny: toolRefusal(v, cli) };
336 const cwd = await io.cwd();
337 // the gate before any process: no loop, no node (status says so itself)
338 if (v.verb !== 'status' && !(await hasGate(io, cwd, true))) return { deny: noLoop(cwd, v.verb) };
339 let session = '';
340 let now = 0;
341 try { session = (await io.sessionId()) || ''; now = await io.now(); } catch { /* the record goes without them */ }
342 // a call of this tool is a tool call of the session too: its sign of life kept fresh
343 await refreshAlive(io, mod, session);
344 if (!READ_ONLY_VERBS.includes(v.verb)) {
345 const owner = await readOwner(io, cwd);
346 if (owner === null) return { deny: `perseveranza tool: the loop's owner could not be read (.perseveranza/state.json), so ${v.verb} is not run through the tool. Run it as a shell command with Bash: ${cli} ${v.argv.join(' ')}. Nothing was run.` };
347 if (owner && owner !== session) return { deny: `perseveranza tool: the loop in this folder belongs to session ${short(owner)}, not to this one (${short(session)}): the tool does not act on another session's loop. Do not touch .perseveranza/; if the user wants this session to take it over, they type /pf resume --takeover. Nothing was run.` };
348 const why = await askGuard(io, mod, cwd, session, e.tool, { verb: v.verb });
349 if (why) return { deny: why };
350 }
351 scheduleActivity(io, mod, noteTool(mod, e, now, session), now);
352 const r = await runCli(io, mod, cwd, toolRun(v), 'tool');
353 if (r.error) return { deny: `perseveranza tool: the ${v.verb} verb could not run (${r.error}). Run it as a shell command instead: ${cli} ${v.argv.join(' ')}` };
354 return { result: formatRun(v.verb, r) };
355 } catch (err) {
356 return { deny: `perseveranza tool failed (${errText(err)}). Run the verb as a shell command instead: node <perseveranza>/src/cli/perseveranza.mjs <verb> ...` };
357 }
358}
359
360// command.run of /pf -> { text, exitCode }. Never throws.
361export async function serveCommand(io, mod, e) {
362 try {
363 const c = commandRun(isObj(e) ? e.args : '');
364 if (c.help) return { text: commandHelp(), exitCode: 0 };
365 if (c.error) return { text: `perseveranza: ${c.error}. Nothing was run.\n${commandHelp()}`, exitCode: 2 };
366 // what starts something or ends a loop runs only for what the user typed
367 const origin = isObj(e) && isObj(e.origin) && typeof e.origin.kind === 'string' ? e.origin.kind : '';
368 const sensitive = ORIGIN_VERBS.includes(c.verb) || (c.verb === 'resume' && c.argv.includes('--takeover'));
369 if (sensitive && !USER_ORIGINS.includes(origin)) return { text: `perseveranza: /pf ${c.verb} runs only when the user types it (this run came from ${origin ? `"${origin.slice(0, 40)}"` : 'an unknown origin'}). Nothing was run.`, exitCode: 1 };
370 const cwd = await io.cwd();
371 if (GATED_COMMAND_VERBS.includes(c.verb) && !(await hasGate(io, cwd, true))) return { text: noLoop(cwd, c.verb), exitCode: 1 };
372 const extra = {};
373 if (c.verb === 'arm') {
374 // arm checks that the mod is alive in this session: it is (this is the mod), so its sign
375 // of life is written first, and arm gets the session id Claude Code gives Bash
376 let session = '';
377 try { session = (await io.sessionId()) || ''; } catch { /* arm says what it misses */ }
378 if (session) { await writeAlive(io, mod, session, cwd); extra.CLAUDE_CODE_SESSION_ID = session; }
379 }
380 const r = await runCli(io, mod, cwd, c, 'command', extra);
381 if (r.error) return { text: `perseveranza: the ${c.verb} verb could not run (${r.error}). From a terminal: ${cliCommand(io.root())} ${c.verb} ...`, exitCode: 1 };
382 return { text: formatRun(c.verb, r), exitCode: exitCodeOf(r) };
383 } catch (err) {
384 return { text: `perseveranza: the command failed (${errText(err)}).`, exitCode: 1 };
385 }
386}
387
388// The mod's sign of life for `arm` (src/shell/mod-alive.mjs): <home>/mod-alive/<session>.json,
389// home as Node's: PERSEVERANZA_HOME, else ~/.perseveranza (USERPROFILE, else HOME: the order of
390// os.homedir() on Windows). Written with $.fs.write (no process); if that is refused, through
391// the bridge. Only where the mod can drive a loop: $.process.run is "CLI only" (the types), and
392// no call says whether this session has it, so the surface does (cliSurface). -> true when written.
393export const SESSION_ID_RE = /^[A-Za-z0-9][A-Za-z0-9_-]{0,127}$/;
394export async function aliveHome(io) {
395 const get = async (f) => { try { const v = await f(); return typeof v === 'string' ? v.trim() : ''; } catch { return ''; } };
396 const own = await get(io.perseveranzaHome);
397 if (own) return own.replace(/[\\/]+$/, '');
398 const base = (await get(io.userProfile)) || (await get(io.home));
399 return base ? `${base.replace(/[\\/]+$/, '')}/.perseveranza` : '';
400}
401
402// Does this session run where $.process.run is? The CLI: `claude` draws on the terminal
403// ($.session.surfaces() lists 'terminal' first), `claude -p` and the SDK draw nowhere (empty).
404// A session drawn only by the Desktop app or the VS Code extension is not the CLI the types
405// promise $.process.run to: no sign of life there, so `arm` keeps the CLI's words (shell mode).
406// A surfaces() that fails says nothing: no sign of life either (arm refuses, --no-mod-check).
407export async function cliSurface(io) {
408 try {
409 const s = await io.surfaces();
410 return Array.isArray(s) && (s.length === 0 || s.includes('terminal'));
411 } catch { return false; }
412}
413
414// opts.bridge false: no fallback through the bridge (a refresh never starts a process)
415export async function writeAlive(io, mod, session, cwd, opts = {}) {
416 if (typeof session !== 'string' || !SESSION_ID_RE.test(session)) return false;
417 if (!(await cliSurface(io))) {
418 try { await io.log(`perseveranza: no sign of life for session ${session.slice(0, 8)}: this session is not drawn by the CLI (Desktop, VS Code), where the mod cannot run the loop; \`arm\` here keeps the CLI's words (--no-mod-check)`); } catch { /* best effort */ }
419 return false;
420 }
421 let at = 0;
422 try { at = await io.now(); } catch { /* 0 */ }
423 const info = { session, at, claudeCode: mod.hello ? mod.hello.claudeCode : undefined, plugin: mod.pluginName || undefined, cwd: typeof cwd === 'string' ? cwd.slice(0, 300) : undefined };
424 const home = await aliveHome(io);
425 if (home) {
426 try { await io.write(`${home}/mod-alive/${session}.json`, JSON.stringify(info)); mod.alive[session] = at; return true; } catch { /* the bridge below */ }
427 }
428 if (opts.bridge === false) return false;
429 try {
430 await setup(io, mod);
431 const r = await io.run([mod.node, mod.bridge], { cwd: typeof cwd === 'string' && cwd ? cwd : undefined, stdin: JSON.stringify({ op: 'alive', cwd: cwd || '.', facts: info }), timeoutMs: 8000 });
432 const a = JSON.parse(String(r.stdout || '').trim());
433 if (a && a.ok === true) { mod.alive[session] = at; return true; }
434 } catch { /* said below */ }
435 try { await io.log(`perseveranza: could not write the sign of life of session ${session.slice(0, 8)}: \`arm\` in this session will refuse (--no-mod-check arms anyway)`); } catch { /* best effort */ }
436 return false;
437}
438
439// The sign of life kept fresh while the session works: written at its start only, a long-lived
440// session's file would grow older than every `claude -p` child started after it, and a prune by
441// count would take it (seen by the verification of phase 3: a REPL session lost its file, and
442// its `arm` refused). So every tool call of the session (the Bash call that runs `arm`
443// included: the hook runs before the tool) rewrites it, at most once per ALIVE_REFRESH_MS, with
444// $.fs.write only (no process). A prune never takes a file younger than a day (mod-alive.mjs),
445// so a session that called a tool in the last day always has its file. Never throws.
446export const ALIVE_REFRESH_MS = 10 * 60 * 1000;
447export async function refreshAlive(io, mod, session) {
448 try {
449 if (typeof session !== 'string' || !SESSION_ID_RE.test(session)) return false;
450 const now = await io.now();
451 const last = mod.alive[session];
452 if (typeof last === 'number' && now - last < ALIVE_REFRESH_MS) return false;
453 // tried once per window, written or not (a Desktop session would log at every call)
454 mod.alive[session] = now;
455 let cwd;
456 try { cwd = await io.cwd(); } catch { /* the file goes without it */ }
457 return await writeAlive(io, mod, session, cwd, { bridge: false });
458 } catch { return false; }
459}
460
461// The signs of life pile up (one per session, `claude -p` children included) and the mod cannot
462// delete a file ($.fs has no unlink): at a session start it lists them ($.fs.list, no process)
463// and, only past ALIVE_PRUNE_AT files or with one older than ALIVE_MAX_AGE_MS, asks the bridge
464// to prune (op 'alive', facts.prune: its own files only, the newest ALIVE_KEEP and every file
465// younger than ALIVE_PROTECT_MS kept, never this
466// session's). Once per process. Never throws. -> true when the bridge was asked.
467export const ALIVE_PRUNE_AT = 200;
468export const ALIVE_MAX_AGE_MS = 30 * 24 * 3600 * 1000;
469export async function pruneAliveIfNeeded(io, mod, session, cwd) {
470 if (mod.alivePruned) return false;
471 mod.alivePruned = true;
472 try {
473 const home = await aliveHome(io);
474 if (!home) return false;
475 const entries = await io.list(`${home}/mod-alive`);
476 if (!Array.isArray(entries)) return false;
477 let now = 0;
478 try { now = await io.now(); } catch { return false; }
479 const own = entries.filter((x) => isObj(x) && x.kind === 'file' && typeof x.name === 'string' && /\.json$/.test(x.name) && SESSION_ID_RE.test(x.name.slice(0, -5)));
480 const old = own.some((x) => typeof x.mtimeMs === 'number' && x.mtimeMs > 0 && now - x.mtimeMs > ALIVE_MAX_AGE_MS);
481 if (own.length <= ALIVE_PRUNE_AT && !old) return false;
482 await setup(io, mod);
483 await io.run([mod.node, mod.bridge], { cwd: typeof cwd === 'string' && cwd ? cwd : undefined, stdin: JSON.stringify({ op: 'alive', cwd: cwd || '.', facts: { prune: true, session } }), timeoutMs: 8000 });
484 return true;
485 } catch { return false; }
486}
487src/core/subagents.mjs 89 lines1// The loop's subagents seen from the mod (Claude Code's in-process hooks): which model a
2// pf-* subagent runs on, and whether a judge about to stop left the verdict the loop asked
3// for. Pure, and no `node:` import: the mod imports this module and has no Node API.
4//
5// routeModel(state, subagentType) -> 'haiku'|'sonnet'|'opus'|null
6// subagentVerdictCheck(state, agentType, artifacts, opts) -> { ok, reason, file }
7//
8// `artifacts` has the shape of the Stop hook's ctx.artifacts: { review, verify, verifyLenses:
9// { <lens>: text } }, each a file's text or null when the file is not there.
10
11import { MODEL_ROUTING, loopAgentName, runningLoopAgents, cleanRequestId, LATE_TOLERANCE_MS, lensFileName } from './machine.mjs';
12import { parseReviewVerdict, parseVerifyVerdict } from './verdicts.mjs';
13import { COMPLEXITIES, singleLens } from './state.mjs';
14
15export { loopAgentName, runningLoopAgents };
16
17// A judge is sent back to write its verdict at most this many times per request; then it is
18// let go and the machine's own `missing` outcome is the net.
19export const MAX_VERDICT_ASKS = 2;
20
21const ROUTES = { 'pf-reviewer': 'review', 'pf-verifier': 'verify', 'pf-executor': 'execute' };
22
23// The model of a loop subagent by the recorded complexity (medium when unknown); null for
24// any other agent, the advisor included (its model is an option of the run, not a route).
25export function routeModel(state, subagentType) {
26 const route = ROUTES[loopAgentName(subagentType)];
27 if (!route) return null;
28 const c = state && COMPLEXITIES.includes(state.complexity) ? state.complexity : 'medium';
29 return MODEL_ROUTING[route][c] || null;
30}
31
32// One verdict file against the current request, by the machine's rule: an id must be the
33// current one; without an id the file must not predate the request (one second of tolerance).
34// -> 'valid' | 'missing' | 'stale' | 'malformed: <error>'
35function judgeFile(state, text, at, parse) {
36 if (text == null) return 'missing';
37 const v = parse(text);
38 if (!v.ok) return `malformed: ${v.error}`;
39 const id = cleanRequestId(v.requestId);
40 if (state.verdictRequestId && id) return id === state.verdictRequestId ? 'valid' : 'stale';
41 const t = Number(at) || 0;
42 const requested = Number(state.verdictRequestedAt) || 0;
43 return requested > 0 && t > 0 && t + LATE_TOLERANCE_MS < requested ? 'stale' : 'valid';
44}
45
46// opts: { askedTimes: how many times this judge was already sent back for this request,
47// artifactAt: the files' mtimes (same shape as artifacts), lens: the lens this
48// verifier was given, when the caller knows it }
49// A judge stopping outside its phase, or with no request id to answer, is none of this
50// check's business (ok). In a round by
51// lenses without `lens`, a verifier cannot be told from its siblings: any valid file of the
52// round lets it go (the machine still asks for the lenses that are missing).
53export function subagentVerdictCheck(state, agentType, artifacts, opts = {}) {
54 const name = loopAgentName(agentType);
55 if (name !== 'pf-reviewer' && name !== 'pf-verifier') return { ok: true, reason: 'not-a-judge', file: null };
56 const s = state && typeof state === 'object' ? state : {};
57 const awaited = name === 'pf-reviewer' ? 'review' : 'final-verify';
58 if (s.phase !== awaited) return { ok: true, reason: 'not-awaited', file: null };
59 // no request id (a state from before the ids): nothing to bind a verdict to, so nothing to
60 // send the judge back for (no prompt ever hands out an empty id)
61 if (!s.verdictRequestId) return { ok: true, reason: 'no-request', file: null };
62 const a = artifacts && typeof artifacts === 'object' ? artifacts : {};
63 const at = opts.artifactAt && typeof opts.artifactAt === 'object' ? opts.artifactAt : {};
64 const asked = Math.max(0, Number(opts.askedTimes) || 0);
65
66 // [file, result] candidates in the order they answer the request
67 let checks;
68 if (name === 'pf-reviewer') checks = [['review.json', judgeFile(s, a.review, at.review, parseReviewVerdict)]];
69 else {
70 const raw = a.verifyLenses && typeof a.verifyLenses === 'object' ? a.verifyLenses : {};
71 const rawAt = at.verifyLenses && typeof at.verifyLenses === 'object' ? at.verifyLenses : {};
72 const lens = (l) => [lensFileName(l), judgeFile(s, raw[l], rawAt[l], parseVerifyVerdict)];
73 const main = ['verify.json', judgeFile(s, a.verify, at.verify, parseVerifyVerdict)];
74 const lenses = Array.isArray(s.verdictLenses) ? s.verdictLenses : [];
75 if (singleLens(lenses)) checks = [main, lens('general')];
76 else if (typeof opts.lens === 'string' && lenses.includes(opts.lens)) {
77 const own = lens(opts.lens);
78 // verify.json covers a lens that wrote no file at all (the machine's rule)
79 checks = own[1] === 'missing' ? [own, main] : [own];
80 } else checks = [...lenses.map(lens), main];
81 }
82 const good = checks.find(([, r]) => r === 'valid');
83 if (good) return { ok: true, reason: 'valid', file: good[0] };
84 // the most telling problem: a file that is there but wrong, before a file that is not there
85 const bad = checks.find(([, r]) => r !== 'missing') || checks[0];
86 if (asked >= MAX_VERDICT_ASKS) return { ok: true, reason: `asked-enough: ${bad[1]}`, file: bad[0] };
87 return { ok: false, reason: bad[1], file: bad[0] };
88}
89src/core/state.mjs 409 lines1// Loop state: schema v2, defaults, normalisation and migration from the v1 flat layout.
2// Pure: no filesystem access. The shell reads/writes the JSON; this module says what it means.
3//
4// Ownership (the contract the v1 comments described, now visible in the shape):
5// phase, counters, flags, owner, usage, tree -> written only by the Stop hook
6// signals, lastTest -> written only by the verbs
7// options, limits -> written only by `arm` (and adaptive budget once)
8// with the exceptions VERB_OWNED lists (complexity, resume's counters and owner release, the
9// test command, the watchdog's interruption). rev: incremented by every save, whoever
10// writes, so a writer can tell that the state changed under it (shell/state-file.mjs).
11
12export const SCHEMA_VERSION = 2;
13export const PHASES = ['plan', 'implement', 'review', 'cleanup', 'final-verify', 'git-finish'];
14export const COMPLEXITIES = ['low', 'medium', 'high'];
15export const DEFAULT_MAX_ITERATIONS = 25;
16export const DEFAULT_MAX_RETRIES = 3;
17// The lenses of the final verification: `general` is the single verifier of old; the others
18// split the adversarial mandate among verifiers that run side by side (arm --verifiers).
19export const LENSES = ['general', 'correctness', 'security', 'tests'];
20export const AUTO_LENSES_HIGH = ['correctness', 'security', 'tests'];
21// the rejected rounds whose findings the next final verification rechecks
22export const MAX_PRIOR_VERIFIES = 3;
23// the rejected reviews of the current step, handed to the advisor of the fix
24export const MAX_PRIOR_REVIEWS = 10;
25// The internal advisor (agents/pf-advisor.md): a second opinion that never routes the loop.
26export const DEFAULT_ADVISOR_MODEL = 'opus';
27// a model name as the Agent tool takes it (opus, sonnet, claude-opus-4-1, opus[1m]...)
28export const ADVISOR_MODEL_RE = /^[A-Za-z0-9][A-Za-z0-9._:[\]-]{0,63}$/;
29export const normalizeAdvisorModel = (v) => (typeof v === 'string' && ADVISOR_MODEL_RE.test(v.trim()) ? v.trim() : DEFAULT_ADVISOR_MODEL);
30
31// Known lenses only, each once, in the order given. -> [] when nothing is left.
32export function normalizeLenses(raw) {
33 if (!Array.isArray(raw)) return [];
34 const out = [];
35 for (const l of raw) {
36 const k = String(l ?? '').trim().toLowerCase();
37 if (LENSES.includes(k) && !out.includes(k)) out.push(k);
38 }
39 return out;
40}
41
42// The lenses a final verification round asks for: the ones armed, else by complexity.
43export function effectiveLenses(s) {
44 const chosen = normalizeLenses(s && s.options && s.options.verifiers);
45 if (chosen.length) return chosen;
46 return s && s.complexity === 'high' ? [...AUTO_LENSES_HIGH] : ['general'];
47}
48
49// One verifier writing verify.json, as before the lenses: the round the machine reads alone.
50export const singleLens = (lenses) => !Array.isArray(lenses) || !lenses.length || (lenses.length === 1 && lenses[0] === 'general');
51
52export function defaultState(overrides = {}) {
53 const s = {
54 schemaVersion: SCHEMA_VERSION,
55 task: '',
56 phase: 'plan',
57 complexity: 'medium',
58 options: {
59 commitSteps: false,
60 gitFinish: true,
61 gitPush: true,
62 approvePlan: false,
63 testCmd: null,
64 externals: [],
65 lang: 'it',
66 // final verification lenses chosen at arm; null = automatic (by complexity, per round)
67 verifiers: null,
68 // the internal advisor at the plan and from the 2nd fix (arm --advisor, --advisor-model)
69 advisor: true,
70 advisorModel: DEFAULT_ADVISOR_MODEL,
71 // how the instructions name the verbs: 'tool' (the mod's `perseveranza` tool, set by
72 // `arm` when the mod is alive in the arming session) or 'shell' (the CLI command)
73 loopMode: 'shell',
74 // paths whose UNTRACKED files the work-tree fingerprint and the git finish leave out, beside
75 // the known tool state (git.mjs VOLATILE_PATHS): arm --ignore, validated there
76 fingerprintIgnore: [],
77 },
78 // staleGates: final passes in a row that did not cover the current tree (pass-stale)
79 // subagentWaits: stops answered with subagent-running (a pf-* subagent still running,
80 // seen only by the mod) for the current verdict request; a new request or phase resets it
81 // quietStops: stops in a row that asked for no new work (subagent-running, missing, idle)
82 counters: { iterations: 0, retries: 0, finalFails: 0, staleGates: 0, subagentWaits: 0, quietStops: 0 },
83 limits: { maxIterations: DEFAULT_MAX_ITERATIONS, maxIterationsExplicit: false, maxRetries: DEFAULT_MAX_RETRIES, maxTokens: null },
84 usage: { inputTokens: 0, outputTokens: 0, cacheReadTokens: 0, cacheCreationTokens: 0, source: null },
85 // resumedAt: when `resume` closed a pause (ms); consumed by the next fire so the gap
86 // it journals is marked as a human's pause, not a dead session
87 // interrupted: written by the watchdog when it kills and restores the session; the next
88 // Stop runs a read-only reconciliation before anything else
89 signals: { lastReport: 'none', claimedDone: false, paused: false, resumedAt: 0, interrupted: null },
90 flags: { repeated: false, cleanedOnce: false, planPresented: false, reconcileAsked: false },
91 lastTest: null,
92 baselineDirty: [],
93 // releasedFrom: the previous owner after `resume --takeover`, until the next fire claims
94 // transcriptPath: the session transcript (a sign of life); claudePid/claudeStartedAt: the
95 // Claude Code process driving the loop, for the watchdog's kill-and-restore
96 owner: { sessionId: null, lastFireAt: 0, releasedFrom: null, releasedAt: 0, transcriptPath: null, claudePid: 0, claudeStartedAt: null },
97 // the work tree as the hook last saw it: lets it notice a stop that changed nothing
98 tree: { fingerprint: null, iteration: 0 },
99 // when the phase that awaits a verdict (review, final-verify) was entered: a verdict file
100 // written before that instant answers an earlier request, not this one
101 verdictRequestedAt: 0,
102 // Correlation token copied into the reviewer/verifier verdict. Unlike file timestamps it
103 // also rejects an old agent that finishes after a replacement request was issued.
104 verdictRequestId: null,
105 // the code fingerprint the request pointed at (null: not computable, e.g. outside git): a
106 // final pass closes only the tree it judged
107 verdictTree: null,
108 // the lenses the current final verification round asked for (fixed when it is requested)
109 verdictLenses: [],
110 // the kept findings (verify-<n>.json) of the last rejected rounds: rechecked by the next
111 priorVerifies: [],
112 // the kept rejections (review-<n>.json) of the current step: the advisor of the fix reads
113 // every attempt that already failed, not only the last one
114 priorReviews: [],
115 // the files of the mod's usage inbox (.perseveranza/usage-inbox/) already added to usage
116 usageInboxSeen: [],
117 armedAt: null,
118 engineVersion: null,
119 rev: 0,
120 };
121 return deepMerge(s, overrides);
122}
123
124function deepMerge(base, patch) {
125 if (!patch || typeof patch !== 'object' || Array.isArray(patch)) return base;
126 const out = { ...base };
127 for (const [k, v] of Object.entries(patch)) {
128 if (v && typeof v === 'object' && !Array.isArray(v) && base[k] && typeof base[k] === 'object' && !Array.isArray(base[k])) {
129 out[k] = deepMerge(base[k], v);
130 } else if (v !== undefined) {
131 out[k] = v;
132 }
133 }
134 return out;
135}
136
137const num = (v, def) => (Number.isFinite(Number(v)) ? Number(v) : def);
138const bool = (v, def) => (typeof v === 'boolean' ? v : def);
139const USAGE_KEYS = ['inputTokens', 'outputTokens', 'cacheReadTokens', 'cacheCreationTokens'];
140// A token count is a finite non-negative integer. A value too large to count (1e308, a
141// string of digits, Infinity) saturates at MAX_TOKENS instead of falling back to 0: tokens
142// already spent never vanish from the budget. NaN, a negative or a non-number is 0.
143export const MAX_TOKENS = Number.MAX_SAFE_INTEGER;
144// one delta of the mod (one flush, one agent, one field) is never more than this: past it the
145// value is clamped and the caller journals it (usage-clamped)
146export const MAX_TOKEN_DELTA = 1e12;
147// A string counts only when it is plain decimal digits (with an optional fraction): what a
148// writer may quote of a JSON number. No sign, exponent, hex (0x10), binary (0b11), octal or
149// whitespace: Number() would read those, and a count that came out of a file must not.
150const DECIMAL = /^\d+(\.\d+)?$/;
151export function tokenCount(v) {
152 const n = typeof v === 'number' ? v : typeof v === 'string' && DECIMAL.test(v) ? Number(v) : NaN;
153 if (Number.isNaN(n) || n <= 0) return 0;
154 return n >= MAX_TOKENS ? MAX_TOKENS : Math.floor(n);
155}
156const tokenCounts = (o) => Object.fromEntries(USAGE_KEYS.map((k) => [k, tokenCount(o && o[k])]));
157const isObj = (v) => v && typeof v === 'object' && !Array.isArray(v);
158const MAX_USAGE_AGENTS = 16;
159// Stops a pf-* subagent still running may hold the loop (subagent-running) for one verdict
160// request (one phase, where none is pending): past it the usual logic applies, so the loop
161// never spends Claude Code's cap of text-only continuations on waiting (see MAX_QUIET_STOPS
162// in machine.mjs for the cap across waits and missing outcomes).
163export const MAX_SUBAGENT_WAITS = 3;
164
165// saturating: a sum never becomes Infinity (which a later read would turn into 0)
166const addCounts = (a, b) => Object.fromEntries(USAGE_KEYS.map((k) => [k, Math.min(MAX_TOKENS, (a ? a[k] : 0) + (b ? b[k] : 0))]));
167const spendOf = (c) => c.inputTokens + c.outputTokens;
168
169// The rows of a mod reading, folded to the cap: main, the biggest spenders, and the rest
170// summed under "other" (never dropped: the budget must see every token).
171function foldAgents(rows) {
172 const byKey = new Map();
173 for (const [k, c] of rows) byKey.set(k, addCounts(byKey.get(k), c));
174 const main = byKey.get('main');
175 const other = byKey.get('other');
176 const rest = [...byKey].filter(([k]) => k !== 'main' && k !== 'other').sort((a, b) => spendOf(b[1]) - spendOf(a[1]));
177 const room = MAX_USAGE_AGENTS - (main ? 1 : 0);
178 const fits = rest.length + (other ? 1 : 0) <= room;
179 const kept = fits ? rest : rest.slice(0, room - 1);
180 let folded = other || null;
181 if (!fits) for (const [, c] of rest.slice(room - 1)) folded = addCounts(folded, c);
182 return Object.fromEntries([...(main ? [['main', main]] : []), ...kept, ...(folded ? [['other', folded]] : [])]);
183}
184
185// state.usage: the totals the budget reads, plus the split by agent kind and
186// the subagents' share. Coerced here, not trusted: it comes from files Claude Code writes.
187// source 'mod': the mod measures every request by agent id, so byAgent IS the reading and the
188// totals are at least its sum (the budget counts every agent's tokens).
189export function normalizeUsage(raw) {
190 const u = isObj(raw) ? raw : {};
191 const out = { ...tokenCounts(u), source: typeof u.source === 'string' ? u.source.slice(0, 40) : null };
192 if (isObj(u.byAgent)) {
193 const rows = Object.entries(u.byAgent).filter(([, v]) => isObj(v)).map(([k, v]) => [String(k).slice(0, 80), tokenCounts(v)]);
194 if (out.source === 'mod') {
195 const sum = rows.reduce((acc, [, c]) => addCounts(acc, c), null);
196 if (sum) for (const k of USAGE_KEYS) out[k] = Math.max(out[k], sum[k]);
197 out.byAgent = foldAgents(rows);
198 } else out.byAgent = Object.fromEntries(rows.slice(0, MAX_USAGE_AGENTS));
199 }
200 if (isObj(u.subagents)) out.subagents = { files: Math.max(0, num(u.subagents.files, 0)), ...tokenCounts(u.subagents) };
201 if (u.partial === true) out.partial = true;
202 return out;
203}
204
205// The tokens the mod measured since its last flush ({ <agentId|'main'>: counts }) added to
206// the reading in state.usage. A reading of another source (the transcripts, before the mod
207// drove the run) becomes the "main" row it started from: nothing already spent is lost.
208// Every field of the delta is a count (tokenCount) clamped to MAX_TOKEN_DELTA; each clamp is
209// pushed to `clamped` ({ agent, field, value }) for the caller to journal. The totals only
210// grow: a hostile delta adds 0, never subtracts.
211// -> a normalized usage with source 'mod'
212export function mergeModUsage(prev, delta, clamped = null) {
213 const p = normalizeUsage(prev);
214 const base = new Map(p.source === 'mod' && p.byAgent ? Object.entries(p.byAgent)
215 : spendOf(p) || p.cacheReadTokens || p.cacheCreationTokens ? [['main', tokenCounts(p)]] : []);
216 let added = null;
217 for (const [k, v] of Object.entries(isObj(delta) ? delta : {})) {
218 if (!isObj(v)) continue;
219 const key = String(k).slice(0, 80) || 'main';
220 const c = Object.fromEntries(USAGE_KEYS.map((f) => {
221 const n = tokenCount(v[f]);
222 if (n > MAX_TOKEN_DELTA) {
223 if (Array.isArray(clamped)) clamped.push({ agent: key, field: f, value: String(v[f]).slice(0, 40) });
224 return [f, MAX_TOKEN_DELTA];
225 }
226 return [f, n];
227 }));
228 base.set(key, addCounts(base.get(key), c));
229 added = addCounts(added, c);
230 }
231 const subs = [...base].filter(([k]) => k !== 'main');
232 const subTotals = subs.reduce((acc, [, c]) => addCounts(acc, c), null);
233 return normalizeUsage({
234 ...addCounts(tokenCounts(p), added),
235 source: 'mod',
236 byAgent: Object.fromEntries(base),
237 ...(subTotals ? { subagents: { files: subs.length, ...subTotals } } : {}),
238 });
239}
240
241// Fill defaults and coerce types on a v2 object. Never throws on odd input.
242export function normalizeState(raw) {
243 const s = defaultState(raw && typeof raw === 'object' ? raw : {});
244 const defaults = defaultState();
245 for (const key of ['counters', 'limits', 'usage', 'signals', 'flags', 'options', 'owner', 'tree']) {
246 if (!s[key] || typeof s[key] !== 'object' || Array.isArray(s[key])) s[key] = defaults[key];
247 }
248 s.schemaVersion = SCHEMA_VERSION;
249 s.task = String(s.task ?? '');
250 if (!PHASES.includes(s.phase)) s.phase = 'plan';
251 if (!COMPLEXITIES.includes(s.complexity)) s.complexity = 'medium';
252 s.counters.iterations = Math.max(0, num(s.counters.iterations, 0));
253 s.counters.retries = Math.max(0, num(s.counters.retries, 0));
254 s.counters.finalFails = Math.max(0, num(s.counters.finalFails, 0));
255 s.counters.staleGates = Math.max(0, num(s.counters.staleGates, 0));
256 s.counters.subagentWaits = Math.max(0, num(s.counters.subagentWaits, 0));
257 s.counters.quietStops = Math.max(0, num(s.counters.quietStops, 0));
258 s.limits.maxIterations = num(s.limits.maxIterations, DEFAULT_MAX_ITERATIONS);
259 if (s.limits.maxIterations < 1) s.limits.maxIterations = DEFAULT_MAX_ITERATIONS;
260 s.limits.maxIterationsExplicit = bool(s.limits.maxIterationsExplicit, false);
261 s.limits.maxRetries = num(s.limits.maxRetries, DEFAULT_MAX_RETRIES);
262 if (s.limits.maxRetries < 1) s.limits.maxRetries = DEFAULT_MAX_RETRIES;
263 s.limits.maxTokens = s.limits.maxTokens == null ? null : Math.max(0, num(s.limits.maxTokens, 0)) || null;
264 s.usage = normalizeUsage(s.usage);
265 s.signals.lastReport = ['pass', 'fail'].includes(s.signals.lastReport) ? s.signals.lastReport : 'none';
266 s.signals.claimedDone = bool(s.signals.claimedDone, false);
267 s.signals.paused = bool(s.signals.paused, false);
268 s.signals.resumedAt = Math.max(0, num(s.signals.resumedAt, 0));
269 if (s.signals.interrupted && typeof s.signals.interrupted === 'object') {
270 const i = s.signals.interrupted;
271 s.signals.interrupted = { at: i.at ?? null, silentMs: Math.max(0, num(i.silentMs, 0)), phase: typeof i.phase === 'string' ? i.phase : '', pending: Array.isArray(i.pending) ? i.pending.map(String).slice(0, 10) : [] };
272 } else s.signals.interrupted = null;
273 s.flags.reconcileAsked = bool(s.flags.reconcileAsked, false);
274 s.flags.repeated = bool(s.flags.repeated, false);
275 s.flags.cleanedOnce = bool(s.flags.cleanedOnce, false);
276 s.flags.planPresented = bool(s.flags.planPresented, false);
277 s.options.commitSteps = bool(s.options.commitSteps, false);
278 s.options.gitFinish = bool(s.options.gitFinish, true);
279 s.options.gitPush = bool(s.options.gitPush, true);
280 s.options.approvePlan = bool(s.options.approvePlan, false);
281 s.options.testCmd = s.options.testCmd ? String(s.options.testCmd) : null;
282 s.options.externals = Array.isArray(s.options.externals) ? s.options.externals.map(String) : [];
283 s.options.lang = typeof s.options.lang === 'string' && s.options.lang ? s.options.lang : 'it';
284 const verifiers = normalizeLenses(s.options.verifiers);
285 s.options.verifiers = verifiers.length ? verifiers : null;
286 s.options.advisor = bool(s.options.advisor, true);
287 s.options.advisorModel = normalizeAdvisorModel(s.options.advisorModel);
288 s.options.loopMode = s.options.loopMode === 'tool' ? 'tool' : 'shell';
289 // plain relative paths only (git.mjs ignorePath validates them again where they are used)
290 s.options.fingerprintIgnore = Array.isArray(s.options.fingerprintIgnore) ? s.options.fingerprintIgnore.filter((p) => typeof p === 'string' && p && p.length <= 200).slice(0, 50) : [];
291 s.baselineDirty = Array.isArray(s.baselineDirty) ? s.baselineDirty.map(String) : [];
292 if (s.lastTest && typeof s.lastTest === 'object') {
293 s.lastTest = {
294 cmd: String(s.lastTest.cmd ?? ''),
295 exitCode: num(s.lastTest.exitCode, 1),
296 iteration: num(s.lastTest.iteration, -1),
297 at: s.lastTest.at ?? null,
298 fingerprint: s.lastTest.fingerprint ?? null,
299 // the same snapshot without documentation files (null on states written before 2.1)
300 codeFingerprint: s.lastTest.codeFingerprint ?? null,
301 // names of the tests that failed, as far as the runner's output could be parsed
302 failed: Array.isArray(s.lastTest.failed) ? s.lastTest.failed.map(String) : [],
303 };
304 } else s.lastTest = null;
305 s.verdictRequestedAt = Math.max(0, num(s.verdictRequestedAt, 0));
306 s.verdictRequestId = typeof s.verdictRequestId === 'string' && s.verdictRequestId ? s.verdictRequestId : null;
307 s.verdictTree = typeof s.verdictTree === 'string' && s.verdictTree ? s.verdictTree : null;
308 s.verdictLenses = normalizeLenses(s.verdictLenses);
309 s.priorVerifies = Array.isArray(s.priorVerifies)
310 ? s.priorVerifies.map(String).filter((n) => /^verify-\d+\.json$/.test(n)).slice(-MAX_PRIOR_VERIFIES)
311 : [];
312 s.priorReviews = Array.isArray(s.priorReviews)
313 ? s.priorReviews.map(String).filter((n) => /^review-\d+\.json$/.test(n)).slice(-MAX_PRIOR_REVIEWS)
314 : [];
315 s.usageInboxSeen = Array.isArray(s.usageInboxSeen) ? s.usageInboxSeen.map(String).filter((n) => /^[\w-]+\.json$/.test(n)).slice(-5000) : [];
316 s.rev = Math.max(0, Math.floor(num(s.rev, 0)));
317 s.tree.fingerprint = typeof s.tree.fingerprint === 'string' && s.tree.fingerprint ? s.tree.fingerprint : null;
318 s.tree.iteration = Math.max(0, num(s.tree.iteration, 0));
319 s.owner.sessionId = typeof s.owner.sessionId === 'string' && s.owner.sessionId ? s.owner.sessionId : null;
320 s.owner.lastFireAt = Math.max(0, num(s.owner.lastFireAt, 0));
321 s.owner.releasedFrom = typeof s.owner.releasedFrom === 'string' && s.owner.releasedFrom ? s.owner.releasedFrom : null;
322 s.owner.releasedAt = Math.max(0, num(s.owner.releasedAt, 0));
323 s.owner.transcriptPath = typeof s.owner.transcriptPath === 'string' && s.owner.transcriptPath ? s.owner.transcriptPath : null;
324 s.owner.claudePid = Math.max(0, num(s.owner.claudePid, 0));
325 s.owner.claudeStartedAt = typeof s.owner.claudeStartedAt === 'string' && s.owner.claudeStartedAt ? s.owner.claudeStartedAt : null;
326 return s;
327}
328
329// The fields the verbs (and the watchdog) write, as paths. The Stop saves the state it read at
330// its start plus its own decisions: a field of this list that changed on disk since that read
331// was changed by a verb meanwhile, and the verb's value wins (mergeVerbFields).
332export const VERB_OWNED = [
333 'signals.lastReport', 'signals.claimedDone', 'signals.paused', 'signals.resumedAt', 'signals.interrupted',
334 'lastTest', 'complexity', 'options.testCmd',
335 'counters.retries', 'counters.finalFails', 'counters.staleGates',
336 'flags.repeated', 'flags.reconcileAsked',
337 'owner.sessionId', 'owner.releasedFrom', 'owner.releasedAt',
338];
339const getPath = (o, p) => p.split('.').reduce((v, k) => (v && typeof v === 'object' ? v[k] : undefined), o);
340const setPath = (o, p, val) => { const ks = p.split('.'); const last = ks.pop(); let t = o; for (const k of ks) { if (!t[k] || typeof t[k] !== 'object') t[k] = {}; t = t[k]; } t[last] = val; };
341const same = (a, b) => JSON.stringify(a ?? null) === JSON.stringify(b ?? null);
342// ours: the state about to be saved; start: the state read at the start; disk: the state on
343// disk now. -> { state, taken: [paths taken from disk] }. Pure; never throws on odd input.
344// The limit: it compares VALUES, not writes. A verb that wrote the value the field already had
345// at the start is invisible, and the Stop's value stands: `resume` resetting counters.retries
346// to 0 while it was 0 at the start and the Stop raised it to 1 keeps the 1, though the reset
347// came later. A per-field write counter would tell them apart; it is not there yet, and the
348// cost is one retry counted once more.
349export function mergeVerbFields(ours, start, disk) {
350 const out = JSON.parse(JSON.stringify(ours));
351 const taken = [];
352 if (!disk || typeof disk !== 'object' || !start || typeof start !== 'object') return { state: out, taken };
353 for (const p of VERB_OWNED) {
354 const d = getPath(disk, p);
355 if (!same(d, getPath(start, p))) { setPath(out, p, d === undefined ? null : JSON.parse(JSON.stringify(d))); taken.push(p); }
356 }
357 return { state: out, taken };
358}
359
360// Did only the fields a verb owns (and the rev) change from start to disk? true: a verb or the
361// watchdog wrote meanwhile (mergeVerbFields keeps it); false: another Stop saved (or the state
362// is not comparable), and its whole state must not be overwritten. Both normalized states.
363export function onlyVerbChanges(start, disk) {
364 if (!start || !disk || typeof start !== 'object' || typeof disk !== 'object') return false;
365 const strip = (s) => { const c = JSON.parse(JSON.stringify(s)); for (const p of VERB_OWNED) setPath(c, p, null); c.rev = 0; return JSON.stringify(c); };
366 return strip(start) === strip(disk);
367}
368
369// A v1 state is a flat object with `phase` and no schemaVersion.
370export function isV1State(raw) {
371 return !!raw && typeof raw === 'object' && raw.schemaVersion == null && typeof raw.phase === 'string';
372}
373
374// Map the v1 flat layout onto v2. Unknown fields are dropped; defaults fill the rest.
375export function migrateV1(raw) {
376 return normalizeState({
377 task: raw.task,
378 phase: raw.phase,
379 complexity: raw.complexity,
380 options: {
381 commitSteps: raw.commitSteps,
382 gitFinish: raw.gitFinish,
383 gitPush: raw.gitPush,
384 approvePlan: raw.approvePlan,
385 testCmd: raw.testCmd ?? null,
386 externals: raw.externals,
387 },
388 counters: { iterations: raw.iterations, retries: raw.retries, finalFails: raw.finalFails },
389 limits: { maxIterations: raw.max, maxIterationsExplicit: true, maxRetries: raw.maxRetries },
390 signals: { lastReport: raw.lastReport, claimedDone: raw.claimedDone, paused: raw.paused },
391 flags: { repeated: raw.repeated, cleanedOnce: raw.cleanedOnce, planPresented: raw.planPresented },
392 lastTest: raw.lastTest ?? null,
393 baselineDirty: raw.baselineDirty,
394 owner: { sessionId: raw.sessionId ?? null, lastFireAt: raw.lastFireAt },
395 });
396}
397
398// Load any raw JSON value into a v2 state.
399// Returns { state, migrated } or { state: null, error } when the input is not a loop state at all.
400export function loadState(raw) {
401 if (!raw || typeof raw !== 'object' || Array.isArray(raw)) return { state: null, error: 'not an object' };
402 if (isV1State(raw)) return { state: migrateV1(raw), migrated: true };
403 if (typeof raw.phase !== 'string') return { state: null, error: 'missing phase' };
404 return { state: normalizeState(raw), migrated: false };
405}
406
407// The exit ramp: phases after the work is declared complete.
408export const EXIT_RAMP = ['cleanup', 'final-verify', 'git-finish'];
409hooks/lib/gate.js 112 lines1// The bridge and the gate, from the mod.
2//
3// Every read and write of the loop's files goes through the bridge, src/shell/mod-bridge.mjs,
4// a Node helper run with $.process.run (one JSON request on stdin, one JSON answer on stdout).
5// $.process.run is CLI only: the mod drives the loop in `claude` and `claude -p`, not in the
6// desktop app or the VS Code extension. The exceptions, all small: the existence checks of
7// the loop's state files ($.fs.exists: a project without a loop never pays for a node
8// process), the owner read by a failed stop and by the tool, the reconciliation flag read by
9// the guard of a tool call (tool.js needsGuard: no process per Edit/Write/Bash), and the fault
10// marker (writeFault below).
11//
12// The flushes (activity, tokens) run one at a time through a serial queue: a slow flush and
13// the next debounce never write side by side. A stop, a judge's check and the read-only
14// operations do not wait behind them: the bridge is safe with overlapping calls (the inbox of
15// the tokens, atomic state writes, see PIANO-MOD "Stato fase 1"), and a stop must not spend
16// its hook's own time waiting for a heartbeat.
17
18export const GATE = '.perseveranza';
19// a fail-open hook that had to skip is journaled this many times per hook per process, then counted
20export const MAX_SKIP_NOTES = 3;
21
22// Where node and the bridge are: PERSEVERANZA_NODE (an absolute path, for a node that is not
23// on the PATH), else `node`; the bridge inside this plugin's folder.
24export async function setup(io, mod) {
25 if (mod.bridge) return;
26 let node = '';
27 try { node = (await io.nodeEnv()) || ''; } catch { /* unset */ }
28 mod.node = node.trim() || 'node';
29 mod.bridge = `${String(io.root()).replace(/[\\/]+$/, '')}/src/shell/mod-bridge.mjs`;
30}
31
32// Is a loop armed in cwd? The bridge's own rule (stop-core's DORMANT check): `state.json` is
33// there, or only its pending copy (a save cut short) and the run was not disarmed (no
34// `state.disarmed.json`, no `state.disarmed.mark`). A `.perseveranza/` folder alone is NOT a
35// loop: `~/.perseveranza` is perseveranza's own config and archive, and a project keeps the
36// folder after a run. `unknown` is what to assume when the check itself fails: true where a
37// wrong "no" would skip the loop's guard, false where a wrong "yes" would block a session.
38export const STATE_FILES = { state: 'state.json', pending: 'state.json.pending', retained: 'state.disarmed.json', mark: 'state.disarmed.mark' };
39export const FAULT_FILE = 'mod-fault.json';
40export const gateFile = (cwd, name) => `${String(cwd).replace(/[\\/]+$/, '')}/${GATE}/${name}`;
41
42export async function hasGate(io, cwd, unknown) {
43 if (typeof cwd !== 'string' || !cwd) return unknown;
44 try {
45 if ((await io.exists(gateFile(cwd, STATE_FILES.state))) === true) return true;
46 if ((await io.exists(gateFile(cwd, STATE_FILES.pending))) !== true) return false;
47 return (await io.exists(gateFile(cwd, STATE_FILES.retained))) !== true && (await io.exists(gateFile(cwd, STATE_FILES.mark))) !== true;
48 } catch { return unknown; }
49}
50
51// The owner session of the loop in cwd, read without the bridge (the one other read the mod
52// does, for a .catch whose bridge failed): '' unclaimed, a session id, or null (unreadable).
53export async function readOwner(io, cwd) {
54 try {
55 const s = JSON.parse(await io.read(gateFile(cwd, STATE_FILES.state)));
56 const id = s && s.owner && typeof s.owner === 'object' ? s.owner.sessionId : null;
57 return typeof id === 'string' ? id : '';
58 } catch { return null; }
59}
60
61// A durable trace when the bridge cannot even journal (node unreachable): one small file of its
62// own, .perseveranza/mod-fault.json, written with $.fs.write (never a file the shell writes).
63// `status` shows it; the next stop that reaches the bridge journals it (mod-fault) and removes
64// it; `arm` reports a leftover one. Never throws.
65export async function writeFault(io, cwd, fault) {
66 try { await io.write(gateFile(cwd, FAULT_FILE), JSON.stringify(fault)); return true; } catch { return false; }
67}
68
69// One request -> { answer } | { error } (no answer: the bridge did not start, timed out,
70// crashed, or printed something that is not JSON). Never throws.
71export async function call(io, mod, op, { cwd, event = {}, facts = {}, timeoutMs }) {
72 try {
73 await setup(io, mod);
74 const r = await io.run([mod.node, mod.bridge], { cwd, stdin: JSON.stringify({ op, cwd, event, facts }), timeoutMs });
75 let answer = null;
76 try { answer = JSON.parse(String(r.stdout || '').trim()); } catch { /* below */ }
77 if (!answer || typeof answer !== 'object' || Array.isArray(answer)) {
78 return { error: `no answer from the bridge (exit ${r.exitCode}${r.stderr ? `: ${String(r.stderr).trim().slice(0, 200)}` : ''})` };
79 }
80 return { answer };
81 } catch (e) {
82 return { error: `no answer from the bridge: ${String((e && e.message) || e).slice(0, 200)}` };
83 }
84}
85
86// The serial queue: task() starts once every task queued before it settled.
87export function enqueue(mod, task) {
88 const run = mod.tail.then(task, task);
89 mod.tail = run.then(() => undefined, () => undefined);
90 return run;
91}
92
93// A flush on its way, from its first step (before it takes anything out of memory) to its
94// last: mod.inflight holds it, and a stop waits for what is there (stop.js).
95export function track(mod, run) {
96 const p = Promise.resolve().then(run);
97 mod.inflight.add(p);
98 const done = () => { mod.inflight.delete(p); };
99 p.then(done, done);
100 return p;
101}
102
103// A fail-open hook that could not do its job: a line in the debug log and, a few times per
104// hook, in the journal (mod-hook-skipped). Never throws.
105export async function skipped(io, mod, cwd, hook, error, session = '') {
106 const n = (mod.skips[hook] || 0) + 1;
107 mod.skips[hook] = n;
108 try { await io.log(`perseveranza: ${hook} skipped (${error})`); } catch { /* best effort */ }
109 if (n > MAX_SKIP_NOTES) return;
110 await call(io, mod, 'journal', { cwd, event: { session_id: session }, facts: { lines: [{ type: 'mod-hook-skipped', hook, error: String(error).slice(0, 300), count: n }] }, timeoutMs: 5000 });
111}
112