SLOPSHOPPER

rd-sub

Watchdog test fixture, not a mod to install. release-delivery probe: notes appended into a subagent during a tool call and after its answer, then a resume

newguardtimer
A shopper browsing a rack in a slop shop
README

watchdog

CI

A second model reviews each step that Claude Code takes and sends it short notes while it works: nit, concern or blocker.

A watchdog flags a planted bug and nudges Claude, which fixes it; the band card marks the note as maybe outdated and opens to the whole note, /watchdog status shows the review cost, and a later review retracts the note and raises a held concern: Claude reported a result it never ran

Requirements

Claude Code 2.1.290 or later. Watchdog is a mod: a plugin whose code Claude Code runs inside your session. The npm stable channel (2.1.285 on 2026-10-06) has no mods. Below 2.1.290 the plugin shows unsupported. Desktop support starts when Claude.app bundles Claude Code 2.1.290 or later.

Check claude --version first. If it is below 2.1.290, move to the npm latest channel: npm install -g @anthropic-ai/claude-code@latest.

Quick start

/plugin marketplace add matteoantoci/claude-plugins
/plugin install watchdog@matteoantoci-plugins

Then run /watchdog on (reviews are off until you do) and ask Claude for a small change. The note shows as a watchdog: [concern] … line in the transcript and as a card, one line with its first sentence above the prompt box. Click the card's ▸, or press ctrl+x tab and then its letter (a, b, c), to read the whole note with its watchdog, age and state; Esc gives the focus back to the prompt. Run /watchdog status to see each watchdog's reviews, notes, tokens and cost. A card names its watchdog when you run two or more.

At its right end a card shows only what needs a look: the subagent type for a note on a subagent; the state while the note has not reached Claude yet, nudge pending (the plugin starts a turn so that Claude reads it), held or aside (Claude reads it with your next prompt); nothing once it is steered (Claude reads it after its next tool result) or nudged. When Claude edited files after the review read its update, the card says outdated? N edits (the open card says may be outdated: N edits since), and Claude reads the same mark with the note. The next review of the same watchdog sees the edits and the note; when the note no longer holds, it retracts it: the card goes, and a note that waits never reaches Claude. A blocker that may be outdated and came after Claude's reply waits as held for that review before it nudges, so Claude does not go after a bug it already fixed.

What runs on your machine

  • The mod runs inside Claude Code with your permissions. Its code is in plugins/watchdog/hooks/.
  • It reads the WATCHDOG.json and WATCHDOG.md files, the session's memory files (such as CLAUDE.md), each update of the agent you work with and, when CLAUDE_WATCHDOG is set in a claude -p run, your project and local settings.
  • It sends each update to the review model, as an agent that Claude Code runs on your account, and puts the notes into your session. /watchdog on also sends one 1-token request for each model, to check that it exists. Apart from these model requests through Claude Code, the mod makes no network calls: its code never calls $.http.fetch or fetch.
  • Its only file write is the dump, under <config>/watchdog/dumps/ (<config> is $CLAUDE_CONFIG_DIR or ~/.claude). It keeps its notes and review state in Claude Code's session state and plugin store. In the terminal, /watchdog dump also copies the dump text to the clipboard.
  • By default a reviewer gets Read, Grep and Glob. A project WATCHDOG.json can grant no more; only <config>/WATCHDOG.json can grant other tools and mcp__* tools. Bash, Edit, Write, NotebookEdit, Agent, SendMessage, AskUserQuestion and ToolSearch are always refused. A reviewer never asks you for a permission.
  • The mod allows its own review spawn (the Agent call of a watchdog:* type) when Claude Code would ask, so no dialog or Auto-mode classifier sees it. A permission rule that denies Agent still wins.
  • A project WATCHDOG.json or WATCHDOG.md sets the number of reviewers, their model, effort and instructions. In a repo you did not write, read these files before /watchdog on. Once on, /watchdog status lists each watchdog with its model, effort and file.

Cost and off switch

  • Each review is one more agent, on opus with medium effort by default (in the demo: 3 reviews, 37.2k tokens, $0.07). The built-in "You should know" mod, when on, runs its own side agent too; turn it off in /plugin to pay for one only.
  • /watchdog off stops reviews for this session. /plugin uninstall watchdog@matteoantoci-plugins removes the plugin.
  • Settings, in /plugin (Installed, Watchdog, Configure options) or /config: onByDefault (default false) turns reviews on in each new interactive session. immuneTurns (0 to 5, default 3) is the number of turns after a nudge before the next nudge for a concern; a nudge is a turn that the plugin starts so that Claude reads a note that came after its reply.
  • Each nudge is one more turn of Claude. After each of your prompts the plugin sends at most 1 nudge for concerns and 2 for blockers (a concern that comes with a blocker rides along); a later note waits for your next prompt. /watchdog status shows both counts, for example nudge 1/1 · blocker 0/2.

Commands

  • /watchdog or /watchdog status: each watchdog's state, reviews, notes, tokens and cost, and the session totals. For a state such as halted, see docs/failures.md.
  • /watchdog on and /watchdog off: turn reviews on or off for this session.
  • /watchdog dump and /watchdog dump raw: write the review log to a file (raw adds the review prompts).

Configure

A WATCHDOG.json in your project or in ~/.claude sets the watchdogs. The load order, every key and the tool grants are in docs/configuration.md. This file adds a second reviewer to the default one:

{ "watchdogs": [{ "name": "default" }, { "name": "security", "model": "sonnet", "effort": "high" }] }

Limitations

  • Notes are advice: the agent may reject one. A review runs in the background, so a note can come after the step; the outdated mark above counts every edit since the review, whatever file it touched.
  • Reviews run only on Anthropic models. claude -p needs CLAUDE_WATCHDOG=on and has no nudge and no cards: see docs/headless.md.
  • The cost comes from the plugin's own price table (plugins/watchdog/hooks/prices.ts); a model not in it shows $?.

Development

npm install sets up the tools and the pre-commit hook. npm run check runs the 6 checks of pre-commit and CI: rules, fmt:check, lint, typecheck, validate and test. See docs/plugin-dev.md. Before a release and before a bump of the pinned Claude Code version, run the live probe by hand, on a real model and login: scripts/live-probe/README.md.

License

Apache-2.0

Source 1 files
hooks/sub.mjs 122 lines
1// Release-probe mod (spec §16.5 late subagent notes; research/smoke-subagent-delivery.md:142,225): it appends one
2// note into a subagent while its `sleep` Bash call runs, one after the agent's final answer (its last turn.step,
3// after next) and one at its turn.complete; once the main turn after that run ends, it resumes the finished agent
4// with $.session.send, as the SendMessage tool resumes one (d.ts SessionSendAddress).
5const FILE = '__LOG__/rd-sub.jsonl';
6const SLEEP = /sleep\s+\d/u;
7const INFLIGHT =
8  '<watchdog-notes>\n<note watchdog="probe" severity="concern">Code word RDINFLIGHT. Use it as WORD in your next ' +
9  'Bash call.</note>\n</watchdog-notes>';
10const LATE =
11  '<watchdog-notes>\n<note watchdog="probe" severity="concern">Late code word RDLATE.</note>\n</watchdog-notes>';
12const RESUME =
13  'RDRESUME. No earlier task is pending. Reply with exactly NOTES=<each code word that a watchdog note gave you, ' +
14  'comma-separated> or NOTES=NONE. Do not use tools.';
15const RESUME_DELAY_MS = 1500;
16const lines = [];
17let writing = null;
18let isDirty = false;
19let agentId = null;
20let runs = 0;
21let isLateStepDone = false;
22let isLateCompleteDone = false;
23let isResumeSent = false;
24
25// Rewrites the whole log file; a write asked while one runs is folded into one more pass (as the observer does).
26const flush = ($) => {
27  if (writing) {
28    isDirty = true;
29    return writing;
30  }
31  writing = (async () => {
32    do {
33      isDirty = false;
34      await $.fs.write(FILE, `${lines.join('\n')}\n`);
35    } while (isDirty);
36  })()
37    .catch(() => undefined)
38    .finally(() => {
39      writing = null;
40    });
41  return writing;
42};
43
44const record = ($, data) => {
45  lines.push(JSON.stringify({ t: Date.now(), ...data }));
46  return flush($);
47};
48
49const errorText = (error) => String(error?.message ?? error).slice(0, 180);
50
51// The append the spec rules out for a subagent (§11.3): `$.session.append({ agentId })`. A `{ deny }` or a
52// rejection (research: "no running loop is <id>") is the outcome.
53const append = async ($, text) => {
54  try {
55    const appended = await $.session.append({ agentId, message: { type: 'user', content: [{ type: 'text', text }] } });
56    return appended && typeof appended === 'object' && 'deny' in appended ? `deny ${appended.deny}` : 'ok';
57  } catch (error) {
58    return `throw ${errorText(error)}`;
59  }
60};
61
62// §16.5: the resume of a finished agent from its JSONL, forced, so it does not wait on the model's choice.
63const resume = async ($) => {
64  try {
65    const sent = await $.session.send({ to: { agentId }, text: RESUME });
66    await record($, { kind: 'resume', agentId, outcome: sent?.isDelivered ? 'delivered' : `refused ${sent?.reason}` });
67  } catch (error) {
68    await record($, { kind: 'resume', agentId, outcome: `throw ${errorText(error)}` });
69  }
70};
71
72export const register = (on) => {
73  // smoke-subagent-delivery.md:225: the first note goes in while the agent's first `sleep` Bash call runs.
74  on('tool.call', async ($, e, next) => {
75    if (!e.agentId || e.tool !== 'Bash' || agentId !== null || !SLEEP.test(String(e.command ?? ''))) {
76      return next(e);
77    }
78    agentId = e.agentId;
79    const pending = next(e);
80    pending.catch(() => undefined);
81    const outcome = await append($, INFLIGHT);
82    await record($, { kind: 'inflight', agentId, outcome });
83    return pending;
84  });
85  // smoke-subagent-delivery.md:225: after the model's final answer (a step with no tool use) the append still
86  // lands in the JSONL, and the claim is that no request follows it.
87  on('turn.step', async function* ($, e, next) {
88    const result = yield* next(e);
89    if (agentId === null || e.agentId !== agentId) {
90      return result;
91    }
92    const tools = result?.toolUses?.length ?? 0;
93    const answer = String(result?.answer ?? '').slice(0, 200);
94    await record($, { kind: 'step', run: runs + 1, index: e.index, tools, answer });
95    if (tools === 0 && runs === 0 && !isLateStepDone) {
96      isLateStepDone = true;
97      await record($, { kind: 'late-step', agentId, outcome: await append($, LATE) });
98    }
99    return result;
100  });
101  on('turn.complete', async ($, e, next) => {
102    if (agentId !== null && e.agentId === agentId) {
103      // research 3B/3C: by the agent's turn.complete its loop is gone, so this append is refused.
104      if (runs === 0 && !isLateCompleteDone) {
105        isLateCompleteDone = true;
106        await record($, { kind: 'late-complete', agentId, outcome: await append($, `${LATE}\nRDLATE-COMPLETE`) });
107      }
108      const result = await next(e);
109      runs += 1;
110      const answer = String(e.answer ?? '').slice(0, 300);
111      await record($, { kind: 'complete', agentId, run: runs, reason: e.reason ?? null, answer });
112      return result;
113    }
114    const result = await next(e);
115    if (!e.agentId && agentId !== null && runs >= 1 && !isResumeSent) {
116      isResumeSent = true;
117      $.clock.after(RESUME_DELAY_MS, () => resume($));
118    }
119    return result;
120  });
121};
122