Test wrapper for the core chat harness mod.

<a href="https://hermitd.dev"><img src="https://img.shields.io/badge/website-hermitd.dev-black.svg" alt="Website" /></a> <a href="LICENSE"><img src="https://img.shields.io/badge/license-MIT-blue.svg" alt="MIT License" /></a> <a href="https://code.claude.com/docs/en/plugins"><img src="https://img.shields.io/badge/Claude%20Code-plugin-orange.svg" alt="Claude Code Plugin" /></a> <a href="plugins/hermitd/CHANGELOG.md"><img src="https://img.shields.io/badge/version-1.4.11-green.svg" alt="Version 1.4.11" /></a> <img src="https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/gtapps/hermitd/_gh_traffic_stats/.github/badges/clones.json" alt="Downloads" /> <img src="https://img.shields.io/badge/PRs-welcome-brightgreen.svg" alt="PRs Welcome" /> <a href="https://discord.gg/54sJqAxhUh"><img src="https://img.shields.io/badge/Discord-Join-5865F2?logo=discord&logoColor=white" alt="Join" /></a>
Run an always-on Claude Code agent on your own machine, for you or your team. Start sessions from your chat in any project it can reach, on the model and effort you choose, and steer them through the agent or take over in the Claude app or CLI.
Use it from your terminal or the Claude app via Remote Control, or connect Discord, Telegram, iMessage, or a custom Claude Code channel.
Give it ongoing responsibilities: maintain research, monitor systems, run routines, and follow up on unfinished work. Between requests, it checks those responsibilities, carries progress across sessions, and reaches you when something needs attention.
Run it on your Claude subscription and extend it with your own MCP servers, skills, and plugins.
It's Claude Code. hermitd is a plugin for the official Claude Code CLI, which runs as Anthropic ships it. New Claude models, features, and fixes work as soon as Claude Code supports them.
Coming from claude-code-hermit? See how to migrate.
<img src="plugins/hermitd/assets/cover.png" alt="Always-on Claude Code agent" />
<a id="quick-start"></a>
Choose one installation method below. Run it from the folder where you want your agent, empty or existing. Uses your Claude subscription on Linux, macOS, or Windows via WSL2. See prerequisites.
With Claude Code 2.1.292+ and Bun 1.4+ installed:
claude plugin install hermitd --marketplace gtapps/hermitd --scope local
claude "/hermitd:hatch"
Prepares Claude Code, Bun, and tmux, installs the plugin, and launches setup:
curl -fsSL https://gtapps.github.io/hermitd/install.sh | bash
Both options set up the agent in this folder. Hatch guides you through the agent’s purpose and preferences, then shows how to start it. Choose Quick for defaults you can adjust later.
After setup, follow the printed next steps to start your agent.
Run in a persistent tmux session:
hermitd start
Requires tmux. The watchdog recovers failed sessions while your machine stays on. Claude Code's /sandbox is recommended for unattended use. To connect a chat, run /hermitd:channel-setup as directed by the setup handoff.
Run the guided setup in Claude Code:
/hermitd:docker-setup
Builds and starts the container, then walks you through authentication and channel pairing. Requires Docker Compose v2.
Customize the container. Ask the agent to add tools, packages, or services to its Docker setup. For example: “Add ffmpeg to the container.”
Optional Docker security controls cover local-network access, DNS policy, resource limits, and plugin installation auditing.
raw/ into maintained knowledge in compiled/, alongside Claude Code's auto memory. /recall searches past sessions, knowledge, proposals, and captured channel conversations.Part of your project channel. With passive mode, the agent saves incoming group messages to look back on later, and wakes when someone you allow @mentions it. It also remembers instructions for that channel.
<a id="configure-it"></a>
Tune from a terminal with /hermit-settings, or change permitted settings from a trusted Discord or Telegram chat. Every write is validated and recorded in a redacted audit ledger; /hermit-settings history [setting] shows what changed. Some of the settings available:
| Key | Default / options (default bold) |
|---|---|
agent_name | your assistant's name |
operator_profile | primary-chat audience: technical / non-technical |
timezone | detected during setup; fallback UTC |
language | detected during setup; fallback en |
escalation | how much it does before asking: conservative / balanced / autonomous |
model | session model: sonnet |
permission_mode | how freely the unattended agent acts: auto |
AGENT_HOOK_PROFILE | guardrail profile: minimal / standard (interactive) / strict (always-on) |
channels | Discord / Telegram / iMessage / third-party channel plugins (+ allowed_users) |
channels.primary | which channel gets outbound pings |
channels.<name>.maintainer_channel_id | optional separate chat for technical alerts, diagnostics, and usage details |
push_notifications | native/mobile push on alerts: true |
remote | remote control; false also requires approval for cross-machine peer messages; true |
ask_gate | route unattended questions to a paired channel: true |
budget | optional daily / weekly / monthly caps; alert or binding pause action |
artifacts | dashboard / proposals / weekly review: dashboard and proposals enabled |
heartbeat.enabled | timed idle sweeps: true |
heartbeat.every | idle sweep cadence: 30m |
heartbeat.active_hours | active window: 08:00–23:00 |
routines | persistent routines managed via /hermit-routines |
monitors | persistent background watches managed via /watch |
scheduled_checks | session-triggered skills at task completion |
reflection.graduation_min_sessions | proposal recurrence bar: 1 |
quality_gate.tier | post-change cleanup spend: budget / balanced / quality |
knowledge.compiled_budget_chars | fresh/resumed startup catalog budget: 2500 |
knowledge.raw_retention_days | raw/ retention: 14 |
knowledge.working_set_warn | warn above N compiled docs: 20 |
auto_session | auto-start session on boot: true |
boot_skill / shutdown_skill | custom boot / teardown skill |
context_hygiene.clear | safe-boundary context clear: enabled, quiet 1h, max age 24h, minimum 20,000 tokens |
context_hygiene.compact | compact long-running active context: enabled, 100000 compactible tokens / 4h cooldown |
CLAUDE_AUTOCOMPACT_PCT_OVERRIDE | auto-compact at % of context: 65 |
MAX_THINKING_TOKENS | thinking-token cap per turn: 10000 |
watchdog.scheduler_enabled | OS scheduler for the watchdog tick: true on tmux always-on (auto-installed at boot); false or hermitd-watchdog uninstall opts out |
watchdog.enabled | recovery/restart tier: false until first scheduler registration (or /docker-setup); hygiene still runs |
Full schema in the Config Reference
Artifacts. The agent uses Claude Code Artifacts to provide an interactive dashboard and custom pages generated on demand that you can view, interact with, and share. Ask it to build your own personalized agent dashboard.
Ask for an update from your terminal or connected chat:
| Command | What it gives you |
|---|---|
/brief | Current status and a summary of recent work. |
/recall | Search past sessions, knowledge, proposals, and captured conversations. |
/hermit-health | Alerts, routines, channels, blockers, and recent learnings. |
/hermit-doctor | Diagnostics for the installation, runtime, scheduling, credentials, and permissions. |
/hermit-evolution | Cost trends, proposal activity, routines, and what the agent has produced over time. |
/cost-reflect | A breakdown of usage by token type, session, and what triggered the work. |
/hermit-dashboard-design | A dashboard designed around what your agent actually tracks. |
The agent reviews evidence from its work and operation. Durable lessons go to memory; non-trivial ideas that would change its behavior are verified, deduplicated, and brought to you for approval.
Work produces evidence
│
▼
Reflect when due
│
▼
Verify and deduplicate
│
┌────┴────┐
▼ ▼
Remember Propose
a lesson a change
│
▼
You approve?
│ │
no yes
│ │
No change Implement
│
▼
Verify result
│
▼
Future evidence
Reflection runs at eligible task or session pauses, daily, and after routines configured to reflect. Approved changes can start now, become a task, or be left for manual implementation. Proposals are resolved when verification passes or later evidence shows the problem is gone.
| Ask | What it does |
|---|---|
“Start X in ~/code/api on Opus at high effort, in its own worktree.” | Starts a session in any project, on your model and effort. |
“Tell me when my session in ~/code/web finishes.” | Watches your other Claude Code sessions. |
!pause · !snooze 2h · !model sonnet · !effort high | Claude Code controls, straight from chat. |
“/later check tomorrow whether those errors have returned.” | Checks whether a fix held up over time. |
| “When I ask for a status update, include blockers.” | Remembers instructions for that channel. |
| “What else could you be doing for me?” | Proposes new capabilities from your work and its tools. |
Quiet heartbeat checks, skipped routines, and passive chat capture use no model tokens. Work, evaluations, and replies consume usage; context management keeps conversation history bounded.
event: a file, a log line, a webhook, a schedule
│
▼
precheck (plain script, no model tokens)
│
├── nothing changed ──▶ back to sleep
│
▼ something to do
Claude Code turn on your machine
│
├── needs a decision ──▶ asks you in chat
│
▼
result in your chat + usage recorded
/cost-reflect.See budgets and routine scheduling for configuration and scheduler fallback behavior.
Reach the running agent through your connected channels or Claude Code Remote Control. You can also start separate sessions for additional work:
/spawn-session, the agent launches a local Claude Code helper in the project, or in another folder with --cwd, and relays its status when it becomes idle. Claude Code isolates the helper's edits in a Git worktree unless the project sets worktree.bgIsolation to none; pass --worktree to give it one from the start. Set the helper’s model and effort with options such as --model sonnet --effort high./rc-gate, the agent manages a Remote Control server on your machine or server. While the gate is open, you can spawn new Claude Code sessions from the Claude app, using your local files and tools. Each session gets its own Git worktree, while the agent keeps running.Both session-spawning paths require a Git workspace. Remote Control requires a Claude sign-in through /login on the machine running the agent.
Watch other sessions. Through Claude Code cross-session messaging, ask the agent to watch a local Claude Code session, including one you started interactively, and notify you when it next becomes idle.
Claude Code controls from chat. Use !model sonnet, !effort high, !advisor opus, !compact, !clear, !doctor (alias !checkup), and !permission-mode auto directly from your connected chat. Control the agent’s work with !pause, !resume, and !snooze 2h. Use /when-done-switch-to --model sonnet to switch automatically at the end of the current turn.
Optional plugins that add domain tools and workflows to your agent.
You can run separate agents for different responsibilities, each with its own working state, knowledge, and routines. See Creating Your Own Hermit.
External orchestration. Other agents and tools can check the agent’s status, health, and recent work through its MCP interface, and request a wake when needed.
Sign-in renewal from chat. When your agent’s Claude sign-in needs renewing, use /relogin from your connected chat. Open the link in your browser, sign in, and send the code back in chat.
Scheduled backups. Optional backups preserve the agent’s knowledge, session reports, settings, and Claude Code memory in Git, with an optional private remote copy. Backups run without model tokens.
Claude Code 2.1.287 blocks plugins named claude*, so this project moved from claude-code-hermit to hermitd. Before migrating, update every registered agent in the Claude config directory to core 1.4.8 and stop all of them, including Docker agents. Run once on the host:
curl -fsSL https://gtapps.github.io/hermitd/migrate.sh | bash
The migration records every agent before replacing the marketplace, moves project state to .hermit/, refreshes launchers and permissions, and rebuilds Docker images. Customized managed files receive .bak copies. If interrupted, rerun the same command to resume. Disabled plugin installs are reported and are not reinstalled.
Follow the printed start command for each agent, then run /hermitd:hermit-evolve. Pending later commands that still use old paths are reported for re-arming.
Run hermitd update from the project folder, or hermitd update <name> from anywhere. Docker updates refresh the host core first, then the container.
hermitd list shows registered and discovered agents on this host, including stopped and missing projects. hermitd status [name] shows transport, execution and its age, open and waiting tasks, and the first runnable task. Both support --json. Listing never removes entries; hermitd prune removes missing projects.
Use hermitd start|stop|restart|attach [name] for lifecycle commands, hermitd pause [name] on|off|snooze <duration>|status, hermitd watchdog [name] run|install|uninstall, or hermitd run [name] <script> [args] for maintenance. Names match the project folder or agent name; with no name, the nearest project above the current folder is used.
See the Upgrade guide for details.
<a id="tips--tuning"></a>
Join the Discord community for setup help and discussion. See CONTRIBUTING.md for reporting bugs or contributing.
Andrej Karpathy inspired the raw/ → compiled/ knowledge system.
Independent project, not affiliated with Anthropic.
register.ts 250 lines1import type { EngineInterface, Register } from 'claude-code';
2
3type Command = { command: string; arg: string | null };
4type Outcome = Command & { status: 'ok' | 'failed' | 'unknown'; text: string };
5type Request = {
6 commands: Command[];
7 by?: string;
8 reply_to?: { source: string; chat_id: string };
9 requested_at?: string;
10};
11type Decision = Request & {
12 decision: 'pass' | 'run' | 'refuse' | 'ok' | 'send_failed';
13 reason?: string;
14 text?: string;
15 silent?: boolean;
16};
17const DEADLINE_MS = 120_000;
18
19// Only a cheap shape filter here. The bridge owns envelope parsing, addressing,
20// authorization and residency; ordinary prompts never start a process.
21export function isCommandPrompt(text: string): boolean {
22 const body = /^\s*<channel\b[^>]*>([\s\S]*)<\/channel>\s*$/.exec(text)?.[1] ?? '';
23 return /^(?:(?:<@!?\d+>|@\S+)\s*)?!(?:model|effort|compact|clear|advisor|doctor|checkup)(?:@[^\s]+)?(?:\s|$)/i.test(body.trim());
24}
25
26export function classifyStdout(command: string, text: string): 'ok' | 'failed' {
27 const success = command === '/model' ? /^Set model to\b/i.test(text)
28 : command === '/effort' ? /^Set effort level to\b/i.test(text)
29 : command === '/advisor' ? /^(?:Advisor set to\b|Advisor (?:disabled|turned off)\b)/i.test(text)
30 : command === '/clear' && text === '';
31 return success ? 'ok' : 'failed';
32}
33
34async function bridge($: EngineInterface, verb: string, payload = ''): Promise<Decision> {
35 const result = await $.process.run([
36 'bash', `${$.plugin.root}/scripts/hermitd-exec.sh`, 'harness-mod', verb,
37 await $.session.id(), payload,
38 ], { cwd: await $.session.cwd() });
39 if (result.exitCode !== 0 || !result.stdout.trim()) throw new Error(`Harness bridge ${verb} failed`);
40 return JSON.parse(result.stdout);
41}
42
43const queue: Request[] = [];
44let request: Request | undefined;
45let outcomes: Outcome[] = [];
46let active: {
47 command: Command;
48 resolved: boolean;
49 evidence?: { status: 'ok' | 'failed'; text: string };
50 stdoutStarted: boolean;
51 expectedModel: string | null;
52 timer: { cancel(): void };
53} | undefined;
54let approval: string | null = null;
55let claiming = false;
56let dispatching = false;
57
58async function finalize($: EngineInterface, input: unknown) {
59 const result = await bridge($, 'finalize', JSON.stringify(input));
60 if (result.decision === 'send_failed') {
61 await $.prompt.submit({ text: `Relay this harness command result once to ${JSON.stringify(result.reply_to)}: ${JSON.stringify(result.text)}. Do not run the command again.` });
62 }
63}
64
65async function finish($: EngineInterface, outcome: Outcome) {
66 if (!active || !request) return;
67 active.timer.cancel();
68 active = undefined;
69 approval = null;
70 outcomes.push(outcome);
71 try {
72 // A dependent effort leg is never run after a failed or uncertain model leg.
73 if (outcome.status !== 'ok' || outcomes.length === request.commands.length) {
74 const completed = request;
75 const results = outcomes;
76 request = undefined;
77 outcomes = [];
78 await finalize($, { ...completed, outcomes: results });
79 }
80 } finally {
81 // A failed reply must not strand the requests queued behind this one.
82 schedule($);
83 }
84}
85
86function observe($: EngineInterface) {
87 if (!active?.resolved || !active.evidence) return;
88 const current = active;
89 // Finish outside the observing hook, so the next command cannot re-enter it.
90 $.clock.after(0, () => {
91 if (active === current) return finish($, { ...current.command, ...current.evidence! });
92 });
93}
94
95async function dispatch($: EngineInterface) {
96 if (active || dispatching) return;
97 request ??= queue.shift();
98 if (!request) return;
99 const command = request.commands[outcomes.length];
100 if (command.command === '/doctor') return runDoctor($, request);
101 const current = {
102 command, resolved: false, stdoutStarted: false, expectedModel: null as string | null,
103 evidence: undefined as { status: 'ok' | 'failed'; text: string } | undefined,
104 timer: $.clock.after(DEADLINE_MS, () => {
105 if (active === current) return finish($, { ...command, status: 'unknown', text: '' });
106 }),
107 };
108 active = current;
109 dispatching = true;
110 approval = command.command === '/model' ? command.arg : null;
111 try {
112 const running = $.command.run({ command: command.command.slice(1), args: command.arg ?? '' });
113 // An ack failure leaves the request for a later claim; it says nothing about this run.
114 if (request.requested_at && outcomes.length === 0) await bridge($, 'ack', request.requested_at).catch(() => undefined);
115 await running;
116 if (active !== current) return;
117 current.resolved = true;
118 observe($);
119 } catch (error) {
120 if (active === current) await finish($, { ...command, status: 'failed', text: String(error) });
121 } finally {
122 dispatching = false;
123 if (!active) schedule($);
124 }
125}
126
127// /doctor is a prompt-type command: run resolves at submit and prints no stdout, and its
128// own turn replies to the chat through the skill-relay marker written first.
129async function runDoctor($: EngineInterface, doctor: Request) {
130 request = undefined;
131 dispatching = true;
132 try {
133 // The bridge answers `pass` rather than exiting non-zero when it cannot record the target.
134 if ((await bridge($, 'relay', JSON.stringify(doctor))).decision !== 'ok') throw new Error('Reply target not recorded');
135 await $.command.run({ command: 'doctor', args: '' });
136 } catch (error) {
137 await finalize($, { ...doctor, outcomes: [{ ...doctor.commands[0], status: 'failed', text: String(error) }] }).catch(() => undefined);
138 } finally {
139 dispatching = false;
140 schedule($);
141 }
142}
143
144function schedule($: EngineInterface) {
145 $.clock.after(0, () => dispatch($));
146}
147
148export const register: Register = on => {
149 on('prompt.submit', async ($, e, next) => {
150 if (e.origin.kind !== 'channel' || !isCommandPrompt(e.text)) return next(e);
151 const result = await bridge($, 'intake', e.text);
152 if (result.decision === 'pass') return next(e);
153 if (result.decision === 'run') {
154 queue.push(result);
155 schedule($);
156 } else if (result.decision === 'refuse' && !result.silent) {
157 $.clock.after(0, () => finalize($, result));
158 }
159 return { drop: 'Harness command handled by the native executor' };
160 });
161
162 on('classic.PreModelSwitch', async ($, e, next) => {
163 const result = await next(e);
164 if (result.block || result.permissionDecision === 'deny' || result.permissionDecision === 'ask') return result;
165 if (approval !== null && e.requested_model === approval) {
166 if (active) active.expectedModel = typeof e.to_model === 'string' ? e.to_model : null;
167 approval = null;
168 return { permissionDecision: 'allow' };
169 }
170 return result;
171 });
172
173 on('classic.PostModelSwitch', async ($, e, next) => {
174 const result = await next(e);
175 if (active?.command.command === '/model' && active.expectedModel !== null && e.to_model === active.expectedModel) {
176 active.evidence = { status: 'ok', text: `Model switched to ${active.command.arg}` };
177 observe($);
178 }
179 return result;
180 });
181
182 on('session.append', async ($, e, next) => {
183 const result = await next(e);
184 if (!active || e.door !== 'command' || e.origin.kind !== 'plugin' || e.origin.name !== $.plugin.name) return result;
185 const content = e.message?.content;
186 const text = typeof content === 'string' ? content : Array.isArray(content)
187 ? content.filter((part: { type: string }) => part.type === 'text').map((part: { text: string }) => part.text).join('\n') : '';
188 const dispatched = /<command-name>([^<]+)<\/command-name>/.exec(text);
189 if (dispatched) {
190 const args = /<command-args>([\s\S]*?)<\/command-args>/.exec(text)?.[1] ?? '';
191 active.stdoutStarted = dispatched[1] === active.command.command && args === (active.command.arg ?? '');
192 }
193 const stdout = /<local-command-stdout>([\s\S]*?)<\/local-command-stdout>/.exec(text);
194 if (!stdout || !active.stdoutStarted) return result;
195 const command = active.command.command;
196 if (command === '/compact') return result;
197 const observed = stdout[1];
198 const status = classifyStdout(command, observed);
199 // A cleared context is reported from classic.SessionStart, once the resident id is current.
200 if (command === '/clear' && status === 'ok') return result;
201 active.evidence = { status, text: observed };
202 observe($);
203 return result;
204 });
205
206 on('session.compact', async ($, e, next) => {
207 const result = await next(e);
208 if (active?.command.command === '/compact' && typeof result.tokensBefore === 'number' && typeof result.tokensAfter === 'number') {
209 active.evidence = { status: 'ok', text: `Compacted ${result.tokensBefore} to ${result.tokensAfter} tokens` };
210 observe($);
211 }
212 return result;
213 });
214
215 on('session.start', async ($, e, next) => {
216 const result = await next(e);
217 await bridge($, 'loaded');
218 return result;
219 });
220
221 // A /clear changes the session id without another session.start. The plugin's own
222 // SessionStart command hook restamps the resident id; next(e) resolves after it, while
223 // session.end and $.command.run resolve before it starts.
224 on('classic.SessionStart', async ($, e, next) => {
225 const result = await next(e);
226 if (e.source === 'clear') {
227 if (active?.command.command === '/clear') {
228 active.evidence = { status: 'ok', text: 'Context cleared' };
229 observe($);
230 }
231 $.clock.after(0, () => bridge($, 'loaded'));
232 }
233 return result;
234 });
235
236 on('turn.complete', async ($, e, next) => {
237 const result = await next(e);
238 if (e.agentId || e.reason !== 'answer' || claiming || active || request || queue.length) return result;
239 claiming = true;
240 try {
241 const claimed = await bridge($, 'claim');
242 if (claimed.decision === 'run') {
243 queue.push(claimed);
244 schedule($);
245 }
246 } finally { claiming = false; }
247 return result;
248 });
249};
250