awsh for Claude Code: an `aw` subagent type whose model steps are answered by a session on the local awsh harness daemon, so any harness and any backend the…

<!-- aither-header:start GENERATED from the ecosystem registry. Edits here are overwritten; change the registry instead. -->
Docs · Source · pip install awdk · The Aither World
The Aither World is an operating system for agents — a Linux you can hand to one, the runtimes it works in, and the tools it works with. awnix is the Linux underneath it; awdk is one of its 67 bricks — each installs on its own, runs offline, and needs no account.
Start here: Point it at a backend you already pay for and run one agent loop.
<!-- aither-header:end -->
<!-- mcp-name: io.github.Aitherium/awdk -->
3 lines of code. Any backend. Local or cloud. Zero lock-in.
Aither ADK is a Python SDK + CLI for building AI agents that run on your hardware — a single helpful agent or a coordinated fleet that delegates work to each other. Agents get tools, persistent knowledge-graph memory, safety filtering, and effort-based model routing out of the box. Swap the LLM backend at runtime — your GPU, Ollama, llama.cpp, or any cloud API — same code, same agents.
pip install awdk
adk quickstart # auto-detect hardware, set up inference
adk init my-agent && cd my-agent && python agent.py
The package is awdk; the command is adk (awdk is the same command, and python -m adk always works). Windows, "adk is not recognized"? pip put the command in a Scripts folder that is not on PATH — its warning names the folder. Add it once, then open a new terminal:
$s = python -c "import sysconfig; print(sysconfig.get_path('scripts'))"
[Environment]::SetEnvironmentVariable('Path', "$([Environment]::GetEnvironmentVariable('Path','User'));$s", 'User')
| You have… | Run this | You get |
|---|---|---|
| Nothing — not even Python | one-line installer (below) | isolated env + first-run wizard |
| No GPU, no API key | adk bonsai-local | Bonsai running free, offline, on CPU — pulls ~300MB image, serves on :8090 |
| A GPU (6 GB+) | adk quickstart | auto-detected vLLM/Ollama, models pulled, ready to chat |
| Just an API key | adk quickstart --cloud | cloud inference (Anthropic / OpenAI / DeepSeek) |
| A whole LAN of machines | adk deploy grid | multi-machine effort-routed inference |
The no-Python one-liner — sets up an isolated environment (via uv) and launches the wizard:
# macOS / Linux
curl -fsSL https://aitherium.com/install.sh | sh
# Windows
powershell -ExecutionPolicy ByPass -c "irm https://aitherium.com/install.ps1 | iex"
Then, whichever path you took:
adk start # chat with your agent (zero config)
adk doctor # something wrong? this names it
Using an AI coding agent (Claude Code, Cursor, Copilot)? Paste the Agent Setup Prompt into your session — it walks the agent through install, auth, inference, and the path from zero to fleet. There's also
llms.txt/llms-full.txtfor tools that ingest those.
Everything in the ADK hangs off five ideas:
AitherAgent("aither"). One object: await agent.chat("...") is the whole API. It has a persona, tools, and memory.ask_agent tool. One YAML file, one adk-serve command, and you have an orchestrator delegating to specialists.If you only remember one thing: agent.chat() is the agent. Everything else is configuration.
| I want to… | Read this | |
|---|---|---|
| Drive Claude Code / Codex / OpenCode / Aider from one shell | AWSH-OMNISHELL-PLAYBOOK.md — install → detect → daemon → UI, verified end to end | |
| Build a real agent or publish a pack | docs/AGENT_DEV_GUIDE.md — the golden path + gotcha checklist | |
| Self-host the full managed-agent experience | QUICKSTART_SELF_HOSTED.md — adk onboard --quick | |
| Operate a self-hosted node long-term | docs/SELF_HOSTING_RUNBOOK.md | |
| Run inference across several machines | GRID_SETUP.md | |
| Wire up a specific LLM provider | docs/providers/ — DeepSeek, Kimi, OpenAI-compatible, local AitherOS | |
| Give my agent a persistent identity/persona | docs/PERSONA.md · `adk soul import | export` |
| Understand the world-model layer | docs/WORLD_MODEL.md | |
| Connect agents across machines (relay) | docs/AITHERRELAY_GUIDE.md | |
| Run a private, local-only companion | PRIVATE_COMPANION.md | |
| Reach my own agent from my phone (Aither Hearth) | docs/agent-home.md — adk home serve, channels, approvals, receipts | |
| See working code | examples/ — five runnable scripts | |
| See what changed | CHANGELOG.md | |
| Browse rendered docs | aitherium.github.io/awdk |
Aither agents speak three protocols for seamless integration with external systems:
Connect your agent to JetBrains, Zed, VS Code, or any ACP-compatible editor over JSON-RPC 2.0 stdio.
adk acp serve # Serve your agent to an editor
acp (registered in adk.harnesses.registry)Map remote A2A agents (Google A2A v0.3.0 compatible) as room participants with full task lifecycle visibility.
from adk.a2a_adapter import A2AAdapter
adapter = A2AAdapter(room_id="main", remote_agent_id="foo")
adapter.on_task_submitted("task_001", "what is AI?")
adapter.on_task_working("task_001", "thinking...")
adapter.on_task_completed("task_001", "AI is...")
adk.a2a_adapter.A2AAdaptera2a.s (submit), a2a.u (update), a2a.d (done)a2a — remote agents appear with their own identity in roomsServe agent-generated RenderBlocks (server-driven UI: tables, forms, charts, approval gates) via the MCP resource protocol using ui:// URIs.
from adk.mcp_ui_resources import RenderBlocksMCPServer, create_table_block, create_scores_block
server = RenderBlocksMCPServer()
blocks = [
create_table_block(columns=["Issue", "Severity"], rows=[[...], [...]]),
create_scores_block({"security": 0.92, "style": 0.78}),
]
uri = server.from_agent_response("reviewer", "task_123", blocks)
# uri -> "ui://agent/reviewer/task_123"
adk.mcp_ui_resources.RenderBlocksMCPServerapplication/vnd.aitheros.renderblocks+jsonui://aw packages — three questions adk can ask about a repositoryadk is the agent runtime; three small, independent packages give it the facts it would otherwise have to guess at. Each answers a different question, each installs on its own, and none of the three requires the others:
| Package | Knows | The question it answers |
|---|---|---|
awgraph | what the code is, and what depends on what | Where is this symptom coming from? |
awgit | what changed, and who is editing it | Is this an in-flight edit someone else owns? |
awrelay | who found what, and who still needs to hear it | Who do I tell? |
pip install awgraph awgit awrelay # or any one of them, alone
Used together, an agent can find a symptom with awgraph, check whether it is an in-flight edit with awgit, and tell the agent already working that file with awrelay — three questions a solo grep-and-guess loop cannot ask at all. The failure they remove is not "the agent was wrong"; it is two agents editing the same file without knowing, and a finding that died in a transcript nobody read.
Each publishes an aither-manifest.json beside its page, and each page renders the others live from those manifests — a project whose manifest is missing shows as unknown rather than silently disappearing: awgraph · awgit · awrelay.
Just want it working? → AWSH-OMNISHELL-PLAYBOOK.md. Install → detect → start the daemon → use it, with verified output at each step. The step people miss is that the harness daemon has to be running: without it the desktop app reports "No harnesses reported by the daemon yet", which reads as a missing feature rather than a stopped process.
Your agent can delegate a task to another coding agent's real product — not a reimplementation of it against the raw API.
That distinction is the whole design. Rebuilding Claude Code's behaviour yourself means inheriting none of its skills, hooks or account handling, and then chasing a product that ships faster than you can track it. So the ADK resolves the real binary on PATH (honouring PATHEXT, so the Windows .cmd shim works), runs it headless with an explicit tool scope, feeds the prompt over stdin — never argv, which is visible in the process table — gives each run its own config dir so concurrent subagents can't corrupt one another's state, and tears down the process tree on timeout.
adk shell harnesses # what can this machine drive, and how to get the rest
adk shell new --harness claude
adk shell send <id> "refactor the retry logic in billing/"
adk shell attach <id> # watch it work
adk shell kill <id> # teardown
adk shell harnesses on a typical box:
ID INSTALLED TRANSPORT DESCRIPTION
claude yes structured-bidi Anthropic Claude Code — bidirectional stream-json, full tool use
gemini yes oneshot-per-turn Google Gemini CLI — one process per turn, stream-json output
terminal yes pty-stream A real shell on this host behind a pseudo-terminal (pwsh/bash)
sandbox NO pty-stream A real Linux TTY inside a dev-workspace container
-> Install Docker Desktop
acp yes structured-bidi JSON-RPC 2.0 stdio harness for JetBrains/Zed/VS Code editors
codex NO oneshot-per-turn OpenAI Codex CLI — one process per turn (codex exec --json)
-> npm i -g @openai/codex
aider NO oneshot-per-turn Aider — pair-programming CLI (one process per turn)
-> pip install aider-install && aider-install
opencode NO oneshot-per-turn OpenCode — open-source coding agent (one process per turn)
-> npm i -g opencode-ai
Ten harnesses are declared; the ones you haven't installed say so and tell you the command. It never silently pretends the world is Claude-only — a harness you don't have is a missing install, not a missing feature, and the difference is printed rather than guessed at.
A per-agent runner does not scale — you end up with claude_runner.py, codex_runner.py, gemini_runner.py, each drifting. So a harness is a row:
HarnessSpec(
id = "codex",
label = "OpenAI Codex CLI",
transport = Transport.ONESHOT_PER_TURN,
binary = "codex",
version_argv = ["--version"],
install_hint = "npm i -g @openai/codex",
json_lines = True,
build_argv = lambda spec, launch: [spec.binary, "exec", "--json", launch.prompt],
)
Four transports cover every agent CLI shipping today: structured-bidi (a persistent bidirectional stream-json session), oneshot-per-turn (a fresh process per turn), pty-stream (a real TTY behind a pseudo-terminal), and http-stream (a remote agent over SSE). Adding an eleventh harness is a table entry, not a new module.
A subagent is launched with an explicit allow-list, and the runner re-validates it fail-closed rather than trusting the caller:
from adk.claude_runner import ClaudeRunner, RunScope
runner = ClaudeRunner()
scope = RunScope(allowed_tools=["Read", "Grep", "Glob"]) # read-only
rec = runner.submit(task="audit error handling in ./api", scope=scope)
rec = runner.get(rec.run_id) # queued | running | completed | failed | cancelled
print(rec.result_text) # one task out, one answer back
runner.kill(rec.run_id) # teardown, whole process tree
The scope becomes --allowedTools on the real CLI, so a subagent asked to audit code cannot write to your disk — enforced by the product you delegated to, not by a prompt asking it nicely.
adk quickstart detects your hardware, pulls the right models, configures backends, and gets you chatting:
pip install awdk
adk quickstart # local GPU: detect → pull models → serve
adk quickstart --cloud # no GPU: enter an API key (Anthropic / OpenAI / DeepSeek)
adk start # start chatting
Either way you get the full harness: tools, skills, memory, and multi-agent coordination.
Want the full self-hosted, managed-agent experience (local LLM → customize a pack → enroll your machine → manage it from the portal)? See QUICKSTART_SELF_HOSTED.md —
adk onboard --quickdoes it in one command.
import asyncio
from adk import AitherAgent
async def main():
agent = AitherAgent("aither") # auto-detects vLLM/Ollama on localhost
response = await agent.chat("Hello! What can you help me with?")
print(response.content)
asyncio.run(main())
The package ships one ready agent — aither, the orchestrator. Add specialists by installing a ready-made pack, or by defining your own. Any agent can then call any other through the built-in ask_agent tool.
# install a ready-made specialist (web research)
adk install pack:openclaw
# define a fleet — the shipped orchestrator + an installed pack + your own agent — and serve it
cat > fleet.yaml <<'YAML'
orchestrator: aither
agents:
- identity: aither # ships with the package
- identity: openclaw # installed above
- name: reviewer # your own — just give it a prompt
system_prompt: "You review code for bugs and security issues."
YAML
adk-serve --fleet fleet.yaml --port 8080
Earn Aitherium tokens by contributing compute to the community embedding pool:
adk volunteer enroll # register as a volunteer (tenant from adk login)
adk volunteer serve # download the embedding model & start llama-server
adk volunteer start # loop: claim batches → embed → submit → earn tokens
Reputation, verified batches and earnings show in the Volunteer Compute panel of the tenant workspace (dgg.aitherium.com) and in adk volunteer status.
| Locked appliances | Aither ADK |
|---|---|
| Their hardware, their cloud | Your hardware, your rules |
| 1 AI assistant | Build a fleet — start with aither, add ready-made packs or your own; they delegate to each other |
| Their model picks | Any model — route by effort level automatically |
| Data on their servers | Data stays on your machine |
| Closed system, monthly fee | Open-core (BSL-1.1) — free, runs entirely on your box |
| Locked to one provider | Runtime backend switching — swap LLM mid-session |
| Cloud-only reasoning | Hybrid reasoning — local orchestration + cloud deep thinking |
adk home runs one personal agent on your machine that answers only you, on the chat apps you already use. The same CLI is installed as aither-hearth.
adk home init --name pip # ~/.aither/agent-home: persona, model, memory
adk home model --byo anthropic # or --local ollama | llamacpp | bonsai
adk home model --check
adk home signin # Sign in with Aitherium
export HEARTH_TELEGRAM_TOKEN=... # a Telegram bot token from @BotFather
adk home serve --channels telegram --pair # prints a 6-digit code: DM it to the bot
adk home serve answers you on the relay, Telegram, Discord, Slack, email, WhatsApp and SMS: every channel whose credentials are in the environment (adk home channels shows which), or exactly the ones in --channels. Pair another channel by sending pair <channel> from one that is already paired.yes <code> on the channel the request arrived on.adk home receipts --verify exits 0 intact, 1 tampered, 2 cannot judge. adk home trust status shows what is enforced.adk home signin, a workspace admin connects a Google account at api.aitherium.com/admin?tab=connections (admin-only); the agent can then read your agenda and mail, and add to them only after an approval. Microsoft 365 is not available yet.adk home say "...", adk home events and /hearth in adk-shell talk to the running serve over 127.0.0.1 instead of starting a second agent.Everything above is free. The paid agent-home pack adds learning that carries across game sessions and more than one agent at a time. Full guide: docs/agent-home.md.
No GPU. No API key. No account. Nothing leaves your machine.
Bonsai is Aitherium's family of ultra-compact models built to make agents sovereign by default — they run on hardware everyone already owns. The 1-bit Bonsai-27B runs on a plain CPU with 4 GB of RAM; Bonsai-4B runs in 2 GB (Android via Termux, Raspberry Pi Zero). Agents on Bonsai get the full harness — tool calling, memory, safety, fleets — not a demo mode.
adk bonsai-local # one command: Docker pulls the image + serves Bonsai-27B on :8090
adk --backend bonsai-local # point your agents at it
Why this matters, concretely:
@tool functions, ask_agent delegation, and pack skills as the big models.When you outgrow it, effort routing lets you keep Bonsai for the cheap calls and send only the hard ones somewhere bigger — see hybrid profiles.
Three packs added in 3.2.0. Each exists because of something the platform's chat models structurally cannot do.
Providers stopped returning raw reasoning. The recovery, from Oh My Pi's externalThinking (MIT), needs no jailbreak: turn the model's native reasoning channel off, then give it a tool whose only parameter is a string described as a private scratchpad. It keeps reasoning — into the tool call, which the API returns in plaintext. What comes back is the model's own shorthand, not a written-for-an-audience summary.
from adk.packs.omp_thinking import reconcile, deep_think_directive
model = {"api": "anthropic-messages", "reasoning": True,
"thinking_requires_effort": True, "thinking_suppress_when_off": True}
reconcile(agent._tools, model) # arms `deep_think` only if the model can take it
print(deep_think_directive(8)["directive"]) # the effort number, aimed at the scratchpad
Two things this pack refuses to do, both deliberate:
reconcile() must run on every swap. Arming it once at startup is correct right up until someone changes models.hooks/register.ts 727 lines1import type { EngineInterface, On, TurnStepChunk, TurnStepResult, TurnUsage } from 'claude-code'
2
3/*
4 * awsh for Claude Code: the `aw` subagent type.
5 *
6 * The step-replacement shape (a real subagent loop whose model request is
7 * answered from elsewhere, chained through a no-op tool across the engine's
8 * hook budget, signed with the model that actually answered) is adapted from
9 * pi-agent-for-claude by Fazal Ali, MIT. This module talks HTTP to the awsh
10 * harness daemon instead of shelling out to a CLI, so it needs no sh, no /tmp
11 * and no tmux, and runs wherever the daemon does.
12 */
13
14/** How often a running session's events are read for new ones. */
15const STREAM_POLL_MS = 150
16
17/**
18 * How long one step streams before handing off to the next. The engine gives a
19 * hook call 10 s and, past it, finishes the step with the real model; the
20 * daemon's session runs on its own, so a longer turn spans several steps,
21 * chained through PROGRESS_TOOL.
22 */
23const STEP_BUDGET_MS = 7_000
24
25/**
26 * The no-op tool a step ends on while the harness is still working: its call
27 * makes the engine run another step, which carries on streaming the same turn.
28 */
29const PROGRESS_TOOL = 'aw_progress'
30
31/** How long a follow-up step waits for its message to reach the transcript. */
32const FOLLOW_UP_WAIT_MS = 5_000
33const POLL_MS = 200
34
35/** How long a new session may take to leave `starting` before it is a failure. */
36const START_WAIT_MS = 6_000
37
38/**
39 * The tool a subagent delivers its final report through, where the session
40 * runs the hand-back contract. A report left as plain text is bounced back
41 * with `[handback-send-enforce]`.
42 */
43const HANDBACK_TOOL = 'SubagentHandback'
44
45/**
46 * User messages the engine writes into a subagent's transcript itself. They
47 * are not the person's follow-ups, so the harness is never sent them.
48 */
49const ENGINE_PREFIXES = [
50 '<system-reminder>',
51 '[handback-send-enforce]',
52 '[Your previous response had no visible output',
53]
54
55/** The harness a spawn gets when neither its prompt nor mods.json names one. */
56const FALLBACK_HARNESS = 'opencode'
57
58/**
59 * How deep a chain of mod-started Claude Code sessions may get. The daemon's
60 * `adk/harnesses/mod.py` holds the same number.
61 */
62const MAX_DEPTH = 2
63
64/** How many notice/raw lines a run keeps to explain a failed turn with. */
65const NOISE_LINES = 6
66
67/** Session states in which the daemon's harness is still producing a turn. */
68const WORKING_STATES = ['starting', 'busy']
69
70/** Where the daemon is and how to talk to it. */
71export type Daemon = { base: string; token: string; home: string }
72
73/** What a spawn asked to be run on. Every field is optional. */
74export type Route = { harness?: string; backend?: string; model?: string; agent?: string }
75
76/** One turn on a daemon session, as it streams across the steps it spans. */
77export type Run = {
78 /** The daemon session answering this agent. */
79 session: string
80 /** The harness that session runs. */
81 harness: string
82 /** This run's number for the agent, which keeps tool call ids unique. */
83 id: number
84 /** The last event `seq` read. */
85 since: number
86 /** How many reads in a row came back empty with the session not working. */
87 quiet: number
88 /** Every piece of text shown so far, across steps. */
89 shown: string
90 /** The text of the turn's answer. */
91 report: string
92 /** Whether the turn has ended. */
93 ended: boolean
94 /** Whether it ended on a failure this module reports, not on an answer. */
95 failed: boolean
96 /** Whether the answer must be handed back through HANDBACK_TOOL. */
97 handback: boolean
98 /** How many steps have handed off to the next through PROGRESS_TOOL. */
99 steps: number
100 /** The model the daemon reported for this session. */
101 model: string
102 /** The last few notice/raw lines: a failed turn's only account of itself. */
103 noise: string[]
104 /** Tokens used since the last step reported them. */
105 usage: { i: number; o: number; cr: number; cw: number }
106}
107
108/**
109 * Registers the `aw` subagent type.
110 *
111 * The spawn runs as it always does, so the engine starts a real subagent loop
112 * with its own id, transcript and entry in `$.agent.list()`; only the model
113 * request inside that loop is replaced, by a turn on an awsh daemon session.
114 * Each loop keeps one daemon session, so a follow-up continues the same
115 * conversation, and the session is visible and steerable from awsh the whole
116 * time (its owner is `claude-code:<agent id>`).
117 *
118 * @param on the engine's registrar
119 */
120export function register(on: On) {
121 const seeds = new Map<string, string>()
122 const consumed = new Map<string, number>()
123 // Loops whose last step delivered the answer through a tool call, with the
124 // text the step after it ends on.
125 const delivered = new Map<string, string>()
126 // PROGRESS_TOOL's full name as the engine registered it, `mcp__<plugin>__…`.
127 let progressTool = ''
128 const lastAnswers = new Map<string, string>()
129 const runs = new Map<string, Run>()
130 // The daemon session each loop talks to, and the route it was opened on.
131 const sessions = new Map<string, { id: string; harness: string }>()
132
133 on('session.start', async ($, e, next) => {
134 const started = await next(e)
135 try {
136 ;({ tool: progressTool } = await $.tool.register({
137 name: PROGRESS_TOOL,
138 description: 'Internal to aw subagents: marks an awsh turn still in progress. Never call it.',
139 }))
140 } catch {
141 // The plugin shares its name with an `awsh` MCP server in the session's
142 // config, and the engine refuses to replace that server. Leave the
143 // session alone; aw subagents report the collision instead of chaining.
144 progressTool = ''
145 }
146 return started
147 })
148
149 on('tool.call', async ($, e, next) => {
150 if (e.tool !== progressTool) return next(e)
151 if (e.agentId === undefined || !runs.has(e.agentId)) {
152 return { deny: `${PROGRESS_TOOL} is internal to aw subagents.` }
153 }
154 return { result: 'the awsh harness is still working.' }
155 })
156
157 on('agent.spawn', async ($, e, next) => {
158 const started = await next(e)
159 if (isAw(e.subagentType) && started.agentId) {
160 seeds.set(started.agentId, e.prompt)
161 }
162 return started
163 })
164
165 on('session.end', async ($, e, next) => {
166 // A daemon session must not outlive the Claude Code session that opened it.
167 const daemon = await daemonOf($).catch(() => undefined)
168 if (daemon !== undefined) {
169 for (const { id } of sessions.values()) {
170 await call($, daemon, 'DELETE', `/sessions/${id}`).catch(() => undefined)
171 }
172 }
173 sessions.clear()
174 return next(e)
175 })
176
177 on('turn.step', async function* ($, e, next) {
178 // An aw subagent's loop, keyed by its id; any other loop is the model's.
179 const agentId = e.agentId
180 const seed = agentId === undefined ? undefined : seeds.get(agentId)
181 if (agentId === undefined || seed === undefined) {
182 return yield* next(e)
183 }
184 const deadline = (await $.clock.now()) + STEP_BUDGET_MS
185
186 // The step after a delivery is that tool's result coming back: the answer
187 // is delivered, so the loop ends here without asking the harness again.
188 const done = delivered.get(agentId)
189 if (done !== undefined) {
190 delivered.delete(agentId)
191 return yield* respond(e, [
192 { kind: 'text', index: 0, text: done },
193 { kind: 'stop', stopReason: 'end_turn', usage: null },
194 ], done, [])
195 }
196
197 let run = runs.get(agentId)
198 if (run === undefined) {
199 const already = consumed.get(agentId) ?? 0
200 const { prompt, count, handback, bounced } = await promptOf($, agentId, seed, already)
201 consumed.set(agentId, count)
202 if (bounced || prompt === undefined) {
203 // A bounce is the engine refusing a plain-text report: hand back the
204 // answer already given rather than asking again.
205 const text = bounced
206 ? lastAnswers.get(agentId) ?? ''
207 : "awsh: no new message reached this agent's transcript."
208 return yield* finish(e, agentId, text, text, 1, handback || bounced, [{ kind: 'text', index: 0, text }])
209 }
210 const started = await start($, agentId, prompt, count, handback, sessions.get(agentId))
211 if ('failure' in started) {
212 // Never throw: a throw hands the step to the model beneath, a Claude
213 // subagent with Claude's tools, and that answer would read as awsh's.
214 const text = started.failure
215 return yield* finish(e, agentId, text, text, 1, handback, [{ kind: 'text', index: 0, text }])
216 }
217 run = started.run
218 sessions.set(agentId, { id: run.session, harness: run.harness })
219 runs.set(agentId, run)
220 }
221
222 const { shown, blocks, lastKind } = yield* pump($, run, deadline, next.signal)
223 if (!run.ended && progressTool === '') {
224 runs.delete(agentId)
225 const text = shown + '\n\nawsh: turn outran one step and cannot chain: an MCP server named '
226 + '"awsh" in this session\'s config blocked registering aw_progress.'
227 return yield* finish(e, agentId, text, text, 1, run.handback, [{ kind: 'text', index: 0, text }])
228 }
229 if (!run.ended) {
230 const input = {}
231 return yield* respond(e, [
232 { kind: 'tool', index: blocks, id: `toolu_aw_${idOf(agentId)}_${run.id}_${run.steps++}`, name: progressTool },
233 { kind: 'input', index: blocks, json: JSON.stringify(input) },
234 { kind: 'stop', stopReason: 'tool_use', usage: takeUsage(run) },
235 ], shown, [{ name: progressTool, input }])
236 }
237
238 runs.delete(agentId)
239 // The hand-back reminder can reach the transcript after the run's first
240 // step read it; checking again now saves a bounce.
241 const handback = run.handback || (await transcriptOf($, agentId)).handback
242 const answer = run.report.trim() || run.shown
243 // Named from the daemon's own record of the session: a model asked what it
244 // is often answers wrongly, and the parent has nothing else to check it by.
245 // A failure is this module's account, not a model's answer: unsigned, so
246 // the parent's rule ("no signature, not awsh's work") holds for it too.
247 const report = run.failed
248 ? answer
249 : `${answer}\n\n— answered by awsh, ${run.harness}/${run.model || 'unreported'}`
250 // The parent reads the last text block, so it must hold the whole answer.
251 // A run spanning several steps streamed its answer across them and gets it
252 // again whole; one whose answer streamed whole here gets just the
253 // signature, onto that same block.
254 const whole = run.steps > 0 && !shown.includes(answer)
255 const extra = whole ? report : report.slice(answer.length)
256 const into = !whole && lastKind === 'text' ? blocks - 1 : blocks
257 const tail: TurnStepChunk[] = extra ? [{ kind: 'text', index: Math.max(into, 0), text: extra }] : []
258 return yield* finish(e, agentId, shown + extra, report, Math.max(into, 0) + (extra ? 1 : 0), handback, tail, takeUsage(run))
259 })
260
261 /**
262 * Ends a step on the answer: plain text, or text then the HANDBACK_TOOL call
263 * that delivers it where the session runs the hand-back contract.
264 */
265 async function* finish(
266 e: { turnId: string; index: number },
267 agentId: string,
268 shown: string,
269 report: string,
270 blocks: number,
271 handback: boolean,
272 chunks: TurnStepChunk[],
273 usage: TurnUsage | null = null,
274 ) {
275 lastAnswers.set(agentId, report)
276 if (!handback) {
277 return yield* respond(e, [...chunks, { kind: 'stop', stopReason: 'end_turn', usage }], shown, [])
278 }
279 delivered.set(agentId, 'Report delivered.')
280 const input = { message: report }
281 const id = `toolu_aw_${idOf(agentId)}_${HANDBACK_TOOL}_${consumed.get(agentId) ?? 0}`
282 return yield* respond(e, [
283 ...chunks,
284 { kind: 'tool', index: blocks, id, name: HANDBACK_TOOL },
285 { kind: 'input', index: blocks, json: JSON.stringify(input) },
286 { kind: 'stop', stopReason: 'tool_use', usage },
287 ], shown, [{ name: HANDBACK_TOOL, input }])
288 }
289}
290
291/** An agent id as a tool-call id may spell it. */
292export function idOf(agentId: string) {
293 return agentId.replace(/[^A-Za-z0-9_]/g, '_')
294}
295
296/**
297 * The tokens a run used since the last step reported them, as this step's
298 * usage under the model the daemon named; null before it named one.
299 */
300export function takeUsage(run: Run | undefined): TurnUsage | null {
301 if (!run?.model) return null
302 const { i, o, cr, cw } = run.usage
303 run.usage = { i: 0, o: 0, cr: 0, cw: 0 }
304 return {
305 model: run.model,
306 input_tokens: i,
307 output_tokens: o,
308 cache_read_input_tokens: cr,
309 cache_creation_input_tokens: cw,
310 }
311}
312
313/** Whether a spawn names this plugin's type, plain or plugin-qualified. */
314export function isAw(subagentType: string) {
315 return subagentType === 'aw' || subagentType.endsWith(':aw')
316}
317
318/**
319 * Takes the `aw-harness:` / `aw-backend:` / `aw-model:` / `aw-agent:` lines out
320 * of a prompt: the way a caller routes a spawn, since the Agent tool's own
321 * `model` names Claude models only. Only the first line of each kind counts.
322 */
323export function routeOf(text: string): { prompt: string; route: Route } {
324 const route: Route = {}
325 let prompt = text
326 const keys = [['harness', 'harness'], ['backend', 'backend'], ['model', 'model'], ['agent', 'agent']] as const
327 for (const [word, field] of keys) {
328 const line = prompt.match(new RegExp(`^[ \\t]*aw-${word}:[ \\t]*([A-Za-z0-9_.:\\/@+-]+)[ \\t]*$`, 'im'))
329 if (!line) continue
330 route[field] = line[1]
331 prompt = prompt.replace(line[0], '')
332 }
333 return { prompt: prompt.trim(), route }
334}
335
336/**
337 * Yields a step's chunks and returns the result the engine expects of it.
338 *
339 * @param e the step being answered
340 * @param chunks what the person watches stream, in order
341 * @param answer the step's visible text
342 * @param toolUses the tool calls the chunks made
343 */
344export async function* respond(
345 e: { turnId: string; index: number },
346 chunks: TurnStepChunk[],
347 answer: string,
348 toolUses: TurnStepResult['toolUses'],
349): AsyncGenerator<TurnStepChunk, TurnStepResult> {
350 for (const chunk of chunks) yield chunk
351 const stop = chunks.at(-1)
352 const stopReason = stop?.kind === 'stop' ? stop.stopReason : 'end_turn'
353 const usage = stop?.kind === 'stop' ? stop.usage : null
354 return { turnId: e.turnId, index: e.index, answer, toolUses, stopReason, usage }
355}
356
357/** Reads a text file, or '' when it is absent or unreadable. */
358export async function readText($: EngineInterface, path: string): Promise<string> {
359 try {
360 const got: unknown = await $.fs.read(path)
361 if (typeof got === 'string') return got
362 const text = (got as { text?: unknown } | undefined)?.text
363 return typeof text === 'string' ? text : ''
364 } catch {
365 return ''
366 }
367}
368
369/**
370 * Finds the daemon: `AITHER_HARNESS_HOST`/`PORT` and the bearer file
371 * `AITHER_HARNESS_TOKEN_FILE`, the same three names every other daemon client
372 * reads, defaulting to 127.0.0.1:8362 and ~/.aither/harness_token. The token is
373 * read on every call rather than cached, since the daemon rewrites it.
374 */
375export async function daemonOf($: EngineInterface): Promise<Daemon> {
376 const home = ((await $.env.get('USERPROFILE')) || (await $.env.get('HOME')) || '').replace(/\\/g, '/')
377 const host = (await $.env.get('AITHER_HARNESS_HOST')) || '127.0.0.1'
378 const port = (await $.env.get('AITHER_HARNESS_PORT')) || '8362'
379 const file = (await $.env.get('AITHER_HARNESS_TOKEN_FILE')) || `${home}/.aither/harness_token`
380 const token = (await readText($, file)).trim()
381 return { base: `http://${host}:${port}`, token, home }
382}
383
384/**
385 * One call to the daemon. Never throws: a failure comes back as status 0 with
386 * the reason, so the caller can end the step on it as text.
387 */
388export async function call(
389 $: EngineInterface,
390 daemon: Daemon,
391 method: string,
392 path: string,
393 body?: unknown,
394): Promise<{ status: number; json: any; detail: string }> {
395 try {
396 const got = await $.http.fetch(`${daemon.base}${path}`, {
397 method,
398 headers: {
399 Authorization: `Bearer ${daemon.token}`,
400 ...(body === undefined ? {} : { 'Content-Type': 'application/json' }),
401 },
402 ...(body === undefined ? {} : { body: JSON.stringify(body) }),
403 })
404 let json: any
405 try {
406 json = JSON.parse(got.text)
407 } catch {
408 json = undefined
409 }
410 const detail = typeof json?.detail === 'string' ? json.detail : got.ok ? '' : got.text.slice(0, 400)
411 return { status: got.status, json, detail }
412 } catch (err) {
413 return { status: 0, json: undefined, detail: String(err) }
414 }
415}
416
417/**
418 * The spawn defaults from `~/.aither/mods.json` (`{ "aw": { "harness": … } }`),
419 * the file `awsettings --domain mods` syncs. Absent or malformed, no defaults.
420 */
421export async function defaultsOf($: EngineInterface, daemon: Daemon): Promise<Route & { permission_mode?: string }> {
422 try {
423 const aw = JSON.parse(await readText($, `${daemon.home}/.aither/mods.json`))?.aw
424 return aw !== null && typeof aw === 'object' ? aw : {}
425 } catch {
426 return {}
427 }
428}
429
430/**
431 * Opens (or reuses) the agent's daemon session and submits one turn to it.
432 *
433 * @param prompt what to send the harness, routing lines still in it
434 * @param id this run's number for the agent
435 * @param handback whether the answer must be handed back
436 * @param open the session a previous turn of this agent opened, if any
437 */
438export async function start(
439 $: EngineInterface,
440 agentId: string,
441 prompt: string,
442 id: number,
443 handback: boolean,
444 open: { id: string; harness: string } | undefined,
445): Promise<{ run: Run } | { failure: string }> {
446 const daemon = await daemonOf($)
447 if (!daemon.token) {
448 return { failure: 'awsh: no harness token found, so the daemon cannot be reached. Start it once (`adk harness serve`) - it writes ~/.aither/harness_token.' }
449 }
450 // A Claude Code the mod started can start the mod again; the daemon tells
451 // each one how deep it is, and the chain stops here rather than at a bill.
452 const depth = Number((await $.env.get('AITHER_AW_DEPTH')) || 0)
453 if (depth >= MAX_DEPTH) {
454 return { failure: `awsh: refused - this Claude Code session is already ${depth} aw spawns deep (the limit is ${MAX_DEPTH}). Do the work here instead of delegating it again.` }
455 }
456 const asked = routeOf(prompt)
457 let session = open
458 // A follow-up that names another harness gets a new session; one that names
459 // none continues the conversation it is part of.
460 if (session !== undefined && asked.route.harness !== undefined && asked.route.harness !== session.harness) {
461 await call($, daemon, 'DELETE', `/sessions/${session.id}`)
462 session = undefined
463 }
464 if (session === undefined) {
465 const defaults = await defaultsOf($, daemon)
466 const harness = asked.route.harness ?? defaults.harness ?? FALLBACK_HARNESS
467 const made = await call($, daemon, 'POST', '/sessions', {
468 harness,
469 cwd: await $.session.cwd(),
470 model_profile: asked.route.backend ?? defaults.backend ?? '',
471 model: asked.route.model ?? defaults.model ?? '',
472 target: asked.route.agent ?? defaults.agent ?? '',
473 permission_mode: defaults.permission_mode ?? '',
474 title: `aw ${harness} (Claude Code subagent)`,
475 // The depth rides in the owner because that is the one free-text field a
476 // session create already carries to the child's environment.
477 owner: `claude-code:${agentId}@d${depth + 1}`,
478 })
479 if (made.status !== 200 || typeof made.json?.id !== 'string') {
480 const why = made.status === 0 ? `the daemon at ${daemon.base} did not answer (${made.detail})` : `${made.status} ${made.detail}`
481 return { failure: `awsh: could not open a ${harness} session: ${why}` }
482 }
483 session = { id: made.json.id, harness }
484 // A harness that cannot start says so within moments; submitting into a
485 // session still `starting` is refused with a 409 that names nothing.
486 for (let waited = 0; waited < START_WAIT_MS; waited += POLL_MS) {
487 const info = await call($, daemon, 'GET', `/sessions/${session.id}`)
488 if (info.json?.state !== 'starting') break
489 await $.clock.sleep(POLL_MS)
490 }
491 }
492 const before = await call($, daemon, 'GET', `/sessions/${session.id}/events?since=999999999`)
493 const sent = await call($, daemon, 'POST', `/sessions/${session.id}/submit`, { text: asked.prompt, submit: true })
494 if (sent.status !== 200) {
495 const events = await call($, daemon, 'GET', `/sessions/${session.id}/events?since=0`)
496 const errors = (events.json?.events ?? [])
497 .filter((ev: any) => ev.kind === 'error')
498 .map((ev: any) => ev.text)
499 .join('\n')
500 return { failure: `awsh: the ${session.harness} session refused the turn: ${sent.status} ${sent.detail}\n${errors}`.trim() }
501 }
502 return {
503 run: {
504 session: session.id,
505 harness: session.harness,
506 id,
507 since: Number(before.json?.last_seq ?? 0),
508 quiet: 0,
509 shown: '',
510 report: '',
511 ended: false,
512 failed: false,
513 handback,
514 steps: 0,
515 model: '',
516 noise: [],
517 usage: { i: 0, o: 0, cr: 0, cw: 0 },
518 },
519 }
520}
521
522/**
523 * Streams a run's new events as this step's chunks, until the turn ends or the
524 * step's `deadline` passes: text as text, thinking as thinking, and each tool
525 * the harness starts as a one-line `▸ tool: args` note in the text.
526 *
527 * Never throws: see `start`. An aborted step interrupts the harness, since the
528 * daemon's session would otherwise carry on with a turn nobody is reading.
529 *
530 * @param deadline when this step must stop streaming, in `$.clock` ms
531 * @param signal the step's abort signal
532 * @returns the text this step showed, how many content blocks it used, and
533 * the kind of the last one
534 */
535export async function* pump(
536 $: EngineInterface,
537 run: Run,
538 deadline: number,
539 signal: AbortSignal,
540): AsyncGenerator<TurnStepChunk, { shown: string; blocks: number; lastKind?: 'text' | 'thinking' }> {
541 let shown = ''
542 let index = -1
543 let kind: 'text' | 'thinking' | undefined
544 let failure = ''
545 const daemon = await daemonOf($)
546
547 try {
548 while (!run.ended && !signal.aborted && (await $.clock.now()) < deadline) {
549 const got = await call($, daemon, 'GET', `/sessions/${run.session}/events?since=${run.since}`)
550 if (got.status !== 200) {
551 failure = `awsh: lost the ${run.harness} session: ${got.status} ${got.detail}`
552 break
553 }
554 const events: any[] = got.json?.events ?? []
555 for (const event of events) {
556 run.since = Math.max(run.since, Number(event.seq ?? 0))
557 let pieceKind: 'text' | 'thinking' = 'text'
558 let text = ''
559 if (event.kind === 'text.delta') {
560 text = event.text ?? ''
561 run.report += text
562 } else if (event.kind === 'thinking.delta') {
563 pieceKind = 'thinking'
564 text = event.text ?? ''
565 } else if (event.kind === 'tool.call') {
566 const args = JSON.stringify(event.data?.input ?? {})
567 text = `\n▸ ${event.tool}: ${args.length > 100 ? `${args.slice(0, 100)}…` : args}\n`
568 } else if (event.kind === 'error') {
569 text = `\n[${run.harness} error] ${event.text ?? ''}\n`
570 run.report += text
571 } else if (event.kind === 'usage') {
572 const u = event.data?.usage ?? {}
573 run.usage.i += Number(u.input_tokens ?? u.input ?? 0)
574 run.usage.o += Number(u.output_tokens ?? u.output ?? 0)
575 run.usage.cr += Number(u.cache_read_input_tokens ?? 0)
576 run.usage.cw += Number(u.cache_creation_input_tokens ?? 0)
577 } else if (event.kind === 'notice' || event.kind === 'raw') {
578 // Never shown as they arrive (a CLI's stderr is mostly noise), but a
579 // failed turn's reason is only ever here, and so is the model name
580 // of a harness the daemon cannot ask (`> build · provider/model`).
581 const line = String(event.text ?? '').replace(/\u001b\[[0-9;]*m/g, '').trim()
582 const named = line.match(/^>\s*\S+\s+·\s+(\S+)$/)
583 if (named) run.model = named[1]
584 else if (line) run.noise = [...run.noise, line].slice(-NOISE_LINES)
585 } else if (event.kind === 'turn.completed') {
586 // A harness that streams nothing still carries its answer here.
587 if (!run.report.trim() && event.text) {
588 text = event.text
589 run.report = text
590 } else if (event.data?.is_error && !run.report.trim()) {
591 failure = `awsh: the ${run.harness} turn failed (exit ${event.data?.exit_code ?? '?'}):\n${[...new Set(run.noise)].join('\n')}`
592 }
593 run.ended = true
594 } else if (event.kind === 'session.exited') {
595 run.ended = true
596 }
597 if (!text) continue
598 if (pieceKind !== kind) {
599 kind = pieceKind
600 index += 1
601 }
602 if (pieceKind === 'text') shown += text
603 yield { kind: pieceKind, index, text }
604 }
605 // A session that stopped working without a closing event still ended.
606 const working = WORKING_STATES.includes(got.json?.state)
607 run.quiet = events.length === 0 && !working ? run.quiet + 1 : 0
608 if (run.quiet >= 3) run.ended = true
609 if (!run.ended) await $.clock.sleep(STREAM_POLL_MS)
610 }
611 } catch (err) {
612 failure = `awsh: the ${run.harness} run failed: ${String(err)}`
613 }
614
615 if (signal.aborted && !run.ended) {
616 run.ended = true
617 await call($, daemon, 'POST', `/sessions/${run.session}/interrupt`)
618 }
619 if (run.ended && !failure && !(run.shown + shown).trim()) {
620 failure = `awsh: the ${run.harness} session ended with no answer.`
621 }
622 if (failure) {
623 run.ended = true
624 run.failed = true
625 shown += failure
626 run.report = failure
627 index += 1
628 kind = 'text'
629 yield { kind: 'text', index, text: failure }
630 }
631 if (run.ended) {
632 const info = await call($, daemon, 'GET', `/sessions/${run.session}`)
633 const binding = info.json?.model_binding
634 run.model = info.json?.reported_model || binding?.model || binding?.profile || run.model
635 }
636 run.shown += shown
637 return { shown, blocks: index + 1, lastKind: kind }
638}
639
640/**
641 * Decides what to send the harness this step, and whether its answer must be
642 * handed back through HANDBACK_TOOL.
643 *
644 * The first step sends the spawn's prompt. A later one waits for a person's
645 * message past the first `already` in `agentId`'s own transcript and sends the
646 * newest; the engine can start the step before it writes the message there,
647 * so this polls briefly.
648 *
649 * @param seed the prompt the spawn was given
650 * @param already how many of the person's messages the harness has been sent
651 */
652export async function promptOf($: EngineInterface, agentId: string, seed: string, already: number) {
653 for (let waited = 0; ; waited += POLL_MS) {
654 const { texts, handback, bounced } = await transcriptOf($, agentId)
655 if (already === 0) return { prompt: seed, count: Math.max(texts.length, 1), handback, bounced: false }
656 if (texts.length > already) return { prompt: texts.at(-1), count: texts.length, handback, bounced: false }
657 if (bounced) return { prompt: undefined, count: already, handback: true, bounced }
658 if (waited >= FOLLOW_UP_WAIT_MS) return { prompt: undefined, count: already, handback, bounced }
659 await $.clock.sleep(POLL_MS)
660 }
661}
662
663/**
664 * Where Claude Code keeps `agentId`'s transcript:
665 * `~/.claude/projects/<project>/<session id>/subagents/agent-<id>.jsonl`.
666 *
667 * Leans on Claude Code's on-disk layout, not an API: `$.session.messages()`
668 * reads the main loop's transcript only. The project directory is the launch
669 * directory with every non-alphanumeric turned into `-`; when that guess
670 * misses, every project directory is tried for this session's id.
671 */
672export async function transcriptPathOf($: EngineInterface, agentId: string): Promise<string | undefined> {
673 const { home } = await daemonOf($)
674 const projects = `${home}/.claude/projects`
675 const sessionId = await $.session.id()
676 const tail = `${sessionId}/subagents/agent-${agentId}.jsonl`
677 const guess = `${projects}/${(await $.session.cwd()).replace(/[^A-Za-z0-9]/g, '-')}/${tail}`
678 if (await $.fs.exists(guess).catch(() => false)) return guess
679 const listed: unknown = await $.fs.list(projects).catch(() => [])
680 for (const entry of Array.isArray(listed) ? listed : []) {
681 const name = typeof entry === 'string' ? entry : (entry as { name?: string })?.name
682 if (!name) continue
683 const path = `${projects}/${name}/${tail}`
684 if (await $.fs.exists(path).catch(() => false)) return path
685 }
686 return undefined
687}
688
689/**
690 * Reads `agentId`'s transcript: the text of every message the person (or the
691 * spawn) sent it, oldest first; whether the engine asked it to hand its
692 * report back through HANDBACK_TOOL; and whether the newest message is the
693 * engine bouncing a report that was not.
694 */
695export async function transcriptOf($: EngineInterface, agentId: string) {
696 const path = await transcriptPathOf($, agentId)
697 const raw = path === undefined ? '' : await readText($, path)
698 const all = raw
699 .split('\n')
700 .flatMap((line) => {
701 try {
702 return [JSON.parse(line)]
703 } catch {
704 return [] // the line Claude Code is still writing
705 }
706 })
707 .filter((entry) => entry.type === 'user')
708 .map((entry) => textOf(entry.message?.content))
709 .filter(Boolean)
710 return {
711 texts: all.filter((text) => !ENGINE_PREFIXES.some((prefix) => text.startsWith(prefix))),
712 handback: all.some((text) => text.includes(HANDBACK_TOOL)),
713 bounced: all.at(-1)?.startsWith('[handback-send-enforce]') ?? false,
714 }
715}
716
717/** Joins a message's text: a plain string, or the text blocks of a block list. */
718export function textOf(content: unknown): string {
719 if (typeof content === 'string') return content.trim()
720 if (!Array.isArray(content)) return ''
721 return content
722 .filter((block) => block?.type === 'text')
723 .map((block) => block.text)
724 .join('\n')
725 .trim()
726}
727